User Tools

Site Tools


provenance:design:ownership_resolution

This is an old revision of the document!


Provenance: Design:Ownership resolution

Working notes behind Ownership resolution. Every query with its population and denominator, the report script and its unedited output, what was folded and what the fold could not reach, the quotes checked against the source, the external sources verified and rejected, and the judgement calls. Corpus-wide selection and extraction caveats are on Corpus and are not restated here.

This page is a carve-out, and most of its audit trail lives elsewhere on purpose. The content page was split out of webXray on 2026-09-11. The scripts that produce its live-database figures, their unedited output, the 175-row adjudication table and the inter-rater material were published in 2026-08 and 2026-09 under the webXray ids and keep those ids, because they are cited from the published record and moving them would break links for no gain:

What Where
The three-list comparison scripts (owner_dbs.py, owner_adjudication.py), their unedited output, the fold residues, the Wayback log, the 2026-08-17 and 2026-09-05 run tables webxray
owner_sample.py, owner_random_sample.py, the full 175-row adjudication with every source, the citation re-fetch, the inter-rater draw and kappa random_sample
The corpus population of webXray itself (the 15 papers, the hand role map), report_webxray.mjs webxray
This page the split itself, the new corpus script and its output, the reproduction record of 2026-09-11, the defect that run found, the reviewer log

No ~~DISCUSSION~~ block: comments belong on the content page. Citations use the same {[citekey]} keys and the same shared Bibliography; this page adds no bibliography entries of its own.

The run

Date 2026-09-11
Corpus at the time data/extract/run1/extractions.jsonl, 5,859 papers, 7 venues, 2010–2026. 5,855 have a readable paper.cols.txt.
What was done Created Ownership resolution; rewrote webXray around what was left; repointed Crawler; added a row to Design; turned the Roadmap row blue.
New code scripts/report_ownership_resolution.mjs (346 lines), committed with its output.
Code changed scripts/report_webxray.mjs — a denominator defect, section C below.
Models Opus 5 wrote the pages and the script. Review layer: three focused Sonnet passes and one Fable generic pass, section I.
Accidental exposure or mistakes caught none in this sitting beyond the defect in C, which was found by the new script disagreeing with the old one.

Scope decision: why a separate page, when the last sitting decided against one

This reverses a recorded decision, so the reversal is recorded too. webxray §“Scope decision” says, on 2026-08-17:

A separate design:ownership_resolution page with webXray as a stub. Rejected: it would leave the wiki's existing red link pointing at a stub, and the material does not split cleanly — webXray's frozen list is the best available illustration of what goes wrong.

Both halves of that reasoning have since stopped holding:

  • “webXray as a stub” was the wrong alternative to compare against. There is enough webXray-specific material for a real page — the three-licence history across four states, the availability census, the domain_owners.json schema, the frozen 2016 PSL, the Wayback timeline, and 15 corpus papers with a hand role map. What is left after the carve-out is 27 KB, not a stub.
  • The red-link problem inverted. On 2026-09-07 the Roadmap queued design:ownership_resolution as a promised page and scripts/sitemap.mjs began gating on it, so from that date not writing it was the dangling promise. The id was fixed by that decision and is used unchanged.
  • The discoverability argument is the item's whole point and was never answered. A reader asking “how do I attribute a third-party domain to a company?” does not search for a dead tool. Both earlier sittings recorded this as the right thing to do; neither did it.

Alternatives considered this sitting and rejected:

Alternative Why not
Leave it, and add a redirect or a prominent pointer on webXray A pointer does not fix search, does not fix the namespace (programming: is instruments; this is a design decision), and leaves the currency claim — which source to use now — filed under the historical one.
Broaden Requests instead That page answers “is this request a tracker?”, a different question with different tooling (filter lists, not owner databases). The two pages cross-link.
Broaden IP classification It is the network-layer sibling (AS-to-organisation) and is already large; the domain layer has different sources, different failure modes and different licences. Cross-linked both ways instead.
Move the random_sample code appendix under provenance:design: too Rejected. Its id is cited from the published content page and from the webXray provenance page; the gain is cosmetic and the cost is broken links. Recorded here so the next sitting does not “tidy” it.
Re-derive the coverage census against a 2026-09-11 snapshot Rejected. The 28-row adjudication is pinned to the 2026-08-17 frame; re-deriving would orphan it. The figures moved unchanged, with their frame dates, and the page carries the two-frames box. This is a carve-out, not a refresh — see the drift check in section B.

The carve-out map

What moved, what stayed, and what had to be repointed. Nothing was re-derived in the move.

Section on the old page Went to
The tool: architecture, and why you cannot install it stayed on webXray
The ownership database (schema, tree, frozen PSL) stayed — it describes webXray's file
How it compares to Tracker Radar and Disconnect moved, as The live sources, compared
Coverage: how much of the third-party surface moved, with the three traps split into their own sub-section
Two lists disagree: error or different question? moved, split: the two axes were promoted to the top of the new page as The question has two axes, the agreement tables became When two lists disagree
Reading all three at once (the code) moved
Which list is right, when they disagree? moved
How often is each list right? A random sample moved, with the inter-rater material given its own sub-section
Where these figures come from, and how to redo them moved, folded into the new page's methodology section
Choosing a resolution source now moved
Assembling the pipeline moved
What to report in a paper moved
Use in publications: the 15 webXray papers, role table stayed
Use in publications: the 136-paper sweep, the Tracker Radar year table moved — they are landscape figures, not webXray figures
Wayback timeline stayed, and was promoted from a methodology bullet to its own section; it was 900 words inside a bullet list

Day-one drift, found and fixed rather than left:

Where What it said after the move Fixed to
webXray intro bullet 2 quoted 69.9% / 93.9% / 2.1% / 17.2% / 7.0% inline as if the page still carried the measurement one sentence keeping 69.9% and 2.1% with a link, and an explicit “nothing on this page restates them”
webXray, The ownership database “§\”Two lists disagree\“ below measures what happens when you forget that” — a same-page reference to a section that had left repointed to [[Design:Ownership resolution#When two lists disagree]], with the finding (root resolution makes agreement worse) stated inline so the sentence still says something
webXray, Related pages listed Requests, Cookies, IP classification — all of which were there for the ownership material new page took those; webXray's list now leads with the new page and adds Policies for policyXray
webXray methodology listed nine owner_* scripts it no longer publishes figures from one bullet saying where they went and that their provenance ids are unchanged
Crawler, webXray bullet “see webXray for what survives of it: the ownership database, and how it compares to Tracker Radar and Disconnect today” — that comparison had moved repointed to the new page; see section H

A site-wide inbound-link sweep found four more, on pages nobody would have thought to check. Grepping every cached page for a link to programming:crawler:webxray rather than only the pages this sitting edited:

Page What it said Fixed to
Requests a Related-pages bullet linking the webXray page under the link text “webXray and domain-to-company ownership” — “once a request is flagged, this is how to answer whose it is” repointed to Ownership resolution, with a note that it moved
Filter lists the same bullet, near-verbatim repointed the same way
Legal enforcement “See webXray for the ownership databases and how much they disagree” — in the bullet on identifying a controller under Reg. 2025/2518 repointed, and extended to name the accuracy rates, which is what that bullet actually needs
Programming (namespace index) the webXray row read “Domain-to-company ownership lists, and what remains of the tool” — a description of the page, now false rewritten to describe the tool page and point the ownership question at the new page

Two of those four used the link text “webXray and domain-to-company ownership”, i.e. they were already treating the webXray page as the ownership page — which is the item's complaint, restated by the wiki itself. A carve-out's drift is not confined to the pages you edited; sweep for inbound links before calling it done. One inbound reference was deliberately left: provenance:design's generated page inventory records a size and a description as of an earlier date, and it is a dated snapshot rather than a live pointer.

A. Corpus queries

All over data/extract/run1/extractions.jsonl (5,859 papers) and data/fulltext/<year>/<venue>/<slug>/paper.cols.txt, whitespace collapsed before matching so a term broken across a PDF column boundary still matches. Every count is a paper count. The script is scripts/report_ownership_resolution.mjs, published in full in section J with its unedited output in section K.

# Question Population / denominator Answer
A1 How many papers have a readable full-text rendering at all? all 5,859 5,855 (99.9%). The four without one can only ever be sweep misses, so every count below is a floor.
A1b …and how many of those are damaged? the 5,855 0 empty, 0 under 2 KB, but 90 carry a NUL byte. The script reads them as JavaScript strings and regexes them, so NULs do not affect it. A grep-based sweep would have silently skipped all 90grep treats a NUL-carrying file as binary and, depending on the build, prints nothing and exits non-zero. None of the 136 union members is one of the 90, checked explicitly, so no published count depends on this — but the next person writing a sweep should not use grep.
A2 How many papers name each ownership resource? all 5,859 PSL 101, Disconnect (list sense) 74, Tracker Radar 32, WhoTracks.me 23, Crunchbase 23, webXray 15
A3 How many of those are inside the crawling population? crawled = 1,120 (crawlConfig present, or studyTypes includes automated-web-crawl) PSL 46, Disconnect 65, Tracker Radar 28, WhoTracks.me 19, Crunchbase 11, webXray 13
A4 How many name any of the five ownership resources? all 5,859 136 (2.3%)
A5 …and how many of those 136 are in the 1,120? crawled = 1,120 107, i.e. 9.6% of that population. Not 12.1% — see section C.
A6 Does adding the PSL change the union? all 5,859 136 → 221. The PSL is deliberately excluded: it answers “what is the registrable domain”, not “whose is it”.
A7 Is ownership resolution growing? each year's own paper count rose to ~3.5% in 2020–2022, has sat at 2.5–2.7% since. Full table on the page; 2010–2015 contributes one paper in total and is omitted rather than padded with zeros.
A8 Which venues? each venue's own paper count PETS 9.6%, IMC 2.7%, TheWebConf 2.3%, IEEE S&P 1.4%, USENIX 1.3%, CCS 1.3%, NDSS 1.1%
A9 What does the extraction schema see, against the full text? all 5,859 schema 89, full text 136, both 82, full-text-only 54, schema-only 7 — residue printed in full below
A10 Of the 15 webXray papers, how many say which version? the 8 that used the crawler or the list 2 say anything; 1 names a commit ([1Matte, Célestin; Bielova, Nataliia; Santos, Cristiana Teixeira (2020): "Do Cookie Banners Respect my Choice? Measuring Legal Compliance of Banners from IAB Europe's Transparency and Consent Framework", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]). Inherited from report_webxray.mjs, re-run and unchanged.
A11 What does “Tracker Radar” mean in the 32 papers that name it? the 32, hand-mapped 11 ownership dataset, 9 tracker/category database, 9 the Collector crawler, 2 citation, 1 compared. Inherited, re-run and unchanged.

A.12 Probe widths: every probe run at two widths, and both printed

A narrowing probe that returns more hits than the loose one, or hits the loose one does not contain, is a broken probe. The script asserts both and throws rather than printing. Both widths are on the record because the gap is itself the claim:

Resource                 Tight probe  Loose probe  Status
-----------------------  -----------  -----------  ----------------------------------------------------------------
webXray                  15           233          hand-verified: all 15, in report_webxray.mjs
Tracker Radar            32           108          hand-verified: all 32, in report_webxray.mjs
WhoTracks.me             23           124          upper bound
Crunchbase               23           23           upper bound
Disconnect (list sense)  74           700          upper bound; the loose width is the reason this one is tightened
Public Suffix List       101          157          upper bound

The patterns themselves are in a code block and not in a table, because a DokuWiki table cell cannot hold a pipe and every one of these regexes is an alternation. \| is not an escape and wrapping the pattern in DokuWiki's nowiki delimiters does not help either: the cell splits regardless.

webXray             tight  /webx[\s-]?ray/i
                    loose  /webx[\s-]?ray|libert/i
                    15 -> 233. "Libert" is a common surname and a French word.
                    The loose width is a sanity bound, not a candidate set.

Tracker Radar       tight  /tracker[\s.-]?radar/i
                    loose  /tracker[\s.-]?radar|duckduckgo/i
                    32 -> 108. Most DuckDuckGo mentions are the search engine
                    or the browser, not the dataset.

WhoTracks.me        tight  /whotracks/i
                    loose  /whotracks|who ?tracks ?\.? ?me|ghostery/i
                    23 -> 124. Ghostery is mostly the extension.

Crunchbase          tight  /crunchbase/i
                    loose  /crunchbase|crunch base/i
                    23 -> 23. NO GAP: nobody spells it with a space. This
                    row's count is not an artefact of the pattern.

Disconnect          tight  /disconnect'?s? (entit|list|block|black|tracking)|disconnect\.me|entities\.json/i
                    loose  /disconnect/i
                    74 -> 700. "Disconnect" is an ordinary English word and a
                    CSP-literature technical term. This is the probe that
                    needed tightening, and the published row is the tight one.

Public Suffix List  tight  /public suffix/i
                    loose  /public suffix|etld\+?1|effective top-?level domain/i
                    101 -> 157. The extra 56 are papers that use the concept
                    without naming the list.

What this does not establish. A tight probe with no gap is not a probe with full recall — it is a probe whose count is stable under widening in the one direction tried. A paper that resolves domain ownership from a source none of these five names, or names none of them at all, is invisible to every width. See section F.

A.13 The schema-only residue, printed in full

Seven papers whose tools[] / classification[] / otherToolsMentioned[] name one of the five resources but whose full text does not match the tight sweep. Printed rather than dropped, because an invisible residue is a residue nobody looks at:

  IEEE-SP/2024/to-auth-or-not-to-auth-a-comparative-analysis-of-the-pre-and-post-login-security
      named: Disconnect Tracker Protection List
  PETS/2022/on-dark-patterns-and-manipulation-of-website-publishers-by-cmps
      named: Disconnect | Disconnect tracking filter list
  PETS/2024/generalizable-active-privacy-choice-designing-a-graphical-user-interface-for-glo
      named: Disconnect Tracker Protection lists
  PETS/2024/support-personas-a-concept-for-tailored-support-of-users-of-privacy-enhancing-te
      named: Disconnect
  USENIX/2019/canvas-fast-and-inexpensive-automotive-network-mapping
      named: physical ECU access and disconnection
  WWW/2018/the-cost-of-digital-advertisement-comparing-user-and-advertiser-views
      named: Disconnect | Disconnect blacklist
  WWW/2023/who-funds-misinformation-a-systematic-analysis-of-the-ad-related-profit-routines
      named: Disconnect

Six of the seven are Disconnect, and they are the tightened pattern's misses: the extractor wrote “Disconnect Tracker Protection List” or bare “Disconnect” into a field, where the paper's own prose says something the list-sense pattern does not match. The seventh is a homonym the extractor invented — “physical ECU access and disconnection” in an automotive CAN-bus paper is not the Disconnect list. The tight pattern's cost is therefore about five papers, all Disconnect, all in the direction of undercounting, and the page's Disconnect row should be read as 74 rather than 74-to-79 only because the union is an upper bound on use anyway. The alternative — publishing the loose 700 — is not a trade worth making.

A.14 Denominators used on the page, stated once

Figure on the page Denominator Never
the six per-resource rows 5,859 for the “share of corpus” column; 1,120 for the “share of crawled” column, with the numerator restricted to that population 5,859 for both
136 / 2.3% 5,859
107 / 9.6% 1,120, numerator restricted to it 136 ÷ 1,120
per-year shares that year's own paper count 5,859
per-venue shares that venue's own paper count 5,859, and never the count alone — the venues differ in size by a factor of three
15 / 13 webXray papers 5,859 and 1,120 respectively
everything about the three lists the domain frames: 32,369 (2026-08-17) or 32,337 (2026-09-05) registrable domains 5,859; these are not corpus figures at all

B. Reproduction record, 2026-09-11

Every script whose figures the new page inherits was re-run before the move, against the same pinned inputs, and diffed against its committed output. A page whose report script no longer runs is a page whose numbers cannot be refreshed.

Script Invocation Result
report_webxray.mjs node scripts/report_webxray.mjs –quotes byte-identical to the committed report_webxray-output.txt before the fix in section C; regenerated after it
owner_dbs.py python3 scripts/owner_dbs.py –cache out/webxray/cache –disagreements 25 reproduces the 2026-08-17 frame exactly: 32,369 domains; 669 / 5,581 / 2,268; 58.6% / 84.3% / 80.5%; 15,651 merged rows; 19 non-hostname rows dropped; 6,941 ICANN and 3,290 PRIVATE rules; 782 of 827 owner names untouched by the fold, 16 of the 45 changed losing a legal suffix
owner_random_sample.py python3 scripts/owner_random_sample.py –sample out/owner_sample.json –rows out/adj_rows_scored.json –table –wiki byte-identical to out/owner_random_sample-output.txt
owner_selfname_census.py python3 scripts/owner_selfname_census.py reproduces 0.8% / 0.7% / 3.2% (25 of 3,215; 250 of 38,368; 250 of 7,850)
report_ownership_resolution.mjs node scripts/report_ownership_resolution.mjs new; section K

A drift check the carve-out made cheap. The new script re-derives the sweep counts from patterns written out independently and then cross-checks nine figures against report_webxray-output.txt. All nine agree. That cross-check is inside the script and throws, so a future run cannot publish two pages that disagree with each other:

Figure                   this script  report_webxray.mjs  verdict
-----------------------  -----------  ------------------  -------
webXray                  15           15                  agree
Tracker Radar            32           32                  agree
WhoTracks.me             23           23                  agree
Crunchbase               23           23                  agree
Disconnect (list sense)  74           74                  agree
Public Suffix List       101          101                 agree
union of the five        136          136                 agree
union ∩ crawled          107          107                 agree
crawled denominator      1120         1120                agree

C. A defect the cross-check found: a numerator from one population, a denominator from another

This is the one substantive correction of the sitting, and it was published for three and a half weeks.

report_webxray.mjs printed, at two places:

  ${pct(ftPapers.length, crawled.length)} of the ${crawled.length} that ran a crawl
  ${pct(anyOwner.length, crawled.length)} of the ${crawled.length} that crawled

ftPapers and anyOwner are sweeps over the whole corpus. Dividing either by crawled.length pairs a corpus-wide numerator with the crawling denominator: the share is of a population the numerator was not drawn from.

Published Correct Why the error is invisible
webXray papers, share of the 1,120 1.3% (15 ÷ 1,120) 1.2% (13 ÷ 1,120) 13 of the 15 are in fact crawling papers, so the two figures differ by one decimal and nothing looks wrong
any ownership resource, share of the 1,120 12.1% (136 ÷ 1,120) 9.6% (107 ÷ 1,120) a 2.5-point error on the more load-bearing figure

Both figures were on the live webXray page, and the 12.1% was one of the figures moving to the new page — which is how it was caught: the new script computed the intersection because it had no reason not to, and the cross-check refused to pass. No guard on the old page could have found it. The number-guard family checks that a figure on a page matches the script's output; the script's output was the wrong number, and the page matched it faithfully.

The fix adds a named helper so the mistake cannot recur silently:

const crawledKeys = new Set(crawled.map(key));
// A share of `crawled` needs a numerator drawn from `crawled`.
const inCrawled = (ks) => ks.filter((k) => crawledKeys.has(k));

report_webxray-output.txt was regenerated and both pages corrected in the same sitting. The webXray page carries a footnote at the changed figure rather than silently restating it.

D. Folding, and what it could not reach

This page's corpus figures need no name fold: every count is a regex over full text or a set membership, and there are no free-text names being aggregated. That is stated rather than assumed, because it is unusual for a page on this wiki.

The live-list figures do fold, and the fold and its residue are published on webxray §B.11. Restated here only in summary, because the content page quotes both halves:

  • Legal-form suffixes only (Inc, LLC, GmbH, S.A.S, …), plus punctuation normalisation. Deliberately no synonyms, so the “agree” columns are a lower bound on real agreement and the “neither” column an upper bound on real disagreement.
  • Residue: the fold leaves 782 of webXray's 827 owner names untouched (94.6%). Of the 45 it changes, only 16 lose a legal-form suffix; the other 29 change on punctuation alone — 56.com, AT&T, Ask.com, Bootstrap_China, Clearstream.TV, Cm_browser, Dictionary.com, Dun & Bradstreet, and the rest are in the published output.
  • What no fold reaches: “Google” against “Alphabet”, “Xandr” against “Microsoft Corporation”. Those are the granularity and vintage axes, and they are the page's subject rather than a normalisation failure. There is no principled synonym table for them, which is why the page reports the agreement figures as bounds and adjudicates a sample instead.

E. Quotes checked against the source

Checked on 2026-09-11 against data/fulltext/<year>/<venue>/<slug>/paper.cols.txt, whitespace collapsed and quotation marks and dashes normalised, with paper.norm.txt and paper.txt as fallbacks.

Paper Quote (head) Verdict
[2Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)] “the product of years of detective work” found in .cols
[2Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)] “because webxray's database of domain ownership primarily contains major ad networks…” NOT contiguous — see below
[3Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] “frequently miss connections among two hostnames” found
[3Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] “twitchcdn.net” found
[4Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] “preferring those that have been updated most recently” found
[4Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] “does not aim at providing full transparency on the organizations behind each domain” found
[1Matte, Célestin; Bielova, Nataliia; Santos, Cristiana Teixeira (2020): "Do Cookie Banners Respect my Choice? Measuring Legal Compliance of Banners from IAB Europe's Transparency and Consent Framework", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] “WebXRay commit 04c3c8e8” found
[5Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)] “Redacted For Privacy” / “Domains By Proxy” / “Whois Guard” found, and the surrounding sentence re-read — see below

The Libert coverage quote is real but is not one contiguous string in any rendering of the PDF. It is the page's longest quotation and it is load-bearing for the coverage section, so it was chased rather than dropped. In paper.cols.txt, paper.norm.txt and paper.txt alike, the two-column layout interleaves a table caption through the sentence:

...because webxray's database of Table 1: Third-Party Prevalence, SSL Use, and First-Party
domain ownership primarily contains major ad networks rather Disclosure than small clients,
and policyxray only searches for identified †Denotes Company has Consumer Services parties,
variability in the long-tail of trackers may not have an outsized effect on overall findings
related to disclosure. Nonetheless, Company % Tracked % SSL % Disclosed it is important to
point out that the number of parties being searched Google † 82.81 80.35 38.29 for is fewer
than the total number of parties present.

Every clause is verbatim and in this order; the interleaved strings are “Table 1: Third-Party Prevalence, SSL Use, and First-Party Disclosure”, “†Denotes Company has Consumer Services” and the table's own header and first row. The page quotes the sentence as the author wrote it and carries a footnote saying exactly this, so a reader running the same grep does not conclude the quote is invented. This is the class of defect a plain substring check reports as a miss and a careless run reports as a fabrication.

One claim was corrected by reading the source. The old page said [5Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)] “resolved organisations for 411 of 762 mobile-specific trackers using CrunchBase, webXray's list, TLS certificates and WHOIS in that order”. The paper's own sentence is narrower:

we successfully retrieved the organization information of 411 trackers from either the CrunchBase database, Tim Libert's library, or TLS certificate, and retrieved the organization […]

The 411 comes from the first three sources; WHOIS contributed a further 251 (“among those 251 trackers whose organization information was retrived from the WHOIS record”). The four sources were tried in that order, but the 411 is not a figure for all four. The new page states it as 411 from the first three and 251 from WHOIS. The consequence for the page's actual point is unchanged: the five registrar-privacy strings total 88, against the largest real company in the top-ten table at 48 (Adobe), and the paper says so itself — “these five organizations cannot represent the real organizations of those trackers”.

F. External sources: verified, and rejected

Repository and vendor state was re-fetched on 2026-09-11, not recalled. Training data is stale by construction for this kind of claim.

Claim on the page Primary source, fetched 2026-09-11 Verdict
Tracker Radar current; last commit on main 2026-08-28; newest release tag 2026.08.28 api.github.com/repos/duckduckgo/tracker-radar/commits?per_page=1a736b501, 2026-08-28T15:25:14Z; /releases2026.08.28, 2026.07.27, 2026.06.08, none flagged prerelease holds
Tracker Radar's pushed_at reads later than main because of unmerged branches pushed_at = 2026-09-02T21:21:31Z against main at 2026-08-28 holds, and the page now names both dates
Disconnect current; last commit 2026-08-28 commits?per_page=14b592c28, 2026-09-05T20:21:14Z moved. Page updated to 2026-09-05 and the table header re-dated to 2026-09-11
Ghostery trackerdb current; last commit 2026-09-01 commits?per_page=1aa82e6d9, 2026-09-01T14:40:47Z holds
Disconnect's headline “14,332 verified domains and entity mappings” disconnect.me/trackerprotection, HTTP 200, string present holds
github.com/timlib/webXray is HTTP 404; the account has 0 public repositories 404; api.github.com/users/timlib → 200, public_repos: 0, company: webXray.ai holds
github.com/timlib/webXray_Domain_Owner_List is HTTP 404 404 holds
No PyPI package pypi.org/pypi/webxray/json → 404 holds
thezedwards/webXray last pushed 2021-03-04, PolyForm Strict pushed_at 2021-03-04T23:52:47Z; GitHub reports NOASSERTION holds
peterjoles/webXray is MIT, last pushed 2023-03-12, 36 commits ahead spdx_id: MIT, pushed_at 2023-03-12T20:15:26Z holds
webxray.org is a placeholder reading “Public interest projects for the interested public.” fetched with a browser User-Agent: HTTP 200, that exact line and nothing else holds
webxray.ai is a live commercial product HTTP 200 holds

Sources rejected, recorded so the next run does not re-add them:

Rejected Why
GitHub's license.spdx_id field for Ghostery trackerdb and Disconnect reports NOASSERTION for both, which is a statement about GitHub's detector and not about the licence. The licences were read from package.json and LICENSE respectively.
GitHub's pushed_at as “last activity” it counts unmerged automation branches. For Tracker Radar it reads four days later than main. Read the branch.
Any figure for Ghostery trackerdb coverage or accuracy none was measured. Its .eno + patterns shape needs a parser nobody could check against the other three. The page says this is an omission rather than a judgement, and lists it as an open question.
Vendor marketing counts as file sizes Disconnect's “14,332 verified domains” is a site headline; entities.json holds 7,850 domains. The page quotes both and says which is which.
Anything recalled rather than fetched about these repositories training data predates all four dates above.

G. What could not be established

  • Ghostery trackerdb is unmeasured. It is a live fourth option and the page gives numbers for two of the three live lists. Closing this needs a parser for the .eno + patterns model whose choices are checkable against the other three — a sitting's work, not a footnote. Filed as an open question on the content page.
  • Recall of the candidate set is unknown. The 136 is a union of six patterns over full text. A paper that resolves domain ownership from a source none of the five names — a bespoke WHOIS pipeline, a commercial feed, an internal list — matches nothing at any width. The page reports 136 as an upper bound on use and says nothing about recall, because nothing here measures it.
  • The share of the 136 that actually resolved an owner is not measured. Only two rows (webXray's 15 and Tracker Radar's 32) carry a hand verdict per paper. Doing the same for Disconnect's 74 is the obvious next increment and was not attempted this sitting.
  • No accuracy figure exists for the tail. Every rate is conditional on the entry being adjudicable, and 18–30% of each list's draws were not. Whether unadjudicable domains differ systematically between lists is exactly what would decide whether Disconnect's lead survives, and a random sample cannot answer it.
  • The coverage census has no neutral denominator. Tracker Radar's own crawl output is the frame, which flatters Tracker Radar. There is no independent census of third-party domains to use instead, and the page says so at every coverage figure rather than pretending otherwise.
  • Whether a hierarchy built from a current list would beat both. webXray's parent_id tree is the only hierarchy any of these files carries and it is frozen; nobody has tried grafting a current corporate-structure source onto Tracker Radar's entities. Filed as an open question.
  • Two adjudication rows from 2026-08-17 remain unresolved (the acquisition date behind 360yield.com, and fwmrm.net's post-2026-spinoff status). Recorded as unresolved rather than guessed; they are excluded from the 28.

H. Corrections owed to Programming:Crawler, and their real state

The task brief for this sitting said Crawler's comparison row “still says webXray's crawler is PhantomJS (historically)”. It does not, and has not since 2026-08-17. The brief's account of its own state was stale; recorded here because a later sitting reading the brief would look for a correction that is already made.

Correction State on 2026-09-11, before this sitting
The comparison row's automation column already fixed. Revision 1786953266 reads “PhantomJS in 2015; consumer Chrome over raw CDP in the last public version”
“stale third-party mirrors, the newest last pushed in 2015 and targets PhantomJS” already fixed in the same revision
The <WRAP todo> asking where webXray is developed already closed in the same revision. The page's remaining <WRAP todo> is three unrelated open questions and was correctly left alone
The licence and most-complete-copy claim already fixed on 2026-09-05, revision 1788636934 — the item's own note says so

The one correction that was owed, and that this sitting created, is the carve-out's own drift: that bullet ended “see webXray for what survives of it: the ownership database, and how it compares to Tracker Radar and Disconnect today”, and the comparison had just left that page. Repointed to Ownership resolution in revision listed below. Nothing else on Crawler was touched: its webXray figure of 7 is the tools[] schema count for a table whose whole column is schema counts, and changing it to the sweep's 15 would make that one cell incomparable with its neighbours.

I. Reviewers

All four were told explicitly that the author's context may not be exhaustive, and all four were handed the page text, the report script, its unedited output and these notes. The three focused passes ran in parallel; the generic pass ran after their findings were applied.

Status at the time of writing: the three focused passes had not returned. They were spawned in parallel with the page text, scripts/report_ownership_resolution.mjs, its unedited output, scripts/report_webxray-output.txt and these notes, against the bundle in review_own/. This section is a stub and the next sitting should fill it or re-run them — it is the one part of this provenance page that is incomplete, and it is recorded as incomplete rather than left to look finished.

Pass Model Brief Returned?
figures versus script Sonnet re-run every script, check every figure, check figures outside the edited window, mutation-test the new script not yet
citations and quotes Sonnet every {[key]} resolves; every quote verbatim in the cited paper across .cols / .norm / .txt; every attributed claim supported; industry claims to a primary source not yet
external currency Sonnet every repository, licence, vendor and version claim re-fetched as of 2026-09-11 not yet
generic, no checklist Fable whatever the focused three were not looking for; reviews this page too not run — it is sequenced after the other three

What the author checked without them, so the gap is bounded rather than open:

Check Result
All four published scripts re-run against pinned inputs report_webxray.mjs and owner_random_sample.py byte-identical to their committed output; owner_dbs.py and owner_selfname_census.py reproduce every figure quoted (section B)
owner_adjudication.py –table reproduces the 28-row table exactly: 1/18/4/1/4, 8/12/6/2/0, 26/0/2/0/0, and the three error rows (1rx.io, jsdelivr.net, stackadapt.com)
domain_owners.json schema table on webXray, every cell reproduces: 319/827 parent_id, depths 508/211/80/24/2/2, 761 uses, 803 platforms, 826 country, 104 trade_groups, 591 policy URLs, 132 GDPR, 4 CCPA, 13 opt-out, 34 crunchbase_id, 52 health-segment, 263 notes, 275 aliases, 69 languages, 2,175 of 3,215
Mutation test of report_ownership_resolution.mjs four mutations, four failures. Swapping tight/loose on Disconnect → “tight (700) > loose (74)”. Making loose disjoint from tight at equal size → “23 tight hits are not in the loose set” — the containment guard catches what a size comparison alone would miss. Injecting a fake union member → cross-check throws. Replacing the crawled intersection with the full union → cross-check throws. No guard is vacuous.
Arithmetic audit of every ratio on the content page all consistent; one error found — “seven times NDSS's rate” is 8.4×, corrected
Rendered DOM of both content pages 18 references for 18 keys on the new page, 17 for 17 on webXray; 12 and 7 tables; 9 and 4 WRAP boxes; every in-page and cross-page anchor resolves; one broken anchor found — the webXray availability table pointed at #Methodology and limitations of these figures for a Wayback timeline that had become its own section
Site-wide inbound-link sweep four pages found and fixed (section above)
Table cells containing a pipe one found in this page's own draft: a nowiki-wrapped [[a|b]] literal in a table cell. A parsed link's pipe is safe in a cell (verified in the rendered DOM); a nowiki-wrapped one is not, because the cell splits before the nowiki is honoured. Rewritten as prose, and the A.12 regex table was moved into a code block for the same reason.
Full-text corpus quality 5,855 of 5,859 renderings present, 0 empty, 90 carrying a NUL byte — see A1b

J. scripts/report_ownership_resolution.mjs, in full

report_ownership_resolution.mjs
// Every corpus figure on Design:Ownership resolution, with its denominator.
//
//   node scripts/report_ownership_resolution.mjs > scripts/report_ownership_resolution-output.txt
//   node scripts/report_ownership_resolution.mjs --wiki    # DokuWiki tables
//
// This page was carved out of Programming:Crawler:webXray on 2026-09-11. The
// corpus-side figures it inherited were computed by scripts/report_webxray.mjs,
// whose population is "papers that name webXray". That is the wrong population
// for a page about ownership resolution in general, so this script re-derives
// the landscape figures from the corpus independently and then CROSS-CHECKS the
// overlapping rows against report_webxray.mjs's committed output. Two
// implementations agreeing is a stronger audit trail than one copied number.
//
// Three things this script does that a naive sweep would not:
//
// 1. Every probe is run at TWO widths and the script asserts tight <= loose AND
//    tight subset-of loose. A narrowing probe that returns more hits, or hits
//    the loose one does not contain, is a broken probe, not a finding. The
//    Disconnect probe is the one that needs this: plain /disconnect/i is a
//    common English word and the CSP literature uses it as a technical term.
//
// 2. Every count is a PAPER count over data/fulltext/<year>/<venue>/<slug>/
//    paper.cols.txt with whitespace collapsed, because a PDF line break inside
//    "Public Suffix List" silently undercounts it.
//
// 3. Papers with no full text are counted and printed, so the sweep's own
//    denominator is visible rather than assumed to be 5,859.
//
// Every row here is an UPPER BOUND on use unless the output says hand-verified:
// a full-text match is a mention, not a use. The two rows that ARE hand
// verified (webXray, Tracker Radar) carry their verdicts in report_webxray.mjs.
 
import fs from 'node:fs';
import path from 'node:path';
import { dataRoot, loadExtractions, POPULATIONS, pct, table, wikiTable } from './lib.mjs';
 
const WIKI = process.argv.includes('--wiki');
const rows = loadExtractions();
const key = (p) => `${p.venue}/${p.year}/${p.slug}`;
const FT = path.join(dataRoot(), 'fulltext');
 
const out = [];
const P = (s = '') => out.push(s);
const heading = (s) => {
  P('');
  P(`=== ${s} ===`);
  P('');
};
const T = (headers, body) => P(WIKI ? wikiTable(headers, body) : table(headers, body));
 
// --- full text, whitespace collapsed -------------------------------------
// Streamed, not cached: the corpus is 5,859 papers and holding every
// paper.cols.txt in a Map exhausts the default V8 heap. One pass, every
// pattern tested against each paper's text, then the text is dropped.
function readText(k) {
  const [venue, year, slug] = k.split('/');
  const f = `${FT}/${year}/${venue}/${slug}/paper.cols.txt`;
  return fs.existsSync(f) ? fs.readFileSync(f, 'utf8').replace(/\s+/g, ' ') : null;
}
 
// Which papers have full text at all, and the hits for every pattern, in one
// pass. PATTERNS is filled below before this runs.
function sweepAll(patterns) {
  const hits = new Map([...patterns.keys()].map((n) => [n, []]));
  const haveText = new Set();
  for (const p of rows) {
    const k = key(p);
    const t = readText(k);
    if (t === null) continue;
    haveText.add(k);
    for (const [name, re] of patterns) if (re.test(t)) hits.get(name).push(k);
  }
  for (const v of hits.values()) v.sort();
  return { hits, haveText };
}
 
// --- the probes, each at two widths --------------------------------------
// `loose` must be a strict superset of `tight` by construction; the script
// checks that it is in fact one and throws if not.
const PROBES = [
  {
    name: 'webXray',
    tight: /webx[\s-]?ray/i,
    loose: /webx[\s-]?ray|libert/i,
    note: 'hand-verified: all 15, in report_webxray.mjs',
  },
  {
    name: 'Tracker Radar',
    tight: /tracker[\s.-]?radar/i,
    loose: /tracker[\s.-]?radar|duckduckgo/i,
    note: 'hand-verified: all 32, in report_webxray.mjs',
  },
  {
    name: 'WhoTracks.me',
    tight: /whotracks/i,
    loose: /whotracks|who ?tracks ?\.? ?me|ghostery/i,
    note: 'upper bound',
  },
  {
    name: 'Crunchbase',
    tight: /crunchbase/i,
    loose: /crunchbase|crunch base/i,
    note: 'upper bound',
  },
  {
    name: 'Disconnect (list sense)',
    tight: /disconnect'?s? (entit|list|block|black|tracking)|disconnect\.me|entities\.json/i,
    loose: /disconnect/i,
    note: 'upper bound; the loose width is the reason this one is tightened',
  },
  {
    name: 'Public Suffix List',
    tight: /public suffix/i,
    loose: /public suffix|etld\+?1|effective top-?level domain/i,
    note: 'upper bound',
  },
];
 
// One pass over the full text: every probe, both widths, plus the set of papers
// that have a readable rendering at all.
const PATTERNS = new Map();
for (const pr of PROBES) {
  PATTERNS.set(`${pr.name}|tight`, pr.tight);
  PATTERNS.set(`${pr.name}|loose`, pr.loose);
}
const { hits: SWEPT, haveText: HAVE_TEXT } = sweepAll(PATTERNS);
 
const withText = rows.filter((p) => HAVE_TEXT.has(key(p)));
const crawled = rows.filter(POPULATIONS.crawled);
 
P('==============================================================================');
P('Design:Ownership resolution — every corpus figure, with its denominator');
P('==============================================================================');
P('');
P(`Corpus: ${rows.length} extracted papers, 7 venues (CCS, IMC, NDSS, PETS, USENIX`);
P('Security, TheWebConf, IEEE S&P), 2010–2026. EuroS&P, ACSAC, RAID, AsiaCCS, CHI');
P('and SOUPS are absent, so every figure here is a claim about those seven venues.');
P('');
P(`Papers with a readable paper.cols.txt: ${withText.length} of ${rows.length} (${pct(withText.length, rows.length)}).`);
P(`Papers that ran a crawl (crawlConfig present or studyTypes includes automated-web-crawl): ${crawled.length}.`);
P('The sweep denominator is the full corpus; papers with no full text can only');
P('ever be misses, so every sweep count below is a floor on the true mention count.');
 
// --- A. probe widths ------------------------------------------------------
heading('A. Probe widths: tight must be contained in loose');
 
const results = new Map();
const widthRows = [];
for (const pr of PROBES) {
  const tight = SWEPT.get(`${pr.name}|tight`);
  const loose = SWEPT.get(`${pr.name}|loose`);
  const tightSet = new Set(tight);
  const looseSet = new Set(loose);
  const escaped = tight.filter((k) => !looseSet.has(k));
  if (tight.length > loose.length) {
    throw new Error(`probe ${pr.name}: tight (${tight.length}) > loose (${loose.length}) — the narrowing probe is broken`);
  }
  if (escaped.length > 0) {
    throw new Error(`probe ${pr.name}: ${escaped.length} tight hits are not in the loose set: ${escaped.slice(0, 5).join(', ')}`);
  }
  results.set(pr.name, tightSet);
  widthRows.push([pr.name, String(tight.length), String(loose.length), pr.note]);
}
T(['Resource', 'Tight probe', 'Loose probe', 'Status'], widthRows);
P('');
P('Both widths are printed because the gap is the claim. Disconnect at the loose');
P('width matches 700 papers and almost none of them mean the list, which is why');
P('the published row is the tight one and why the page says so. A row whose two');
P('widths are close is a row whose count is not an artefact of the pattern.');
 
// --- B. the published landscape table ------------------------------------
heading('B. The ownership-resolution landscape (the page\'s main table)');
 
const landscape = PROBES.map((pr) => {
  const s = results.get(pr.name);
  const inCrawled = crawled.filter((p) => s.has(key(p))).length;
  return [pr.name, String(s.size), pct(s.size, rows.length), String(inCrawled), pct(inCrawled, crawled.length), pr.note.startsWith('hand-verified') ? 'yes' : 'no — upper bound'];
})
  .sort((a, b) => Number(b[1]) - Number(a[1]));
T(['Resource named in full text', 'Papers', `Share of ${rows.length}`, `of which crawled`, `Share of ${crawled.length}`, 'Hand-verified?'], landscape);
 
// --- C. the union ---------------------------------------------------------
heading('C. The union: how widespread is ownership resolution at all?');
 
// The union is over the five OWNERSHIP resources. The Public Suffix List is
// deliberately excluded: it answers "what is the registrable domain", not
// "whose is it", and including it would nearly double the union on a resource
// that every crawl paper touches for unrelated reasons.
const UNION_NAMES = ['webXray', 'Tracker Radar', 'WhoTracks.me', 'Crunchbase', 'Disconnect (list sense)'];
const union = new Set();
for (const n of UNION_NAMES) for (const k of results.get(n)) union.add(k);
P(`Union of ${UNION_NAMES.join(' / ')}:`);
P(`  ${union.size} papers — ${pct(union.size, rows.length)} of the ${rows.length}-paper corpus,`);
const unionCrawled = crawled.filter((p) => union.has(key(p)));
P(`  and ${pct(unionCrawled.length, crawled.length)} of the ${crawled.length} that ran a crawl (${unionCrawled.length} papers).`);
P('');
P('The Public Suffix List is NOT in the union: it answers "what is the registrable');
P('domain", not "whose is it". Adding it would take the union to ' +
  (() => { const u2 = new Set(union); for (const k of results.get('Public Suffix List')) u2.add(k); return u2.size; })() +
  ' on a resource');
P('every crawling paper touches for unrelated reasons.');
P('');
P('Each member is an upper bound on USE, so the union is an upper bound too:');
P('136 papers mention one of these; far fewer resolved an owner with one.');
 
// --- D. per-year adoption -------------------------------------------------
heading('D. Per-year: is ownership resolution growing?');
 
const years = [...new Set(rows.map((p) => p.year))].sort();
const yearRows = years.map((y) => {
  const inYear = rows.filter((p) => p.year === y);
  const hit = inYear.filter((p) => union.has(key(p)));
  const star = y >= 2025 ? '*' : '';
  return [`${y}${star}`, String(inYear.length), String(hit.length), pct(hit.length, inYear.length)];
});
T(['Year', 'Corpus papers', 'Naming an ownership resource', 'Share'], yearRows);
P('');
P('* 2025 and 2026 are the provisional corpus edge: CCS 2026 and IMC 2026 have not');
P('  been held, and IEEE S&P / WWW 2026 abstracts are not in OpenAlex, so selection');
P('  under-samples them BY CONSTRUCTION. Do not read 2026 as a complete year.');
P('  The share, not the count, is the readable column for those two rows.');
 
// --- E. venue shape -------------------------------------------------------
heading('E. Venue shape of the union');
 
const venues = [...new Set(rows.map((p) => p.venue))].sort();
const venueRows = venues.map((v) => {
  const inVenue = rows.filter((p) => p.venue === v);
  const hit = inVenue.filter((p) => union.has(key(p)));
  return [v, String(inVenue.length), String(hit.length), pct(hit.length, inVenue.length)];
}).sort((a, b) => Number(b[2]) - Number(a[2]));
T(['Venue', 'Corpus papers', 'Naming an ownership resource', 'Share of venue'], venueRows);
P('');
P('Read the SHARE column, not the count: the venues differ in size by a factor of');
P('three. IMC and PETS are where this work lands; the shares are small everywhere,');
P('which is the honest headline — ownership resolution is a step inside a paper,');
P('not a paper topic, and most papers that do it do not say how.');
 
// --- F. the schema's own view --------------------------------------------
heading('F. What the extraction schema sees, against the sweep');
 
// classification[].groundTruthSource / resourceName naming an ownership list.
const SCHEMA_RE = /webx[\s-]?ray|tracker[\s.-]?radar|whotracks|disconnect|crunchbase/i;
const schemaPapers = new Set();
for (const p of rows) {
  const names = [
    ...p.tools.map((t) => t.name),
    ...p.classification.map((c) => c.resourceName),
    ...p.classification.map((c) => c.groundTruthSource),
    ...p.otherToolsMentioned.map((t) => t.name),
  ].filter((s) => typeof s === 'string');
  if (names.some((n) => SCHEMA_RE.test(n))) schemaPapers.add(key(p));
}
const both = [...union].filter((k) => schemaPapers.has(k));
P(`Papers whose tools[] / classification[] / otherToolsMentioned[] name one of the`);
P(`five resources:            ${schemaPapers.size}`);
P(`Papers whose FULL TEXT names one:              ${union.size}`);
P(`In both:                                       ${both.length}`);
P(`Full text only (schema misses):                ${union.size - both.length}`);
P(`Schema only (no full-text match):              ${schemaPapers.size - both.length}`);
P('');
P('The schema-only residue is printed rather than dropped:');
const schemaOnly = [...schemaPapers].filter((k) => !union.has(k)).sort();
for (const k of schemaOnly) {
  const p = rows.find((r) => key(r) === k);
  const hits = [
    ...p.tools.map((t) => t.name),
    ...p.classification.map((c) => c.resourceName),
    ...p.classification.map((c) => c.groundTruthSource),
    ...p.otherToolsMentioned.map((t) => t.name),
  ].filter((s) => typeof s === 'string' && SCHEMA_RE.test(s));
  P(`  ${k}`);
  P(`      named: ${[...new Set(hits)].join(' | ')}`);
  P(`      full text present: ${HAVE_TEXT.has(k)}`);
}
P('');
P('This is why the page publishes the full-text sweep and not the schema field:');
P('the schema fires on a fraction of the papers that name these lists, because');
P('most mentions are in related work or a reference list rather than in a tools');
P('sentence the extractor reads as a tool. Neither signal is "the" population;');
P('both are reported.');
 
// --- G. cross-check against report_webxray.mjs ---------------------------
heading('G. Cross-check against scripts/report_webxray-output.txt');
 
const committed = 'scripts/report_webxray-output.txt';
if (!fs.existsSync(committed)) {
  P(`${committed} not found — cross-check SKIPPED.`);
} else {
  const txt = fs.readFileSync(committed, 'utf8');
  const CHECKS = [
    ['webXray', /^webXray\s+(\d+)\s/m],
    ['Tracker Radar', /^Tracker Radar\s+(\d+)\s/m],
    ['WhoTracks.me', /^WhoTracks\.me\s+(\d+)\s/m],
    ['Crunchbase', /^Crunchbase\s+(\d+)\s/m],
    ['Disconnect (list sense)', /^Disconnect \(list sense\)\s+(\d+)\s/m],
    ['Public Suffix List', /^Public Suffix List\s+(\d+)\s/m],
  ];
  const checkRows = [];
  let bad = 0;
  for (const [name, re] of CHECKS) {
    const m = txt.match(re);
    if (m === null) throw new Error(`cross-check: could not find "${name}" in ${committed}`);
    const theirs = Number(m[1]);
    const mine = results.get(name).size;
    if (theirs !== mine) bad += 1;
    checkRows.push([name, String(mine), String(theirs), theirs === mine ? 'agree' : '*** DISAGREE ***']);
  }
  const um = txt.match(/^\s*(\d+) papers \((\d+\.\d)% of the corpus\)\. (\d+) of them are among the (\d+) that crawled -- (\d+\.\d)% of that population\./m);
  if (um === null) throw new Error(`cross-check: could not find the union line in ${committed}`);
  if (Number(um[1]) !== union.size) bad += 1;
  checkRows.push(['union of the five', String(union.size), um[1], Number(um[1]) === union.size ? 'agree' : '*** DISAGREE ***']);
  if (Number(um[3]) !== unionCrawled.length) bad += 1;
  checkRows.push(['union ∩ crawled', String(unionCrawled.length), um[3], Number(um[3]) === unionCrawled.length ? 'agree' : '*** DISAGREE ***']);
  if (Number(um[4]) !== crawled.length) bad += 1;
  checkRows.push(['crawled denominator', String(crawled.length), um[4], Number(um[4]) === crawled.length ? 'agree' : '*** DISAGREE ***']);
  T(['Figure', 'this script', 'report_webxray.mjs', 'verdict'], checkRows);
  P('');
  if (bad > 0) {
    throw new Error(`${bad} figures disagree between the two independent implementations — do not publish either.`);
  }
  P('All rows agree. The two scripts share no code beyond lib.mjs and the');
  P('whitespace-collapse convention; the patterns were written out separately.');
}
 
// --- H. the page's non-corpus figures, listed so they are not mistaken ----
heading('H. Every number on the page that does NOT come from this corpus');
 
P('Listed so a reader auditing the page knows which script to re-run, and so');
P('that a refresh of this script is never mistaken for a refresh of the page.');
P('');
T(['Figure family', 'Script', 'Snapshot it is pinned to'], [
  ['List shape (827 / 19,148 / 1,887 owners; 3,215 / 38,368 / 7,850 domains)', 'owner_dbs.py --cache out/webxray/cache', '2026-08-17'],
  ['Coverage census (669 / 5,581 / 2,268 of 32,369; the prevalence weights; the rank slices)', 'owner_dbs.py --cache out/webxray/cache', '2026-08-17'],
  ['Agreement between pairs (612 / 464 / 1,538 and the "neither" columns)', 'owner_dbs.py --cache out/webxray/cache', '2026-08-17'],
  ['The 28-row hand adjudication of high-prevalence disagreements', 'owner_adjudication.py --table', '2026-08-17'],
  ['Accuracy rates (69.9 / 73.7 / 94.4%, the encounter-weighted column, the CIs)', 'owner_random_sample.py --sample out/owner_sample.json --rows out/adj_rows_scored.json', '2026-09-05'],
  ['Coverage counts inside the random-sample section (664 / 5,566 / 2,264 of 32,337)', 'owner_sample.py --cache out/owner_cache', '2026-09-05'],
  ['Self-named entity shares (0.8% / 0.7% / 3.2%)', 'owner_selfname_census.py', '2026-09-05'],
  ['Inter-rater kappa (+0.400, +0.680)', 'owner_irr_kappa.py', '2026-09-11'],
  ['Repository state, licences, Wayback dates', 'hand checks + wayback_webxray.sh', '2026-08-17 / 2026-09-05'],
]);
P('');
P('Two of the three lists change weekly. Re-run before citing.');
 
console.log(out.join('\n'));

K. Its unedited output

Run as node scripts/report_ownership_resolution.mjs.

report_ownership_resolution-output.txt
==============================================================================
Design:Ownership resolution — every corpus figure, with its denominator
==============================================================================
 
Corpus: 5859 extracted papers, 7 venues (CCS, IMC, NDSS, PETS, USENIX
Security, TheWebConf, IEEE S&P), 2010–2026. EuroS&P, ACSAC, RAID, AsiaCCS, CHI
and SOUPS are absent, so every figure here is a claim about those seven venues.
 
Papers with a readable paper.cols.txt: 5855 of 5859 (99.9%).
Papers that ran a crawl (crawlConfig present or studyTypes includes automated-web-crawl): 1120.
The sweep denominator is the full corpus; papers with no full text can only
ever be misses, so every sweep count below is a floor on the true mention count.
 
=== A. Probe widths: tight must be contained in loose ===
 
Resource                 Tight probe  Loose probe  Status
-----------------------  -----------  -----------  ----------------------------------------------------------------
webXray                  15           233          hand-verified: all 15, in report_webxray.mjs
Tracker Radar            32           108          hand-verified: all 32, in report_webxray.mjs
WhoTracks.me             23           124          upper bound
Crunchbase               23           23           upper bound
Disconnect (list sense)  74           700          upper bound; the loose width is the reason this one is tightened
Public Suffix List       101          157          upper bound
 
Both widths are printed because the gap is the claim. Disconnect at the loose
width matches 700 papers and almost none of them mean the list, which is why
the published row is the tight one and why the page says so. A row whose two
widths are close is a row whose count is not an artefact of the pattern.
 
=== B. The ownership-resolution landscape (the page's main table) ===
 
Resource named in full text  Papers  Share of 5859  of which crawled  Share of 1120  Hand-verified?
---------------------------  ------  -------------  ----------------  -------------  ----------------
Public Suffix List           101     1.7%           46                4.1%           no — upper bound
Disconnect (list sense)      74      1.3%           65                5.8%           no — upper bound
Tracker Radar                32      0.5%           28                2.5%           yes
WhoTracks.me                 23      0.4%           19                1.7%           no — upper bound
Crunchbase                   23      0.4%           11                1.0%           no — upper bound
webXray                      15      0.3%           13                1.2%           yes
 
=== C. The union: how widespread is ownership resolution at all? ===
 
Union of webXray / Tracker Radar / WhoTracks.me / Crunchbase / Disconnect (list sense):
  136 papers — 2.3% of the 5859-paper corpus,
  and 9.6% of the 1120 that ran a crawl (107 papers).
 
The Public Suffix List is NOT in the union: it answers "what is the registrable
domain", not "whose is it". Adding it would take the union to 221 on a resource
every crawling paper touches for unrelated reasons.
 
Each member is an upper bound on USE, so the union is an upper bound too:
136 papers mention one of these; far fewer resolved an owner with one.
 
=== D. Per-year: is ownership resolution growing? ===
 
Year   Corpus papers  Naming an ownership resource  Share
-----  -------------  ----------------------------  -----
2010   119            0                             0.0%
2011   116            1                             0.9%
2012   151            0                             0.0%
2013   125            0                             0.0%
2014   166            0                             0.0%
2015   190            0                             0.0%
2016   182            2                             1.1%
2017   231            4                             1.7%
2018   254            5                             2.0%
2019   402            9                             2.2%
2020   404            15                            3.7%
2021   379            13                            3.4%
2022   546            19                            3.5%
2023   719            19                            2.6%
2024   690            17                            2.5%
2025*  770            21                            2.7%
2026*  415            11                            2.7%
 
* 2025 and 2026 are the provisional corpus edge: CCS 2026 and IMC 2026 have not
  been held, and IEEE S&P / WWW 2026 abstracts are not in OpenAlex, so selection
  under-samples them BY CONSTRUCTION. Do not read 2026 as a complete year.
  The share, not the count, is the readable column for those two rows.
 
=== E. Venue shape of the union ===
 
Venue    Corpus papers  Naming an ownership resource  Share of venue
-------  -------------  ----------------------------  --------------
PETS     510            49                            9.6%
USENIX   1410           19                            1.3%
WWW      843            19                            2.3%
IMC      638            17                            2.7%
CCS      990            13                            1.3%
IEEE-SP  767            11                            1.4%
NDSS     701            8                             1.1%
 
Read the SHARE column, not the count: the venues differ in size by a factor of
three. IMC and PETS are where this work lands; the shares are small everywhere,
which is the honest headline — ownership resolution is a step inside a paper,
not a paper topic, and most papers that do it do not say how.
 
=== F. What the extraction schema sees, against the sweep ===
 
Papers whose tools[] / classification[] / otherToolsMentioned[] name one of the
five resources:            89
Papers whose FULL TEXT names one:              136
In both:                                       82
Full text only (schema misses):                54
Schema only (no full-text match):              7
 
The schema-only residue is printed rather than dropped:
  IEEE-SP/2024/to-auth-or-not-to-auth-a-comparative-analysis-of-the-pre-and-post-login-security
      named: Disconnect Tracker Protection List
      full text present: true
  PETS/2022/on-dark-patterns-and-manipulation-of-website-publishers-by-cmps
      named: Disconnect | Disconnect tracking filter list
      full text present: true
  PETS/2024/generalizable-active-privacy-choice-designing-a-graphical-user-interface-for-glo
      named: Disconnect Tracker Protection lists
      full text present: true
  PETS/2024/support-personas-a-concept-for-tailored-support-of-users-of-privacy-enhancing-te
      named: Disconnect
      full text present: true
  USENIX/2019/canvas-fast-and-inexpensive-automotive-network-mapping
      named: physical ECU access and disconnection
      full text present: true
  WWW/2018/the-cost-of-digital-advertisement-comparing-user-and-advertiser-views
      named: Disconnect | Disconnect blacklist
      full text present: true
  WWW/2023/who-funds-misinformation-a-systematic-analysis-of-the-ad-related-profit-routines
      named: Disconnect
      full text present: true
 
This is why the page publishes the full-text sweep and not the schema field:
the schema fires on a fraction of the papers that name these lists, because
most mentions are in related work or a reference list rather than in a tools
sentence the extractor reads as a tool. Neither signal is "the" population;
both are reported.
 
=== G. Cross-check against scripts/report_webxray-output.txt ===
 
Figure                   this script  report_webxray.mjs  verdict
-----------------------  -----------  ------------------  -------
webXray                  15           15                  agree
Tracker Radar            32           32                  agree
WhoTracks.me             23           23                  agree
Crunchbase               23           23                  agree
Disconnect (list sense)  74           74                  agree
Public Suffix List       101          101                 agree
union of the five        136          136                 agree
union ∩ crawled          107          107                 agree
crawled denominator      1120         1120                agree
 
All rows agree. The two scripts share no code beyond lib.mjs and the
whitespace-collapse convention; the patterns were written out separately.
 
=== H. Every number on the page that does NOT come from this corpus ===
 
Listed so a reader auditing the page knows which script to re-run, and so
that a refresh of this script is never mistaken for a refresh of the page.
 
Figure family                                                                             Script                                                                                 Snapshot it is pinned to
----------------------------------------------------------------------------------------  -------------------------------------------------------------------------------------  ------------------------
List shape (827 / 19,148 / 1,887 owners; 3,215 / 38,368 / 7,850 domains)                  owner_dbs.py --cache out/webxray/cache                                                 2026-08-17
Coverage census (669 / 5,581 / 2,268 of 32,369; the prevalence weights; the rank slices)  owner_dbs.py --cache out/webxray/cache                                                 2026-08-17
Agreement between pairs (612 / 464 / 1,538 and the "neither" columns)                     owner_dbs.py --cache out/webxray/cache                                                 2026-08-17
The 28-row hand adjudication of high-prevalence disagreements                             owner_adjudication.py --table                                                          2026-08-17
Accuracy rates (69.9 / 73.7 / 94.4%, the encounter-weighted column, the CIs)              owner_random_sample.py --sample out/owner_sample.json --rows out/adj_rows_scored.json  2026-09-05
Coverage counts inside the random-sample section (664 / 5,566 / 2,264 of 32,337)          owner_sample.py --cache out/owner_cache                                                2026-09-05
Self-named entity shares (0.8% / 0.7% / 3.2%)                                             owner_selfname_census.py                                                               2026-09-05
Inter-rater kappa (+0.400, +0.680)                                                        owner_irr_kappa.py                                                                     2026-09-11
Repository state, licences, Wayback dates                                                 hand checks + wayback_webxray.sh                                                       2026-08-17 / 2026-09-05
 
Two of the three lists change weekly. Re-run before citing.
  • Ownership resolution — the page these notes are for.
  • webxray — the live-list scripts, the fold residues, the Wayback log, the 2026-08-17 and 2026-09-05 run tables.
  • random_sample — the 175-row adjudication, the estimator, the inter-rater draw and kappa.
  • Corpus — corpus-level selection and extraction caveats.

References

[1]
Matte, Célestin; Bielova, Nataliia; Santos, Cristiana Teixeira (2020): "Do Cookie Banners Respect my Choice? Measuring Legal Compliance of Banners from IAB Europe's Transparency and Consent Framework", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[2]
Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)
[3]
Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[4]
Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[5]
Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)
provenance/design/ownership_resolution.1789149855.txt.gz · Last modified: by karel.kubicek.claude