This is an old revision of the document!
Table of Contents
Provenance: Design:Ownership resolution
Working notes behind Ownership resolution. Every query with its population and denominator, the report script and its unedited output, what was folded and what the fold could not reach, the quotes checked against the source, the external sources verified and rejected, and the judgement calls. Corpus-wide selection and extraction caveats are on Corpus and are not restated here.
This page is a carve-out, and most of its audit trail lives elsewhere on purpose. The content page was split out of webXray on 2026-09-11. The scripts that produce its live-database figures, their unedited output, the 175-row adjudication table and the inter-rater material were published in 2026-08 and 2026-09 under the webXray ids and keep those ids, because they are cited from the published record and moving them would break links for no gain:
| What | Where |
|---|---|
The three-list comparison scripts (owner_dbs.py, owner_adjudication.py), their unedited output, the fold residues, the Wayback log, the 2026-08-17 and 2026-09-05 run tables | webxray |
owner_sample.py, owner_random_sample.py, the full 175-row adjudication with every source, the citation re-fetch, the inter-rater draw and kappa | random_sample |
The corpus population of webXray itself (the 15 papers, the hand role map), report_webxray.mjs | webxray |
| This page | the split itself, the new corpus script and its output, the reproduction record of 2026-09-11, the defect that run found, the reviewer log |
No ~~DISCUSSION~~ block: comments belong on the content page. Citations use the same {[citekey]} keys and the same shared Bibliography; this page adds no bibliography entries of its own.
The run
| Date | 2026-09-11 |
| Corpus at the time | data/extract/run1/extractions.jsonl, 5,859 papers, 7 venues, 2010–2026. 5,855 have a readable paper.cols.txt. |
| What was done | Created Ownership resolution; rewrote webXray around what was left; repointed Crawler; added a row to Design; turned the Roadmap row blue. |
| New code | scripts/report_ownership_resolution.mjs (346 lines), committed with its output. |
| Code changed | scripts/report_webxray.mjs — a denominator defect, section C below. |
| Models | Opus 5 wrote the pages and the script. Review layer: three focused Sonnet passes and one Fable generic pass, section I. |
| Accidental exposure or mistakes caught | none in this sitting beyond the defect in C, which was found by the new script disagreeing with the old one. |
Scope decision: why a separate page, when the last sitting decided against one
This reverses a recorded decision, so the reversal is recorded too. webxray §“Scope decision” says, on 2026-08-17:
A separatedesign:ownership_resolutionpage with webXray as a stub. Rejected: it would leave the wiki's existing red link pointing at a stub, and the material does not split cleanly — webXray's frozen list is the best available illustration of what goes wrong.
Both halves of that reasoning have since stopped holding:
- “webXray as a stub” was the wrong alternative to compare against. There is enough webXray-specific material for a real page — the three-licence history across four states, the availability census, the
domain_owners.jsonschema, the frozen 2016 PSL, the Wayback timeline, and 15 corpus papers with a hand role map. What is left after the carve-out is 27 KB, not a stub. - The red-link problem inverted. On 2026-09-07 the Roadmap queued
design:ownership_resolutionas a promised page andscripts/sitemap.mjsbegan gating on it, so from that date not writing it was the dangling promise. The id was fixed by that decision and is used unchanged. - The discoverability argument is the item's whole point and was never answered. A reader asking “how do I attribute a third-party domain to a company?” does not search for a dead tool. Both earlier sittings recorded this as the right thing to do; neither did it.
Alternatives considered this sitting and rejected:
| Alternative | Why not |
|---|---|
| Leave it, and add a redirect or a prominent pointer on webXray | A pointer does not fix search, does not fix the namespace (programming: is instruments; this is a design decision), and leaves the currency claim — which source to use now — filed under the historical one. |
| Broaden Requests instead | That page answers “is this request a tracker?”, a different question with different tooling (filter lists, not owner databases). The two pages cross-link. |
| Broaden IP classification | It is the network-layer sibling (AS-to-organisation) and is already large; the domain layer has different sources, different failure modes and different licences. Cross-linked both ways instead. |
Move the random_sample code appendix under provenance:design: too | Rejected. Its id is cited from the published content page and from the webXray provenance page; the gain is cosmetic and the cost is broken links. Recorded here so the next sitting does not “tidy” it. |
| Re-derive the coverage census against a 2026-09-11 snapshot | Rejected. The 28-row adjudication is pinned to the 2026-08-17 frame; re-deriving would orphan it. The figures moved unchanged, with their frame dates, and the page carries the two-frames box. This is a carve-out, not a refresh — see the drift check in section B. |
The carve-out map
What moved, what stayed, and what had to be repointed. Nothing was re-derived in the move.
| Section on the old page | Went to |
|---|---|
| The tool: architecture, and why you cannot install it | stayed on webXray |
| The ownership database (schema, tree, frozen PSL) | stayed — it describes webXray's file |
| How it compares to Tracker Radar and Disconnect | moved, as The live sources, compared |
| Coverage: how much of the third-party surface | moved, with the three traps split into their own sub-section |
| Two lists disagree: error or different question? | moved, split: the two axes were promoted to the top of the new page as The question has two axes, the agreement tables became When two lists disagree |
| Reading all three at once (the code) | moved |
| Which list is right, when they disagree? | moved |
| How often is each list right? A random sample | moved, with the inter-rater material given its own sub-section |
| Where these figures come from, and how to redo them | moved, folded into the new page's methodology section |
| Choosing a resolution source now | moved |
| Assembling the pipeline | moved |
| What to report in a paper | moved |
| Use in publications: the 15 webXray papers, role table | stayed |
| Use in publications: the 136-paper sweep, the Tracker Radar year table | moved — they are landscape figures, not webXray figures |
| Wayback timeline | stayed, and was promoted from a methodology bullet to its own section; it was 900 words inside a bullet list |
Day-one drift, found and fixed rather than left:
| Where | What it said after the move | Fixed to |
|---|---|---|
| webXray intro bullet 2 | quoted 69.9% / 93.9% / 2.1% / 17.2% / 7.0% inline as if the page still carried the measurement | one sentence keeping 69.9% and 2.1% with a link, and an explicit “nothing on this page restates them” |
webXray, The ownership database | “§\”Two lists disagree\“ below measures what happens when you forget that” — a same-page reference to a section that had left | repointed to [[Design:Ownership resolution#When two lists disagree]], with the finding (root resolution makes agreement worse) stated inline so the sentence still says something |
webXray, Related pages | listed Requests, Cookies, IP classification — all of which were there for the ownership material | new page took those; webXray's list now leads with the new page and adds Policies for policyXray |
| webXray methodology | listed nine owner_* scripts it no longer publishes figures from | one bullet saying where they went and that their provenance ids are unchanged |
| Crawler, webXray bullet | “see webXray for what survives of it: the ownership database, and how it compares to Tracker Radar and Disconnect today” — that comparison had moved | repointed to the new page; see section H |
A site-wide inbound-link sweep found four more, on pages nobody would have thought to check. Grepping every cached page for a link to programming:crawler:webxray rather than only the pages this sitting edited:
| Page | What it said | Fixed to |
|---|---|---|
| Requests | a Related-pages bullet linking the webXray page under the link text “webXray and domain-to-company ownership” — “once a request is flagged, this is how to answer whose it is” | repointed to Ownership resolution, with a note that it moved |
| Filter lists | the same bullet, near-verbatim | repointed the same way |
| Legal enforcement | “See webXray for the ownership databases and how much they disagree” — in the bullet on identifying a controller under Reg. 2025/2518 | repointed, and extended to name the accuracy rates, which is what that bullet actually needs |
| Programming (namespace index) | the webXray row read “Domain-to-company ownership lists, and what remains of the tool” — a description of the page, now false | rewritten to describe the tool page and point the ownership question at the new page |
Two of those four used the link text “webXray and domain-to-company ownership”, i.e. they were already treating the webXray page as the ownership page — which is the item's complaint, restated by the wiki itself. A carve-out's drift is not confined to the pages you edited; sweep for inbound links before calling it done. One inbound reference was deliberately left: provenance:design's generated page inventory records a size and a description as of an earlier date, and it is a dated snapshot rather than a live pointer.
A. Corpus queries
All over data/extract/run1/extractions.jsonl (5,859 papers) and data/fulltext/<year>/<venue>/<slug>/paper.cols.txt, whitespace collapsed before matching so a term broken across a PDF column boundary still matches. Every count is a paper count. The script is scripts/report_ownership_resolution.mjs, published in full in section J with its unedited output in section K.
| # | Question | Population / denominator | Answer |
|---|---|---|---|
| A1 | How many papers have a readable full-text rendering at all? | all 5,859 | 5,855 (99.9%). The four without one can only ever be sweep misses, so every count below is a floor. |
| A1b | …and how many of those are damaged? | the 5,855 | 0 empty, 0 under 2 KB, but 90 carry a NUL byte. The script reads them as JavaScript strings and regexes them, so NULs do not affect it. A grep-based sweep would have silently skipped all 90 — grep treats a NUL-carrying file as binary and, depending on the build, prints nothing and exits non-zero. None of the 136 union members is one of the 90, checked explicitly, so no published count depends on this — but the next person writing a sweep should not use grep. |
| A2 | How many papers name each ownership resource? | all 5,859 | PSL 101, Disconnect (list sense) 74, Tracker Radar 32, WhoTracks.me 23, Crunchbase 23, webXray 15 |
| A3 | How many of those are inside the crawling population? | crawled = 1,120 (crawlConfig present, or studyTypes includes automated-web-crawl) | PSL 46, Disconnect 65, Tracker Radar 28, WhoTracks.me 19, Crunchbase 11, webXray 13 |
| A4 | How many name any of the five ownership resources? | all 5,859 | 136 (2.3%) |
| A5 | …and how many of those 136 are in the 1,120? | crawled = 1,120 | 107, i.e. 9.6% of that population. Not 12.1% — see section C. |
| A6 | Does adding the PSL change the union? | all 5,859 | 136 → 221. The PSL is deliberately excluded: it answers “what is the registrable domain”, not “whose is it”. |
| A7 | Is ownership resolution growing? | each year's own paper count | rose to ~3.5% in 2020–2022, has sat at 2.5–2.7% since. Full table on the page; 2010–2015 contributes one paper in total and is omitted rather than padded with zeros. |
| A8 | Which venues? | each venue's own paper count | PETS 9.6%, IMC 2.7%, TheWebConf 2.3%, IEEE S&P 1.4%, USENIX 1.3%, CCS 1.3%, NDSS 1.1% |
| A9 | What does the extraction schema see, against the full text? | all 5,859 | schema 89, full text 136, both 82, full-text-only 54, schema-only 7 — residue printed in full below |
| A10 | Of the 15 webXray papers, how many say which version? | the 8 that used the crawler or the list | 2 say anything; 1 names a commit ([1Matte, Célestin; Bielova, Nataliia; Santos, Cristiana Teixeira (2020): "Do Cookie Banners Respect my Choice? Measuring Legal Compliance of Banners from IAB Europe's Transparency and Consent Framework", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]). Inherited from report_webxray.mjs, re-run and unchanged. |
| A11 | What does “Tracker Radar” mean in the 32 papers that name it? | the 32, hand-mapped | 11 ownership dataset, 9 tracker/category database, 9 the Collector crawler, 2 citation, 1 compared. Inherited, re-run and unchanged. |
A.12 Probe widths: every probe run at two widths, and both printed
A narrowing probe that returns more hits than the loose one, or hits the loose one does not contain, is a broken probe. The script asserts both and throws rather than printing. Both widths are on the record because the gap is itself the claim:
Resource Tight probe Loose probe Status ----------------------- ----------- ----------- ---------------------------------------------------------------- webXray 15 233 hand-verified: all 15, in report_webxray.mjs Tracker Radar 32 108 hand-verified: all 32, in report_webxray.mjs WhoTracks.me 23 124 upper bound Crunchbase 23 23 upper bound Disconnect (list sense) 74 700 upper bound; the loose width is the reason this one is tightened Public Suffix List 101 157 upper bound
The patterns themselves are in a code block and not in a table, because a DokuWiki table cell cannot hold a pipe and every one of these regexes is an alternation. \| is not an escape and wrapping the pattern in DokuWiki's nowiki delimiters does not help either: the cell splits regardless.
webXray tight /webx[\s-]?ray/i
loose /webx[\s-]?ray|libert/i
15 -> 233. "Libert" is a common surname and a French word.
The loose width is a sanity bound, not a candidate set.
Tracker Radar tight /tracker[\s.-]?radar/i
loose /tracker[\s.-]?radar|duckduckgo/i
32 -> 108. Most DuckDuckGo mentions are the search engine
or the browser, not the dataset.
WhoTracks.me tight /whotracks/i
loose /whotracks|who ?tracks ?\.? ?me|ghostery/i
23 -> 124. Ghostery is mostly the extension.
Crunchbase tight /crunchbase/i
loose /crunchbase|crunch base/i
23 -> 23. NO GAP: nobody spells it with a space. This
row's count is not an artefact of the pattern.
Disconnect tight /disconnect'?s? (entit|list|block|black|tracking)|disconnect\.me|entities\.json/i
loose /disconnect/i
74 -> 700. "Disconnect" is an ordinary English word and a
CSP-literature technical term. This is the probe that
needed tightening, and the published row is the tight one.
Public Suffix List tight /public suffix/i
loose /public suffix|etld\+?1|effective top-?level domain/i
101 -> 157. The extra 56 are papers that use the concept
without naming the list.
What this does not establish. A tight probe with no gap is not a probe with full recall — it is a probe whose count is stable under widening in the one direction tried. A paper that resolves domain ownership from a source none of these five names, or names none of them at all, is invisible to every width. See section F.
A.13 The schema-only residue, printed in full
Seven papers whose tools[] / classification[] / otherToolsMentioned[] name one of the five resources but whose full text does not match the tight sweep. Printed rather than dropped, because an invisible residue is a residue nobody looks at:
IEEE-SP/2024/to-auth-or-not-to-auth-a-comparative-analysis-of-the-pre-and-post-login-security
named: Disconnect Tracker Protection List
PETS/2022/on-dark-patterns-and-manipulation-of-website-publishers-by-cmps
named: Disconnect | Disconnect tracking filter list
PETS/2024/generalizable-active-privacy-choice-designing-a-graphical-user-interface-for-glo
named: Disconnect Tracker Protection lists
PETS/2024/support-personas-a-concept-for-tailored-support-of-users-of-privacy-enhancing-te
named: Disconnect
USENIX/2019/canvas-fast-and-inexpensive-automotive-network-mapping
named: physical ECU access and disconnection
WWW/2018/the-cost-of-digital-advertisement-comparing-user-and-advertiser-views
named: Disconnect | Disconnect blacklist
WWW/2023/who-funds-misinformation-a-systematic-analysis-of-the-ad-related-profit-routines
named: Disconnect
Six of the seven are Disconnect, and they are the tightened pattern's misses: the extractor wrote “Disconnect Tracker Protection List” or bare “Disconnect” into a field, where the paper's own prose says something the list-sense pattern does not match. The seventh is a homonym the extractor invented — “physical ECU access and disconnection” in an automotive CAN-bus paper is not the Disconnect list. The tight pattern's cost is therefore about five papers, all Disconnect, all in the direction of undercounting, and the page's Disconnect row should be read as 74 rather than 74-to-79 only because the union is an upper bound on use anyway. The alternative — publishing the loose 700 — is not a trade worth making.
A.14 Denominators used on the page, stated once
| Figure on the page | Denominator | Never |
|---|---|---|
| the six per-resource rows | 5,859 for the “share of corpus” column; 1,120 for the “share of crawled” column, with the numerator restricted to that population | 5,859 for both |
| 136 / 2.3% | 5,859 | — |
| 107 / 9.6% | 1,120, numerator restricted to it | 136 ÷ 1,120 |
| per-year shares | that year's own paper count | 5,859 |
| per-venue shares | that venue's own paper count | 5,859, and never the count alone — the venues differ in size by a factor of three |
| 15 / 13 webXray papers | 5,859 and 1,120 respectively | — |
| everything about the three lists | the domain frames: 32,369 (2026-08-17) or 32,337 (2026-09-05) registrable domains | 5,859; these are not corpus figures at all |
B. Reproduction record, 2026-09-11
Every script whose figures the new page inherits was re-run before the move, against the same pinned inputs, and diffed against its committed output. A page whose report script no longer runs is a page whose numbers cannot be refreshed.
| Script | Invocation | Result |
|---|---|---|
report_webxray.mjs | node scripts/report_webxray.mjs –quotes | byte-identical to the committed report_webxray-output.txt before the fix in section C; regenerated after it |
owner_dbs.py | python3 scripts/owner_dbs.py –cache out/webxray/cache –disagreements 25 | reproduces the 2026-08-17 frame exactly: 32,369 domains; 669 / 5,581 / 2,268; 58.6% / 84.3% / 80.5%; 15,651 merged rows; 19 non-hostname rows dropped; 6,941 ICANN and 3,290 PRIVATE rules; 782 of 827 owner names untouched by the fold, 16 of the 45 changed losing a legal suffix |
owner_random_sample.py | python3 scripts/owner_random_sample.py –sample out/owner_sample.json –rows out/adj_rows_scored.json –table –wiki | byte-identical to out/owner_random_sample-output.txt |
owner_selfname_census.py | python3 scripts/owner_selfname_census.py | reproduces 0.8% / 0.7% / 3.2% (25 of 3,215; 250 of 38,368; 250 of 7,850) |
report_ownership_resolution.mjs | node scripts/report_ownership_resolution.mjs | new; section K |
A drift check the carve-out made cheap. The new script re-derives the sweep counts from patterns written out independently and then cross-checks nine figures against report_webxray-output.txt. All nine agree. That cross-check is inside the script and throws, so a future run cannot publish two pages that disagree with each other:
Figure this script report_webxray.mjs verdict ----------------------- ----------- ------------------ ------- webXray 15 15 agree Tracker Radar 32 32 agree WhoTracks.me 23 23 agree Crunchbase 23 23 agree Disconnect (list sense) 74 74 agree Public Suffix List 101 101 agree union of the five 136 136 agree union ∩ crawled 107 107 agree crawled denominator 1120 1120 agree
C. A defect the cross-check found: a numerator from one population, a denominator from another
This is the one substantive correction of the sitting, and it was published for three and a half weeks.
report_webxray.mjs printed, at two places:
${pct(ftPapers.length, crawled.length)} of the ${crawled.length} that ran a crawl
${pct(anyOwner.length, crawled.length)} of the ${crawled.length} that crawled
ftPapers and anyOwner are sweeps over the whole corpus. Dividing either by crawled.length pairs a corpus-wide numerator with the crawling denominator: the share is of a population the numerator was not drawn from.
| Published | Correct | Why the error is invisible | |
|---|---|---|---|
| webXray papers, share of the 1,120 | 1.3% (15 ÷ 1,120) | 1.2% (13 ÷ 1,120) | 13 of the 15 are in fact crawling papers, so the two figures differ by one decimal and nothing looks wrong |
| any ownership resource, share of the 1,120 | 12.1% (136 ÷ 1,120) | 9.6% (107 ÷ 1,120) | a 2.5-point error on the more load-bearing figure |
Both figures were on the live webXray page, and the 12.1% was one of the figures moving to the new page — which is how it was caught: the new script computed the intersection because it had no reason not to, and the cross-check refused to pass. No guard on the old page could have found it. The number-guard family checks that a figure on a page matches the script's output; the script's output was the wrong number, and the page matched it faithfully.
The fix adds a named helper so the mistake cannot recur silently:
const crawledKeys = new Set(crawled.map(key)); // A share of `crawled` needs a numerator drawn from `crawled`. const inCrawled = (ks) => ks.filter((k) => crawledKeys.has(k));
report_webxray-output.txt was regenerated and both pages corrected in the same sitting. The webXray page carries a footnote at the changed figure rather than silently restating it.
D. Folding, and what it could not reach
This page's corpus figures need no name fold: every count is a regex over full text or a set membership, and there are no free-text names being aggregated. That is stated rather than assumed, because it is unusual for a page on this wiki.
The live-list figures do fold, and the fold and its residue are published on webxray §B.11. Restated here only in summary, because the content page quotes both halves:
- Legal-form suffixes only (
Inc,LLC,GmbH,S.A.S, …), plus punctuation normalisation. Deliberately no synonyms, so the “agree” columns are a lower bound on real agreement and the “neither” column an upper bound on real disagreement. - Residue: the fold leaves 782 of webXray's 827 owner names untouched (94.6%). Of the 45 it changes, only 16 lose a legal-form suffix; the other 29 change on punctuation alone —
56.com,AT&T,Ask.com,Bootstrap_China,Clearstream.TV,Cm_browser,Dictionary.com,Dun & Bradstreet, and the rest are in the published output. - What no fold reaches: “Google” against “Alphabet”, “Xandr” against “Microsoft Corporation”. Those are the granularity and vintage axes, and they are the page's subject rather than a normalisation failure. There is no principled synonym table for them, which is why the page reports the agreement figures as bounds and adjudicates a sample instead.
E. Quotes checked against the source
Checked on 2026-09-11 against data/fulltext/<year>/<venue>/<slug>/paper.cols.txt, whitespace collapsed and quotation marks and dashes normalised, with paper.norm.txt and paper.txt as fallbacks.
| Paper | Quote (head) | Verdict |
|---|---|---|
| [2Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)] | “the product of years of detective work” | found in .cols |
| [2Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)] | “because webxray's database of domain ownership primarily contains major ad networks…” | NOT contiguous — see below |
| [3Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] | “frequently miss connections among two hostnames” | found |
| [3Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] | “twitchcdn.net” | found |
| [4Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] | “preferring those that have been updated most recently” | found |
| [4Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] | “does not aim at providing full transparency on the organizations behind each domain” | found |
| [1Matte, Célestin; Bielova, Nataliia; Santos, Cristiana Teixeira (2020): "Do Cookie Banners Respect my Choice? Measuring Legal Compliance of Banners from IAB Europe's Transparency and Consent Framework", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] | “WebXRay commit 04c3c8e8” | found |
| [5Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)] | “Redacted For Privacy” / “Domains By Proxy” / “Whois Guard” | found, and the surrounding sentence re-read — see below |
The Libert coverage quote is real but is not one contiguous string in any rendering of the PDF. It is the page's longest quotation and it is load-bearing for the coverage section, so it was chased rather than dropped. In paper.cols.txt, paper.norm.txt and paper.txt alike, the two-column layout interleaves a table caption through the sentence:
...because webxray's database of Table 1: Third-Party Prevalence, SSL Use, and First-Party domain ownership primarily contains major ad networks rather Disclosure than small clients, and policyxray only searches for identified †Denotes Company has Consumer Services parties, variability in the long-tail of trackers may not have an outsized effect on overall findings related to disclosure. Nonetheless, Company % Tracked % SSL % Disclosed it is important to point out that the number of parties being searched Google † 82.81 80.35 38.29 for is fewer than the total number of parties present.
Every clause is verbatim and in this order; the interleaved strings are “Table 1: Third-Party Prevalence, SSL Use, and First-Party Disclosure”, “†Denotes Company has Consumer Services” and the table's own header and first row. The page quotes the sentence as the author wrote it and carries a footnote saying exactly this, so a reader running the same grep does not conclude the quote is invented. This is the class of defect a plain substring check reports as a miss and a careless run reports as a fabrication.
One claim was corrected by reading the source. The old page said [5Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)] “resolved organisations for 411 of 762 mobile-specific trackers using CrunchBase, webXray's list, TLS certificates and WHOIS in that order”. The paper's own sentence is narrower:
we successfully retrieved the organization information of 411 trackers from either the CrunchBase database, Tim Libert's library, or TLS certificate, and retrieved the organization […]
The 411 comes from the first three sources; WHOIS contributed a further 251 (“among those 251 trackers whose organization information was retrived from the WHOIS record”). The four sources were tried in that order, but the 411 is not a figure for all four. The new page states it as 411 from the first three and 251 from WHOIS. The consequence for the page's actual point is unchanged: the five registrar-privacy strings total 88, against the largest real company in the top-ten table at 48 (Adobe), and the paper says so itself — “these five organizations cannot represent the real organizations of those trackers”.
F. External sources: verified, and rejected
Repository and vendor state was re-fetched on 2026-09-11, not recalled. Training data is stale by construction for this kind of claim.
| Claim on the page | Primary source, fetched 2026-09-11 | Verdict |
|---|---|---|
Tracker Radar current; last commit on main 2026-08-28; newest release tag 2026.08.28 | api.github.com/repos/duckduckgo/tracker-radar/commits?per_page=1 → a736b501, 2026-08-28T15:25:14Z; /releases → 2026.08.28, 2026.07.27, 2026.06.08, none flagged prerelease | holds |
Tracker Radar's pushed_at reads later than main because of unmerged branches | pushed_at = 2026-09-02T21:21:31Z against main at 2026-08-28 | holds, and the page now names both dates |
| Disconnect current; last commit 2026-08-28 | commits?per_page=1 → 4b592c28, 2026-09-05T20:21:14Z | moved. Page updated to 2026-09-05 and the table header re-dated to 2026-09-11 |
Ghostery trackerdb current; last commit 2026-09-01 | commits?per_page=1 → aa82e6d9, 2026-09-01T14:40:47Z | holds |
| Disconnect's headline “14,332 verified domains and entity mappings” | disconnect.me/trackerprotection, HTTP 200, string present | holds |
github.com/timlib/webXray is HTTP 404; the account has 0 public repositories | 404; api.github.com/users/timlib → 200, public_repos: 0, company: webXray.ai | holds |
github.com/timlib/webXray_Domain_Owner_List is HTTP 404 | 404 | holds |
| No PyPI package | pypi.org/pypi/webxray/json → 404 | holds |
thezedwards/webXray last pushed 2021-03-04, PolyForm Strict | pushed_at 2021-03-04T23:52:47Z; GitHub reports NOASSERTION | holds |
peterjoles/webXray is MIT, last pushed 2023-03-12, 36 commits ahead | spdx_id: MIT, pushed_at 2023-03-12T20:15:26Z | holds |
webxray.org is a placeholder reading “Public interest projects for the interested public.” | fetched with a browser User-Agent: HTTP 200, that exact line and nothing else | holds |
webxray.ai is a live commercial product | HTTP 200 | holds |
Sources rejected, recorded so the next run does not re-add them:
| Rejected | Why |
|---|---|
GitHub's license.spdx_id field for Ghostery trackerdb and Disconnect | reports NOASSERTION for both, which is a statement about GitHub's detector and not about the licence. The licences were read from package.json and LICENSE respectively. |
GitHub's pushed_at as “last activity” | it counts unmerged automation branches. For Tracker Radar it reads four days later than main. Read the branch. |
Any figure for Ghostery trackerdb coverage or accuracy | none was measured. Its .eno + patterns shape needs a parser nobody could check against the other three. The page says this is an omission rather than a judgement, and lists it as an open question. |
| Vendor marketing counts as file sizes | Disconnect's “14,332 verified domains” is a site headline; entities.json holds 7,850 domains. The page quotes both and says which is which. |
| Anything recalled rather than fetched about these repositories | training data predates all four dates above. |
G. What could not be established
- Ghostery
trackerdbis unmeasured. It is a live fourth option and the page gives numbers for two of the three live lists. Closing this needs a parser for the.eno+patternsmodel whose choices are checkable against the other three — a sitting's work, not a footnote. Filed as an open question on the content page. - Recall of the candidate set is unknown. The 136 is a union of six patterns over full text. A paper that resolves domain ownership from a source none of the five names — a bespoke WHOIS pipeline, a commercial feed, an internal list — matches nothing at any width. The page reports 136 as an upper bound on use and says nothing about recall, because nothing here measures it.
- The share of the 136 that actually resolved an owner is not measured. Only two rows (webXray's 15 and Tracker Radar's 32) carry a hand verdict per paper. Doing the same for Disconnect's 74 is the obvious next increment and was not attempted this sitting.
- No accuracy figure exists for the tail. Every rate is conditional on the entry being adjudicable, and 18–30% of each list's draws were not. Whether unadjudicable domains differ systematically between lists is exactly what would decide whether Disconnect's lead survives, and a random sample cannot answer it.
- The coverage census has no neutral denominator. Tracker Radar's own crawl output is the frame, which flatters Tracker Radar. There is no independent census of third-party domains to use instead, and the page says so at every coverage figure rather than pretending otherwise.
- Whether a hierarchy built from a current list would beat both. webXray's
parent_idtree is the only hierarchy any of these files carries and it is frozen; nobody has tried grafting a current corporate-structure source onto Tracker Radar's entities. Filed as an open question. - Two adjudication rows from 2026-08-17 remain unresolved (the acquisition date behind
360yield.com, andfwmrm.net's post-2026-spinoff status). Recorded as unresolved rather than guessed; they are excluded from the 28.
H. Corrections owed to Programming:Crawler, and their real state
The task brief for this sitting said Crawler's comparison row “still says webXray's crawler is PhantomJS (historically)”. It does not, and has not since 2026-08-17. The brief's account of its own state was stale; recorded here because a later sitting reading the brief would look for a correction that is already made.
| Correction | State on 2026-09-11, before this sitting |
|---|---|
| The comparison row's automation column | already fixed. Revision 1786953266 reads “PhantomJS in 2015; consumer Chrome over raw CDP in the last public version” |
| “stale third-party mirrors, the newest last pushed in 2015 and targets PhantomJS” | already fixed in the same revision |
The <WRAP todo> asking where webXray is developed | already closed in the same revision. The page's remaining <WRAP todo> is three unrelated open questions and was correctly left alone |
| The licence and most-complete-copy claim | already fixed on 2026-09-05, revision 1788636934 — the item's own note says so |
The one correction that was owed, and that this sitting created, is the carve-out's own drift: that bullet ended “see webXray for what survives of it: the ownership database, and how it compares to Tracker Radar and Disconnect today”, and the comparison had just left that page. Repointed to Ownership resolution in revision listed below. Nothing else on Crawler was touched: its webXray figure of 7 is the tools[] schema count for a table whose whole column is schema counts, and changing it to the sweep's 15 would make that one cell incomparable with its neighbours.
I. Reviewers
All four were told explicitly that the author's context may not be exhaustive, and all four were handed the page text, the report script, its unedited output and these notes. The three focused passes ran in parallel; the generic pass ran after their findings were applied.
All three focused passes returned and every finding was applied; the generic pass was not run. Each was handed the page text, scripts/report_ownership_resolution.mjs, its unedited output, scripts/report_webxray-output.txt and these notes, against the bundle in review_own/. Every finding below was re-verified by the author before being accepted — a reviewer's report is a lead, not a result.
| Pass | Model | Returned | Findings | Accepted | Rejected |
|---|---|---|---|---|---|
| external currency | Sonnet | yes | 1 substantive of ~40 checks | 1 | 0 |
| citations and quotes | Sonnet | yes | 1 substantive of 22 citekeys, 11 quotes, 13 attributed claims, 7 vendor sources and all 15 rows of the webXray table | 1 | 0 |
| figures versus script | Sonnet | yes | 2 substantive, both inside a published script rather than on a page | 2 | 0 |
| generic, no checklist | Fable | not run — sequenced after the other three | — | — | — |
Accepted: the LLM row was an overclaim (external currency)
The Choosing a resolution source now table read “emerging, and not yet at the domain layer”. The reviewer found arXiv:2606.20868, Can LLMs Reason About Brand Ownership?, submitted 2026-06-18 — four models evaluated on domain-to-brand attribution over 36 heavily-phished brands.
Verified independently before accepting: the abstract was fetched (HTTP 200) and read. The paper is real, the date is right, and the task genuinely is domain-layer ownership attribution. Two qualifications the reviewer did not make, and which changed how it went on the page:
- Its question is phishing and squatting defence — “is this domain the brand's own?” — not third-party tracker attribution. 36 brands is not a coverage claim.
- Its result cuts the other way. Models enumerate a brand's domains at up to 82% precision from memory, but on ownership verification without external tools macro F1 is at most 0.37, rising by up to 0.65 with WHOIS. So the honest update is not “LLMs are arriving, consider them” but “the one published attempt at this task fails without retrieval” — which strengthens the row's existing advice rather than softening it.
The corpus-scoped half of the sentence (“no corpus paper applies this to third-party domain ownership”) was not changed: it was and is true, and an arXiv preprint is outside the seven venues. Cited as a footnote with its arXiv id rather than a {[key]}, so no bibliography entry and no cache purge.
Accepted: an undocumented column splice, of the opposite kind (citations and quotes)
The page footnotes the [2Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)] coverage sentence as non-contiguous in every rendering. The reviewer found that [6Selmo, Carlos; Carisimo, Esteban; Bustamante, Fabián E.; Alvarez-Hamelin, J. Ignacio (2025): "Learning AS-to-Organization Mappings with Borges", in: Proceedings of the 2025 ACM Internet Measurement Conference, pp. 120-133. (DOI)]'s “an Internet shaped by constant mergers, rebrandings, and regional variation” has the inverse problem, unflagged.
Verified independently: the string constant mergers does not appear in paper.cols.txt at all — nor in paper.norm.txt or paper.txt — while pypdf over paper.pdf returns the sentence fully contiguous. So the quote is correct and the corpus's own rendering is the thing that is wrong. Footnoted to say so, because a reader checking the quote against the corpus would otherwise conclude it was invented.
This is the two-renderings rule in both directions on one page: Libert fails in every rendering and is real; Borges fails in the column-repaired rendering and is real. A quote-check against a single rendering would have reported one false FAIL and, had the Libert footnote not already existed, one apparent fabrication.
Accepted: a published script's own ALLOW list carried the fold it warns against (figures versus script)
This is the most valuable finding of the sitting, and section B of this page had just asserted the opposite.
report_webxray.mjs ends with a section Z, Every number on the page that does NOT come from this corpus, which exists so check_page_numbers.mjs can tell a non-corpus figure from a corpus one. Five of its lines were hardcoded constants, and they were the ICANN+PRIVATE fold with exact-key lookup — the variant owner_dbs.py's own output labels <- THE TRAP and both content pages warn against by name, because it manufactures a coverage hole out of googleapis.com:
| Section Z said | The published fold gives |
|---|---|
| 45,525 registrable domains | 32,369 |
| webXray 657 (1.4%), TR 5,539 (12.2%), Disconnect 2,277 (5.0%) | 669 (2.1%), 5,581 (17.2%), 2,268 (7.0%) |
| weighted 54.5% / 79.3% / 75.8% | 58.6% / 84.3% / 80.5% |
| top 100: 70% / 95% / 91% | 71% / 98% / 94% |
| 197 of 601, 216 of 461, 584 of 1,532 | 198 of 612, 217 of 464, 584 of 1,538 |
| “16,396 of 47,836 rows are hostnames” | the label-count heuristic the page says “answers neither question; do not use it” — replaced with the ICANN fold's 15,651 |
Verified independently before accepting: each figure was matched against the live owner_dbs.py run in section B, and grepped for on both content pages. None appears as prose on a live page, so nothing a reader saw was wrong — but the ALLOW list of a number guard is exactly where a wrong figure does its damage silently, by blessing the wrong value if it ever reaches a page. Corrected, and the corrected lines carry a comment naming the trap.
What this says about section B of this page. B calls report_webxray.mjs's output “byte-identical” to its committed copy, and it was — before and after. Byte-identity proves a script matches its own last run. It proves nothing about whether the content is still true, and here the script had been faithfully reproducing a stale constant. A reproduction record is not a correctness record, and this page should not have implied otherwise.
Accepted: a guard that loses its own warning (figures versus script)
report_webxray.mjs buffered every line and flushed once at the bottom. Its section-A guard —
if (roleResidue.length || roleExtra.length) { P('FAILURE: the hand map and the sweep disagree. Re-read the new papers before publishing any figure below.'); }
— called P() and let the script continue. Section B then dereferences ROLE.get(k).role on the now-missing key and throws, before the single console.log at the bottom has run.
Reproduced by deleting one ROLE entry: stdout 0 bytes, a bare TypeError stack, and the FAILURE line that exists for precisely that case nowhere at all. The guard fires and is then thrown away. Fixed with a fail() helper that flushes the buffer, repeats the message on stderr and exits 1; re-running the same mutation now prints 1,283 bytes ending in the named paper and the FAILURE line, and exits 1. On a healthy run the output is byte-identical to before, so the fix is behaviour-neutral where it should be.
What the reviewers confirmed, which is also a result
- All 22 distinct citekeys across the three pages resolve to exactly one bibliography entry; no collisions anywhere in the 993-entry file.
- All 11 quoted strings and 13 attributed non-quote claims verified verbatim or in substance, including the self-contradictory [3Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] sentence this page declines to derive a percentage from, and the [5Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)] 411/251 split this sitting corrected.
- All 15 rows of the webXray publication table verified against source text — the reviewer was asked for five and did fifteen.
- Every figure on both content pages independently re-derived from a live re-run of all six scripts, including the inter-rater kappa, whose invocation the reviewer had to reconstruct because it was not in the brief. Both published code snippets (
owner_lookup.pyandresolve()) executed against the cached JSON and their output matched the page character for character. - Three further mutations of
report_ownership_resolution.mjsby the reviewer, including a re-creation of this sitting's own denominator bug — all caught, none vacuous. - Every repository, licence, URL, version, dependency and acquisition claim on all three pages re-fetched and holding to the day, including the 2020-06-24 Disconnect relicensing commits, PolyForm Strict still having no SPDX id, and
lxml/psycopg2-binarywheels still stopping at CPython 3.9 against a 2025-10-31 Python 3.9 EOL.
Noted, not acted on
| Observation | Disposition |
|---|---|
| Today's live PSL has 6,950 ICANN / 3,375 PRIVATE rules against the 6,941 / 3,290 in the 2026-08-17 snapshot | Not a defect. The figure is snapshot-pinned and the page says the PSL moves. Re-deriving it would orphan the adjudication — the same reason the whole 2026-08-17 frame is frozen. |
The reviewer's Wayback CDX pull put timlib/webXray's last archived 200 at 2022-12-30 where the page says 2023-03-31 | Left as published. The reviewer attributed the gap to CDX collapse-window settings and did not claim the page was wrong. It is a bound either way and both bounds support the same conclusion. Flagged here so a future sitting can settle it rather than rediscover it. |
libert2015_invisible and karaj2018whotracksme could not be quote-checked — neither paper is in the corpus | Correct and expected. Neither is quoted; both are cited only as identifiers. |
| The figures reviewer could not verify section H's claim that Crawler was already corrected, having no access to that page's history | Fair, and the claim stands. Revisions 1786953266 and 1788636934 were read from node scripts/dw.mjs history programming:crawler and the live page text quoted in section H. |
| It could not verify who produced “rater 2” in the inter-rater work, taking the label from the script's own output | A real limit of that measurement, and it is not this page's to fix. Recorded on random_sample where the draw was made. |
| The Sánchez-Rola 2021-versus-2022 venue-year discrepancy | Left as published, with the page's existing italic note that the corpus files it under 2022 while Crossref and the paper's header say 2021. |
Still owed
- The generic (Fable) pass was never run, so nothing has read these pages looking for what the three checklists were not looking for — overstated claims, structure, a page that does not answer its own question. It is the one pass still owed, and it was also meant to review this page.
report_webxray.mjs's section Z is a hand-maintained ALLOW list of constants and that is what let it go stale. It should be derived fromowner_dbs.py's output rather than retyped, or checked against it by a guard that throws. Not done this sitting: it needs the two scripts to share a format, which is a change to both.- The same buffered-output defect may exist in other
report_*.mjsscripts on this wiki. Onlyreport_webxray.mjsandreport_ownership_resolution.mjswere checked.
What the author checked without them, so the gap is bounded rather than open:
| Check | Result |
|---|---|
| All four published scripts re-run against pinned inputs | report_webxray.mjs and owner_random_sample.py byte-identical to their committed output; owner_dbs.py and owner_selfname_census.py reproduce every figure quoted (section B) |
owner_adjudication.py –table | reproduces the 28-row table exactly: 1/18/4/1/4, 8/12/6/2/0, 26/0/2/0/0, and the three error rows (1rx.io, jsdelivr.net, stackadapt.com) |
domain_owners.json schema table on webXray, every cell | reproduces: 319/827 parent_id, depths 508/211/80/24/2/2, 761 uses, 803 platforms, 826 country, 104 trade_groups, 591 policy URLs, 132 GDPR, 4 CCPA, 13 opt-out, 34 crunchbase_id, 52 health-segment, 263 notes, 275 aliases, 69 languages, 2,175 of 3,215 |
Mutation test of report_ownership_resolution.mjs | four mutations, four failures. Swapping tight/loose on Disconnect → “tight (700) > loose (74)”. Making loose disjoint from tight at equal size → “23 tight hits are not in the loose set” — the containment guard catches what a size comparison alone would miss. Injecting a fake union member → cross-check throws. Replacing the crawled intersection with the full union → cross-check throws. No guard is vacuous. |
| Arithmetic audit of every ratio on the content page | all consistent; one error found — “seven times NDSS's rate” is 8.4×, corrected |
| Rendered DOM of both content pages | 18 references for 18 keys on the new page, 17 for 17 on webXray; 12 and 7 tables; 9 and 4 WRAP boxes; every in-page and cross-page anchor resolves; one broken anchor found — the webXray availability table pointed at #Methodology and limitations of these figures for a Wayback timeline that had become its own section |
| Site-wide inbound-link sweep | four pages found and fixed (section above) |
| Table cells containing a pipe | one found in this page's own draft: a nowiki-wrapped [[a|b]] literal in a table cell. A parsed link's pipe is safe in a cell (verified in the rendered DOM); a nowiki-wrapped one is not, because the cell splits before the nowiki is honoured. Rewritten as prose, and the A.12 regex table was moved into a code block for the same reason. |
| Full-text corpus quality | 5,855 of 5,859 renderings present, 0 empty, 90 carrying a NUL byte — see A1b |
J. scripts/report_ownership_resolution.mjs, in full
- report_ownership_resolution.mjs
// Every corpus figure on Design:Ownership resolution, with its denominator. // // node scripts/report_ownership_resolution.mjs > scripts/report_ownership_resolution-output.txt // node scripts/report_ownership_resolution.mjs --wiki # DokuWiki tables // // This page was carved out of Programming:Crawler:webXray on 2026-09-11. The // corpus-side figures it inherited were computed by scripts/report_webxray.mjs, // whose population is "papers that name webXray". That is the wrong population // for a page about ownership resolution in general, so this script re-derives // the landscape figures from the corpus independently and then CROSS-CHECKS the // overlapping rows against report_webxray.mjs's committed output. Two // implementations agreeing is a stronger audit trail than one copied number. // // Three things this script does that a naive sweep would not: // // 1. Every probe is run at TWO widths and the script asserts tight <= loose AND // tight subset-of loose. A narrowing probe that returns more hits, or hits // the loose one does not contain, is a broken probe, not a finding. The // Disconnect probe is the one that needs this: plain /disconnect/i is a // common English word and the CSP literature uses it as a technical term. // // 2. Every count is a PAPER count over data/fulltext/<year>/<venue>/<slug>/ // paper.cols.txt with whitespace collapsed, because a PDF line break inside // "Public Suffix List" silently undercounts it. // // 3. Papers with no full text are counted and printed, so the sweep's own // denominator is visible rather than assumed to be 5,859. // // Every row here is an UPPER BOUND on use unless the output says hand-verified: // a full-text match is a mention, not a use. The two rows that ARE hand // verified (webXray, Tracker Radar) carry their verdicts in report_webxray.mjs. import fs from 'node:fs'; import path from 'node:path'; import { dataRoot, loadExtractions, POPULATIONS, pct, table, wikiTable } from './lib.mjs'; const WIKI = process.argv.includes('--wiki'); const rows = loadExtractions(); const key = (p) => `${p.venue}/${p.year}/${p.slug}`; const FT = path.join(dataRoot(), 'fulltext'); const out = []; const P = (s = '') => out.push(s); const heading = (s) => { P(''); P(`=== ${s} ===`); P(''); }; const T = (headers, body) => P(WIKI ? wikiTable(headers, body) : table(headers, body)); // --- full text, whitespace collapsed ------------------------------------- // Streamed, not cached: the corpus is 5,859 papers and holding every // paper.cols.txt in a Map exhausts the default V8 heap. One pass, every // pattern tested against each paper's text, then the text is dropped. function readText(k) { const [venue, year, slug] = k.split('/'); const f = `${FT}/${year}/${venue}/${slug}/paper.cols.txt`; return fs.existsSync(f) ? fs.readFileSync(f, 'utf8').replace(/\s+/g, ' ') : null; } // Which papers have full text at all, and the hits for every pattern, in one // pass. PATTERNS is filled below before this runs. function sweepAll(patterns) { const hits = new Map([...patterns.keys()].map((n) => [n, []])); const haveText = new Set(); for (const p of rows) { const k = key(p); const t = readText(k); if (t === null) continue; haveText.add(k); for (const [name, re] of patterns) if (re.test(t)) hits.get(name).push(k); } for (const v of hits.values()) v.sort(); return { hits, haveText }; } // --- the probes, each at two widths -------------------------------------- // `loose` must be a strict superset of `tight` by construction; the script // checks that it is in fact one and throws if not. const PROBES = [ { name: 'webXray', tight: /webx[\s-]?ray/i, loose: /webx[\s-]?ray|libert/i, note: 'hand-verified: all 15, in report_webxray.mjs', }, { name: 'Tracker Radar', tight: /tracker[\s.-]?radar/i, loose: /tracker[\s.-]?radar|duckduckgo/i, note: 'hand-verified: all 32, in report_webxray.mjs', }, { name: 'WhoTracks.me', tight: /whotracks/i, loose: /whotracks|who ?tracks ?\.? ?me|ghostery/i, note: 'upper bound', }, { name: 'Crunchbase', tight: /crunchbase/i, loose: /crunchbase|crunch base/i, note: 'upper bound', }, { name: 'Disconnect (list sense)', tight: /disconnect'?s? (entit|list|block|black|tracking)|disconnect\.me|entities\.json/i, loose: /disconnect/i, note: 'upper bound; the loose width is the reason this one is tightened', }, { name: 'Public Suffix List', tight: /public suffix/i, loose: /public suffix|etld\+?1|effective top-?level domain/i, note: 'upper bound', }, ]; // One pass over the full text: every probe, both widths, plus the set of papers // that have a readable rendering at all. const PATTERNS = new Map(); for (const pr of PROBES) { PATTERNS.set(`${pr.name}|tight`, pr.tight); PATTERNS.set(`${pr.name}|loose`, pr.loose); } const { hits: SWEPT, haveText: HAVE_TEXT } = sweepAll(PATTERNS); const withText = rows.filter((p) => HAVE_TEXT.has(key(p))); const crawled = rows.filter(POPULATIONS.crawled); P('=============================================================================='); P('Design:Ownership resolution — every corpus figure, with its denominator'); P('=============================================================================='); P(''); P(`Corpus: ${rows.length} extracted papers, 7 venues (CCS, IMC, NDSS, PETS, USENIX`); P('Security, TheWebConf, IEEE S&P), 2010–2026. EuroS&P, ACSAC, RAID, AsiaCCS, CHI'); P('and SOUPS are absent, so every figure here is a claim about those seven venues.'); P(''); P(`Papers with a readable paper.cols.txt: ${withText.length} of ${rows.length} (${pct(withText.length, rows.length)}).`); P(`Papers that ran a crawl (crawlConfig present or studyTypes includes automated-web-crawl): ${crawled.length}.`); P('The sweep denominator is the full corpus; papers with no full text can only'); P('ever be misses, so every sweep count below is a floor on the true mention count.'); // --- A. probe widths ------------------------------------------------------ heading('A. Probe widths: tight must be contained in loose'); const results = new Map(); const widthRows = []; for (const pr of PROBES) { const tight = SWEPT.get(`${pr.name}|tight`); const loose = SWEPT.get(`${pr.name}|loose`); const tightSet = new Set(tight); const looseSet = new Set(loose); const escaped = tight.filter((k) => !looseSet.has(k)); if (tight.length > loose.length) { throw new Error(`probe ${pr.name}: tight (${tight.length}) > loose (${loose.length}) — the narrowing probe is broken`); } if (escaped.length > 0) { throw new Error(`probe ${pr.name}: ${escaped.length} tight hits are not in the loose set: ${escaped.slice(0, 5).join(', ')}`); } results.set(pr.name, tightSet); widthRows.push([pr.name, String(tight.length), String(loose.length), pr.note]); } T(['Resource', 'Tight probe', 'Loose probe', 'Status'], widthRows); P(''); P('Both widths are printed because the gap is the claim. Disconnect at the loose'); P('width matches 700 papers and almost none of them mean the list, which is why'); P('the published row is the tight one and why the page says so. A row whose two'); P('widths are close is a row whose count is not an artefact of the pattern.'); // --- B. the published landscape table ------------------------------------ heading('B. The ownership-resolution landscape (the page\'s main table)'); const landscape = PROBES.map((pr) => { const s = results.get(pr.name); const inCrawled = crawled.filter((p) => s.has(key(p))).length; return [pr.name, String(s.size), pct(s.size, rows.length), String(inCrawled), pct(inCrawled, crawled.length), pr.note.startsWith('hand-verified') ? 'yes' : 'no — upper bound']; }) .sort((a, b) => Number(b[1]) - Number(a[1])); T(['Resource named in full text', 'Papers', `Share of ${rows.length}`, `of which crawled`, `Share of ${crawled.length}`, 'Hand-verified?'], landscape); // --- C. the union --------------------------------------------------------- heading('C. The union: how widespread is ownership resolution at all?'); // The union is over the five OWNERSHIP resources. The Public Suffix List is // deliberately excluded: it answers "what is the registrable domain", not // "whose is it", and including it would nearly double the union on a resource // that every crawl paper touches for unrelated reasons. const UNION_NAMES = ['webXray', 'Tracker Radar', 'WhoTracks.me', 'Crunchbase', 'Disconnect (list sense)']; const union = new Set(); for (const n of UNION_NAMES) for (const k of results.get(n)) union.add(k); P(`Union of ${UNION_NAMES.join(' / ')}:`); P(` ${union.size} papers — ${pct(union.size, rows.length)} of the ${rows.length}-paper corpus,`); const unionCrawled = crawled.filter((p) => union.has(key(p))); P(` and ${pct(unionCrawled.length, crawled.length)} of the ${crawled.length} that ran a crawl (${unionCrawled.length} papers).`); P(''); P('The Public Suffix List is NOT in the union: it answers "what is the registrable'); P('domain", not "whose is it". Adding it would take the union to ' + (() => { const u2 = new Set(union); for (const k of results.get('Public Suffix List')) u2.add(k); return u2.size; })() + ' on a resource'); P('every crawling paper touches for unrelated reasons.'); P(''); P('Each member is an upper bound on USE, so the union is an upper bound too:'); P('136 papers mention one of these; far fewer resolved an owner with one.'); // --- D. per-year adoption ------------------------------------------------- heading('D. Per-year: is ownership resolution growing?'); const years = [...new Set(rows.map((p) => p.year))].sort(); const yearRows = years.map((y) => { const inYear = rows.filter((p) => p.year === y); const hit = inYear.filter((p) => union.has(key(p))); const star = y >= 2025 ? '*' : ''; return [`${y}${star}`, String(inYear.length), String(hit.length), pct(hit.length, inYear.length)]; }); T(['Year', 'Corpus papers', 'Naming an ownership resource', 'Share'], yearRows); P(''); P('* 2025 and 2026 are the provisional corpus edge: CCS 2026 and IMC 2026 have not'); P(' been held, and IEEE S&P / WWW 2026 abstracts are not in OpenAlex, so selection'); P(' under-samples them BY CONSTRUCTION. Do not read 2026 as a complete year.'); P(' The share, not the count, is the readable column for those two rows.'); // --- E. venue shape ------------------------------------------------------- heading('E. Venue shape of the union'); const venues = [...new Set(rows.map((p) => p.venue))].sort(); const venueRows = venues.map((v) => { const inVenue = rows.filter((p) => p.venue === v); const hit = inVenue.filter((p) => union.has(key(p))); return [v, String(inVenue.length), String(hit.length), pct(hit.length, inVenue.length)]; }).sort((a, b) => Number(b[2]) - Number(a[2])); T(['Venue', 'Corpus papers', 'Naming an ownership resource', 'Share of venue'], venueRows); P(''); P('Read the SHARE column, not the count: the venues differ in size by a factor of'); P('three. IMC and PETS are where this work lands; the shares are small everywhere,'); P('which is the honest headline — ownership resolution is a step inside a paper,'); P('not a paper topic, and most papers that do it do not say how.'); // --- F. the schema's own view -------------------------------------------- heading('F. What the extraction schema sees, against the sweep'); // classification[].groundTruthSource / resourceName naming an ownership list. const SCHEMA_RE = /webx[\s-]?ray|tracker[\s.-]?radar|whotracks|disconnect|crunchbase/i; const schemaPapers = new Set(); for (const p of rows) { const names = [ ...p.tools.map((t) => t.name), ...p.classification.map((c) => c.resourceName), ...p.classification.map((c) => c.groundTruthSource), ...p.otherToolsMentioned.map((t) => t.name), ].filter((s) => typeof s === 'string'); if (names.some((n) => SCHEMA_RE.test(n))) schemaPapers.add(key(p)); } const both = [...union].filter((k) => schemaPapers.has(k)); P(`Papers whose tools[] / classification[] / otherToolsMentioned[] name one of the`); P(`five resources: ${schemaPapers.size}`); P(`Papers whose FULL TEXT names one: ${union.size}`); P(`In both: ${both.length}`); P(`Full text only (schema misses): ${union.size - both.length}`); P(`Schema only (no full-text match): ${schemaPapers.size - both.length}`); P(''); P('The schema-only residue is printed rather than dropped:'); const schemaOnly = [...schemaPapers].filter((k) => !union.has(k)).sort(); for (const k of schemaOnly) { const p = rows.find((r) => key(r) === k); const hits = [ ...p.tools.map((t) => t.name), ...p.classification.map((c) => c.resourceName), ...p.classification.map((c) => c.groundTruthSource), ...p.otherToolsMentioned.map((t) => t.name), ].filter((s) => typeof s === 'string' && SCHEMA_RE.test(s)); P(` ${k}`); P(` named: ${[...new Set(hits)].join(' | ')}`); P(` full text present: ${HAVE_TEXT.has(k)}`); } P(''); P('This is why the page publishes the full-text sweep and not the schema field:'); P('the schema fires on a fraction of the papers that name these lists, because'); P('most mentions are in related work or a reference list rather than in a tools'); P('sentence the extractor reads as a tool. Neither signal is "the" population;'); P('both are reported.'); // --- G. cross-check against report_webxray.mjs --------------------------- heading('G. Cross-check against scripts/report_webxray-output.txt'); const committed = 'scripts/report_webxray-output.txt'; if (!fs.existsSync(committed)) { P(`${committed} not found — cross-check SKIPPED.`); } else { const txt = fs.readFileSync(committed, 'utf8'); const CHECKS = [ ['webXray', /^webXray\s+(\d+)\s/m], ['Tracker Radar', /^Tracker Radar\s+(\d+)\s/m], ['WhoTracks.me', /^WhoTracks\.me\s+(\d+)\s/m], ['Crunchbase', /^Crunchbase\s+(\d+)\s/m], ['Disconnect (list sense)', /^Disconnect \(list sense\)\s+(\d+)\s/m], ['Public Suffix List', /^Public Suffix List\s+(\d+)\s/m], ]; const checkRows = []; let bad = 0; for (const [name, re] of CHECKS) { const m = txt.match(re); if (m === null) throw new Error(`cross-check: could not find "${name}" in ${committed}`); const theirs = Number(m[1]); const mine = results.get(name).size; if (theirs !== mine) bad += 1; checkRows.push([name, String(mine), String(theirs), theirs === mine ? 'agree' : '*** DISAGREE ***']); } const um = txt.match(/^\s*(\d+) papers \((\d+\.\d)% of the corpus\)\. (\d+) of them are among the (\d+) that crawled -- (\d+\.\d)% of that population\./m); if (um === null) throw new Error(`cross-check: could not find the union line in ${committed}`); if (Number(um[1]) !== union.size) bad += 1; checkRows.push(['union of the five', String(union.size), um[1], Number(um[1]) === union.size ? 'agree' : '*** DISAGREE ***']); if (Number(um[3]) !== unionCrawled.length) bad += 1; checkRows.push(['union ∩ crawled', String(unionCrawled.length), um[3], Number(um[3]) === unionCrawled.length ? 'agree' : '*** DISAGREE ***']); if (Number(um[4]) !== crawled.length) bad += 1; checkRows.push(['crawled denominator', String(crawled.length), um[4], Number(um[4]) === crawled.length ? 'agree' : '*** DISAGREE ***']); T(['Figure', 'this script', 'report_webxray.mjs', 'verdict'], checkRows); P(''); if (bad > 0) { throw new Error(`${bad} figures disagree between the two independent implementations — do not publish either.`); } P('All rows agree. The two scripts share no code beyond lib.mjs and the'); P('whitespace-collapse convention; the patterns were written out separately.'); } // --- H. the page's non-corpus figures, listed so they are not mistaken ---- heading('H. Every number on the page that does NOT come from this corpus'); P('Listed so a reader auditing the page knows which script to re-run, and so'); P('that a refresh of this script is never mistaken for a refresh of the page.'); P(''); T(['Figure family', 'Script', 'Snapshot it is pinned to'], [ ['List shape (827 / 19,148 / 1,887 owners; 3,215 / 38,368 / 7,850 domains)', 'owner_dbs.py --cache out/webxray/cache', '2026-08-17'], ['Coverage census (669 / 5,581 / 2,268 of 32,369; the prevalence weights; the rank slices)', 'owner_dbs.py --cache out/webxray/cache', '2026-08-17'], ['Agreement between pairs (612 / 464 / 1,538 and the "neither" columns)', 'owner_dbs.py --cache out/webxray/cache', '2026-08-17'], ['The 28-row hand adjudication of high-prevalence disagreements', 'owner_adjudication.py --table', '2026-08-17'], ['Accuracy rates (69.9 / 73.7 / 94.4%, the encounter-weighted column, the CIs)', 'owner_random_sample.py --sample out/owner_sample.json --rows out/adj_rows_scored.json', '2026-09-05'], ['Coverage counts inside the random-sample section (664 / 5,566 / 2,264 of 32,337)', 'owner_sample.py --cache out/owner_cache', '2026-09-05'], ['Self-named entity shares (0.8% / 0.7% / 3.2%)', 'owner_selfname_census.py', '2026-09-05'], ['Inter-rater kappa (+0.400, +0.680)', 'owner_irr_kappa.py', '2026-09-11'], ['Repository state, licences, Wayback dates', 'hand checks + wayback_webxray.sh', '2026-08-17 / 2026-09-05'], ]); P(''); P('Two of the three lists change weekly. Re-run before citing.'); console.log(out.join('\n'));
K. Its unedited output
Run as node scripts/report_ownership_resolution.mjs.
- report_ownership_resolution-output.txt
============================================================================== Design:Ownership resolution — every corpus figure, with its denominator ============================================================================== Corpus: 5859 extracted papers, 7 venues (CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P), 2010–2026. EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS are absent, so every figure here is a claim about those seven venues. Papers with a readable paper.cols.txt: 5855 of 5859 (99.9%). Papers that ran a crawl (crawlConfig present or studyTypes includes automated-web-crawl): 1120. The sweep denominator is the full corpus; papers with no full text can only ever be misses, so every sweep count below is a floor on the true mention count. === A. Probe widths: tight must be contained in loose === Resource Tight probe Loose probe Status ----------------------- ----------- ----------- ---------------------------------------------------------------- webXray 15 233 hand-verified: all 15, in report_webxray.mjs Tracker Radar 32 108 hand-verified: all 32, in report_webxray.mjs WhoTracks.me 23 124 upper bound Crunchbase 23 23 upper bound Disconnect (list sense) 74 700 upper bound; the loose width is the reason this one is tightened Public Suffix List 101 157 upper bound Both widths are printed because the gap is the claim. Disconnect at the loose width matches 700 papers and almost none of them mean the list, which is why the published row is the tight one and why the page says so. A row whose two widths are close is a row whose count is not an artefact of the pattern. === B. The ownership-resolution landscape (the page's main table) === Resource named in full text Papers Share of 5859 of which crawled Share of 1120 Hand-verified? --------------------------- ------ ------------- ---------------- ------------- ---------------- Public Suffix List 101 1.7% 46 4.1% no — upper bound Disconnect (list sense) 74 1.3% 65 5.8% no — upper bound Tracker Radar 32 0.5% 28 2.5% yes WhoTracks.me 23 0.4% 19 1.7% no — upper bound Crunchbase 23 0.4% 11 1.0% no — upper bound webXray 15 0.3% 13 1.2% yes === C. The union: how widespread is ownership resolution at all? === Union of webXray / Tracker Radar / WhoTracks.me / Crunchbase / Disconnect (list sense): 136 papers — 2.3% of the 5859-paper corpus, and 9.6% of the 1120 that ran a crawl (107 papers). The Public Suffix List is NOT in the union: it answers "what is the registrable domain", not "whose is it". Adding it would take the union to 221 on a resource every crawling paper touches for unrelated reasons. Each member is an upper bound on USE, so the union is an upper bound too: 136 papers mention one of these; far fewer resolved an owner with one. === D. Per-year: is ownership resolution growing? === Year Corpus papers Naming an ownership resource Share ----- ------------- ---------------------------- ----- 2010 119 0 0.0% 2011 116 1 0.9% 2012 151 0 0.0% 2013 125 0 0.0% 2014 166 0 0.0% 2015 190 0 0.0% 2016 182 2 1.1% 2017 231 4 1.7% 2018 254 5 2.0% 2019 402 9 2.2% 2020 404 15 3.7% 2021 379 13 3.4% 2022 546 19 3.5% 2023 719 19 2.6% 2024 690 17 2.5% 2025* 770 21 2.7% 2026* 415 11 2.7% * 2025 and 2026 are the provisional corpus edge: CCS 2026 and IMC 2026 have not been held, and IEEE S&P / WWW 2026 abstracts are not in OpenAlex, so selection under-samples them BY CONSTRUCTION. Do not read 2026 as a complete year. The share, not the count, is the readable column for those two rows. === E. Venue shape of the union === Venue Corpus papers Naming an ownership resource Share of venue ------- ------------- ---------------------------- -------------- PETS 510 49 9.6% USENIX 1410 19 1.3% WWW 843 19 2.3% IMC 638 17 2.7% CCS 990 13 1.3% IEEE-SP 767 11 1.4% NDSS 701 8 1.1% Read the SHARE column, not the count: the venues differ in size by a factor of three. IMC and PETS are where this work lands; the shares are small everywhere, which is the honest headline — ownership resolution is a step inside a paper, not a paper topic, and most papers that do it do not say how. === F. What the extraction schema sees, against the sweep === Papers whose tools[] / classification[] / otherToolsMentioned[] name one of the five resources: 89 Papers whose FULL TEXT names one: 136 In both: 82 Full text only (schema misses): 54 Schema only (no full-text match): 7 The schema-only residue is printed rather than dropped: IEEE-SP/2024/to-auth-or-not-to-auth-a-comparative-analysis-of-the-pre-and-post-login-security named: Disconnect Tracker Protection List full text present: true PETS/2022/on-dark-patterns-and-manipulation-of-website-publishers-by-cmps named: Disconnect | Disconnect tracking filter list full text present: true PETS/2024/generalizable-active-privacy-choice-designing-a-graphical-user-interface-for-glo named: Disconnect Tracker Protection lists full text present: true PETS/2024/support-personas-a-concept-for-tailored-support-of-users-of-privacy-enhancing-te named: Disconnect full text present: true USENIX/2019/canvas-fast-and-inexpensive-automotive-network-mapping named: physical ECU access and disconnection full text present: true WWW/2018/the-cost-of-digital-advertisement-comparing-user-and-advertiser-views named: Disconnect | Disconnect blacklist full text present: true WWW/2023/who-funds-misinformation-a-systematic-analysis-of-the-ad-related-profit-routines named: Disconnect full text present: true This is why the page publishes the full-text sweep and not the schema field: the schema fires on a fraction of the papers that name these lists, because most mentions are in related work or a reference list rather than in a tools sentence the extractor reads as a tool. Neither signal is "the" population; both are reported. === G. Cross-check against scripts/report_webxray-output.txt === Figure this script report_webxray.mjs verdict ----------------------- ----------- ------------------ ------- webXray 15 15 agree Tracker Radar 32 32 agree WhoTracks.me 23 23 agree Crunchbase 23 23 agree Disconnect (list sense) 74 74 agree Public Suffix List 101 101 agree union of the five 136 136 agree union ∩ crawled 107 107 agree crawled denominator 1120 1120 agree All rows agree. The two scripts share no code beyond lib.mjs and the whitespace-collapse convention; the patterns were written out separately. === H. Every number on the page that does NOT come from this corpus === Listed so a reader auditing the page knows which script to re-run, and so that a refresh of this script is never mistaken for a refresh of the page. Figure family Script Snapshot it is pinned to ---------------------------------------------------------------------------------------- ------------------------------------------------------------------------------------- ------------------------ List shape (827 / 19,148 / 1,887 owners; 3,215 / 38,368 / 7,850 domains) owner_dbs.py --cache out/webxray/cache 2026-08-17 Coverage census (669 / 5,581 / 2,268 of 32,369; the prevalence weights; the rank slices) owner_dbs.py --cache out/webxray/cache 2026-08-17 Agreement between pairs (612 / 464 / 1,538 and the "neither" columns) owner_dbs.py --cache out/webxray/cache 2026-08-17 The 28-row hand adjudication of high-prevalence disagreements owner_adjudication.py --table 2026-08-17 Accuracy rates (69.9 / 73.7 / 94.4%, the encounter-weighted column, the CIs) owner_random_sample.py --sample out/owner_sample.json --rows out/adj_rows_scored.json 2026-09-05 Coverage counts inside the random-sample section (664 / 5,566 / 2,264 of 32,337) owner_sample.py --cache out/owner_cache 2026-09-05 Self-named entity shares (0.8% / 0.7% / 3.2%) owner_selfname_census.py 2026-09-05 Inter-rater kappa (+0.400, +0.680) owner_irr_kappa.py 2026-09-11 Repository state, licences, Wayback dates hand checks + wayback_webxray.sh 2026-08-17 / 2026-09-05 Two of the three lists change weekly. Re-run before citing.
Related
- Ownership resolution — the page these notes are for.
- webxray — the live-list scripts, the fold residues, the Wayback log, the 2026-08-17 and 2026-09-05 run tables.
- random_sample — the 175-row adjudication, the estimator, the inter-rater draw and kappa.
- Corpus — corpus-level selection and extraction caveats.
References
- [1]
- Matte, Célestin; Bielova, Nataliia; Santos, Cristiana Teixeira (2020): "Do Cookie Banners Respect my Choice? Measuring Legal Compliance of Banners from IAB Europe's Transparency and Consent Framework", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
- [2]
- Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)
- [3]
- Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
- [4]
- Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
- [5]
- Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)
- [6]
- Selmo, Carlos; Carisimo, Esteban; Bustamante, Fabián E.; Alvarez-Hamelin, J. Ignacio (2025): "Learning AS-to-Organization Mappings with Borges", in: Proceedings of the 2025 ACM Internet Measurement Conference, pp. 120-133. (DOI)
