User Tools

Site Tools


provenance:design:ownership_resolution

This is an old revision of the document!


Provenance: Design:Ownership resolution

Working notes behind Ownership resolution. Every query with its population and denominator, the report script and its unedited output, what was folded and what the fold could not reach, the quotes checked against the source, the external sources verified and rejected, and the judgement calls. Corpus-wide selection and extraction caveats are on Corpus and are not restated here.

This page is a carve-out, and most of its audit trail lives elsewhere on purpose. The content page was split out of webXray on 2026-09-11. The scripts that produce its live-database figures, their unedited output, the 175-row adjudication table and the inter-rater material were published in 2026-08 and 2026-09 under the webXray ids and keep those ids, because they are cited from the published record and moving them would break links for no gain:

What Where
The three-list comparison scripts (owner_dbs.py, owner_adjudication.py), their unedited output, the fold residues, the Wayback log, the 2026-08-17 and 2026-09-05 run tables webxray
owner_sample.py, owner_random_sample.py, the full 175-row adjudication with every source, the citation re-fetch, the inter-rater draw and kappa random_sample
The corpus population of webXray itself (the 15 papers, the hand role map), report_webxray.mjs webxray
This page the split itself, the new corpus script and its output, the reproduction record of 2026-09-11, the defect that run found, the reviewer log

No ~~DISCUSSION~~ block: comments belong on the content page. Citations use the same {[citekey]} keys and the same shared Bibliography; this page adds no bibliography entries of its own.

Updated 2026-09-11 (second sitting the same day). The 40 sampled domains the 2026-09-05 pass could not adjudicate were re-adjudicated against harder sourcing and 28 of them settled, which changed the content page's headline table, its error bullet, its residue bullet and — the consequential one — its Fisher table, where the one comparison that survived Bonferroni no longer does. Section L below is that pass in full: its queries, its new source kinds, the citations it broke and fixed, the 12 rows still unresolved with every route tried, and its own reviewer log. Everything above section L is the 2026-09-05 sitting as it stood and is not rewritten, except the query table in section K, which names the script and rows file the live figures now come from.

The run

Date 2026-09-11
Corpus at the time data/extract/run1/extractions.jsonl, 5,859 papers, 7 venues, 2010–2026. 5,855 have a readable paper.cols.txt.
What was done Created Ownership resolution; rewrote webXray around what was left; repointed Crawler; added a row to Design; turned the Roadmap row blue.
New code scripts/report_ownership_resolution.mjs (346 lines), committed with its output.
Code changed scripts/report_webxray.mjs — a denominator defect, section C below.
Models Opus 5 wrote the pages and the script. Review layer: three focused Sonnet passes and one Fable generic pass — all four ran, section I. The 2026-08-17 and 2026-09-05 ownership adjudications this page inherits were also made by language models, not by a human expert; that disclosure belongs on the content page and section I records that the carve-out dropped it.
Accidental exposure or mistakes caught no exposure. Nine defects, all caught before or during review and all listed in C and I: the denominator in C; a stale ALLOW list and an unreachable guard in report_webxray.mjs; a dropped LLM-adjudicator disclosure; an asserted p-value that was never computed; a raw-versus-post-stratified rate description; three arithmetic slips; four stale inbound links; a broken anchor; a pipe in a table cell. The most serious two — the disclosure and the p-value — were found by the generic pass, which is the pass with no checklist.

Scope decision: why a separate page, when the last sitting decided against one

This reverses a recorded decision, so the reversal is recorded too. webxray §“Scope decision” says, on 2026-08-17:

A separate design:ownership_resolution page with webXray as a stub. Rejected: it would leave the wiki's existing red link pointing at a stub, and the material does not split cleanly — webXray's frozen list is the best available illustration of what goes wrong.

Both halves of that reasoning have since stopped holding:

  • “webXray as a stub” was the wrong alternative to compare against. There is enough webXray-specific material for a real page — the three-licence history across four states, the availability census, the domain_owners.json schema, the frozen 2016 PSL, the Wayback timeline, and 15 corpus papers with a hand role map. What is left after the carve-out is 27 KB, not a stub.
  • The red-link problem inverted. On 2026-09-07 the Roadmap queued design:ownership_resolution as a promised page and scripts/sitemap.mjs began gating on it, so from that date not writing it was the dangling promise. The id was fixed by that decision and is used unchanged.
  • The discoverability argument is the item's whole point and was never answered. A reader asking “how do I attribute a third-party domain to a company?” does not search for a dead tool. Both earlier sittings recorded this as the right thing to do; neither did it.

Alternatives considered this sitting and rejected:

Alternative Why not
Leave it, and add a redirect or a prominent pointer on webXray A pointer does not fix search, does not fix the namespace (programming: is instruments; this is a design decision), and leaves the currency claim — which source to use now — filed under the historical one.
Broaden Requests instead That page answers “is this request a tracker?”, a different question with different tooling (filter lists, not owner databases). The two pages cross-link.
Broaden IP classification It is the network-layer sibling (AS-to-organisation) and is already large; the domain layer has different sources, different failure modes and different licences. Cross-linked both ways instead.
Move the random_sample code appendix under provenance:design: too Rejected. Its id is cited from the published content page and from the webXray provenance page; the gain is cosmetic and the cost is broken links. Recorded here so the next sitting does not “tidy” it.
Re-derive the coverage census against a 2026-09-11 snapshot Rejected. The 28-row adjudication is pinned to the 2026-08-17 frame; re-deriving would orphan it. The figures moved unchanged, with their frame dates, and the page carries the two-frames box. This is a carve-out, not a refresh — see the drift check in section B.

The carve-out map

What moved, what stayed, and what had to be repointed. Nothing was re-derived in the move.

Section on the old page Went to
The tool: architecture, and why you cannot install it stayed on webXray
The ownership database (schema, tree, frozen PSL) stayed — it describes webXray's file
How it compares to Tracker Radar and Disconnect moved, as The live sources, compared
Coverage: how much of the third-party surface moved, with the three traps split into their own sub-section
Two lists disagree: error or different question? moved, split: the two axes were promoted to the top of the new page as The question has two axes, the agreement tables became When two lists disagree
Reading all three at once (the code) moved
Which list is right, when they disagree? moved
How often is each list right? A random sample moved, with the inter-rater material given its own sub-section
Where these figures come from, and how to redo them moved, folded into the new page's methodology section
Choosing a resolution source now moved
Assembling the pipeline moved
What to report in a paper moved
Use in publications: the 15 webXray papers, role table stayed
Use in publications: the 136-paper sweep, the Tracker Radar year table moved — they are landscape figures, not webXray figures
Wayback timeline stayed, and was promoted from a methodology bullet to its own section; it was 900 words inside a bullet list

Day-one drift, found and fixed rather than left:

Where What it said after the move Fixed to
webXray intro bullet 2 quoted 69.9% / 93.9% / 2.1% / 17.2% / 7.0% inline as if the page still carried the measurement one sentence keeping 69.9% and 2.1% with a link, and an explicit “nothing on this page restates them”
webXray, The ownership database “§\”Two lists disagree\“ below measures what happens when you forget that” — a same-page reference to a section that had left repointed to [[Design:Ownership resolution#When two lists disagree]], with the finding (root resolution makes agreement worse) stated inline so the sentence still says something
webXray, Related pages listed Requests, Cookies, IP classification — all of which were there for the ownership material new page took those; webXray's list now leads with the new page and adds Policies for policyXray
webXray methodology listed nine owner_* scripts it no longer publishes figures from one bullet saying where they went and that their provenance ids are unchanged
the 28-row adjudication paragraph the old page said the adjudicators were “Sonnet sub-agents working to a published brief, single-rated”. The rewrite dropped the clause, leaving “adjudicated one by one against primary sources — company newsrooms, SEC filings” — which reads as human expert work. This was accepted as a blocking finding on 2026-09-05 and the carve-out silently reversed it six days later restored, and the same disclosure added to the random-sample section, which never carried it
Crawler, webXray bullet “see webXray for what survives of it: the ownership database, and how it compares to Tracker Radar and Disconnect today” — that comparison had moved repointed to the new page; see section H

A site-wide inbound-link sweep found four more, on pages nobody would have thought to check. Grepping every cached page for a link to programming:crawler:webxray rather than only the pages this sitting edited:

Page What it said Fixed to
Requests a Related-pages bullet linking the webXray page under the link text “webXray and domain-to-company ownership” — “once a request is flagged, this is how to answer whose it is” repointed to Ownership resolution, with a note that it moved
Filter lists the same bullet, near-verbatim repointed the same way
Legal enforcement “See webXray for the ownership databases and how much they disagree” — in the bullet on identifying a controller under Reg. 2025/2518 repointed, and extended to name the accuracy rates, which is what that bullet actually needs
Programming (namespace index) the webXray row read “Domain-to-company ownership lists, and what remains of the tool” — a description of the page, now false rewritten to describe the tool page and point the ownership question at the new page

Two of those four used the link text “webXray and domain-to-company ownership”, i.e. they were already treating the webXray page as the ownership page — which is the item's complaint, restated by the wiki itself. A carve-out's drift is not confined to the pages you edited; sweep for inbound links before calling it done. One inbound reference was deliberately left: provenance:design's generated page inventory records a size and a description as of an earlier date, and it is a dated snapshot rather than a live pointer.

A. Corpus queries

All over data/extract/run1/extractions.jsonl (5,859 papers) and data/fulltext/<year>/<venue>/<slug>/paper.cols.txt, whitespace collapsed before matching so a term broken across a PDF column boundary still matches. Every count is a paper count. The script is scripts/report_ownership_resolution.mjs, published in full in section J with its unedited output in section K.

# Question Population / denominator Answer
A1 How many papers have a readable full-text rendering at all? all 5,859 5,855 (99.9%). The four without one can only ever be sweep misses, so every count below is a floor.
A1b …and how many of those are damaged? the 5,855 0 empty, 0 under 2 KB, but 90 carry a NUL byte. The script reads them as JavaScript strings and regexes them, so NULs do not affect it. A grep-based sweep would have silently skipped all 90grep treats a NUL-carrying file as binary and, depending on the build, prints nothing and exits non-zero. None of the 136 union members is one of the 90, checked explicitly, so no published count depends on this — but the next person writing a sweep should not use grep.
A2 How many papers name each ownership resource? all 5,859 PSL 101, Disconnect (list sense) 74, Tracker Radar 32, WhoTracks.me 23, Crunchbase 23, webXray 15
A3 How many of those are inside the crawling population? crawled = 1,120 (crawlConfig present, or studyTypes includes automated-web-crawl) PSL 46, Disconnect 65, Tracker Radar 28, WhoTracks.me 19, Crunchbase 11, webXray 13
A4 How many name any of the five ownership resources? all 5,859 136 (2.3%)
A5 …and how many of those 136 are in the 1,120? crawled = 1,120 107, i.e. 9.6% of that population. Not 12.1% — see section C.
A6 Does adding the PSL change the union? all 5,859 136 → 221. The PSL is deliberately excluded: it answers “what is the registrable domain”, not “whose is it”.
A7 Is ownership resolution growing? each year's own paper count rose to ~3.5% in 2020–2022, has sat at 2.5–2.7% since. Full table on the page; 2010–2015 contributes one paper in total and is omitted rather than padded with zeros.
A8 Which venues? each venue's own paper count PETS 9.6%, IMC 2.7%, TheWebConf 2.3%, IEEE S&P 1.4%, USENIX 1.3%, CCS 1.3%, NDSS 1.1%
A9 What does the extraction schema see, against the full text? all 5,859 schema 89, full text 136, both 82, full-text-only 54, schema-only 7 — residue printed in full below
A10 Of the 15 webXray papers, how many say which version? the 8 that used the crawler or the list 2 say anything; 1 names a commit ([1Matte, Célestin; Bielova, Nataliia; Santos, Cristiana Teixeira (2020): "Do Cookie Banners Respect my Choice? Measuring Legal Compliance of Banners from IAB Europe's Transparency and Consent Framework", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]). Inherited from report_webxray.mjs, re-run and unchanged.
A11 What does “Tracker Radar” mean in the 32 papers that name it? the 32, hand-mapped 11 ownership dataset, 9 tracker/category database, 9 the Collector crawler, 2 citation, 1 compared. Inherited, re-run and unchanged.

A.12 Probe widths: every probe run at two widths, and both printed

A narrowing probe that returns more hits than the loose one, or hits the loose one does not contain, is a broken probe. The script asserts both and throws rather than printing. Both widths are on the record because the gap is itself the claim:

Resource                 Tight probe  Loose probe  Status
-----------------------  -----------  -----------  ----------------------------------------------------------------
webXray                  15           233          hand-verified: all 15, in report_webxray.mjs
Tracker Radar            32           108          hand-verified: all 32, in report_webxray.mjs
WhoTracks.me             23           124          upper bound
Crunchbase               23           23           upper bound
Disconnect (list sense)  74           700          upper bound; the loose width is the reason this one is tightened
Public Suffix List       101          157          upper bound

The patterns themselves are in a code block and not in a table, because a DokuWiki table cell cannot hold a pipe and every one of these regexes is an alternation. \| is not an escape and wrapping the pattern in DokuWiki's nowiki delimiters does not help either: the cell splits regardless.

webXray             tight  /webx[\s-]?ray/i
                    loose  /webx[\s-]?ray|libert/i
                    15 -> 233. "Libert" is a common surname and a French word.
                    The loose width is a sanity bound, not a candidate set.

Tracker Radar       tight  /tracker[\s.-]?radar/i
                    loose  /tracker[\s.-]?radar|duckduckgo/i
                    32 -> 108. Most DuckDuckGo mentions are the search engine
                    or the browser, not the dataset.

WhoTracks.me        tight  /whotracks/i
                    loose  /whotracks|who ?tracks ?\.? ?me|ghostery/i
                    23 -> 124. Ghostery is mostly the extension.

Crunchbase          tight  /crunchbase/i
                    loose  /crunchbase|crunch base/i
                    23 -> 23. NO GAP: nobody spells it with a space. This
                    row's count is not an artefact of the pattern.

Disconnect          tight  /disconnect'?s? (entit|list|block|black|tracking)|disconnect\.me|entities\.json/i
                    loose  /disconnect/i
                    74 -> 700. "Disconnect" is an ordinary English word and a
                    CSP-literature technical term. This is the probe that
                    needed tightening, and the published row is the tight one.

Public Suffix List  tight  /public suffix/i
                    loose  /public suffix|etld\+?1|effective top-?level domain/i
                    101 -> 157. The extra 56 are papers that use the concept
                    without naming the list.

What this does not establish. A tight probe with no gap is not a probe with full recall — it is a probe whose count is stable under widening in the one direction tried. A paper that resolves domain ownership from a source none of these five names, or names none of them at all, is invisible to every width. See section F.

A.13 The schema-only residue, printed in full

Seven papers whose tools[] / classification[] / otherToolsMentioned[] name one of the five resources but whose full text does not match the tight sweep. Printed rather than dropped, because an invisible residue is a residue nobody looks at:

  IEEE-SP/2024/to-auth-or-not-to-auth-a-comparative-analysis-of-the-pre-and-post-login-security
      named: Disconnect Tracker Protection List
  PETS/2022/on-dark-patterns-and-manipulation-of-website-publishers-by-cmps
      named: Disconnect | Disconnect tracking filter list
  PETS/2024/generalizable-active-privacy-choice-designing-a-graphical-user-interface-for-glo
      named: Disconnect Tracker Protection lists
  PETS/2024/support-personas-a-concept-for-tailored-support-of-users-of-privacy-enhancing-te
      named: Disconnect
  USENIX/2019/canvas-fast-and-inexpensive-automotive-network-mapping
      named: physical ECU access and disconnection
  WWW/2018/the-cost-of-digital-advertisement-comparing-user-and-advertiser-views
      named: Disconnect | Disconnect blacklist
  WWW/2023/who-funds-misinformation-a-systematic-analysis-of-the-ad-related-profit-routines
      named: Disconnect

Six of the seven are Disconnect, and they are the tightened pattern's misses: the extractor wrote “Disconnect Tracker Protection List” or bare “Disconnect” into a field, where the paper's own prose says something the list-sense pattern does not match. The seventh is a homonym the extractor invented — “physical ECU access and disconnection” in an automotive CAN-bus paper is not the Disconnect list. The tight pattern's cost is therefore about five papers, all Disconnect, all in the direction of undercounting, and the page's Disconnect row should be read as 74 rather than 74-to-79 only because the union is an upper bound on use anyway. The alternative — publishing the loose 700 — is not a trade worth making.

A.14 Denominators used on the page, stated once

Figure on the page Denominator Never
the six per-resource rows 5,859 for the “share of corpus” column; 1,120 for the “share of crawled” column, with the numerator restricted to that population 5,859 for both
136 / 2.3% 5,859
107 / 9.6% 1,120, numerator restricted to it 136 ÷ 1,120
per-year shares that year's own paper count 5,859
per-venue shares that venue's own paper count 5,859, and never the count alone — the venues differ in size by a factor of three
15 / 13 webXray papers 5,859 and 1,120 respectively
everything about the three lists the domain frames: 32,369 (2026-08-17) or 32,337 (2026-09-05) registrable domains 5,859; these are not corpus figures at all

B. Reproduction record, 2026-09-11

Every script whose figures the new page inherits was re-run before the move, against the same pinned inputs, and diffed against its committed output. A page whose report script no longer runs is a page whose numbers cannot be refreshed.

Script Invocation Result
report_webxray.mjs node scripts/report_webxray.mjs –quotes byte-identical to the committed report_webxray-output.txt before the fix in section C; regenerated after it
owner_dbs.py python3 scripts/owner_dbs.py –cache out/webxray/cache –disagreements 25 reproduces the 2026-08-17 frame exactly: 32,369 domains; 669 / 5,581 / 2,268; 58.6% / 84.3% / 80.5%; 15,651 merged rows; 19 non-hostname rows dropped; 6,941 ICANN and 3,290 PRIVATE rules; 782 of 827 owner names untouched by the fold, 16 of the 45 changed losing a legal suffix
owner_random_sample.py python3 scripts/owner_random_sample.py –sample out/owner_sample.json –rows out/adj_rows_scored.json –table –wiki byte-identical to out/owner_random_sample-output.txt
owner_selfname_census.py python3 scripts/owner_selfname_census.py reproduces 0.8% / 0.7% / 3.2% (25 of 3,215; 250 of 38,368; 250 of 7,850)
owner_adjudication.py python3 scripts/owner_adjudication.py –table reproduces the 28-row table exactly: 1/18/4/1/4, 8/12/6/2/0, 26/0/2/0/0, and the three rows scored as error rather than staleness
owner_irr_kappa.py invocation reconstructed by the figures reviewer, since it was missing from the brief reproduces pooled kappa +0.400 (+0.241, +0.551), per-list +0.027 / +0.478 / +0.460, settled-only +0.680
report_ownership_resolution.mjs node scripts/report_ownership_resolution.mjs new; section K

A drift check the carve-out made cheap. The new script re-derives the sweep counts from patterns written out independently and then cross-checks nine figures against report_webxray-output.txt. All nine agree. That cross-check is inside the script and throws, so a future run cannot publish two pages that disagree with each other:

Figure                   this script  report_webxray.mjs  verdict
-----------------------  -----------  ------------------  -------
webXray                  15           15                  agree
Tracker Radar            32           32                  agree
WhoTracks.me             23           23                  agree
Crunchbase               23           23                  agree
Disconnect (list sense)  74           74                  agree
Public Suffix List       101          101                 agree
union of the five        136          136                 agree
union ∩ crawled          107          107                 agree
crawled denominator      1120         1120                agree

C. A defect the cross-check found: a numerator from one population, a denominator from another

This is the one substantive correction of the sitting, and it was published for three and a half weeks.

report_webxray.mjs printed, at two places:

  ${pct(ftPapers.length, crawled.length)} of the ${crawled.length} that ran a crawl
  ${pct(anyOwner.length, crawled.length)} of the ${crawled.length} that crawled

ftPapers and anyOwner are sweeps over the whole corpus. Dividing either by crawled.length pairs a corpus-wide numerator with the crawling denominator: the share is of a population the numerator was not drawn from.

Published Correct Why the error is invisible
webXray papers, share of the 1,120 1.3% (15 ÷ 1,120) 1.2% (13 ÷ 1,120) 13 of the 15 are in fact crawling papers, so the two figures differ by one decimal and nothing looks wrong
any ownership resource, share of the 1,120 12.1% (136 ÷ 1,120) 9.6% (107 ÷ 1,120) a 2.5-point error on the more load-bearing figure

Both figures were on the live webXray page, and the 12.1% was one of the figures moving to the new page — which is how it was caught: the new script computed the intersection because it had no reason not to, and the cross-check refused to pass. No guard on the old page could have found it. The number-guard family checks that a figure on a page matches the script's output; the script's output was the wrong number, and the page matched it faithfully.

The fix adds a named helper so the mistake cannot recur silently:

const crawledKeys = new Set(crawled.map(key));
// A share of `crawled` needs a numerator drawn from `crawled`.
const inCrawled = (ks) => ks.filter((k) => crawledKeys.has(k));

report_webxray-output.txt was regenerated and both pages corrected in the same sitting. The webXray page carries a footnote at the changed figure rather than silently restating it.

D. Folding, and what it could not reach

This page's corpus figures need no name fold: every count is a regex over full text or a set membership, and there are no free-text names being aggregated. That is stated rather than assumed, because it is unusual for a page on this wiki.

The live-list figures do fold, and the fold and its residue are published on webxray §B.11. Restated here only in summary, because the content page quotes both halves:

  • Legal-form suffixes only (Inc, LLC, GmbH, S.A.S, …), plus punctuation normalisation. Deliberately no synonyms, so the “agree” columns are a lower bound on real agreement and the “neither” column an upper bound on real disagreement.
  • Residue: the fold leaves 782 of webXray's 827 owner names untouched (94.6%). Of the 45 it changes, only 16 lose a legal-form suffix; the other 29 change on punctuation alone — 56.com, AT&T, Ask.com, Bootstrap_China, Clearstream.TV, Cm_browser, Dictionary.com, Dun & Bradstreet, and the rest are in the published output.
  • What no fold reaches: “Google” against “Alphabet”, “Xandr” against “Microsoft Corporation”. Those are the granularity and vintage axes, and they are the page's subject rather than a normalisation failure. There is no principled synonym table for them, which is why the page reports the agreement figures as bounds and adjudicates a sample instead.

E. Quotes checked against the source

Checked on 2026-09-11 against data/fulltext/<year>/<venue>/<slug>/paper.cols.txt, whitespace collapsed and quotation marks and dashes normalised, with paper.norm.txt and paper.txt as fallbacks.

Paper Quote (head) Verdict
[2Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)] “the product of years of detective work” found in .cols
[2Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)] “because webxray's database of domain ownership primarily contains major ad networks…” NOT contiguous — see below
[3Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] “frequently miss connections among two hostnames” found
[3Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] “twitchcdn.net” found
[4Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] “preferring those that have been updated most recently” found
[4Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] “does not aim at providing full transparency on the organizations behind each domain” found
[1Matte, Célestin; Bielova, Nataliia; Santos, Cristiana Teixeira (2020): "Do Cookie Banners Respect my Choice? Measuring Legal Compliance of Banners from IAB Europe's Transparency and Consent Framework", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] “WebXRay commit 04c3c8e8” found
[5Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)] “Redacted For Privacy” / “Domains By Proxy” / “Whois Guard” found, and the surrounding sentence re-read — see below

The Libert coverage quote is real but is not one contiguous string in any rendering of the PDF. It is the page's longest quotation and it is load-bearing for the coverage section, so it was chased rather than dropped. In paper.cols.txt, paper.norm.txt and paper.txt alike, the two-column layout interleaves a table caption through the sentence:

...because webxray's database of Table 1: Third-Party Prevalence, SSL Use, and First-Party
domain ownership primarily contains major ad networks rather Disclosure than small clients,
and policyxray only searches for identified †Denotes Company has Consumer Services parties,
variability in the long-tail of trackers may not have an outsized effect on overall findings
related to disclosure. Nonetheless, Company % Tracked % SSL % Disclosed it is important to
point out that the number of parties being searched Google † 82.81 80.35 38.29 for is fewer
than the total number of parties present.

Every clause is verbatim and in this order; the interleaved strings are “Table 1: Third-Party Prevalence, SSL Use, and First-Party Disclosure”, “†Denotes Company has Consumer Services” and the table's own header and first row. The page quotes the sentence as the author wrote it and carries a footnote saying exactly this, so a reader running the same grep does not conclude the quote is invented. This is the class of defect a plain substring check reports as a miss and a careless run reports as a fabrication.

One claim was corrected by reading the source. The old page said [5Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)] “resolved organisations for 411 of 762 mobile-specific trackers using CrunchBase, webXray's list, TLS certificates and WHOIS in that order”. The paper's own sentence is narrower:

we successfully retrieved the organization information of 411 trackers from either the CrunchBase database, Tim Libert's library, or TLS certificate, and retrieved the organization […]

The 411 comes from the first three sources; WHOIS contributed a further 251 (“among those 251 trackers whose organization information was retrived from the WHOIS record”). The four sources were tried in that order, but the 411 is not a figure for all four. The new page states it as 411 from the first three and 251 from WHOIS. The consequence for the page's actual point is unchanged: the five registrar-privacy strings total 88, against the largest real company in the top-ten table at 48 (Adobe), and the paper says so itself — “these five organizations cannot represent the real organizations of those trackers”.

F. External sources: verified, and rejected

Repository and vendor state was re-fetched on 2026-09-11, not recalled. Training data is stale by construction for this kind of claim.

Claim on the page Primary source, fetched 2026-09-11 Verdict
Tracker Radar current; last commit on main 2026-08-28; newest release tag 2026.08.28 api.github.com/repos/duckduckgo/tracker-radar/commits?per_page=1a736b501, 2026-08-28T15:25:14Z; /releases2026.08.28, 2026.07.27, 2026.06.08, none flagged prerelease holds
Tracker Radar's pushed_at reads later than main because of unmerged branches pushed_at = 2026-09-02T21:21:31Z against main at 2026-08-28 holds, and the page now names both dates
Disconnect current; last commit 2026-08-28 commits?per_page=14b592c28, 2026-09-05T20:21:14Z moved. Page updated to 2026-09-05 and the table header re-dated to 2026-09-11
Ghostery trackerdb current; last commit 2026-09-01 commits?per_page=1aa82e6d9, 2026-09-01T14:40:47Z holds
Disconnect's headline “14,332 verified domains and entity mappings” disconnect.me/trackerprotection, HTTP 200, string present holds
github.com/timlib/webXray is HTTP 404; the account has 0 public repositories 404; api.github.com/users/timlib → 200, public_repos: 0, company: webXray.ai holds
github.com/timlib/webXray_Domain_Owner_List is HTTP 404 404 holds
No PyPI package pypi.org/pypi/webxray/json → 404 holds
thezedwards/webXray last pushed 2021-03-04, PolyForm Strict pushed_at 2021-03-04T23:52:47Z; GitHub reports NOASSERTION holds
peterjoles/webXray is MIT, last pushed 2023-03-12, 36 commits ahead spdx_id: MIT, pushed_at 2023-03-12T20:15:26Z holds
webxray.org is a placeholder reading “Public interest projects for the interested public.” fetched with a browser User-Agent: HTTP 200, that exact line and nothing else holds
webxray.ai is a live commercial product HTTP 200 holds

Sources rejected, recorded so the next run does not re-add them:

Rejected Why
GitHub's license.spdx_id field for Ghostery trackerdb and Disconnect reports NOASSERTION for both, which is a statement about GitHub's detector and not about the licence. The licences were read from package.json and LICENSE respectively.
GitHub's pushed_at as “last activity” it counts unmerged automation branches. For Tracker Radar it reads four days later than main. Read the branch.
Any figure for Ghostery trackerdb coverage or accuracy none was measured. Its .eno + patterns shape needs a parser nobody could check against the other three. The page says this is an omission rather than a judgement, and lists it as an open question.
Vendor marketing counts as file sizes Disconnect's “14,332 verified domains” is a site headline; entities.json holds 7,850 domains. The page quotes both and says which is which.
Anything recalled rather than fetched about these repositories training data predates all four dates above.

G. What could not be established

  • Ghostery trackerdb is unmeasured. It is a live fourth option and the page gives numbers for two of the three live lists. Closing this needs a parser for the .eno + patterns model whose choices are checkable against the other three — a sitting's work, not a footnote. Filed as an open question on the content page.
  • Recall of the candidate set is unknown. The 136 is a union of six patterns over full text. A paper that resolves domain ownership from a source none of the five names — a bespoke WHOIS pipeline, a commercial feed, an internal list — matches nothing at any width. The page reports 136 as an upper bound on use and says nothing about recall, because nothing here measures it.
  • The share of the 136 that actually resolved an owner is not measured. Only two rows (webXray's 15 and Tracker Radar's 32) carry a hand verdict per paper. Doing the same for Disconnect's 74 is the obvious next increment and was not attempted this sitting.
  • ~~No accuracy figure exists for the tail.~~ Partly closed on 2026-09-11, and the answer mattered. This read: “Every rate is conditional on the entry being adjudicable, and 18–30% of each list's draws were not. Whether unadjudicable domains differ systematically between lists is exactly what would decide whether Disconnect's lead survives, and a random sample cannot answer it.” A second pass with harder sourcing settled 28 of the 40, leaving 6.7% / 6.7% / 10.0%. The formal test found no significant difference between recovered and already-eligible rows (p = 0.64 / 1.00 / 0.40, on 7, 9 and 12 rows — almost no power), but Disconnect's estimate fell 94.4% → 89.2% and its lead over webXray stopped surviving Bonferroni. So the answer is: the tail was somewhat worse for the list the worry named, and the worry was worth acting on. See section L. What remains open is the residue's own residue — 12 domains that no register, no registry record and no archived legal page can reach, and nothing bounds those from inside the sample either.
  • The coverage census has no neutral denominator. Tracker Radar's own crawl output is the frame, which flatters Tracker Radar. There is no independent census of third-party domains to use instead, and the page says so at every coverage figure rather than pretending otherwise.
  • Whether a hierarchy built from a current list would beat both. webXray's parent_id tree is the only hierarchy any of these files carries and it is frozen; nobody has tried grafting a current corporate-structure source onto Tracker Radar's entities. Filed as an open question.
  • Two adjudication rows from 2026-08-17 remain unresolved (the acquisition date behind 360yield.com, and fwmrm.net's post-2026-spinoff status). Recorded as unresolved rather than guessed; they are excluded from the 28.

H. Corrections owed to Programming:Crawler, and their real state

The task brief for this sitting said Crawler's comparison row “still says webXray's crawler is PhantomJS (historically)”. It does not, and has not since 2026-08-17. The brief's account of its own state was stale; recorded here because a later sitting reading the brief would look for a correction that is already made.

Correction State on 2026-09-11, before this sitting
The comparison row's automation column already fixed. Revision 1786953266 reads “PhantomJS in 2015; consumer Chrome over raw CDP in the last public version”
“stale third-party mirrors, the newest last pushed in 2015 and targets PhantomJS” already fixed in the same revision
The <WRAP todo> asking where webXray is developed already closed in the same revision. The page's remaining <WRAP todo> is three unrelated open questions and was correctly left alone
The licence and most-complete-copy claim already fixed on 2026-09-05, revision 1788636934 — the item's own note says so

The one correction that was owed, and that this sitting created, is the carve-out's own drift: that bullet ended “see webXray for what survives of it: the ownership database, and how it compares to Tracker Radar and Disconnect today”, and the comparison had just left that page. Repointed to Ownership resolution in revision listed below. Nothing else on Crawler was touched: its webXray figure of 7 is the tools[] schema count for a table whose whole column is schema counts, and changing it to the sweep's 15 would make that one cell incomparable with its neighbours.

I. Reviewers

All four were told explicitly that the author's context may not be exhaustive, and all four were handed the page text, the report script, its unedited output and these notes. The three focused passes ran in parallel; the generic pass ran afterwards, against the pages as the first three had left them.

All four passes returned and every finding was applied. The generic pass — the one with no checklist — found more, and worse, than the three focused passes combined. Each was handed the page text, scripts/report_ownership_resolution.mjs, its unedited output, scripts/report_webxray-output.txt and these notes, against the bundle in review_own/. Every finding below was re-verified by the author before being accepted — a reviewer's report is a lead, not a result.

Pass Model Returned Findings Accepted Rejected
external currency Sonnet yes 1 substantive of ~40 checks 1 0
citations and quotes Sonnet yes 1 substantive of 22 citekeys, 11 quotes, 13 attributed claims, 7 vendor sources and all 15 rows of the webXray table 1 0
figures versus script Sonnet yes 2 substantive, both inside a published script rather than on a page 2 0
generic, no checklist Fable yes 13 substantive, including the two most serious of the sitting 13 2 partial

Accepted: the LLM row was an overclaim (external currency)

The Choosing a resolution source now table read “emerging, and not yet at the domain layer”. The reviewer found arXiv:2606.20868, Can LLMs Reason About Brand Ownership?, submitted 2026-06-18 — four models evaluated on domain-to-brand attribution over 36 heavily-phished brands.

Verified independently before accepting: the abstract was fetched (HTTP 200) and read. The paper is real, the date is right, and the task genuinely is domain-layer ownership attribution. Two qualifications the reviewer did not make, and which changed how it went on the page:

  • Its question is phishing and squatting defence — “is this domain the brand's own?” — not third-party tracker attribution. 36 brands is not a coverage claim.
  • Its result cuts the other way. Models enumerate a brand's domains at up to 82% precision from memory, but on ownership verification without external tools macro F1 is at most 0.37, rising by up to 0.65 with WHOIS. So the honest update is not “LLMs are arriving, consider them” but “the one published attempt at this task fails without retrieval” — which strengthens the row's existing advice rather than softening it.

The corpus-scoped half of the sentence (“no corpus paper applies this to third-party domain ownership”) was not changed: it was and is true, and an arXiv preprint is outside the seven venues. Cited as a footnote with its arXiv id rather than a {[key]}, so no bibliography entry and no cache purge.

Accepted: an undocumented column splice, of the opposite kind (citations and quotes)

The page footnotes the [2Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)] coverage sentence as non-contiguous in every rendering. The reviewer found that [6Selmo, Carlos; Carisimo, Esteban; Bustamante, Fabián E.; Alvarez-Hamelin, J. Ignacio (2025): "Learning AS-to-Organization Mappings with Borges", in: Proceedings of the 2025 ACM Internet Measurement Conference, pp. 120-133. (DOI)]'s “an Internet shaped by constant mergers, rebrandings, and regional variation” has the inverse problem, unflagged.

Verified independently: the string constant mergers does not appear in paper.cols.txt at all — nor in paper.norm.txt or paper.txt — while pypdf over paper.pdf returns the sentence fully contiguous. So the quote is correct and the corpus's own rendering is the thing that is wrong. Footnoted to say so, because a reader checking the quote against the corpus would otherwise conclude it was invented.

This is the two-renderings rule in both directions on one page: Libert fails in every rendering and is real; Borges fails in the column-repaired rendering and is real. A quote-check against a single rendering would have reported one false FAIL and, had the Libert footnote not already existed, one apparent fabrication.

Accepted: a published script's own ALLOW list carried the fold it warns against (figures versus script)

This is the most valuable finding of the sitting, and section B of this page had just asserted the opposite.

report_webxray.mjs ends with a section Z, Every number on the page that does NOT come from this corpus, which exists so check_page_numbers.mjs can tell a non-corpus figure from a corpus one. Five of its lines were hardcoded constants, and they were the ICANN+PRIVATE fold with exact-key lookup — the variant owner_dbs.py's own output labels <- THE TRAP and both content pages warn against by name, because it manufactures a coverage hole out of googleapis.com:

Section Z said The published fold gives
45,525 registrable domains 32,369
webXray 657 (1.4%), TR 5,539 (12.2%), Disconnect 2,277 (5.0%) 669 (2.1%), 5,581 (17.2%), 2,268 (7.0%)
weighted 54.5% / 79.3% / 75.8% 58.6% / 84.3% / 80.5%
top 100: 70% / 95% / 91% 71% / 98% / 94%
197 of 601, 216 of 461, 584 of 1,532 198 of 612, 217 of 464, 584 of 1,538
“16,396 of 47,836 rows are hostnames” the label-count heuristic the page says “answers neither question; do not use it” — replaced with the ICANN fold's 15,651

Verified independently before accepting: each figure was matched against the live owner_dbs.py run in section B, and grepped for on both content pages. None appears as prose on a live page, so nothing a reader saw was wrong — but the ALLOW list of a number guard is exactly where a wrong figure does its damage silently, by blessing the wrong value if it ever reaches a page. Corrected, and the corrected lines carry a comment naming the trap.

What this says about section B of this page. B calls report_webxray.mjs's output “byte-identical” to its committed copy, and it was — before and after. Byte-identity proves a script matches its own last run. It proves nothing about whether the content is still true, and here the script had been faithfully reproducing a stale constant. A reproduction record is not a correctness record, and this page should not have implied otherwise.

Accepted: a guard that loses its own warning (figures versus script)

report_webxray.mjs buffered every line and flushed once at the bottom. Its section-A guard —

if (roleResidue.length || roleExtra.length) {
  P('FAILURE: the hand map and the sweep disagree. Re-read the new papers before publishing any figure below.');
}

— called P() and let the script continue. Section B then dereferences ROLE.get(k).role on the now-missing key and throws, before the single console.log at the bottom has run.

Reproduced by deleting one ROLE entry: stdout 0 bytes, a bare TypeError stack, and the FAILURE line that exists for precisely that case nowhere at all. The guard fires and is then thrown away. Fixed with a fail() helper that flushes the buffer, repeats the message on stderr and exits 1; re-running the same mutation now prints 1,283 bytes ending in the named paper and the FAILURE line, and exits 1. On a healthy run the output is byte-identical to before, so the fix is behaviour-neutral where it should be.

What the reviewers confirmed, which is also a result

  • All 22 distinct citekeys across the three pages resolve to exactly one bibliography entry; no collisions anywhere in the 993-entry file.
  • All 11 quoted strings and 13 attributed non-quote claims verified verbatim or in substance, including the self-contradictory [3Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] sentence this page declines to derive a percentage from, and the [5Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)] 411/251 split this sitting corrected.
  • All 15 rows of the webXray publication table verified against source text — the reviewer was asked for five and did fifteen.
  • Every figure on both content pages independently re-derived from a live re-run of all six scripts, including the inter-rater kappa, whose invocation the reviewer had to reconstruct because it was not in the brief. Both published code snippets (owner_lookup.py and resolve()) executed against the cached JSON and their output matched the page character for character.
  • Three further mutations of report_ownership_resolution.mjs by the reviewer, including a re-creation of this sitting's own denominator bug — all caught, none vacuous.
  • Every repository, licence, URL, version, dependency and acquisition claim on all three pages re-fetched and holding to the day, including the 2020-06-24 Disconnect relicensing commits, PolyForm Strict still having no SPDX id, and lxml/psycopg2-binary wheels still stopping at CPython 3.9 against a 2025-10-31 Python 3.9 EOL.

Accepted, and the worst finding of the sitting: the LLM-adjudicator disclosure was dropped (generic)

webxray's 2026-09-05 reviewer log records, verbatim: “The pages never say the adjudicators were language models … ACCEPTED, blocking. The content page now says 'by Sonnet sub-agents working to a published brief, single-rated'.” The carve-out dropped that sentence, and the new page said only “adjudicated one by one against primary sources — company newsrooms, SEC filings, or the domain's own legal documents”.

Verified before accepting: the clause is in the pre-split page text (live_programming_crawler_webxray.txt), and a grep of the new page for Sonnet, sub-agent and language model returned zero hits. A reader deciding whether to cite 94.4% was being told, by omission, that a person adjudicated it.

Restored in two places — beside the 28-row table and, for the first time, beside the random-sample table, which never carried it even before the split. No reviewer in three sittings had noticed the second gap. A finding accepted as blocking can be undone by a later structural edit, and no guard on this wiki looks for that.

Accepted: a p-value asserted without being computed (generic)

The page said: “Disconnect's lead over webXray is a difference this sample can establish; its lead over Tracker Radar is not, and neither is Tracker Radar's over webXray.” Two p-values were published (0.014, 0.82). The Disconnect-versus-Tracker-Radar pair was never computed — the webXray provenance log shows the original reviewer supplied only the other two, and the sentence generalised from their absence.

Verified: Fisher's exact re-implemented from scratch over the eligible raw counts reproduces the two published values to three decimals (0.0137, 0.820) and gives p = 0.0248 for the untested pair. The sentence was false by the page's own test, directly under the short answer's headline comparison.

Rewritten as a three-row table with all three p-values and a Bonferroni threshold of 0.0167 for three tests on one sample — under which 0.014 survives and 0.025 does not, so the recommendation stands but for a stated reason rather than an invented one. The page now says in its own text that an earlier version asserted this without computing it.

This is the failure mode to remember: an absence of evidence in a reviewer's report was read as evidence of absence. Nothing in the three focused briefs would have caught it — the figures pass checks the page against the script, and this number was in neither.

Accepted: the estimator is post-stratified, and the page described it as unweighted (generic)

“domain-level — … Every entry weighs the same.” It does not: owner_random_sample.py computes R = Σ_h (N_h/N)·p_h over the 24/12/12/12 quartile allocation. That is why 94.4% ≠ 39/42 = 92.9% and 69.9% ≠ 35/49 = 71.4% — and why the Fisher paragraph's raw counts did not match the table above it, which a reader dividing them would have read as an error on the page. Corrected, with the raw counts now stated beside the weighted ones.

Accepted: three arithmetic slips the author's own "arithmetic audit" missed (generic)

Said Is
“A third of the sample could not be settled” 42 of 180 = 23%; the per-list shares in the same bullet are 18 / 22 / 30%
“the intervals are ±13 percentage points at best” “at best” is Disconnect's ±6.6; ±13.8 is the worst. Now “±7 to ±14”
“Two of the three lists change weekly” the page's own table says Tracker Radar regenerates monthly

The author's audit table below claims “all consistent; one error found”. That was wrong, and it is left standing above with this correction beside it rather than quietly rewritten — an audit that reports a clean result it did not earn is worse than no audit.

Accepted: overstatements and structure (generic)

Finding Action
The short answer headlines the two metrics under which the recommendation looks best, while the body argues two other metrics matter more (encounter-weighted accuracy, where the frozen list tops the table; weighted coverage, where the live lists near-tie) short answer now states this in three sentences and says why the recommendation survives it
“read the encounter-weighted column as an ordering rather than as a number” — all three intervals overlap and no pairwise test is run now says it supports neither, and why it is nonetheless the quantity that matters
webXray's inter-rater kappa is +0.027 on 12 rows — chance — and the page stated it without drawing the implication one sentence added: its 69.9% carries rater uncertainty on top of the sampling interval
“rose to about 3.5% around 2020–2022 and has sat between 2.5% and 2.7% since” is a story about noise verified: 2022 vs 2023 p = 0.41, pooled 2020–22 vs 2023–26 p = 0.11. Headline “It is not growing” kept, curve dropped
The “29 of 525 verdicts changed, direction overwhelmingly favourable” discount sat two screens below the table it discounts promoted to a <WRAP important> directly under the accuracy table
“most papers should merge them and say in which order” — the page never said which order now states it: Disconnect first for the name, Tracker Radar as fallback, record which key matched
Wiki-internal QA on the content page (the schema-versus-sweep paragraph, the probe-width paragraph) trimmed to pointers here; the two quote-splice footnotes were kept, because they defend a reader who greps the corpus and finds nothing
Three sentences on webXray describing the pre-split page all three fixed
No reading list for the stated reader three papers, in order, added to the short answer

Partially rejected (generic)

Finding Disposition
Remove the Libert and Selmo quote-splice footnotes as wiki-internal QA Rejected. A reader who checks either quote against this corpus finds nothing and concludes it was invented. The footnotes exist for them, not for us. The other items in that finding were accepted.
“98 em-dashes, 26 'not X' contrasts, nearly every section ends on an aphorism — someone who did the work would let a few sections end on the number” Accepted as fair, not acted on this sitting. It is a real observation about the register and it would take a prose pass, not an edit. Recorded so it is not rediscovered.

Noted, not acted on

Observation Disposition
Today's live PSL has 6,950 ICANN / 3,375 PRIVATE rules against the 6,941 / 3,290 in the 2026-08-17 snapshot Not a defect. The figure is snapshot-pinned and the page says the PSL moves. Re-deriving it would orphan the adjudication — the same reason the whole 2026-08-17 frame is frozen.
The reviewer's Wayback CDX pull put timlib/webXray's last archived 200 at 2022-12-30 where the page says 2023-03-31 Left as published. The reviewer attributed the gap to CDX collapse-window settings and did not claim the page was wrong. It is a bound either way and both bounds support the same conclusion. Flagged here so a future sitting can settle it rather than rediscover it.
libert2015_invisible and karaj2018whotracksme could not be quote-checked — neither paper is in the corpus Correct and expected. Neither is quoted; both are cited only as identifiers.
The figures reviewer could not verify section H's claim that Crawler was already corrected, having no access to that page's history Fair, and the claim stands. Revisions 1786953266 and 1788636934 were read from node scripts/dw.mjs history programming:crawler and the live page text quoted in section H.
It could not verify who produced “rater 2” in the inter-rater work, taking the label from the script's own output A real limit of that measurement, and it is not this page's to fix. Recorded on random_sample where the draw was made.
The Sánchez-Rola 2021-versus-2022 venue-year discrepancy Left as published, with the page's existing italic note that the corpus files it under 2022 while Crossref and the paper's header say 2021.

Still owed

  • A finding accepted as blocking was silently reversed by a later structural edit, and nothing detected it for six days except a reviewer with no checklist. There is no guard on this wiki that re-checks accepted findings against a rewritten page. The cheapest version would be a list of blocking findings per page, with the sentence each one put there, checked as a string.
  • The prose register was flagged and not fixed: 98 em-dashes in 8,700 words, and most sections closing on an aphorism rather than on the number. A prose pass, not an edit.
  • report_webxray.mjs's section Z is a hand-maintained ALLOW list of constants and that is what let it go stale. It should be derived from owner_dbs.py's output rather than retyped, or checked against it by a guard that throws. Not done this sitting: it needs the two scripts to share a format, which is a change to both.
  • The same buffered-output defect may exist in other report_*.mjs scripts on this wiki. Only report_webxray.mjs and report_ownership_resolution.mjs were checked.

What the author checked without them, so the gap is bounded rather than open:

Check Result
All published scripts re-run against pinned inputs they reproduce their committed output — which, as C and I both show, is not evidence the output is right: byte-identity blessed a wrong denominator for three and a half weeks and five wrong constants for longer. A reproduction record is not a correctness record.
owner_adjudication.py –table reproduces the 28-row table exactly: 1/18/4/1/4, 8/12/6/2/0, 26/0/2/0/0, and the three error rows (1rx.io, jsdelivr.net, stackadapt.com)
domain_owners.json schema table on webXray, every cell reproduces: 319/827 parent_id, depths 508/211/80/24/2/2, 761 uses, 803 platforms, 826 country, 104 trade_groups, 591 policy URLs, 132 GDPR, 4 CCPA, 13 opt-out, 34 crunchbase_id, 52 health-segment, 263 notes, 275 aliases, 69 languages, 2,175 of 3,215
Mutation test of report_ownership_resolution.mjs four mutations, four failures. Swapping tight/loose on Disconnect → “tight (700) > loose (74)”. Making loose disjoint from tight at equal size → “23 tight hits are not in the loose set” — the containment guard catches what a size comparison alone would miss. Injecting a fake union member → cross-check throws. Replacing the crawled intersection with the full union → cross-check throws. No guard is vacuous.
Arithmetic audit of every ratio on the content page all consistent; one error found — “seven times NDSS's rate” is 8.4×, corrected
Rendered DOM of both content pages 18 references for 18 keys on the new page, 17 for 17 on webXray; 12 and 7 tables; 9 and 4 WRAP boxes; every in-page and cross-page anchor resolves; one broken anchor found — the webXray availability table pointed at #Methodology and limitations of these figures for a Wayback timeline that had become its own section
Site-wide inbound-link sweep four pages found and fixed (section above)
Table cells containing a pipe one found in this page's own draft: a nowiki-wrapped [[a|b]] literal in a table cell. A parsed link's pipe is safe in a cell (verified in the rendered DOM); a nowiki-wrapped one is not, because the cell splits before the nowiki is honoured. Rewritten as prose, and the A.12 regex table was moved into a code block for the same reason.
Full-text corpus quality 5,855 of 5,859 renderings present, 0 empty, 90 carrying a NUL byte — see A1b

L. The second pass over the residue, 2026-09-11

The 2026-09-05 sample drew 180 times over 175 distinct domains and could not adjudicate 40 of them (22.9%), 42 draws. Every accuracy rate it produced was therefore conditional on the entry being adjudicable, the exclusion rate rose toward the prevalence tail, and it was worst — 30% — for Disconnect, the list the content page recommends. Nothing inside that sample could test whether the excluded rows were systematically worse. This pass re-adjudicated exactly those 40 rows, changing nothing else, against three routes the first pass did not take.

L.1 What was done, and in what order

Step Command What it produced
Build the worksheet python3 - over out/owner_sample.json + out/adj_rows_scored.json out/tail40.json — the 40 domains with each list's claim, the residue membership per list, and the 2026-09-05 note saying why that pass gave up
Mechanical evidence bash scripts/owner_tail_probe.sh out/tail40.json out/tail_probe out/tail_probe.txt (2,471 lines, 2m 29s, 6-way parallel)
Adjudication 5 Sonnet sub-agents × 8 domains, round-robin over the prevalence order so every batch spans the range out/tail_adj/result1..5.json
Gate python3 scripts/owner_tail_merge.py –sample out/owner_sample.json –base out/adj_rows_scored.json –tail out/tail40.json –dir out/tail_adj –out out/adj_rows_tail.json out/adj_rows_tail.json, 0 problems
Citation check python3 scripts/owner_verify_sources.py –rows out/tail_rows_only.json –cache out/tail_verify_cache –out out/tail_verify.json 3 defects in the checker itself — see L.5
Hand corrections python3 scripts/owner_tail_corrections.py –rows out/adj_rows_tail.json –out out/adj_rows_tail_final.json 6 rows: resource ×3, keep ×2, rescore ×1
Re-estimate python3 scripts/owner_random_sample.py –sample out/owner_sample.json –rows out/adj_rows_tail_final.json –table –wiki out/owner_random_sample-tail-output.txt
What changed python3 scripts/report_tail_pass.py … out/report_tail_pass-output.txt

The 2026-09-05 output file is kept, not overwritten. out/owner_random_sample-output.txt is the before; out/owner_random_sample-tail-output.txt is what the page's figures are now checked against, and out/owner_random_sample-tail.diff is the 338-line diff between them.

Reproducing the before is a trap worth recording. out/adj_rows.json and out/adj_rows_final.json both still carry self-named verdicts, and owner_random_sample.py exits fatally on them. The committed 2026-09-05 output is reproduced byte-identically only by –rows out/adj_rows_scored.json –wiki — the –wiki matters, because sections F and G exist only under it.

L.2 The three routes, and which one paid

Route What it is Rows it settled
register-by-number a national company register searched by company/VAT/register number rather than by name — the first pass failed on strings like “Admiral”, “Collective” and “Globo”, which are unsearchable 4
rdap-history the registrar's own port-43 WHOIS, not RDAP. rdap.org returns a thin registry record with no registrant for .com/.net; the registrar's server returns it un-redacted when the registrant has not bought privacy protection, which is a per-registration fact and not a per-registrar one – whois.tucows.com returns a real organisation for some domains and Contact Privacy Inc. Customer 0168237798 for cdnbasket.net. Also the registration and last-changed events, which bound a claim even when the organisation is redacted 9
archived-legal a Wayback capture of the domain's own legal page, fetched with the id_ suffix so the Archive's banner is not in the bytes 2
live-source a live document the first pass did not reach — usually because it looked for a page naming the domain and this pass looked for the entity, or because the page is on a subdomain 13

These counts are after the hand corrections. They differ from the owner_tail_merge.py output published in the appendix, which ran before them: that run reports live-source 12 and rdap-history 10, because the i.ua correction moved one row from rdap-history to live-source. Nothing else moved. The largest single route is live-source – a live document the first pass did not reach, usually because it looked for a page naming the domain where this pass looked for the entity.

The registrar hop is nonetheless the cheapest thing to add to any future pass: it is one extra socket per domain, and it settled rows RDAP alone could not. It produced Applied Technologies Internet SAS for at-o.net, Conversant, Inc. for the three ValueClick-style sync domains, Acoustic, L.P. for pages02.net, Leven Labs, Inc. for both Admiral domains and Google LLC for blogblog.com — none of which RDAP shows. This container has no whois(1) and no apt, so scripts/whois43.py speaks the protocol directly: IANA referral, TLD server, then the registrar's server when the registry is thin.

L.3 The sourcing bar was widened, and by how much

Three source kinds are new in this pass and every row records which it used:

Source kind New? Rows What it is worth
domain-register yes 9 the registry's record of who holds the name, contractually required to be accurate but asserted by the registrant and validated by nobody. Weaker than a company register
legal-doc no 8 the domain's own live privacy policy, terms, imprint or legal-entity footer — the 2026-09-05 bar
parent-site no 4 the acquirer's or parent's own site naming the brand or domain as theirs — the 2026-09-05 bar
register no 3 a national company register record, or an OV/EV certificate's validated O = field — the 2026-09-05 bar
archived-legal-doc yes 2 the domain's own legal document read at a date it existed. Establishes the operator at the capture date, so it needs registry continuity or a live source about the entity to become a claim about today
filing no 1 an SEC or other regulator's filing — the 2026-09-05 bar
newsroom no 1 the company's own press release — the 2026-09-05 bar

So 17 of the 28 recovered rows clear the original bar and 11 rest on a kind the first pass did not use. tls-san — another organisation's certificate on this host — was defined in the brief as corroboration only and is the source of record for 0 rows.

This pass produced its own counter-example to its weakest source kind, and it is kept rather than smoothed away. i.ua's registry record (whois://whois.ua/i.ua) names Digital Ventures LLC. The portal's own user agreement (help.i.ua/agreement/) names a different company as its administration: ТОВ «КЕПРЕЙТ ПАРТНЕРС», Ukrainian register code 33500955. On that domain the registrant of record is not the operator. The row was re-sourced to the legal document, and the finding is why report_tail_pass.py carries section E3: a sensitivity run that drops every recovered row resting on an uncorroborated registrant organisation. Two remain, and dropping both moves webXray +0.0, Tracker Radar −0.6 and Disconnect +0.0 percentage points.

L.4 Every recovered row, with its route and its source

Domain In the residue of Owner settled on Route Source kind Citation re-fetched?
spot.im Tracker Radar Open Web Technologies Ltd. (OpenWeb) live-source parent-site OK
company-target.com webXray Demandbase, Inc. live-source parent-site OK
adgrx.com webXray AdGear Technologies, Inc. (“AdGear”, trading as Samsung Ads; wholly-ow live-source legal-doc FETCHFAIL
app-us1.com Tracker Radar ActiveCampaign, LLC (ActiveCampaign) live-source domain-register OK
govx.com Disconnect GovX, Inc. register-by-number filing OK
opti-digital.com Disconnect Opti Digital SAS (2 Rue des Cortalets, 66400 Céret, France) live-source legal-doc OK
ksearchnet.com Disconnect Klevu Oy (operating subsidiary of Athos Commerce) live-source register OK
kameleoon.io Disconnect Kameleoon SAS (Kameleoon) register-by-number legal-doc OK
gssprt.jp Disconnect Geniee, Inc. rdap-history domain-register OK
acint.net Tracker Radar Poshibalov Evgeny Vasilyevich, operating the self-titled project 'Acin rdap-history domain-register OK
yceml.net webXray Conversant, Inc. (operating brand Epsilon; ultimate parent Publicis Gr rdap-history domain-register OK
travelpayouts.com Tracker Radar Go Travel Un Limited (Hong Kong; trading as Travelpayouts) archived-legal archived-legal-doc OK
blogblog.com webXray Google LLC (Blogger) rdap-history domain-register OK
at-o.net Disconnect Applied Technologies Internet SAS (AT Internet), controlled by Piano S register-by-number register OK
cnevids.com Tracker Radar Condé Nast Entertainment (Advance Publications) live-source legal-doc OK
pages02.net Disconnect Acoustic, L.P. rdap-history domain-register OK
trustpilot.net Tracker Radar Trustpilot A/S (Pilestraede 58, 5th floor, DK-1112 Copenhagen K, Denma live-source legal-doc OK
awltovhc.com webXray Epsilon (d/b/a 'Epsilon PeopleCloud Digital Media Solutions', formerly rdap-history newsroom OK
lduhtrp.net webXray Conversant, Inc. (Epsilon / Publicis Groupe) rdap-history domain-register OK
wishabi.com webXray Flipp Corp. live-source legal-doc OK
sa-as.com Disconnect FoundryCo, Inc. rdap-history domain-register OK
cratecamera.com Tracker Radar Leven Labs, Inc. (DBA Admiral) rdap-history domain-register OK
globo.com Disconnect Globo Comunicação e Participações S.A. (CNPJ 27.865.757/0001-02) live-source legal-doc OK
offshoregeology.com Disconnect Admiral (Leven Labs, Inc.) live-source parent-site OK
km0trk.com Disconnect Good On You Pty Ltd (Good On You) live-source parent-site OK
cjponyparts.com Tracker Radar CJ Pony Parts, Inc. archived-legal archived-legal-doc OK
cedscdn.it Tracker Radar CED Digital & Servizi S.r.l. (Caltagirone Editore group) register-by-number register OK
i.ua Disconnect ТОВ «КЕПРЕЙТ ПАРТНЕРС» (LLC Keprait Partners), Ukrainian register code live-source legal-doc OK

The one non-OK row is adgrx.com: samsungads.ca serves an expired certificate, so curl refuses it and the checker never sees a body. Fetched by hand with –insecure: HTTP 200, 46,160 bytes, and the quoted row is byte-verbatim in it. The checker was deliberately not given an –insecure retry — an ownership citation whose host identity cannot be validated should surface, not be swallowed.

L.5 Three defects the citation check had, found by these rows

owner_verify_sources.py was written for the 2026-09-05 rows and this pass cited kinds of evidence it had never seen. Each defect would have reported a correct citation as a fabrication, and each fix carries a control:

Defect Found on Fix Control
No whois:// scheme, so every registry citation FETCHFAILed 9 rows a port-43 branch calling whois43.ask against the named server, not re-resolved through IANA, so the row is checked against the record it cites the 9 rows now verify OK
body.decode(“utf8”, “replace”) turns every byte of a windows-1251 page into U+FFFD, so a Cyrillic quote can never match help.i.ua/agreement/ a decode() that sniffs the charset the document declares, from the bytes rather than the header, because the cache holds the body alone the quote is found; the same quote with one letter changed is not. The normaliser had already been fixed once for this class of bug, one layer later
A rate-limit stub cached as evidence, giving a permanent false NOTFOUND sa-as.com, blogblog.com refuse to return or cache a WHOIS answer under 400 bytes or one that does not name the queried domain; one retry 20 s later both verify OK on the retry, and the guard was seen to fire before the retry was added

Final tally over the 40 tail rows: FETCHFAIL=1, NOSOURCE=12, OK=27. The 12 NOSOURCE are the rows still unresolved, which have no citation by definition.

L.6 Hand corrections

Six rows a machine flagged and a human then re-checked. keep means the citation holds and the checker or the transport was at fault; resource means the claim holds and the row cited the wrong document for it; rescore means the verdict was wrong on the brief's own rules.

Domain Action Why
app-us1.com resource The cited page is a 404 whose only evidence is a <link> tag, which the quote checker normalises to the empty string and reports as NOQUOTE – so the row's only citation was one the checker cannot see, on a page that does not exist. Re-issued the registrar WHOIS query live (whois.markmonitor.com, port 43): the registrant organisation is un-redacted and names ActiveCampaign, LLC. The verdict does not change; the evidence under it does, from a 404 page to the registry record. Rate effect: none.
blogblog.com resource The row's source URL was prose – 'whois:blogblog.com (registrar WHOIS, whois.markmonitor.com)' – which no client can dereference, so it FETCHFAILed. Rewritten in the whois:<server>/<domain> form the checker parses and re-queried live: 'Registrant Organization: Google LLC'. Same record, same quote, a URL that resolves. Rate effect: none.
ksearchnet.com resource The row cites an OV certificate's validated O= field – which the brief accepts as register-grade – but gave the https:// URL of the host, which now returns HTTP 404, so the checker fetched a 404 body and looked for a certificate subject in it. Repointed to the tls:// form the checker handles. Handshake redone 2026-09-11: subject=C=FI, ST=Uusimaa, O=Klevu Oy, CN=*.ksearchnet.com, issuer Sectigo Public Server Authentication CA OV R36. Rate effect: none.
i.ua rescore Two things were wrong. (1) The row rested on the WHOIS registrant, 'Digital Ventures LLC'. The portal's own user agreement names a different company as its administration – ТОВ «КЕПРЕЙТ ПАРТНЕРС», register code 33500955 – so on this domain the registrant of record is NOT the operator. That is the single most important finding about the domain-register source kind this pass added, and it is kept in the residue notes rather than smoothed away. (2) The adjudicator scored Disconnect's entry 'I.UA' as error because the registrant's name differs from it. But the brief scores a trading brand for the same business as current, and the operator's own agreement calls the property 'порталу I.UA' – the portal I.UA. error means 'never the owner at any time', which this is not. Rescored current. Rate effect: moves Disconnect UP, which is the self-serving direction, so the alternative reading is published beside it and the sensitivity run reports what happens without it.
adgrx.com keep FETCHFAIL 'HTTP 000' is the TLS layer, not the citation: samsungads.ca serves an expired certificate, so curl refuses it and the checker never sees a body. Re-fetched by hand with –insecure on 2026-09-11: HTTP 200, 46,160 bytes, and the quoted row 'adgrx.com</span></td><td>ADGRX_UID' is byte-verbatim in it. The checker is deliberately NOT given an –insecure retry: an ownership citation whose host identity cannot be validated should surface, not be swallowed. Same call the 2026-09-11 inter-rater pass made on its own expired-cert row.
sa-as.com keep NOTFOUND was a cached rate-limit stub, not a bad citation. whois.markmonitor.com answers the fourth rapid query with a record that has no registrant block, and the checker cached it. Re-queried 25 seconds later: 3,357 bytes, 'Registrant Organization: FoundryCo, Inc.' present. owner_verify_sources.py now refuses to cache a whois answer under 400 bytes or one that does not name the queried domain, and blogblog.com hit the same stub. The row still rests on an UNCORROBORATED registrant organisation and is counted as such.

Only one verdict changed, i.ua, and it moves Disconnect up, which is the self-serving direction. The alternative reading is that “I.UA” names no company at all and the row is an error; under it Disconnect's domain-level current would be lower still than the 89.2% now published, so the published figure is the more favourable of the two readings and not the less.

L.7 What is still unresolved, and every route tried

12 distinct domains of the 175, down from 40. These are not rows where the register was not tried. Each note below is the adjudicator's own account of what each of the three routes returned, in full and not truncated – an earlier draft cut this column at 420 characters mid-word, which made “every route tried” a claim a reader could not check.

Two of the twelve are adjudicator limits rather than evidence limits, and saying so matters because they are the two a person could still close. hqseek.com has a live site: it redirects to an adult domain and the adjudicator declined to fetch further, so the row is unresolved because a model stopped, not because the evidence is gone – and its note records that Tracker Radar's claimed owner matches no name the domain ever surfaced, which would be an error rather than an unknown if anyone read the pages. stripst.com is unresolved because Stripchat's legal pages return HTTP 406 to a non-browser client, which is a transport failure, not an absence of documents. The other ten are genuinely evidence-exhausted.

Domain In the residue of What each list claims Why it is still unresolved
1rx.io webXray wx: Blinkx / TR: RhythmOne / DC: Nexxen register-by-number: no company/VAT number available anywhere for 'Blinkx'/'RhythmOne'/'Nexxen' tied to this specific domain, so nothing to look up. rdap-history: RDAP has no data for the .io TLD via rdap.org; registry WHOIS (whois.nic.io) refers to GoDaddy, and the registrant is fully REDACTED (Domains By Proxy, LLC) with no organisation name, created 2015-05-28. archived-legal: root domain 1rx.io has zero Wayback captures ever (checked cdx with and without wildcards); only the ad-serving subdomain a-ams.1rx.io was archived, and only as RTB delivery JS calls, never a privacy/legal page. Also checked SEC EDGAR full-text search for the literal string '1rx.io' (Nexxen/Tremor International is Nasdaq-listed) – 0 hits. No primary source anywhere names this domain, so it stays unresolved, same as the first pass.
agkn.com webXray, Disconnect wx: Neustar Marketing / TR: TransUnion LLC / DC: TransUnion Register-by-number: SEC EDGAR full-text search for the exact phrase “agkn.com” across all filings returned 0 hits; “AdAdvisor” appears only in two old NEUSTAR INC (CIK 1265888) filings from before the 2021 TransUnion acquisition, nothing from TransUnion itself. RDAP/WHOIS history: registrant is fully privacy-proxied (Brandsight Privacy Customer, PO Box, Boise ID) with no organization field populated, at either GoDaddy or the registrar's own WHOIS; only the 2005 creation date is usable and it merely predates the whole ownership chain rather than confirming any link in it. Archived legal pages: Wayback shows a 2016 privacy.html and cookie-sync URLs branded “Neustar AdAdvisor” through 2021, but nothing after, and I could not fetch a current TransUnion page (transunion.com/privacy/media-digital-marketing-solutions, transunion.com/privacy/neustar) to bridge that gap – both returned HTTP 403 (Cloudflare/Akamai bot wall) to curl and to WebFetch alike, so no live source could be confirmed. Leaving unresolved rather than guessing from the secondary-reporting chain.
marphezis.com Tracker Radar TR: Online Media Solutions Ltd. dba Brightcom register-by-number: no identifier available for 'Online Media Solutions Ltd. dba Brightcom' tied to this domain specifically. rdap-history: RDAP/registry WHOIS registrant is fully REDACTED (GoDaddy/Domains By Proxy pattern), registered 2015-07-14, last changed 2026-07-15 (routine renewal, not evidence of transfer). archived-legal: fetched the only 200-status legal-adjacent capture found – it is a GoDaddy domain-parking iframe (mcc.godaddy.com/park/…) from 2016, i.e. the domain was parked/lapsed, not operated by anyone; no privacy/terms page was ever archived. Also checked brightcom.com's own privacy/legal pages (403/404, no mention possible) and SEC EDGAR full-text search for 'marphezis' (0 hits, though Brightcom is TASE-listed, not SEC, so this is weak). Nothing ties the domain to Brightcom or any other real operator; stays unresolved.
cdnbasket.net Tracker Radar, Disconnect TR: Bounce Exchange / DC: cdnbasket.net All three routes came back empty. Register-by-number: no identifying company/VAT number found anywhere for this domain or for Bounce Exchange/Wunderkind. RDAP/WHOIS history: registrant is privacy-proxied today (Contact Privacy Inc., Tucows), last changed 2026-08-15, with no un-redacted historical record available. Archived legal pages: the Wayback CDX index has zero captures of cdnbasket.net at any URL, ever. bounceexchange.com does 302 to wunderkind.co, and Wunderkind's own privacy policy (fetched, quoted above) confirms 'Wunderkind Corporation' as a real live entity, but nothing there or in its cookie policy names cdnbasket.net, so the tie to Bounce Exchange/Wunderkind stays an unconfirmed lead, not a resolution.
mapixl.com Disconnect DC: MarketingArchitects All three routes tried and none produced an identifier for the actual operator, let alone one naming 'MarketingArchitects'. register-by-number: no imprint/legal page exists anywhere to pull a company number from. rdap-history: re-queried whois.godaddy.com by socket myself this session – registrant is 'Domains By Proxy, LLC', an explicit privacy proxy, which the brief excludes as a domain-register source; DNS is Cloudflare-fronted (104.21.x/172.67.x, no MX, TXT is an opaque verification hash) with no organizational signal. archived-legal: Wayback CDX for mapixl.com shows only redirect/revisit captures back to 2019; the one distinct body I could fetch (20230420115444id_) turned out to be a Microsoft Azure AD sign-in page, unrelated to any ad-tech company and not a legal page. Live fetch returns a Cloudflare JS-challenge page (HTTP 403, title 'Attention Required! | Cloudflare'). No route ties this domain to MarketingArchitects or to any other named entity; left unresolved as instructed.
htplayground.com Disconnect DC: htplayground.com register-by-number: Disconnect's own entry is just the domain string, so there is no candidate company name to get an identifier for. rdap-history: rdap.org has no RDAP service for this registrar; registry WHOIS is thin, and the registrar-side WHOIS (whois.registrar.amazon) shows a proxy registrant ('c/o whoisproxy.com', 604 Cameron Street, Alexandria VA – a known privacy-proxy mail-drop address), no real organisation. archived-legal: the domain has real Wayback history, but it is a personal phpBB forum from 2007-2008 ('H.T Playground', self-described as a hangout for Vietnamese Catholic youth-group members, TNTT/Huynh Truong) – not a company, and a CDX search across the whole path space found zero privacy/legal/terms/about captures at any point in the domain's history. Domain is currently fully dead (no TLS handshake, no HTTP response). No route produced a real company; unresolved, same as the first pass.
contentabc.com Tracker Radar TR: Aylo / DC: ContentABC Register-by-number: no imprint, VAT, or company number for either candidate ('Aylo'/Ethical Capital Partners or a literal 'ContentABC') is attached to this domain anywhere I could find, so there is no number to look up. RDAP/WHOIS history: registrant is fully proxied via 'Whois Privacy (enumDNS dba)' in Luxembourg through EuroDNS S.A.; no organization ever disclosed; creation date 2009 does not narrow the candidates. Archived legal pages: the mechanical probe's Wayback CDX shows zero captures for legal pages, root, or the URL inventory. contentabc.com itself does not resolve from this sandbox (HTTP 000), and a web search for “contentabc.com” together with Aylo/MindGeek returned no corroborating primary source. Disconnect's 'ContentABC' is just the domain label capitalised with no independent evidence behind it.
stat-track.com Disconnect DC: StackTrack All three routes were tried. Register-by-number: Disconnect's 'StackTrack' is not a findable real company, and a GitHub-issue claim of Moosend ownership could not be verified – Moosend's own privacy, cookie and terms pages (all fetched) never mention stat-track.com. RDAP/WHOIS history: rdap.verisign.com (queried live, quoted above) returns only a registrar (MarkMonitor) entity with no registrant object at all – full redaction, domain re-registered 2016-02-02. Archived legal pages: Wayback's CDX index only has captures from 2002-2011 of a self-hosted Prosper202 affiliate-tracking install, which entirely predates the 2016 re-registration and cannot speak to who holds the name now; nothing has been captured since.
mmstat.com webXray wx: Alimama / TR: Alibaba Group / DC: AlibabaGroup Register-by-number: tried China's official ICP/beian registry at beian.miit.gov.cn directly – it returned HTTP 521 behind a JS/cookie challenge I could not solve with curl. A third-party ICP lookup aggregator (icplishi.com; not an accepted source, same category as a WHOIS aggregator) suggests the beian filer is an 'Alibaba (China) … Co., Ltd.' entity, but this is an unverified lead, not admissible evidence, so it is not used as a verdict basis. RDAP/WHOIS history: 'Alibaba Cloud Computing (Beijing) Co., Ltd.' in the WHOIS record is the REGISTRAR (whois.verisign-grs.com and grs-whois.hichina.com both confirm this), not a Registrant Organization – no registrant org line is present at all, only province 'zhe jiang' and country CN. Archived legal pages: Wayback shows only sparse root captures from 2004-2016 (mostly 404 after 2016) and zero legal-page captures. Also tried mmstat.com live (HTTP 000, no connection) and Alibaba Group's/Alibaba.com's own live privacy pages (rule.alibaba.com/rule/detail/2034.htm, alibabagroup.com/en-US/privacy-policy) via curl – neither mentions mmstat.com.
collective-media.net webXray wx: Collective / TR: Collective, Inc. / DC: Collective Tried all three routes. Register-by-number: SEC EDGAR full-text search for '“Collective, Inc.” advertising' returns 424 hits but none for the ad-network Collective, Inc. (it was VC-backed/private, never an SEC filer) – no identifier to look up. RDAP/WHOIS history: re-queried whois.godaddy.com this session; Registrant Organization is 'Domains By Proxy, LLC', a proxy, not acceptable per the domain-register rule. Archived legal pages: the CDX inventory has no privacy/terms/about capture; the one 'root page' capture that returns 200 is contaminated with Internet Archive's own Petabox/search-portal HTML rather than the actual site (an IA artifact), so it establishes nothing. webXray/TR/Disconnect's 'Collective' claim (matching the once-real ad network Collective, Inc.) could not be confirmed or refuted for who holds the domain today; stays unresolved.
hqseek.com Tracker Radar TR: Michal Vodrazka All three routes tried. register-by-number: no imprint or identifiable number anywhere; the domain redirects to an adult site (hqporn.pics) which per content policy I did not fetch further. rdap-history: RDAP has no JSON endpoint for this domain; the two-hop WHOIS in the probe shows only registrar DNC Holdings/DirectNIC and registrant 'Jewella Privacy LLC' (a privacy service), no unredacted organization. archived-legal: fetched the earliest available capture with content (2004-04-04, an adult-directory page even then) and its meta tags read 'copyright: Miguel Bain' / 'author: Miguel Bain' – a real named individual, but a DIFFERENT name from Tracker Radar's claimed 'Michal Vodrazka', and there is a 20+ year gap with a privacy-proxied registrant in between, so this cannot be bridged to a verdict about today. Flagging the tension: TR's specific personal-name claim does not match the only name this domain's own pages have ever surfaced.
stripst.com Disconnect DC: Stripchat Register-by-number: no imprint or company number is available anywhere for stripst.com, and nothing on the domain or Stripchat's reachable pages names an operating legal entity, so there is no number to look up. RDAP/WHOIS history: registrant is fully proxied ('Withheld For Privacy LLC', Delaware) via NameCheap; the mechanical probe's RDAP lookup also failed; TLS is a domain-validated Google Trust Services cert with no O= field, so no organization there either. Archived legal pages: the mechanical probe confirms zero Wayback captures of stripst.com itself (root, legal, and inventory all empty). I additionally tried Stripchat's own live /privacy, /terms and /legal pages hoping to find a parent-site mention of stripst.com, but all three returned HTTP 406 (bot-blocked) to curl; stripst.com itself, cdn.stripst.com, and www.stripst.com all return HTTP 000 (no connection) from this sandbox. Disconnect's 'Stripchat' claim is plausible (third-party scan data and CT-log subdomain names like cdn.stripst.com are consistent with Stripchat infrastructure) but remains unverified against a primary source.

Two obstacles are worth naming for anyone attempting this again: mmstat.com would be settled by China's ICP registry at beian.miit.gov.cn, which serves a JavaScript challenge no fetch here could pass; agkn.com by TransUnion's own pages, which return 403 to everything that is not a browser session. The remaining residue is a lower bound on what harder sourcing can reach, not a claim that these domains have no owner.

L.8 What the pass could not establish

  • Whether the 12 still-unresolved rows are worse than the rest is still unbounded, and this pass makes the bound narrower rather than tighter: the rows it recovered are by construction the easier end of the residue, so the test in section B of report_tail_pass.py is a test of the recovered rows, not of the residue. It can show the recovered rows are worse, or fail to; it cannot say anything about what is left. report_tail_pass.py says so in its own docstring and the content page says so in prose.
  • The recovered/existing comparison has almost no power. 7, 9 and 12 recovered rows. A difference would have to be enormous to register, and none of the three p-values is interpretable as evidence of no difference.
  • The adjudicators were again language models, one rater per domain, working to out/tail_adj/INSTRUCTIONS_TAIL.md. No inter-rater agreement was measured for this pass — the +0.400 kappa on the content page is the 2026-09-05 rows. Measuring it would need a second model over the same 40 and was not attempted this sitting.
  • Three sub-agents could not re-issue a port-43 query themselves (this sandbox's curl has no whois:// support and bash has no /dev/tcp) and relied on the probe's already-fetched WHOIS text, disclosing it per row. Every such citation was independently re-fetched by owner_verify_sources.py afterwards, which is what closes that gap.
  • cratecamera.com and sa-as.com rest on an uncorroborated registrant organisation and nothing further could be found for either: Admiral publishes no allowlist of its own domains, and no Foundry page names sa-as.com. They are in the sensitivity run for that reason.

L.9 Reviewers, and what each one was worth

All four told explicitly that the author's context may not be exhaustive, all handed the page text, the scripts, their unedited output and these notes. The three focused passes ran in parallel against a bundle cut at 19:21; the generic pass ran afterwards against the pages as the first three had left them.

Findings accepted

Pass Finding What was done
figures vs script added 14 rows to webXray, 15 to Tracker Radar — the recovered counts are 7 / 9 / 12. 14 and 15 match no quantity in the data Accepted, and it is the author's own fabrication. Corrected.
figures vs script the same sentence mixes conventions: 9 of Disconnect's 12 were current … against 41 of its previous 42. 41/42 is current+granularity; against a current-only numerator the comparable figure is 39 of 42 Accepted; rewritten to name the convention.
figures vs script / citations / generic both guards assert a correct figure is present and neither asserts a superseded one is absent, so the headline box could carry the old rates with both exiting 0 Accepted and closed with code. check_tail_figures.py gained section G; three mutations that reintroduce a superseded rate were added to the harness.
figures vs script the whois:// retry runs on a short response but returns on the first raised exception, contradicting its own comment Accepted; fixed, and the failure path exercised by hand.
citations adj_rows_tail_final.json's i.ua row still carried the adjudicator's superseded “error” reasoning beside a current verdict Accepted; note rewritten, superseded text kept, marked as superseded.
citations offshoregeology.com's note flags a WHOIS line the adjudicator could not re-fetch; this sandbox can Accepted; gap closed in the note, and the reviewer reproduced it independently.
external currency Disconnect's entities.json has not changed since 2026-08-07; the 2026-09-05 commit the page cites touched services.json Accepted, verified independently before acting: commits?path=entities.jsonab6ff5a3a8, 2026-08-07; repo HEAD 4b592c288a, 2026-09-05; the three raw files are byte-identical (412,191 bytes, same MD5). Both places on the content page corrected, with the query in a footnote.
external currency the RDAP-vs-registrar claim is stated as a general rule; it holds only where the registrant has not bought privacy, and that is per-registration, not per-registrar Accepted, reproduced: whois.tucows.com returns Contact Privacy Inc. Customer 0168237798 for cdnbasket.net while returning real organisations for others. §L.2 qualified, and the generator too.

The generic pass, which had no checklist and found the most

It ran after the three focused ones, against the pages as they had left them, and it returned 15 findings. Fourteen were accepted. Two of them changed what the page claims, and both were things all three focused passes and the author's own site-wide sweep had read past.

# Finding What was done
1 The overturned headline is a fact about one metric and the page never said which. All three pairwise tests were published on current only. On current+granularity — which the page itself says is right when the unit is “which company” — both Disconnect comparisons clear Bonferroni, before *and* after, so the second pass overturns nothing there. Worse: report_tail_pass.py tested current+granularity in section B and current-only in E2, each the choice that made its own claim mildest, and disclosed neither Accepted; the most consequential finding of the sitting. Reproduced independently before acting: 52/54 vs 44/56 → *p* = 0.008, 52/54 vs 43/56 → *p* = 0.004; before, 0.010 and 0.004. Sections B and E2 now print both metrics and say why; the Fisher table on the page carries both; the recommendation and the short answer name the metric each claim belongs to
2 “All four errors were found by the second pass” is false. elfsight.com was Tracker Radar's one error before it Accepted. Three are new. Corrected on the content page and on provenance:programming:crawler:webxray
3 The short answer's “no longer clears” implies Disconnect-vs-Tracker-Radar was once significant. It never was — 0.025 before, against a threshold of 0.0167 Accepted; rewritten to say it never cleared it
4 Disconnect's only error, km0trk.com, is a scoring decision on a bare-label entry presented as the same kind of failure as a law-firm attribution — and the rule that separates it from i.ua, scored the other way, is nowhere on the page. The remaining Disconnect residue is enriched for that class, which bears on the worst-case bound Accepted in both halves. The rule is now stated (is the label a brand the operator's own documents use for the property?), with the admission that a reasonable person would score it the other way. The residue enrichment is verified from the §L.7 table — 3 of Disconnect's 6 remaining rows: cdnbasket.net and htplayground.com are the bare domain string, stat-track.com's “StackTrack” is not a findable company — and said on the page next to the bound
5 The published sensitivity run covers 2 of the 11 rows that rest on the new bar, and the i.ua alternative reading that the published correction log promises “the sensitivity run reports” was never computed Accepted. E3 now has three runs: the original 2-row one, the same-bar run that drops all 11 (webXray 69.2%, Tracker Radar 72.1%, Disconnect 89.0% — the result does not depend on the widened bar), and the i.ua alternative (Disconnect 88.6%). All three are on the page and all three are asserted by the figure guard
6 “Indistinguishable for the other two” is exactly the interpretation §L.8 says the p-values cannot bear Accepted; the summary now says nothing here establishes that the tail was worse for any list, and names the metric dependence
7 The residue is described as evidence-exhausted; two rows are adjudicator limits, not evidence limits, and every “why it is still unresolved” cell was cut at 420 characters mid-word, so “every route tried” was not checkable Accepted in both halves. hqseek.com (adjudicator declined a live adult redirect on content policy, and its note records a possible error nobody ran down) and stripst.com (HTTP 406 to every non-browser fetch) are now named as limits of this pass. The truncation is removed from the generator, so §L.7 carries the notes in full
8 The new section never says who adjudicated the second pass Accepted. §I of this page records dropping that disclosure as the worst finding of the previous sitting, so this is a repeat. The content page now says: Sonnet sub-agents, one rater per domain, no inter-rater agreement measured for this pass
9 The largest single movement of the pass is absent from the content page: Tracker Radar encounter-weighted 68.6% → 74.3%, almost all of it one row (spot.im, prevalence 0.011) — the column the page calls “the number a paper's attribution error depends on” Accepted; a row added to the before/after table and a sentence to the prose
10 “29 of the 525 verdicts were changed” predates this sitting's own i.ua rescore Accepted; 30 now, split by date
11 Stale limit bullet on provenance:programming:crawler:webxray: “n is 42–49 per list … 69.9 vs 73.7” Accepted. Now 54–56 and 69.2 vs 73.9, with the Fisher *p*. The sweep could not see it because neither figure carries a % — recorded in the bullet itself
12 §L.2's “the registrar hop did most of the work” disagrees with the owner_tail_merge.py output published two sections below (9/13 against 10/12) Accepted; the difference is the i.ua correction moving one row's route, and §L.2 now says its counts are post-correction and the appendix output is not
13 Two published scripts print different 95% intervals for the same estimate, under a claim of “imported, not reimplemented” Accepted. The intervals are removed from report_tail_pass.py's output rather than reconciled: its bootstrap draws from a shared generator, so an identically seeded re-run legitimately differs. The estimator's own output remains the only source for intervals, and the script now says so
14 Three prose misstatements: “changed the answers more than it changed the rows” (28 of 40 rows changed); “narrower rather than tighter” (synonyms); “the highest and by the widest margin” (unclear) Two accepted, one rejected. The first is rewritten to the actual quantity (23% of draws → 7.8%); the second is rewritten. The third is left: in a bullet that has just given both rates, “highest and by the widest margin” reads plainly, and rewording it costs more than it buys
15 report_tail_pass.py's docstring shows a pre-corrections input file, so copying it reproduces a Disconnect with two errors Accepted; docstring fixed to the file the published run used

What this pass cost and what it bought. It was the only one of the four to question the *choice of metric* rather than the arithmetic within one, and that question changed the page's headline claim from “the second pass overturned Disconnect's lead” to “it overturned a legal-entity-level lead over a list this page calls historical, and cost it nothing at the company level”. The three focused passes checked every number against the script and found the script's own metric switch invisible, because each number was right for the metric its section used. A checklist cannot find a framing defect; this is the strongest evidence so far that the generic slot earns its place.

Findings recorded as already-fixed (stale bundle)

Three of the four passes independently reported that the content page's <WRAP important> “short answer” box still carried 94.4% / 73.7% / 93.9%. It did when their bundle was cut at 19:21. A site-wide sweep (scripts/sweep_moved_figures.py) had found and fixed it at 19:28, along with three more the first patch missed. The finding was right and the bundle was stale — the same trap this wiki has recorded before. Two of the four passes also named the <WRAP todo> that still asked for the thing this sitting did; that one was not yet fixed and has been.

Nothing found in

Fisher's exact (independently reimplemented: 0.031 / 0.083 / 0.831, matching); the eligibility rule (read against the estimator's, equivalent, and asserted at runtime); the worst-case convention; the raw-count arithmetic; the source-kind split (9 / 2 / 17, recomputed from the rows); the reproducibility of every committed output; citekeys (22, all resolving once against a freshly fetched bibliography, none added); all 28 recovered citations re-fetched and verified verbatim by a second party; the three error verdicts, each re-derived from a primary source; the i.ua override, judged substantively right by an independent read of the surrounding agreement text.

J. scripts/report_ownership_resolution.mjs, in full

report_ownership_resolution.mjs
// Every corpus figure on Design:Ownership resolution, with its denominator.
//
//   node scripts/report_ownership_resolution.mjs > scripts/report_ownership_resolution-output.txt
//   node scripts/report_ownership_resolution.mjs --wiki    # DokuWiki tables
//
// This page was carved out of Programming:Crawler:webXray on 2026-09-11. The
// corpus-side figures it inherited were computed by scripts/report_webxray.mjs,
// whose population is "papers that name webXray". That is the wrong population
// for a page about ownership resolution in general, so this script re-derives
// the landscape figures from the corpus independently and then CROSS-CHECKS the
// overlapping rows against report_webxray.mjs's committed output. Two
// implementations agreeing is a stronger audit trail than one copied number.
//
// Three things this script does that a naive sweep would not:
//
// 1. Every probe is run at TWO widths and the script asserts tight <= loose AND
//    tight subset-of loose. A narrowing probe that returns more hits, or hits
//    the loose one does not contain, is a broken probe, not a finding. The
//    Disconnect probe is the one that needs this: plain /disconnect/i is a
//    common English word and the CSP literature uses it as a technical term.
//
// 2. Every count is a PAPER count over data/fulltext/<year>/<venue>/<slug>/
//    paper.cols.txt with whitespace collapsed, because a PDF line break inside
//    "Public Suffix List" silently undercounts it.
//
// 3. Papers with no full text are counted and printed, so the sweep's own
//    denominator is visible rather than assumed to be 5,859.
//
// Every row here is an UPPER BOUND on use unless the output says hand-verified:
// a full-text match is a mention, not a use. The two rows that ARE hand
// verified (webXray, Tracker Radar) carry their verdicts in report_webxray.mjs.
 
import fs from 'node:fs';
import path from 'node:path';
import { dataRoot, loadExtractions, POPULATIONS, pct, table, wikiTable } from './lib.mjs';
 
const WIKI = process.argv.includes('--wiki');
const rows = loadExtractions();
const key = (p) => `${p.venue}/${p.year}/${p.slug}`;
const FT = path.join(dataRoot(), 'fulltext');
 
const out = [];
const P = (s = '') => out.push(s);
const heading = (s) => {
  P('');
  P(`=== ${s} ===`);
  P('');
};
const T = (headers, body) => P(WIKI ? wikiTable(headers, body) : table(headers, body));
 
// --- full text, whitespace collapsed -------------------------------------
// Streamed, not cached: the corpus is 5,859 papers and holding every
// paper.cols.txt in a Map exhausts the default V8 heap. One pass, every
// pattern tested against each paper's text, then the text is dropped.
function readText(k) {
  const [venue, year, slug] = k.split('/');
  const f = `${FT}/${year}/${venue}/${slug}/paper.cols.txt`;
  return fs.existsSync(f) ? fs.readFileSync(f, 'utf8').replace(/\s+/g, ' ') : null;
}
 
// Which papers have full text at all, and the hits for every pattern, in one
// pass. PATTERNS is filled below before this runs.
function sweepAll(patterns) {
  const hits = new Map([...patterns.keys()].map((n) => [n, []]));
  const haveText = new Set();
  for (const p of rows) {
    const k = key(p);
    const t = readText(k);
    if (t === null) continue;
    haveText.add(k);
    for (const [name, re] of patterns) if (re.test(t)) hits.get(name).push(k);
  }
  for (const v of hits.values()) v.sort();
  return { hits, haveText };
}
 
// --- the probes, each at two widths --------------------------------------
// `loose` must be a strict superset of `tight` by construction; the script
// checks that it is in fact one and throws if not.
const PROBES = [
  {
    name: 'webXray',
    tight: /webx[\s-]?ray/i,
    loose: /webx[\s-]?ray|libert/i,
    note: 'hand-verified: all 15, in report_webxray.mjs',
  },
  {
    name: 'Tracker Radar',
    tight: /tracker[\s.-]?radar/i,
    loose: /tracker[\s.-]?radar|duckduckgo/i,
    note: 'hand-verified: all 32, in report_webxray.mjs',
  },
  {
    name: 'WhoTracks.me',
    tight: /whotracks/i,
    loose: /whotracks|who ?tracks ?\.? ?me|ghostery/i,
    note: 'upper bound',
  },
  {
    name: 'Crunchbase',
    tight: /crunchbase/i,
    loose: /crunchbase|crunch base/i,
    note: 'upper bound',
  },
  {
    name: 'Disconnect (list sense)',
    tight: /disconnect'?s? (entit|list|block|black|tracking)|disconnect\.me|entities\.json/i,
    loose: /disconnect/i,
    note: 'upper bound; the loose width is the reason this one is tightened',
  },
  {
    name: 'Public Suffix List',
    tight: /public suffix/i,
    loose: /public suffix|etld\+?1|effective top-?level domain/i,
    note: 'upper bound',
  },
];
 
// One pass over the full text: every probe, both widths, plus the set of papers
// that have a readable rendering at all.
const PATTERNS = new Map();
for (const pr of PROBES) {
  PATTERNS.set(`${pr.name}|tight`, pr.tight);
  PATTERNS.set(`${pr.name}|loose`, pr.loose);
}
const { hits: SWEPT, haveText: HAVE_TEXT } = sweepAll(PATTERNS);
 
const withText = rows.filter((p) => HAVE_TEXT.has(key(p)));
const crawled = rows.filter(POPULATIONS.crawled);
 
P('==============================================================================');
P('Design:Ownership resolution — every corpus figure, with its denominator');
P('==============================================================================');
P('');
P(`Corpus: ${rows.length} extracted papers, 7 venues (CCS, IMC, NDSS, PETS, USENIX`);
P('Security, TheWebConf, IEEE S&P), 2010–2026. EuroS&P, ACSAC, RAID, AsiaCCS, CHI');
P('and SOUPS are absent, so every figure here is a claim about those seven venues.');
P('');
P(`Papers with a readable paper.cols.txt: ${withText.length} of ${rows.length} (${pct(withText.length, rows.length)}).`);
P(`Papers that ran a crawl (crawlConfig present or studyTypes includes automated-web-crawl): ${crawled.length}.`);
P('The sweep denominator is the full corpus; papers with no full text can only');
P('ever be misses, so every sweep count below is a floor on the true mention count.');
 
// --- A. probe widths ------------------------------------------------------
heading('A. Probe widths: tight must be contained in loose');
 
const results = new Map();
const widthRows = [];
for (const pr of PROBES) {
  const tight = SWEPT.get(`${pr.name}|tight`);
  const loose = SWEPT.get(`${pr.name}|loose`);
  const tightSet = new Set(tight);
  const looseSet = new Set(loose);
  const escaped = tight.filter((k) => !looseSet.has(k));
  if (tight.length > loose.length) {
    throw new Error(`probe ${pr.name}: tight (${tight.length}) > loose (${loose.length}) — the narrowing probe is broken`);
  }
  if (escaped.length > 0) {
    throw new Error(`probe ${pr.name}: ${escaped.length} tight hits are not in the loose set: ${escaped.slice(0, 5).join(', ')}`);
  }
  results.set(pr.name, tightSet);
  widthRows.push([pr.name, String(tight.length), String(loose.length), pr.note]);
}
T(['Resource', 'Tight probe', 'Loose probe', 'Status'], widthRows);
P('');
P('Both widths are printed because the gap is the claim. Disconnect at the loose');
P('width matches 700 papers and almost none of them mean the list, which is why');
P('the published row is the tight one and why the page says so. A row whose two');
P('widths are close is a row whose count is not an artefact of the pattern.');
 
// --- B. the published landscape table ------------------------------------
heading('B. The ownership-resolution landscape (the page\'s main table)');
 
const landscape = PROBES.map((pr) => {
  const s = results.get(pr.name);
  const inCrawled = crawled.filter((p) => s.has(key(p))).length;
  return [pr.name, String(s.size), pct(s.size, rows.length), String(inCrawled), pct(inCrawled, crawled.length), pr.note.startsWith('hand-verified') ? 'yes' : 'no — upper bound'];
})
  .sort((a, b) => Number(b[1]) - Number(a[1]));
T(['Resource named in full text', 'Papers', `Share of ${rows.length}`, `of which crawled`, `Share of ${crawled.length}`, 'Hand-verified?'], landscape);
 
// --- C. the union ---------------------------------------------------------
heading('C. The union: how widespread is ownership resolution at all?');
 
// The union is over the five OWNERSHIP resources. The Public Suffix List is
// deliberately excluded: it answers "what is the registrable domain", not
// "whose is it", and including it would nearly double the union on a resource
// that every crawl paper touches for unrelated reasons.
const UNION_NAMES = ['webXray', 'Tracker Radar', 'WhoTracks.me', 'Crunchbase', 'Disconnect (list sense)'];
const union = new Set();
for (const n of UNION_NAMES) for (const k of results.get(n)) union.add(k);
P(`Union of ${UNION_NAMES.join(' / ')}:`);
P(`  ${union.size} papers — ${pct(union.size, rows.length)} of the ${rows.length}-paper corpus,`);
const unionCrawled = crawled.filter((p) => union.has(key(p)));
P(`  and ${pct(unionCrawled.length, crawled.length)} of the ${crawled.length} that ran a crawl (${unionCrawled.length} papers).`);
P('');
P('The Public Suffix List is NOT in the union: it answers "what is the registrable');
P('domain", not "whose is it". Adding it would take the union to ' +
  (() => { const u2 = new Set(union); for (const k of results.get('Public Suffix List')) u2.add(k); return u2.size; })() +
  ' on a resource');
P('every crawling paper touches for unrelated reasons.');
P('');
P('Each member is an upper bound on USE, so the union is an upper bound too:');
P('136 papers mention one of these; far fewer resolved an owner with one.');
 
// --- D. per-year adoption -------------------------------------------------
heading('D. Per-year: is ownership resolution growing?');
 
const years = [...new Set(rows.map((p) => p.year))].sort();
const yearRows = years.map((y) => {
  const inYear = rows.filter((p) => p.year === y);
  const hit = inYear.filter((p) => union.has(key(p)));
  const star = y >= 2025 ? '*' : '';
  return [`${y}${star}`, String(inYear.length), String(hit.length), pct(hit.length, inYear.length)];
});
T(['Year', 'Corpus papers', 'Naming an ownership resource', 'Share'], yearRows);
P('');
P('* 2025 and 2026 are the provisional corpus edge: CCS 2026 and IMC 2026 have not');
P('  been held, and IEEE S&P / WWW 2026 abstracts are not in OpenAlex, so selection');
P('  under-samples them BY CONSTRUCTION. Do not read 2026 as a complete year.');
P('  The share, not the count, is the readable column for those two rows.');
 
// --- E. venue shape -------------------------------------------------------
heading('E. Venue shape of the union');
 
const venues = [...new Set(rows.map((p) => p.venue))].sort();
const venueRows = venues.map((v) => {
  const inVenue = rows.filter((p) => p.venue === v);
  const hit = inVenue.filter((p) => union.has(key(p)));
  return [v, String(inVenue.length), String(hit.length), pct(hit.length, inVenue.length)];
}).sort((a, b) => Number(b[2]) - Number(a[2]));
T(['Venue', 'Corpus papers', 'Naming an ownership resource', 'Share of venue'], venueRows);
P('');
P('Read the SHARE column, not the count: the venues differ in size by a factor of');
P('three. IMC and PETS are where this work lands; the shares are small everywhere,');
P('which is the honest headline — ownership resolution is a step inside a paper,');
P('not a paper topic, and most papers that do it do not say how.');
 
// --- F. the schema's own view --------------------------------------------
heading('F. What the extraction schema sees, against the sweep');
 
// classification[].groundTruthSource / resourceName naming an ownership list.
const SCHEMA_RE = /webx[\s-]?ray|tracker[\s.-]?radar|whotracks|disconnect|crunchbase/i;
const schemaPapers = new Set();
for (const p of rows) {
  const names = [
    ...p.tools.map((t) => t.name),
    ...p.classification.map((c) => c.resourceName),
    ...p.classification.map((c) => c.groundTruthSource),
    ...p.otherToolsMentioned.map((t) => t.name),
  ].filter((s) => typeof s === 'string');
  if (names.some((n) => SCHEMA_RE.test(n))) schemaPapers.add(key(p));
}
const both = [...union].filter((k) => schemaPapers.has(k));
P(`Papers whose tools[] / classification[] / otherToolsMentioned[] name one of the`);
P(`five resources:            ${schemaPapers.size}`);
P(`Papers whose FULL TEXT names one:              ${union.size}`);
P(`In both:                                       ${both.length}`);
P(`Full text only (schema misses):                ${union.size - both.length}`);
P(`Schema only (no full-text match):              ${schemaPapers.size - both.length}`);
P('');
P('The schema-only residue is printed rather than dropped:');
const schemaOnly = [...schemaPapers].filter((k) => !union.has(k)).sort();
for (const k of schemaOnly) {
  const p = rows.find((r) => key(r) === k);
  const hits = [
    ...p.tools.map((t) => t.name),
    ...p.classification.map((c) => c.resourceName),
    ...p.classification.map((c) => c.groundTruthSource),
    ...p.otherToolsMentioned.map((t) => t.name),
  ].filter((s) => typeof s === 'string' && SCHEMA_RE.test(s));
  P(`  ${k}`);
  P(`      named: ${[...new Set(hits)].join(' | ')}`);
  P(`      full text present: ${HAVE_TEXT.has(k)}`);
}
P('');
P('This is why the page publishes the full-text sweep and not the schema field:');
P('the schema fires on a fraction of the papers that name these lists, because');
P('most mentions are in related work or a reference list rather than in a tools');
P('sentence the extractor reads as a tool. Neither signal is "the" population;');
P('both are reported.');
 
// --- G. cross-check against report_webxray.mjs ---------------------------
heading('G. Cross-check against scripts/report_webxray-output.txt');
 
const committed = 'scripts/report_webxray-output.txt';
if (!fs.existsSync(committed)) {
  P(`${committed} not found — cross-check SKIPPED.`);
} else {
  const txt = fs.readFileSync(committed, 'utf8');
  const CHECKS = [
    ['webXray', /^webXray\s+(\d+)\s/m],
    ['Tracker Radar', /^Tracker Radar\s+(\d+)\s/m],
    ['WhoTracks.me', /^WhoTracks\.me\s+(\d+)\s/m],
    ['Crunchbase', /^Crunchbase\s+(\d+)\s/m],
    ['Disconnect (list sense)', /^Disconnect \(list sense\)\s+(\d+)\s/m],
    ['Public Suffix List', /^Public Suffix List\s+(\d+)\s/m],
  ];
  const checkRows = [];
  let bad = 0;
  for (const [name, re] of CHECKS) {
    const m = txt.match(re);
    if (m === null) throw new Error(`cross-check: could not find "${name}" in ${committed}`);
    const theirs = Number(m[1]);
    const mine = results.get(name).size;
    if (theirs !== mine) bad += 1;
    checkRows.push([name, String(mine), String(theirs), theirs === mine ? 'agree' : '*** DISAGREE ***']);
  }
  const um = txt.match(/^\s*(\d+) papers \((\d+\.\d)% of the corpus\)\. (\d+) of them are among the (\d+) that crawled -- (\d+\.\d)% of that population\./m);
  if (um === null) throw new Error(`cross-check: could not find the union line in ${committed}`);
  if (Number(um[1]) !== union.size) bad += 1;
  checkRows.push(['union of the five', String(union.size), um[1], Number(um[1]) === union.size ? 'agree' : '*** DISAGREE ***']);
  if (Number(um[3]) !== unionCrawled.length) bad += 1;
  checkRows.push(['union ∩ crawled', String(unionCrawled.length), um[3], Number(um[3]) === unionCrawled.length ? 'agree' : '*** DISAGREE ***']);
  if (Number(um[4]) !== crawled.length) bad += 1;
  checkRows.push(['crawled denominator', String(crawled.length), um[4], Number(um[4]) === crawled.length ? 'agree' : '*** DISAGREE ***']);
  T(['Figure', 'this script', 'report_webxray.mjs', 'verdict'], checkRows);
  P('');
  if (bad > 0) {
    throw new Error(`${bad} figures disagree between the two independent implementations — do not publish either.`);
  }
  P('All rows agree. The two scripts share no code beyond lib.mjs and the');
  P('whitespace-collapse convention; the patterns were written out separately.');
}
 
// --- H. the page's non-corpus figures, listed so they are not mistaken ----
heading('H. Every number on the page that does NOT come from this corpus');
 
P('Listed so a reader auditing the page knows which script to re-run, and so');
P('that a refresh of this script is never mistaken for a refresh of the page.');
P('');
T(['Figure family', 'Script', 'Snapshot it is pinned to'], [
  ['List shape (827 / 19,148 / 1,887 owners; 3,215 / 38,368 / 7,850 domains)', 'owner_dbs.py --cache out/webxray/cache', '2026-08-17'],
  ['Coverage census (669 / 5,581 / 2,268 of 32,369; the prevalence weights; the rank slices)', 'owner_dbs.py --cache out/webxray/cache', '2026-08-17'],
  ['Agreement between pairs (612 / 464 / 1,538 and the "neither" columns)', 'owner_dbs.py --cache out/webxray/cache', '2026-08-17'],
  ['The 28-row hand adjudication of high-prevalence disagreements', 'owner_adjudication.py --table', '2026-08-17'],
  ['Accuracy rates (69.2 / 73.9 / 89.2%, the encounter-weighted column, the CIs)', 'owner_random_sample.py --sample out/owner_sample.json --rows out/adj_rows_tail_final.json', '2026-09-05 draw, 2026-09-11 second pass'],
  ['What the second pass over the 40 unadjudicable rows changed', 'owner_tail_probe.sh | owner_tail_merge.py | owner_tail_corrections.py | report_tail_pass.py', '2026-09-11'],
  ['Coverage counts inside the random-sample section (664 / 5,566 / 2,264 of 32,337)', 'owner_sample.py --cache out/owner_cache', '2026-09-05'],
  ['Self-named entity shares (0.8% / 0.7% / 3.2%)', 'owner_selfname_census.py', '2026-09-05'],
  ['Inter-rater kappa (+0.400, +0.680)', 'owner_irr_kappa.py', '2026-09-11'],
  ['Repository state, licences, Wayback dates', 'hand checks + wayback_webxray.sh', '2026-08-17 / 2026-09-05'],
]);
P('');
P('Two of the three lists change weekly. Re-run before citing.');
 
console.log(out.join('\n'));

K. Its unedited output

Run as node scripts/report_ownership_resolution.mjs.

report_ownership_resolution-output.txt
==============================================================================
Design:Ownership resolution — every corpus figure, with its denominator
==============================================================================
 
Corpus: 5859 extracted papers, 7 venues (CCS, IMC, NDSS, PETS, USENIX
Security, TheWebConf, IEEE S&P), 2010–2026. EuroS&P, ACSAC, RAID, AsiaCCS, CHI
and SOUPS are absent, so every figure here is a claim about those seven venues.
 
Papers with a readable paper.cols.txt: 5855 of 5859 (99.9%).
Papers that ran a crawl (crawlConfig present or studyTypes includes automated-web-crawl): 1120.
The sweep denominator is the full corpus; papers with no full text can only
ever be misses, so every sweep count below is a floor on the true mention count.
 
=== A. Probe widths: tight must be contained in loose ===
 
Resource                 Tight probe  Loose probe  Status
-----------------------  -----------  -----------  ----------------------------------------------------------------
webXray                  15           233          hand-verified: all 15, in report_webxray.mjs
Tracker Radar            32           108          hand-verified: all 32, in report_webxray.mjs
WhoTracks.me             23           124          upper bound
Crunchbase               23           23           upper bound
Disconnect (list sense)  74           700          upper bound; the loose width is the reason this one is tightened
Public Suffix List       101          157          upper bound
 
Both widths are printed because the gap is the claim. Disconnect at the loose
width matches 700 papers and almost none of them mean the list, which is why
the published row is the tight one and why the page says so. A row whose two
widths are close is a row whose count is not an artefact of the pattern.
 
=== B. The ownership-resolution landscape (the page's main table) ===
 
Resource named in full text  Papers  Share of 5859  of which crawled  Share of 1120  Hand-verified?
---------------------------  ------  -------------  ----------------  -------------  ----------------
Public Suffix List           101     1.7%           46                4.1%           no — upper bound
Disconnect (list sense)      74      1.3%           65                5.8%           no — upper bound
Tracker Radar                32      0.5%           28                2.5%           yes
WhoTracks.me                 23      0.4%           19                1.7%           no — upper bound
Crunchbase                   23      0.4%           11                1.0%           no — upper bound
webXray                      15      0.3%           13                1.2%           yes
 
=== C. The union: how widespread is ownership resolution at all? ===
 
Union of webXray / Tracker Radar / WhoTracks.me / Crunchbase / Disconnect (list sense):
  136 papers — 2.3% of the 5859-paper corpus,
  and 9.6% of the 1120 that ran a crawl (107 papers).
 
The Public Suffix List is NOT in the union: it answers "what is the registrable
domain", not "whose is it". Adding it would take the union to 221 on a resource
every crawling paper touches for unrelated reasons.
 
Each member is an upper bound on USE, so the union is an upper bound too:
136 papers mention one of these; far fewer resolved an owner with one.
 
=== D. Per-year: is ownership resolution growing? ===
 
Year   Corpus papers  Naming an ownership resource  Share
-----  -------------  ----------------------------  -----
2010   119            0                             0.0%
2011   116            1                             0.9%
2012   151            0                             0.0%
2013   125            0                             0.0%
2014   166            0                             0.0%
2015   190            0                             0.0%
2016   182            2                             1.1%
2017   231            4                             1.7%
2018   254            5                             2.0%
2019   402            9                             2.2%
2020   404            15                            3.7%
2021   379            13                            3.4%
2022   546            19                            3.5%
2023   719            19                            2.6%
2024   690            17                            2.5%
2025*  770            21                            2.7%
2026*  415            11                            2.7%
 
* 2025 and 2026 are the provisional corpus edge: CCS 2026 and IMC 2026 have not
  been held, and IEEE S&P / WWW 2026 abstracts are not in OpenAlex, so selection
  under-samples them BY CONSTRUCTION. Do not read 2026 as a complete year.
  The share, not the count, is the readable column for those two rows.
 
=== E. Venue shape of the union ===
 
Venue    Corpus papers  Naming an ownership resource  Share of venue
-------  -------------  ----------------------------  --------------
PETS     510            49                            9.6%
USENIX   1410           19                            1.3%
WWW      843            19                            2.3%
IMC      638            17                            2.7%
CCS      990            13                            1.3%
IEEE-SP  767            11                            1.4%
NDSS     701            8                             1.1%
 
Read the SHARE column, not the count: the venues differ in size by a factor of
three. IMC and PETS are where this work lands; the shares are small everywhere,
which is the honest headline — ownership resolution is a step inside a paper,
not a paper topic, and most papers that do it do not say how.
 
=== F. What the extraction schema sees, against the sweep ===
 
Papers whose tools[] / classification[] / otherToolsMentioned[] name one of the
five resources:            89
Papers whose FULL TEXT names one:              136
In both:                                       82
Full text only (schema misses):                54
Schema only (no full-text match):              7
 
The schema-only residue is printed rather than dropped:
  IEEE-SP/2024/to-auth-or-not-to-auth-a-comparative-analysis-of-the-pre-and-post-login-security
      named: Disconnect Tracker Protection List
      full text present: true
  PETS/2022/on-dark-patterns-and-manipulation-of-website-publishers-by-cmps
      named: Disconnect | Disconnect tracking filter list
      full text present: true
  PETS/2024/generalizable-active-privacy-choice-designing-a-graphical-user-interface-for-glo
      named: Disconnect Tracker Protection lists
      full text present: true
  PETS/2024/support-personas-a-concept-for-tailored-support-of-users-of-privacy-enhancing-te
      named: Disconnect
      full text present: true
  USENIX/2019/canvas-fast-and-inexpensive-automotive-network-mapping
      named: physical ECU access and disconnection
      full text present: true
  WWW/2018/the-cost-of-digital-advertisement-comparing-user-and-advertiser-views
      named: Disconnect | Disconnect blacklist
      full text present: true
  WWW/2023/who-funds-misinformation-a-systematic-analysis-of-the-ad-related-profit-routines
      named: Disconnect
      full text present: true
 
This is why the page publishes the full-text sweep and not the schema field:
the schema fires on a fraction of the papers that name these lists, because
most mentions are in related work or a reference list rather than in a tools
sentence the extractor reads as a tool. Neither signal is "the" population;
both are reported.
 
=== G. Cross-check against scripts/report_webxray-output.txt ===
 
Figure                   this script  report_webxray.mjs  verdict
-----------------------  -----------  ------------------  -------
webXray                  15           15                  agree
Tracker Radar            32           32                  agree
WhoTracks.me             23           23                  agree
Crunchbase               23           23                  agree
Disconnect (list sense)  74           74                  agree
Public Suffix List       101          101                 agree
union of the five        136          136                 agree
union ∩ crawled          107          107                 agree
crawled denominator      1120         1120                agree
 
All rows agree. The two scripts share no code beyond lib.mjs and the
whitespace-collapse convention; the patterns were written out separately.
 
=== H. Every number on the page that does NOT come from this corpus ===
 
Listed so a reader auditing the page knows which script to re-run, and so
that a refresh of this script is never mistaken for a refresh of the page.
 
Figure family                                                                             Script                                                                                       Snapshot it is pinned to
----------------------------------------------------------------------------------------  -------------------------------------------------------------------------------------------  ---------------------------------------
List shape (827 / 19,148 / 1,887 owners; 3,215 / 38,368 / 7,850 domains)                  owner_dbs.py --cache out/webxray/cache                                                       2026-08-17
Coverage census (669 / 5,581 / 2,268 of 32,369; the prevalence weights; the rank slices)  owner_dbs.py --cache out/webxray/cache                                                       2026-08-17
Agreement between pairs (612 / 464 / 1,538 and the "neither" columns)                     owner_dbs.py --cache out/webxray/cache                                                       2026-08-17
The 28-row hand adjudication of high-prevalence disagreements                             owner_adjudication.py --table                                                                2026-08-17
Accuracy rates (69.2 / 73.9 / 89.2%, the encounter-weighted column, the CIs)              owner_random_sample.py --sample out/owner_sample.json --rows out/adj_rows_tail_final.json    2026-09-05 draw, 2026-09-11 second pass
What the second pass over the 40 unadjudicable rows changed                               owner_tail_probe.sh | owner_tail_merge.py | owner_tail_corrections.py | report_tail_pass.py  2026-09-11
Coverage counts inside the random-sample section (664 / 5,566 / 2,264 of 32,337)          owner_sample.py --cache out/owner_cache                                                      2026-09-05
Self-named entity shares (0.8% / 0.7% / 3.2%)                                             owner_selfname_census.py                                                                     2026-09-05
Inter-rater kappa (+0.400, +0.680)                                                        owner_irr_kappa.py                                                                           2026-09-11
Repository state, licences, Wayback dates                                                 hand checks + wayback_webxray.sh                                                             2026-08-17 / 2026-09-05
 
Two of the three lists change weekly. Re-run before citing.
  • Ownership resolution — the page these notes are for.
  • webxray — the live-list scripts, the fold residues, the Wayback log, the 2026-08-17 and 2026-09-05 run tables.
  • random_sample — the 175-row adjudication, the estimator, the inter-rater draw and kappa.
  • residue_pass — the 2026-09-11 second pass over the residue in full: its probe, its brief, its merge gate, its corrections, its report, its guard, its mutation harness and every unedited output, including the 40-domain probe log and the diff between the two estimates.
  • Corpus — corpus-level selection and extraction caveats.

References

[1]
Matte, Célestin; Bielova, Nataliia; Santos, Cristiana Teixeira (2020): "Do Cookie Banners Respect my Choice? Measuring Legal Compliance of Banners from IAB Europe's Transparency and Consent Framework", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[2]
Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)
[3]
Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[4]
Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[5]
Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[6]
Selmo, Carlos; Carisimo, Esteban; Bustamante, Fabián E.; Alvarez-Hamelin, J. Ignacio (2025): "Learning AS-to-Organization Mappings with Borges", in: Proceedings of the 2025 ACM Internet Measurement Conference, pp. 120-133. (DOI)
provenance/design/ownership_resolution.1789166430.txt.gz · Last modified: by karel.kubicek.claude