This is an old revision of the document!
Table of Contents
Provenance: Design:Ownership resolution
Working notes behind Ownership resolution. Every query with its population and denominator, the report script and its unedited output, what was folded and what the fold could not reach, the quotes checked against the source, the external sources verified and rejected, and the judgement calls. Corpus-wide selection and extraction caveats are on Corpus and are not restated here.
This page is a carve-out, and most of its audit trail lives elsewhere on purpose. The content page was split out of webXray on 2026-09-11. The scripts that produce its live-database figures, their unedited output, the 175-row adjudication table and the inter-rater material were published in 2026-08 and 2026-09 under the webXray ids and keep those ids, because they are cited from the published record and moving them would break links for no gain:
| What | Where |
|---|---|
The three-list comparison scripts (owner_dbs.py, owner_adjudication.py), their unedited output, the fold residues, the Wayback log, the 2026-08-17 and 2026-09-05 run tables | webxray |
owner_sample.py, owner_random_sample.py, the full 175-row adjudication with every source, the citation re-fetch, the inter-rater draw and kappa | random_sample |
The corpus population of webXray itself (the 15 papers, the hand role map), report_webxray.mjs | webxray |
| This page | the split itself, the new corpus script and its output, the reproduction record of 2026-09-11, the defect that run found, the reviewer log |
No ~~DISCUSSION~~ block: comments belong on the content page. Citations use the same {[citekey]} keys and the same shared Bibliography; this page adds no bibliography entries of its own.
Updated 2026-09-11 (second sitting the same day). The 40 sampled domains the
2026-09-05 pass could not adjudicate were re-adjudicated against harder sourcing
and 28 of them settled, which changed the content page's headline table, its
error bullet, its residue bullet and — the consequential one — its Fisher
table, where the one comparison that survived Bonferroni no longer does.
Section L below is that pass in full: its queries, its new source kinds, the
citations it broke and fixed, the 12 rows still unresolved with every route
tried, and its own reviewer log. Everything above section L is the 2026-09-05
sitting as it stood and is not rewritten, except the query table in section K,
which names the script and rows file the live figures now come from.
The run
| Date | 2026-09-11 |
| Corpus at the time | data/extract/run1/extractions.jsonl, 5,859 papers, 7 venues, 2010–2026. 5,855 have a readable paper.cols.txt. |
| What was done | Created Ownership resolution; rewrote webXray around what was left; repointed Crawler; added a row to Design; turned the Roadmap row blue. |
| New code | scripts/report_ownership_resolution.mjs (346 lines), committed with its output. |
| Code changed | scripts/report_webxray.mjs — a denominator defect, section C below. |
| Models | Opus 5 wrote the pages and the script. Review layer: three focused Sonnet passes and one Fable generic pass — all four ran, section I. The 2026-08-17 and 2026-09-05 ownership adjudications this page inherits were also made by language models, not by a human expert; that disclosure belongs on the content page and section I records that the carve-out dropped it. |
| Accidental exposure or mistakes caught | no exposure. Nine defects, all caught before or during review and all listed in C and I: the denominator in C; a stale ALLOW list and an unreachable guard in report_webxray.mjs; a dropped LLM-adjudicator disclosure; an asserted p-value that was never computed; a raw-versus-post-stratified rate description; three arithmetic slips; four stale inbound links; a broken anchor; a pipe in a table cell. The most serious two — the disclosure and the p-value — were found by the generic pass, which is the pass with no checklist. |
Scope decision: why a separate page, when the last sitting decided against one
This reverses a recorded decision, so the reversal is recorded too. webxray §“Scope decision” says, on 2026-08-17:
A separatedesign:ownership_resolutionpage with webXray as a stub. Rejected: it would leave the wiki's existing red link pointing at a stub, and the material does not split cleanly — webXray's frozen list is the best available illustration of what goes wrong.
Both halves of that reasoning have since stopped holding:
- “webXray as a stub” was the wrong alternative to compare against. There is enough webXray-specific material for a real page — the three-licence history across four states, the availability census, the
domain_owners.jsonschema, the frozen 2016 PSL, the Wayback timeline, and 15 corpus papers with a hand role map. What is left after the carve-out is 27 KB, not a stub. - The red-link problem inverted. On 2026-09-07 the Roadmap queued
design:ownership_resolutionas a promised page andscripts/sitemap.mjsbegan gating on it, so from that date not writing it was the dangling promise. The id was fixed by that decision and is used unchanged. - The discoverability argument is the item's whole point and was never answered. A reader asking “how do I attribute a third-party domain to a company?” does not search for a dead tool. Both earlier sittings recorded this as the right thing to do; neither did it.
Alternatives considered this sitting and rejected:
| Alternative | Why not |
|---|---|
| Leave it, and add a redirect or a prominent pointer on webXray | A pointer does not fix search, does not fix the namespace (programming: is instruments; this is a design decision), and leaves the currency claim — which source to use now — filed under the historical one. |
| Broaden Requests instead | That page answers “is this request a tracker?”, a different question with different tooling (filter lists, not owner databases). The two pages cross-link. |
| Broaden IP classification | It is the network-layer sibling (AS-to-organisation) and is already large; the domain layer has different sources, different failure modes and different licences. Cross-linked both ways instead. |
Move the random_sample code appendix under provenance:design: too | Rejected. Its id is cited from the published content page and from the webXray provenance page; the gain is cosmetic and the cost is broken links. Recorded here so the next sitting does not “tidy” it. |
| Re-derive the coverage census against a 2026-09-11 snapshot | Rejected. The 28-row adjudication is pinned to the 2026-08-17 frame; re-deriving would orphan it. The figures moved unchanged, with their frame dates, and the page carries the two-frames box. This is a carve-out, not a refresh — see the drift check in section B. |
The carve-out map
What moved, what stayed, and what had to be repointed. Nothing was re-derived in the move.
| Section on the old page | Went to |
|---|---|
| The tool: architecture, and why you cannot install it | stayed on webXray |
| The ownership database (schema, tree, frozen PSL) | stayed — it describes webXray's file |
| How it compares to Tracker Radar and Disconnect | moved, as The live sources, compared |
| Coverage: how much of the third-party surface | moved, with the three traps split into their own sub-section |
| Two lists disagree: error or different question? | moved, split: the two axes were promoted to the top of the new page as The question has two axes, the agreement tables became When two lists disagree |
| Reading all three at once (the code) | moved |
| Which list is right, when they disagree? | moved |
| How often is each list right? A random sample | moved, with the inter-rater material given its own sub-section |
| Where these figures come from, and how to redo them | moved, folded into the new page's methodology section |
| Choosing a resolution source now | moved |
| Assembling the pipeline | moved |
| What to report in a paper | moved |
| Use in publications: the 15 webXray papers, role table | stayed |
| Use in publications: the 136-paper sweep, the Tracker Radar year table | moved — they are landscape figures, not webXray figures |
| Wayback timeline | stayed, and was promoted from a methodology bullet to its own section; it was 900 words inside a bullet list |
Day-one drift, found and fixed rather than left:
| Where | What it said after the move | Fixed to |
|---|---|---|
| webXray intro bullet 2 | quoted 69.9% / 93.9% / 2.1% / 17.2% / 7.0% inline as if the page still carried the measurement | one sentence keeping 69.9% and 2.1% with a link, and an explicit “nothing on this page restates them” |
webXray, The ownership database | “§\”Two lists disagree\“ below measures what happens when you forget that” — a same-page reference to a section that had left | repointed to [[Design:Ownership resolution#When two lists disagree]], with the finding (root resolution makes agreement worse) stated inline so the sentence still says something |
webXray, Related pages | listed Requests, Cookies, IP classification — all of which were there for the ownership material | new page took those; webXray's list now leads with the new page and adds Policies for policyXray |
| webXray methodology | listed nine owner_* scripts it no longer publishes figures from | one bullet saying where they went and that their provenance ids are unchanged |
| the 28-row adjudication paragraph | the old page said the adjudicators were “Sonnet sub-agents working to a published brief, single-rated”. The rewrite dropped the clause, leaving “adjudicated one by one against primary sources — company newsrooms, SEC filings” — which reads as human expert work. This was accepted as a blocking finding on 2026-09-05 and the carve-out silently reversed it six days later | restored, and the same disclosure added to the random-sample section, which never carried it |
| Crawler, webXray bullet | “see webXray for what survives of it: the ownership database, and how it compares to Tracker Radar and Disconnect today” — that comparison had moved | repointed to the new page; see section H |
A site-wide inbound-link sweep found four more, on pages nobody would have thought to check. Grepping every cached page for a link to programming:crawler:webxray rather than only the pages this sitting edited:
| Page | What it said | Fixed to |
|---|---|---|
| Requests | a Related-pages bullet linking the webXray page under the link text “webXray and domain-to-company ownership” — “once a request is flagged, this is how to answer whose it is” | repointed to Ownership resolution, with a note that it moved |
| Filter lists | the same bullet, near-verbatim | repointed the same way |
| Legal enforcement | “See webXray for the ownership databases and how much they disagree” — in the bullet on identifying a controller under Reg. 2025/2518 | repointed, and extended to name the accuracy rates, which is what that bullet actually needs |
| Programming (namespace index) | the webXray row read “Domain-to-company ownership lists, and what remains of the tool” — a description of the page, now false | rewritten to describe the tool page and point the ownership question at the new page |
Two of those four used the link text “webXray and domain-to-company ownership”, i.e. they were already treating the webXray page as the ownership page — which is the item's complaint, restated by the wiki itself. A carve-out's drift is not confined to the pages you edited; sweep for inbound links before calling it done. One inbound reference was deliberately left: provenance:design's generated page inventory records a size and a description as of an earlier date, and it is a dated snapshot rather than a live pointer.
A. Corpus queries
All over data/extract/run1/extractions.jsonl (5,859 papers) and data/fulltext/<year>/<venue>/<slug>/paper.cols.txt, whitespace collapsed before matching so a term broken across a PDF column boundary still matches. Every count is a paper count. The script is scripts/report_ownership_resolution.mjs, published in full in section J with its unedited output in section K.
| # | Question | Population / denominator | Answer |
|---|---|---|---|
| A1 | How many papers have a readable full-text rendering at all? | all 5,859 | 5,855 (99.9%). The four without one can only ever be sweep misses, so every count below is a floor. |
| A1b | …and how many of those are damaged? | the 5,855 | 0 empty, 0 under 2 KB, but 90 carry a NUL byte. The script reads them as JavaScript strings and regexes them, so NULs do not affect it. A grep-based sweep would have silently skipped all 90 — grep treats a NUL-carrying file as binary and, depending on the build, prints nothing and exits non-zero. None of the 136 union members is one of the 90, checked explicitly, so no published count depends on this — but the next person writing a sweep should not use grep. |
| A2 | How many papers name each ownership resource? | all 5,859 | PSL 101, Disconnect (list sense) 74, Tracker Radar 32, WhoTracks.me 23, Crunchbase 23, webXray 15 |
| A3 | How many of those are inside the crawling population? | crawled = 1,120 (crawlConfig present, or studyTypes includes automated-web-crawl) | PSL 46, Disconnect 65, Tracker Radar 28, WhoTracks.me 19, Crunchbase 11, webXray 13 |
| A4 | How many name any of the five ownership resources? | all 5,859 | 136 (2.3%) |
| A5 | …and how many of those 136 are in the 1,120? | crawled = 1,120 | 107, i.e. 9.6% of that population. Not 12.1% — see section C. |
| A6 | Does adding the PSL change the union? | all 5,859 | 136 → 221. The PSL is deliberately excluded: it answers “what is the registrable domain”, not “whose is it”. |
| A7 | Is ownership resolution growing? | each year's own paper count | rose to ~3.5% in 2020–2022, has sat at 2.5–2.7% since. Full table on the page; 2010–2015 contributes one paper in total and is omitted rather than padded with zeros. |
| A8 | Which venues? | each venue's own paper count | PETS 9.6%, IMC 2.7%, TheWebConf 2.3%, IEEE S&P 1.4%, USENIX 1.3%, CCS 1.3%, NDSS 1.1% |
| A9 | What does the extraction schema see, against the full text? | all 5,859 | schema 89, full text 136, both 82, full-text-only 54, schema-only 7 — residue printed in full below |
| A10 | Of the 15 webXray papers, how many say which version? | the 8 that used the crawler or the list | 2 say anything; 1 names a commit ([1Matte, Célestin; Bielova, Nataliia; Santos, Cristiana Teixeira (2020): "Do Cookie Banners Respect my Choice? Measuring Legal Compliance of Banners from IAB Europe's Transparency and Consent Framework", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]). Inherited from report_webxray.mjs, re-run and unchanged. |
| A11 | What does “Tracker Radar” mean in the 32 papers that name it? | the 32, hand-mapped | 11 ownership dataset, 9 tracker/category database, 9 the Collector crawler, 2 citation, 1 compared. Inherited, re-run and unchanged. |
A.12 Probe widths: every probe run at two widths, and both printed
A narrowing probe that returns more hits than the loose one, or hits the loose one does not contain, is a broken probe. The script asserts both and throws rather than printing. Both widths are on the record because the gap is itself the claim:
Resource Tight probe Loose probe Status ----------------------- ----------- ----------- ---------------------------------------------------------------- webXray 15 233 hand-verified: all 15, in report_webxray.mjs Tracker Radar 32 108 hand-verified: all 32, in report_webxray.mjs WhoTracks.me 23 124 upper bound Crunchbase 23 23 upper bound Disconnect (list sense) 74 700 upper bound; the loose width is the reason this one is tightened Public Suffix List 101 157 upper bound
The patterns themselves are in a code block and not in a table, because a DokuWiki table cell cannot hold a pipe and every one of these regexes is an alternation. \| is not an escape and wrapping the pattern in DokuWiki's nowiki delimiters does not help either: the cell splits regardless.
webXray tight /webx[\s-]?ray/i
loose /webx[\s-]?ray|libert/i
15 -> 233. "Libert" is a common surname and a French word.
The loose width is a sanity bound, not a candidate set.
Tracker Radar tight /tracker[\s.-]?radar/i
loose /tracker[\s.-]?radar|duckduckgo/i
32 -> 108. Most DuckDuckGo mentions are the search engine
or the browser, not the dataset.
WhoTracks.me tight /whotracks/i
loose /whotracks|who ?tracks ?\.? ?me|ghostery/i
23 -> 124. Ghostery is mostly the extension.
Crunchbase tight /crunchbase/i
loose /crunchbase|crunch base/i
23 -> 23. NO GAP: nobody spells it with a space. This
row's count is not an artefact of the pattern.
Disconnect tight /disconnect'?s? (entit|list|block|black|tracking)|disconnect\.me|entities\.json/i
loose /disconnect/i
74 -> 700. "Disconnect" is an ordinary English word and a
CSP-literature technical term. This is the probe that
needed tightening, and the published row is the tight one.
Public Suffix List tight /public suffix/i
loose /public suffix|etld\+?1|effective top-?level domain/i
101 -> 157. The extra 56 are papers that use the concept
without naming the list.
What this does not establish. A tight probe with no gap is not a probe with full recall — it is a probe whose count is stable under widening in the one direction tried. A paper that resolves domain ownership from a source none of these five names, or names none of them at all, is invisible to every width. See section F.
A.13 The schema-only residue, printed in full
Seven papers whose tools[] / classification[] / otherToolsMentioned[] name one of the five resources but whose full text does not match the tight sweep. Printed rather than dropped, because an invisible residue is a residue nobody looks at:
IEEE-SP/2024/to-auth-or-not-to-auth-a-comparative-analysis-of-the-pre-and-post-login-security
named: Disconnect Tracker Protection List
PETS/2022/on-dark-patterns-and-manipulation-of-website-publishers-by-cmps
named: Disconnect | Disconnect tracking filter list
PETS/2024/generalizable-active-privacy-choice-designing-a-graphical-user-interface-for-glo
named: Disconnect Tracker Protection lists
PETS/2024/support-personas-a-concept-for-tailored-support-of-users-of-privacy-enhancing-te
named: Disconnect
USENIX/2019/canvas-fast-and-inexpensive-automotive-network-mapping
named: physical ECU access and disconnection
WWW/2018/the-cost-of-digital-advertisement-comparing-user-and-advertiser-views
named: Disconnect | Disconnect blacklist
WWW/2023/who-funds-misinformation-a-systematic-analysis-of-the-ad-related-profit-routines
named: Disconnect
Six of the seven are Disconnect, and they are the tightened pattern's misses: the extractor wrote “Disconnect Tracker Protection List” or bare “Disconnect” into a field, where the paper's own prose says something the list-sense pattern does not match. The seventh is a homonym the extractor invented — “physical ECU access and disconnection” in an automotive CAN-bus paper is not the Disconnect list. The tight pattern's cost is therefore about five papers, all Disconnect, all in the direction of undercounting, and the page's Disconnect row should be read as 74 rather than 74-to-79 only because the union is an upper bound on use anyway. The alternative — publishing the loose 700 — is not a trade worth making.
A.14 Denominators used on the page, stated once
| Figure on the page | Denominator | Never |
|---|---|---|
| the six per-resource rows | 5,859 for the “share of corpus” column; 1,120 for the “share of crawled” column, with the numerator restricted to that population | 5,859 for both |
| 136 / 2.3% | 5,859 | — |
| 107 / 9.6% | 1,120, numerator restricted to it | 136 ÷ 1,120 |
| per-year shares | that year's own paper count | 5,859 |
| per-venue shares | that venue's own paper count | 5,859, and never the count alone — the venues differ in size by a factor of three |
| 15 / 13 webXray papers | 5,859 and 1,120 respectively | — |
| everything about the three lists | the domain frames: 32,369 (2026-08-17) or 32,337 (2026-09-05) registrable domains | 5,859; these are not corpus figures at all |
B. Reproduction record, 2026-09-11
Every script whose figures the new page inherits was re-run before the move, against the same pinned inputs, and diffed against its committed output. A page whose report script no longer runs is a page whose numbers cannot be refreshed.
| Script | Invocation | Result |
|---|---|---|
report_webxray.mjs | node scripts/report_webxray.mjs –quotes | byte-identical to the committed report_webxray-output.txt before the fix in section C; regenerated after it |
owner_dbs.py | python3 scripts/owner_dbs.py –cache out/webxray/cache –disagreements 25 | reproduces the 2026-08-17 frame exactly: 32,369 domains; 669 / 5,581 / 2,268; 58.6% / 84.3% / 80.5%; 15,651 merged rows; 19 non-hostname rows dropped; 6,941 ICANN and 3,290 PRIVATE rules; 782 of 827 owner names untouched by the fold, 16 of the 45 changed losing a legal suffix |
owner_random_sample.py | python3 scripts/owner_random_sample.py –sample out/owner_sample.json –rows out/adj_rows_scored.json –table –wiki | byte-identical to out/owner_random_sample-output.txt |
owner_selfname_census.py | python3 scripts/owner_selfname_census.py | reproduces 0.8% / 0.7% / 3.2% (25 of 3,215; 250 of 38,368; 250 of 7,850) |
owner_adjudication.py | python3 scripts/owner_adjudication.py –table | reproduces the 28-row table exactly: 1/18/4/1/4, 8/12/6/2/0, 26/0/2/0/0, and the three rows scored as error rather than staleness |
owner_irr_kappa.py | invocation reconstructed by the figures reviewer, since it was missing from the brief | reproduces pooled kappa +0.400 (+0.241, +0.551), per-list +0.027 / +0.478 / +0.460, settled-only +0.680 |
report_ownership_resolution.mjs | node scripts/report_ownership_resolution.mjs | new; section K |
A drift check the carve-out made cheap. The new script re-derives the sweep counts from patterns written out independently and then cross-checks nine figures against report_webxray-output.txt. All nine agree. That cross-check is inside the script and throws, so a future run cannot publish two pages that disagree with each other:
Figure this script report_webxray.mjs verdict ----------------------- ----------- ------------------ ------- webXray 15 15 agree Tracker Radar 32 32 agree WhoTracks.me 23 23 agree Crunchbase 23 23 agree Disconnect (list sense) 74 74 agree Public Suffix List 101 101 agree union of the five 136 136 agree union ∩ crawled 107 107 agree crawled denominator 1120 1120 agree
C. A defect the cross-check found: a numerator from one population, a denominator from another
This is the one substantive correction of the sitting, and it was published for three and a half weeks.
report_webxray.mjs printed, at two places:
${pct(ftPapers.length, crawled.length)} of the ${crawled.length} that ran a crawl
${pct(anyOwner.length, crawled.length)} of the ${crawled.length} that crawled
ftPapers and anyOwner are sweeps over the whole corpus. Dividing either by crawled.length pairs a corpus-wide numerator with the crawling denominator: the share is of a population the numerator was not drawn from.
| Published | Correct | Why the error is invisible | |
|---|---|---|---|
| webXray papers, share of the 1,120 | 1.3% (15 ÷ 1,120) | 1.2% (13 ÷ 1,120) | 13 of the 15 are in fact crawling papers, so the two figures differ by one decimal and nothing looks wrong |
| any ownership resource, share of the 1,120 | 12.1% (136 ÷ 1,120) | 9.6% (107 ÷ 1,120) | a 2.5-point error on the more load-bearing figure |
Both figures were on the live webXray page, and the 12.1% was one of the figures moving to the new page — which is how it was caught: the new script computed the intersection because it had no reason not to, and the cross-check refused to pass. No guard on the old page could have found it. The number-guard family checks that a figure on a page matches the script's output; the script's output was the wrong number, and the page matched it faithfully.
The fix adds a named helper so the mistake cannot recur silently:
const crawledKeys = new Set(crawled.map(key)); // A share of `crawled` needs a numerator drawn from `crawled`. const inCrawled = (ks) => ks.filter((k) => crawledKeys.has(k));
report_webxray-output.txt was regenerated and both pages corrected in the same sitting. The webXray page carries a footnote at the changed figure rather than silently restating it.
D. Folding, and what it could not reach
This page's corpus figures need no name fold: every count is a regex over full text or a set membership, and there are no free-text names being aggregated. That is stated rather than assumed, because it is unusual for a page on this wiki.
The live-list figures do fold, and the fold and its residue are published on webxray §B.11. Restated here only in summary, because the content page quotes both halves:
- Legal-form suffixes only (
Inc,LLC,GmbH,S.A.S, …), plus punctuation normalisation. Deliberately no synonyms, so the “agree” columns are a lower bound on real agreement and the “neither” column an upper bound on real disagreement. - Residue: the fold leaves 782 of webXray's 827 owner names untouched (94.6%). Of the 45 it changes, only 16 lose a legal-form suffix; the other 29 change on punctuation alone —
56.com,AT&T,Ask.com,Bootstrap_China,Clearstream.TV,Cm_browser,Dictionary.com,Dun & Bradstreet, and the rest are in the published output. - What no fold reaches: “Google” against “Alphabet”, “Xandr” against “Microsoft Corporation”. Those are the granularity and vintage axes, and they are the page's subject rather than a normalisation failure. There is no principled synonym table for them, which is why the page reports the agreement figures as bounds and adjudicates a sample instead.
E. Quotes checked against the source
Checked on 2026-09-11 against data/fulltext/<year>/<venue>/<slug>/paper.cols.txt, whitespace collapsed and quotation marks and dashes normalised, with paper.norm.txt and paper.txt as fallbacks.
| Paper | Quote (head) | Verdict |
|---|---|---|
| [2Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)] | “the product of years of detective work” | found in .cols |
| [2Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)] | “because webxray's database of domain ownership primarily contains major ad networks…” | NOT contiguous — see below |
| [3Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] | “frequently miss connections among two hostnames” | found |
| [3Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] | “twitchcdn.net” | found |
| [4Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] | “preferring those that have been updated most recently” | found |
| [4Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] | “does not aim at providing full transparency on the organizations behind each domain” | found |
| [1Matte, Célestin; Bielova, Nataliia; Santos, Cristiana Teixeira (2020): "Do Cookie Banners Respect my Choice? Measuring Legal Compliance of Banners from IAB Europe's Transparency and Consent Framework", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] | “WebXRay commit 04c3c8e8” | found |
| [5Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)] | “Redacted For Privacy” / “Domains By Proxy” / “Whois Guard” | found, and the surrounding sentence re-read — see below |
The Libert coverage quote is real but is not one contiguous string in any rendering of the PDF. It is the page's longest quotation and it is load-bearing for the coverage section, so it was chased rather than dropped. In paper.cols.txt, paper.norm.txt and paper.txt alike, the two-column layout interleaves a table caption through the sentence:
...because webxray's database of Table 1: Third-Party Prevalence, SSL Use, and First-Party domain ownership primarily contains major ad networks rather Disclosure than small clients, and policyxray only searches for identified †Denotes Company has Consumer Services parties, variability in the long-tail of trackers may not have an outsized effect on overall findings related to disclosure. Nonetheless, Company % Tracked % SSL % Disclosed it is important to point out that the number of parties being searched Google † 82.81 80.35 38.29 for is fewer than the total number of parties present.
Every clause is verbatim and in this order; the interleaved strings are “Table 1: Third-Party Prevalence, SSL Use, and First-Party Disclosure”, “†Denotes Company has Consumer Services” and the table's own header and first row. The page quotes the sentence as the author wrote it and carries a footnote saying exactly this, so a reader running the same grep does not conclude the quote is invented. This is the class of defect a plain substring check reports as a miss and a careless run reports as a fabrication.
One claim was corrected by reading the source. The old page said [5Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)] “resolved organisations for 411 of 762 mobile-specific trackers using CrunchBase, webXray's list, TLS certificates and WHOIS in that order”. The paper's own sentence is narrower:
we successfully retrieved the organization information of 411 trackers from either the CrunchBase database, Tim Libert's library, or TLS certificate, and retrieved the organization […]
The 411 comes from the first three sources; WHOIS contributed a further 251 (“among those 251 trackers whose organization information was retrived from the WHOIS record”). The four sources were tried in that order, but the 411 is not a figure for all four. The new page states it as 411 from the first three and 251 from WHOIS. The consequence for the page's actual point is unchanged: the five registrar-privacy strings total 88, against the largest real company in the top-ten table at 48 (Adobe), and the paper says so itself — “these five organizations cannot represent the real organizations of those trackers”.
F. External sources: verified, and rejected
Repository and vendor state was re-fetched on 2026-09-11, not recalled. Training data is stale by construction for this kind of claim.
| Claim on the page | Primary source, fetched 2026-09-11 | Verdict |
|---|---|---|
Tracker Radar current; last commit on main 2026-08-28; newest release tag 2026.08.28 | api.github.com/repos/duckduckgo/tracker-radar/commits?per_page=1 → a736b501, 2026-08-28T15:25:14Z; /releases → 2026.08.28, 2026.07.27, 2026.06.08, none flagged prerelease | holds |
Tracker Radar's pushed_at reads later than main because of unmerged branches | pushed_at = 2026-09-02T21:21:31Z against main at 2026-08-28 | holds, and the page now names both dates |
| Disconnect current; last commit 2026-08-28 | commits?per_page=1 → 4b592c28, 2026-09-05T20:21:14Z | moved. Page updated to 2026-09-05 and the table header re-dated to 2026-09-11 |
Ghostery trackerdb current; last commit 2026-09-01 | commits?per_page=1 → aa82e6d9, 2026-09-01T14:40:47Z | holds |
| Disconnect's headline “14,332 verified domains and entity mappings” | disconnect.me/trackerprotection, HTTP 200, string present | holds |
github.com/timlib/webXray is HTTP 404; the account has 0 public repositories | 404; api.github.com/users/timlib → 200, public_repos: 0, company: webXray.ai | holds |
github.com/timlib/webXray_Domain_Owner_List is HTTP 404 | 404 | holds |
| No PyPI package | pypi.org/pypi/webxray/json → 404 | holds |
thezedwards/webXray last pushed 2021-03-04, PolyForm Strict | pushed_at 2021-03-04T23:52:47Z; GitHub reports NOASSERTION | holds |
peterjoles/webXray is MIT, last pushed 2023-03-12, 36 commits ahead | spdx_id: MIT, pushed_at 2023-03-12T20:15:26Z | holds |
webxray.org is a placeholder reading “Public interest projects for the interested public.” | fetched with a browser User-Agent: HTTP 200, that exact line and nothing else | holds |
webxray.ai is a live commercial product | HTTP 200 | holds |
Sources rejected, recorded so the next run does not re-add them:
| Rejected | Why |
|---|---|
GitHub's license.spdx_id field for Ghostery trackerdb and Disconnect | reports NOASSERTION for both, which is a statement about GitHub's detector and not about the licence. The licences were read from package.json and LICENSE respectively. |
GitHub's pushed_at as “last activity” | it counts unmerged automation branches. For Tracker Radar it reads four days later than main. Read the branch. |
Any figure for Ghostery trackerdb coverage or accuracy | none was measured. Its .eno + patterns shape needs a parser nobody could check against the other three. The page says this is an omission rather than a judgement, and lists it as an open question. |
| Vendor marketing counts as file sizes | Disconnect's “14,332 verified domains” is a site headline; entities.json holds 7,850 domains. The page quotes both and says which is which. |
| Anything recalled rather than fetched about these repositories | training data predates all four dates above. |
G. What could not be established
- Ghostery
trackerdbis unmeasured. It is a live fourth option and the page gives numbers for two of the three live lists. Closing this needs a parser for the.eno+patternsmodel whose choices are checkable against the other three — a sitting's work, not a footnote. Filed as an open question on the content page. - Recall of the candidate set is unknown. The 136 is a union of six patterns over full text. A paper that resolves domain ownership from a source none of the five names — a bespoke WHOIS pipeline, a commercial feed, an internal list — matches nothing at any width. The page reports 136 as an upper bound on use and says nothing about recall, because nothing here measures it.
- The share of the 136 that actually resolved an owner is not measured. Only two rows (webXray's 15 and Tracker Radar's 32) carry a hand verdict per paper. Doing the same for Disconnect's 74 is the obvious next increment and was not attempted this sitting.
- ~~No accuracy figure exists for the tail.~~ Partly closed on 2026-09-11, and the answer mattered. This read: “Every rate is conditional on the entry being adjudicable, and 18–30% of each list's draws were not. Whether unadjudicable domains differ systematically between lists is exactly what would decide whether Disconnect's lead survives, and a random sample cannot answer it.” A second pass with harder sourcing settled 28 of the 40, leaving 6.7% / 6.7% / 10.0%. The formal test found no significant difference between recovered and already-eligible rows (p = 0.64 / 1.00 / 0.40, on 7, 9 and 12 rows — almost no power), but Disconnect's estimate fell 94.4% → 89.2% and its lead over webXray stopped surviving Bonferroni. So the answer is: the tail was somewhat worse for the list the worry named, and the worry was worth acting on. See section L. What remains open is the residue's own residue — 12 domains that no register, no registry record and no archived legal page can reach, and nothing bounds those from inside the sample either.
- The coverage census has no neutral denominator. Tracker Radar's own crawl output is the frame, which flatters Tracker Radar. There is no independent census of third-party domains to use instead, and the page says so at every coverage figure rather than pretending otherwise.
- Whether a hierarchy built from a current list would beat both. webXray's
parent_idtree is the only hierarchy any of these files carries and it is frozen; nobody has tried grafting a current corporate-structure source onto Tracker Radar's entities. Filed as an open question. - Two adjudication rows from 2026-08-17 remain unresolved (the acquisition date behind
360yield.com, andfwmrm.net's post-2026-spinoff status). Recorded as unresolved rather than guessed; they are excluded from the 28.
H. Corrections owed to Programming:Crawler, and their real state
The task brief for this sitting said Crawler's comparison row “still says webXray's crawler is PhantomJS (historically)”. It does not, and has not since 2026-08-17. The brief's account of its own state was stale; recorded here because a later sitting reading the brief would look for a correction that is already made.
| Correction | State on 2026-09-11, before this sitting |
|---|---|
| The comparison row's automation column | already fixed. Revision 1786953266 reads “PhantomJS in 2015; consumer Chrome over raw CDP in the last public version” |
| “stale third-party mirrors, the newest last pushed in 2015 and targets PhantomJS” | already fixed in the same revision |
The <WRAP todo> asking where webXray is developed | already closed in the same revision. The page's remaining <WRAP todo> is three unrelated open questions and was correctly left alone |
| The licence and most-complete-copy claim | already fixed on 2026-09-05, revision 1788636934 — the item's own note says so |
The one correction that was owed, and that this sitting created, is the carve-out's own drift: that bullet ended “see webXray for what survives of it: the ownership database, and how it compares to Tracker Radar and Disconnect today”, and the comparison had just left that page. Repointed to Ownership resolution in revision listed below. Nothing else on Crawler was touched: its webXray figure of 7 is the tools[] schema count for a table whose whole column is schema counts, and changing it to the sweep's 15 would make that one cell incomparable with its neighbours.
I. Reviewers
All four were told explicitly that the author's context may not be exhaustive, and all four were handed the page text, the report script, its unedited output and these notes. The three focused passes ran in parallel; the generic pass ran afterwards, against the pages as the first three had left them.
All four passes returned and every finding was applied. The generic pass — the one with no checklist — found more, and worse, than the three focused passes combined. Each was handed the page text, scripts/report_ownership_resolution.mjs, its unedited output, scripts/report_webxray-output.txt and these notes, against the bundle in review_own/. Every finding below was re-verified by the author before being accepted — a reviewer's report is a lead, not a result.
| Pass | Model | Returned | Findings | Accepted | Rejected |
|---|---|---|---|---|---|
| external currency | Sonnet | yes | 1 substantive of ~40 checks | 1 | 0 |
| citations and quotes | Sonnet | yes | 1 substantive of 22 citekeys, 11 quotes, 13 attributed claims, 7 vendor sources and all 15 rows of the webXray table | 1 | 0 |
| figures versus script | Sonnet | yes | 2 substantive, both inside a published script rather than on a page | 2 | 0 |
| generic, no checklist | Fable | yes | 13 substantive, including the two most serious of the sitting | 13 | 2 partial |
Accepted: the LLM row was an overclaim (external currency)
The Choosing a resolution source now table read “emerging, and not yet at the domain layer”. The reviewer found arXiv:2606.20868, Can LLMs Reason About Brand Ownership?, submitted 2026-06-18 — four models evaluated on domain-to-brand attribution over 36 heavily-phished brands.
Verified independently before accepting: the abstract was fetched (HTTP 200) and read. The paper is real, the date is right, and the task genuinely is domain-layer ownership attribution. Two qualifications the reviewer did not make, and which changed how it went on the page:
- Its question is phishing and squatting defence — “is this domain the brand's own?” — not third-party tracker attribution. 36 brands is not a coverage claim.
- Its result cuts the other way. Models enumerate a brand's domains at up to 82% precision from memory, but on ownership verification without external tools macro F1 is at most 0.37, rising by up to 0.65 with WHOIS. So the honest update is not “LLMs are arriving, consider them” but “the one published attempt at this task fails without retrieval” — which strengthens the row's existing advice rather than softening it.
The corpus-scoped half of the sentence (“no corpus paper applies this to third-party domain ownership”) was not changed: it was and is true, and an arXiv preprint is outside the seven venues. Cited as a footnote with its arXiv id rather than a {[key]}, so no bibliography entry and no cache purge.
Accepted: an undocumented column splice, of the opposite kind (citations and quotes)
The page footnotes the [2Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)] coverage sentence as non-contiguous in every rendering. The reviewer found that [6Selmo, Carlos; Carisimo, Esteban; Bustamante, Fabián E.; Alvarez-Hamelin, J. Ignacio (2025): "Learning AS-to-Organization Mappings with Borges", in: Proceedings of the 2025 ACM Internet Measurement Conference, pp. 120-133. (DOI)]'s “an Internet shaped by constant mergers, rebrandings, and regional variation” has the inverse problem, unflagged.
Verified independently: the string constant mergers does not appear in paper.cols.txt at all — nor in paper.norm.txt or paper.txt — while pypdf over paper.pdf returns the sentence fully contiguous. So the quote is correct and the corpus's own rendering is the thing that is wrong. Footnoted to say so, because a reader checking the quote against the corpus would otherwise conclude it was invented.
This is the two-renderings rule in both directions on one page: Libert fails in every rendering and is real; Borges fails in the column-repaired rendering and is real. A quote-check against a single rendering would have reported one false FAIL and, had the Libert footnote not already existed, one apparent fabrication.
Accepted: a published script's own ALLOW list carried the fold it warns against (figures versus script)
This is the most valuable finding of the sitting, and section B of this page had just asserted the opposite.
report_webxray.mjs ends with a section Z, Every number on the page that does NOT come from this corpus, which exists so check_page_numbers.mjs can tell a non-corpus figure from a corpus one. Five of its lines were hardcoded constants, and they were the ICANN+PRIVATE fold with exact-key lookup — the variant owner_dbs.py's own output labels <- THE TRAP and both content pages warn against by name, because it manufactures a coverage hole out of googleapis.com:
| Section Z said | The published fold gives |
|---|---|
| 45,525 registrable domains | 32,369 |
| webXray 657 (1.4%), TR 5,539 (12.2%), Disconnect 2,277 (5.0%) | 669 (2.1%), 5,581 (17.2%), 2,268 (7.0%) |
| weighted 54.5% / 79.3% / 75.8% | 58.6% / 84.3% / 80.5% |
| top 100: 70% / 95% / 91% | 71% / 98% / 94% |
| 197 of 601, 216 of 461, 584 of 1,532 | 198 of 612, 217 of 464, 584 of 1,538 |
| “16,396 of 47,836 rows are hostnames” | the label-count heuristic the page says “answers neither question; do not use it” — replaced with the ICANN fold's 15,651 |
Verified independently before accepting: each figure was matched against the live owner_dbs.py run in section B, and grepped for on both content pages. None appears as prose on a live page, so nothing a reader saw was wrong — but the ALLOW list of a number guard is exactly where a wrong figure does its damage silently, by blessing the wrong value if it ever reaches a page. Corrected, and the corrected lines carry a comment naming the trap.
What this says about section B of this page. B calls report_webxray.mjs's output “byte-identical” to its committed copy, and it was — before and after. Byte-identity proves a script matches its own last run. It proves nothing about whether the content is still true, and here the script had been faithfully reproducing a stale constant. A reproduction record is not a correctness record, and this page should not have implied otherwise.
Accepted: a guard that loses its own warning (figures versus script)
report_webxray.mjs buffered every line and flushed once at the bottom. Its section-A guard —
if (roleResidue.length || roleExtra.length) { P('FAILURE: the hand map and the sweep disagree. Re-read the new papers before publishing any figure below.'); }
— called P() and let the script continue. Section B then dereferences ROLE.get(k).role on the now-missing key and throws, before the single console.log at the bottom has run.
Reproduced by deleting one ROLE entry: stdout 0 bytes, a bare TypeError stack, and the FAILURE line that exists for precisely that case nowhere at all. The guard fires and is then thrown away. Fixed with a fail() helper that flushes the buffer, repeats the message on stderr and exits 1; re-running the same mutation now prints 1,283 bytes ending in the named paper and the FAILURE line, and exits 1. On a healthy run the output is byte-identical to before, so the fix is behaviour-neutral where it should be.
What the reviewers confirmed, which is also a result
- All 22 distinct citekeys across the three pages resolve to exactly one bibliography entry; no collisions anywhere in the 993-entry file.
- All 11 quoted strings and 13 attributed non-quote claims verified verbatim or in substance, including the self-contradictory [3Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] sentence this page declines to derive a percentage from, and the [5Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)] 411/251 split this sitting corrected.
- All 15 rows of the webXray publication table verified against source text — the reviewer was asked for five and did fifteen.
- Every figure on both content pages independently re-derived from a live re-run of all six scripts, including the inter-rater kappa, whose invocation the reviewer had to reconstruct because it was not in the brief. Both published code snippets (
owner_lookup.pyandresolve()) executed against the cached JSON and their output matched the page character for character. - Three further mutations of
report_ownership_resolution.mjsby the reviewer, including a re-creation of this sitting's own denominator bug — all caught, none vacuous. - Every repository, licence, URL, version, dependency and acquisition claim on all three pages re-fetched and holding to the day, including the 2020-06-24 Disconnect relicensing commits, PolyForm Strict still having no SPDX id, and
lxml/psycopg2-binarywheels still stopping at CPython 3.9 against a 2025-10-31 Python 3.9 EOL.
Accepted, and the worst finding of the sitting: the LLM-adjudicator disclosure was dropped (generic)
webxray's 2026-09-05 reviewer log records, verbatim: “The pages never say the adjudicators were language models … ACCEPTED, blocking. The content page now says 'by Sonnet sub-agents working to a published brief, single-rated'.” The carve-out dropped that sentence, and the new page said only “adjudicated one by one against primary sources — company newsrooms, SEC filings, or the domain's own legal documents”.
Verified before accepting: the clause is in the pre-split page text (live_programming_crawler_webxray.txt), and a grep of the new page for Sonnet, sub-agent and language model returned zero hits. A reader deciding whether to cite 94.4% was being told, by omission, that a person adjudicated it.
Restored in two places — beside the 28-row table and, for the first time, beside the random-sample table, which never carried it even before the split. No reviewer in three sittings had noticed the second gap. A finding accepted as blocking can be undone by a later structural edit, and no guard on this wiki looks for that.
Accepted: a p-value asserted without being computed (generic)
The page said: “Disconnect's lead over webXray is a difference this sample can establish; its lead over Tracker Radar is not, and neither is Tracker Radar's over webXray.” Two p-values were published (0.014, 0.82). The Disconnect-versus-Tracker-Radar pair was never computed — the webXray provenance log shows the original reviewer supplied only the other two, and the sentence generalised from their absence.
Verified: Fisher's exact re-implemented from scratch over the eligible raw counts reproduces the two published values to three decimals (0.0137, 0.820) and gives p = 0.0248 for the untested pair. The sentence was false by the page's own test, directly under the short answer's headline comparison.
Rewritten as a three-row table with all three p-values and a Bonferroni threshold of 0.0167 for three tests on one sample — under which 0.014 survives and 0.025 does not, so the recommendation stands but for a stated reason rather than an invented one. The page now says in its own text that an earlier version asserted this without computing it.
This is the failure mode to remember: an absence of evidence in a reviewer's report was read as evidence of absence. Nothing in the three focused briefs would have caught it — the figures pass checks the page against the script, and this number was in neither.
Accepted: the estimator is post-stratified, and the page described it as unweighted (generic)
“domain-level — … Every entry weighs the same.” It does not: owner_random_sample.py computes R = Σ_h (N_h/N)·p_h over the 24/12/12/12 quartile allocation. That is why 94.4% ≠ 39/42 = 92.9% and 69.9% ≠ 35/49 = 71.4% — and why the Fisher paragraph's raw counts did not match the table above it, which a reader dividing them would have read as an error on the page. Corrected, with the raw counts now stated beside the weighted ones.
Accepted: three arithmetic slips the author's own "arithmetic audit" missed (generic)
| Said | Is |
|---|---|
| “A third of the sample could not be settled” | 42 of 180 = 23%; the per-list shares in the same bullet are 18 / 22 / 30% |
| “the intervals are ±13 percentage points at best” | “at best” is Disconnect's ±6.6; ±13.8 is the worst. Now “±7 to ±14” |
| “Two of the three lists change weekly” | the page's own table says Tracker Radar regenerates monthly |
The author's audit table below claims “all consistent; one error found”. That was wrong, and it is left standing above with this correction beside it rather than quietly rewritten — an audit that reports a clean result it did not earn is worse than no audit.
Accepted: overstatements and structure (generic)
| Finding | Action |
|---|---|
| The short answer headlines the two metrics under which the recommendation looks best, while the body argues two other metrics matter more (encounter-weighted accuracy, where the frozen list tops the table; weighted coverage, where the live lists near-tie) | short answer now states this in three sentences and says why the recommendation survives it |
| “read the encounter-weighted column as an ordering rather than as a number” — all three intervals overlap and no pairwise test is run | now says it supports neither, and why it is nonetheless the quantity that matters |
| webXray's inter-rater kappa is +0.027 on 12 rows — chance — and the page stated it without drawing the implication | one sentence added: its 69.9% carries rater uncertainty on top of the sampling interval |
| “rose to about 3.5% around 2020–2022 and has sat between 2.5% and 2.7% since” is a story about noise | verified: 2022 vs 2023 p = 0.41, pooled 2020–22 vs 2023–26 p = 0.11. Headline “It is not growing” kept, curve dropped |
| The “29 of 525 verdicts changed, direction overwhelmingly favourable” discount sat two screens below the table it discounts | promoted to a <WRAP important> directly under the accuracy table |
| “most papers should merge them and say in which order” — the page never said which order | now states it: Disconnect first for the name, Tracker Radar as fallback, record which key matched |
| Wiki-internal QA on the content page (the schema-versus-sweep paragraph, the probe-width paragraph) | trimmed to pointers here; the two quote-splice footnotes were kept, because they defend a reader who greps the corpus and finds nothing |
| Three sentences on webXray describing the pre-split page | all three fixed |
| No reading list for the stated reader | three papers, in order, added to the short answer |
Partially rejected (generic)
| Finding | Disposition |
|---|---|
| Remove the Libert and Selmo quote-splice footnotes as wiki-internal QA | Rejected. A reader who checks either quote against this corpus finds nothing and concludes it was invented. The footnotes exist for them, not for us. The other items in that finding were accepted. |
| “98 em-dashes, 26 'not X' contrasts, nearly every section ends on an aphorism — someone who did the work would let a few sections end on the number” | Accepted as fair, not acted on this sitting. It is a real observation about the register and it would take a prose pass, not an edit. Recorded so it is not rediscovered. |
Noted, not acted on
| Observation | Disposition |
|---|---|
| Today's live PSL has 6,950 ICANN / 3,375 PRIVATE rules against the 6,941 / 3,290 in the 2026-08-17 snapshot | Not a defect. The figure is snapshot-pinned and the page says the PSL moves. Re-deriving it would orphan the adjudication — the same reason the whole 2026-08-17 frame is frozen. |
The reviewer's Wayback CDX pull put timlib/webXray's last archived 200 at 2022-12-30 where the page says 2023-03-31 | Left as published. The reviewer attributed the gap to CDX collapse-window settings and did not claim the page was wrong. It is a bound either way and both bounds support the same conclusion. Flagged here so a future sitting can settle it rather than rediscover it. |
libert2015_invisible and karaj2018whotracksme could not be quote-checked — neither paper is in the corpus | Correct and expected. Neither is quoted; both are cited only as identifiers. |
| The figures reviewer could not verify section H's claim that Crawler was already corrected, having no access to that page's history | Fair, and the claim stands. Revisions 1786953266 and 1788636934 were read from node scripts/dw.mjs history programming:crawler and the live page text quoted in section H. |
| It could not verify who produced “rater 2” in the inter-rater work, taking the label from the script's own output | A real limit of that measurement, and it is not this page's to fix. Recorded on random_sample where the draw was made. |
| The Sánchez-Rola 2021-versus-2022 venue-year discrepancy | Left as published, with the page's existing italic note that the corpus files it under 2022 while Crossref and the paper's header say 2021. |
Still owed
- A finding accepted as blocking was silently reversed by a later structural edit, and nothing detected it for six days except a reviewer with no checklist. There is no guard on this wiki that re-checks accepted findings against a rewritten page. The cheapest version would be a list of blocking findings per page, with the sentence each one put there, checked as a string.
- The prose register was flagged and not fixed: 98 em-dashes in 8,700 words, and most sections closing on an aphorism rather than on the number. A prose pass, not an edit.
report_webxray.mjs's section Z is a hand-maintained ALLOW list of constants and that is what let it go stale. It should be derived fromowner_dbs.py's output rather than retyped, or checked against it by a guard that throws. Not done this sitting: it needs the two scripts to share a format, which is a change to both.- The same buffered-output defect may exist in other
report_*.mjsscripts on this wiki. Onlyreport_webxray.mjsandreport_ownership_resolution.mjswere checked.
What the author checked without them, so the gap is bounded rather than open:
| Check | Result |
|---|---|
| All published scripts re-run against pinned inputs | they reproduce their committed output — which, as C and I both show, is not evidence the output is right: byte-identity blessed a wrong denominator for three and a half weeks and five wrong constants for longer. A reproduction record is not a correctness record. |
owner_adjudication.py –table | reproduces the 28-row table exactly: 1/18/4/1/4, 8/12/6/2/0, 26/0/2/0/0, and the three error rows (1rx.io, jsdelivr.net, stackadapt.com) |
domain_owners.json schema table on webXray, every cell | reproduces: 319/827 parent_id, depths 508/211/80/24/2/2, 761 uses, 803 platforms, 826 country, 104 trade_groups, 591 policy URLs, 132 GDPR, 4 CCPA, 13 opt-out, 34 crunchbase_id, 52 health-segment, 263 notes, 275 aliases, 69 languages, 2,175 of 3,215 |
Mutation test of report_ownership_resolution.mjs | four mutations, four failures. Swapping tight/loose on Disconnect → “tight (700) > loose (74)”. Making loose disjoint from tight at equal size → “23 tight hits are not in the loose set” — the containment guard catches what a size comparison alone would miss. Injecting a fake union member → cross-check throws. Replacing the crawled intersection with the full union → cross-check throws. No guard is vacuous. |
| Arithmetic audit of every ratio on the content page | all consistent; one error found — “seven times NDSS's rate” is 8.4×, corrected |
| Rendered DOM of both content pages | 18 references for 18 keys on the new page, 17 for 17 on webXray; 12 and 7 tables; 9 and 4 WRAP boxes; every in-page and cross-page anchor resolves; one broken anchor found — the webXray availability table pointed at #Methodology and limitations of these figures for a Wayback timeline that had become its own section |
| Site-wide inbound-link sweep | four pages found and fixed (section above) |
| Table cells containing a pipe | one found in this page's own draft: a nowiki-wrapped [[a|b]] literal in a table cell. A parsed link's pipe is safe in a cell (verified in the rendered DOM); a nowiki-wrapped one is not, because the cell splits before the nowiki is honoured. Rewritten as prose, and the A.12 regex table was moved into a code block for the same reason. |
| Full-text corpus quality | 5,855 of 5,859 renderings present, 0 empty, 90 carrying a NUL byte — see A1b |
L. The second pass over the residue, 2026-09-11
The 2026-09-05 sample drew 180 times over 175 distinct domains and could not adjudicate 40 of them (22.9%), 42 draws. Every accuracy rate it produced was therefore conditional on the entry being adjudicable, the exclusion rate rose toward the prevalence tail, and it was worst — 30% — for Disconnect, the list the content page recommends. Nothing inside that sample could test whether the excluded rows were systematically worse. This pass re-adjudicated exactly those 40 rows, changing nothing else, against three routes the first pass did not take.
L.1 What was done, and in what order
| Step | Command | What it produced |
|---|---|---|
| Build the worksheet | python3 - over out/owner_sample.json + out/adj_rows_scored.json | out/tail40.json — the 40 domains with each list's claim, the residue membership per list, and the 2026-09-05 note saying why that pass gave up |
| Mechanical evidence | bash scripts/owner_tail_probe.sh out/tail40.json out/tail_probe | out/tail_probe.txt (2,471 lines, 2m 29s, 6-way parallel) |
| Adjudication | 5 Sonnet sub-agents × 8 domains, round-robin over the prevalence order so every batch spans the range | out/tail_adj/result1..5.json |
| Gate | python3 scripts/owner_tail_merge.py –sample out/owner_sample.json –base out/adj_rows_scored.json –tail out/tail40.json –dir out/tail_adj –out out/adj_rows_tail.json | out/adj_rows_tail.json, 0 problems |
| Citation check | python3 scripts/owner_verify_sources.py –rows out/tail_rows_only.json –cache out/tail_verify_cache –out out/tail_verify.json | 3 defects in the checker itself — see L.5 |
| Hand corrections | python3 scripts/owner_tail_corrections.py –rows out/adj_rows_tail.json –out out/adj_rows_tail_final.json | 6 rows: resource ×3, keep ×2, rescore ×1 |
| Re-estimate | python3 scripts/owner_random_sample.py –sample out/owner_sample.json –rows out/adj_rows_tail_final.json –table –wiki | out/owner_random_sample-tail-output.txt |
| What changed | python3 scripts/report_tail_pass.py … | out/report_tail_pass-output.txt |
The 2026-09-05 output file is kept, not overwritten. out/owner_random_sample-output.txt is the before; out/owner_random_sample-tail-output.txt is what the page's figures are now checked against, and out/owner_random_sample-tail.diff is the 338-line diff between them.
Reproducing the before is a trap worth recording. out/adj_rows.json and out/adj_rows_final.json both still carry self-named verdicts, and owner_random_sample.py exits fatally on them. The committed 2026-09-05 output is reproduced byte-identically only by –rows out/adj_rows_scored.json –wiki — the –wiki matters, because sections F and G exist only under it.
L.2 The three routes, and which one paid
| Route | What it is | Rows it settled |
|---|---|---|
register-by-number | a national company register searched by company/VAT/register number rather than by name — the first pass failed on strings like “Admiral”, “Collective” and “Globo”, which are unsearchable | 4 |
rdap-history | the registrar's own port-43 WHOIS, not RDAP. rdap.org returns a thin registry record with no registrant for .com/.net; the registrar's server returns it un-redacted. Also the registration and last-changed events, which bound a claim even when the organisation is redacted | 9 |
archived-legal | a Wayback capture of the domain's own legal page, fetched with the id_ suffix so the Archive's banner is not in the bytes | 2 |
live-source | a live document the first pass did not reach — usually because it looked for a page naming the domain and this pass looked for the entity, or because the page is on a subdomain | 13 |
The registrar hop is what did most of the work, and it is the cheapest thing
to add to any future pass: it is one extra socket per domain. It produced
Applied Technologies Internet SAS for at-o.net, Conversant, Inc. for the
three ValueClick-style sync domains, Acoustic, L.P. for pages02.net,
Leven Labs, Inc. for both Admiral domains and Google LLC for
blogblog.com — none of which RDAP shows. This container has no whois(1)
and no apt, so scripts/whois43.py speaks the protocol directly: IANA
referral, TLD server, then the registrar's server when the registry is thin.
L.3 The sourcing bar was widened, and by how much
Three source kinds are new in this pass and every row records which it used:
| Source kind | New? | Rows | What it is worth |
|---|---|---|---|
domain-register | yes | 9 | the registry's record of who holds the name, contractually required to be accurate but asserted by the registrant and validated by nobody. Weaker than a company register |
legal-doc | no | 8 | the domain's own live privacy policy, terms, imprint or legal-entity footer — the 2026-09-05 bar |
parent-site | no | 4 | the acquirer's or parent's own site naming the brand or domain as theirs — the 2026-09-05 bar |
register | no | 3 | a national company register record, or an OV/EV certificate's validated O = field — the 2026-09-05 bar |
archived-legal-doc | yes | 2 | the domain's own legal document read at a date it existed. Establishes the operator at the capture date, so it needs registry continuity or a live source about the entity to become a claim about today |
filing | no | 1 | an SEC or other regulator's filing — the 2026-09-05 bar |
newsroom | no | 1 | the company's own press release — the 2026-09-05 bar |
So 17 of the 28 recovered rows clear the original bar and 11 rest on a
kind the first pass did not use. tls-san — another organisation's
certificate on this host — was defined in the brief as corroboration only and
is the source of record for 0 rows.
This pass produced its own counter-example to its weakest source kind, and
it is kept rather than smoothed away. i.ua's registry record
(whois://whois.ua/i.ua) names Digital Ventures LLC. The portal's own
user agreement (help.i.ua/agreement/) names a different company as its
administration: ТОВ «КЕПРЕЙТ ПАРТНЕРС», Ukrainian register code 33500955.
On that domain the registrant of record is not the operator. The row was
re-sourced to the legal document, and the finding is why
report_tail_pass.py carries section E3: a sensitivity run that drops every
recovered row resting on an uncorroborated registrant organisation. Two
remain, and dropping both moves webXray +0.0, Tracker Radar −0.6 and
Disconnect +0.0 percentage points.
L.4 Every recovered row, with its route and its source
| Domain | In the residue of | Owner settled on | Route | Source kind | Citation re-fetched? |
|---|---|---|---|---|---|
spot.im | Tracker Radar | Open Web Technologies Ltd. (OpenWeb) | live-source | parent-site | OK |
company-target.com | webXray | Demandbase, Inc. | live-source | parent-site | OK |
adgrx.com | webXray | AdGear Technologies, Inc. (“AdGear”, trading as Samsung Ads; wholly-ow | live-source | legal-doc | FETCHFAIL |
app-us1.com | Tracker Radar | ActiveCampaign, LLC (ActiveCampaign) | live-source | domain-register | OK |
govx.com | Disconnect | GovX, Inc. | register-by-number | filing | OK |
opti-digital.com | Disconnect | Opti Digital SAS (2 Rue des Cortalets, 66400 Céret, France) | live-source | legal-doc | OK |
ksearchnet.com | Disconnect | Klevu Oy (operating subsidiary of Athos Commerce) | live-source | register | OK |
kameleoon.io | Disconnect | Kameleoon SAS (Kameleoon) | register-by-number | legal-doc | OK |
gssprt.jp | Disconnect | Geniee, Inc. | rdap-history | domain-register | OK |
acint.net | Tracker Radar | Poshibalov Evgeny Vasilyevich, operating the self-titled project 'Acin | rdap-history | domain-register | OK |
yceml.net | webXray | Conversant, Inc. (operating brand Epsilon; ultimate parent Publicis Gr | rdap-history | domain-register | OK |
travelpayouts.com | Tracker Radar | Go Travel Un Limited (Hong Kong; trading as Travelpayouts) | archived-legal | archived-legal-doc | OK |
blogblog.com | webXray | Google LLC (Blogger) | rdap-history | domain-register | OK |
at-o.net | Disconnect | Applied Technologies Internet SAS (AT Internet), controlled by Piano S | register-by-number | register | OK |
cnevids.com | Tracker Radar | Condé Nast Entertainment (Advance Publications) | live-source | legal-doc | OK |
pages02.net | Disconnect | Acoustic, L.P. | rdap-history | domain-register | OK |
trustpilot.net | Tracker Radar | Trustpilot A/S (Pilestraede 58, 5th floor, DK-1112 Copenhagen K, Denma | live-source | legal-doc | OK |
awltovhc.com | webXray | Epsilon (d/b/a 'Epsilon PeopleCloud Digital Media Solutions', formerly | rdap-history | newsroom | OK |
lduhtrp.net | webXray | Conversant, Inc. (Epsilon / Publicis Groupe) | rdap-history | domain-register | OK |
wishabi.com | webXray | Flipp Corp. | live-source | legal-doc | OK |
sa-as.com | Disconnect | FoundryCo, Inc. | rdap-history | domain-register | OK |
cratecamera.com | Tracker Radar | Leven Labs, Inc. (DBA Admiral) | rdap-history | domain-register | OK |
globo.com | Disconnect | Globo Comunicação e Participações S.A. (CNPJ 27.865.757/0001-02) | live-source | legal-doc | OK |
offshoregeology.com | Disconnect | Admiral (Leven Labs, Inc.) | live-source | parent-site | OK |
km0trk.com | Disconnect | Good On You Pty Ltd (Good On You) | live-source | parent-site | OK |
cjponyparts.com | Tracker Radar | CJ Pony Parts, Inc. | archived-legal | archived-legal-doc | OK |
cedscdn.it | Tracker Radar | CED Digital & Servizi S.r.l. (Caltagirone Editore group) | register-by-number | register | OK |
i.ua | Disconnect | ТОВ «КЕПРЕЙТ ПАРТНЕРС» (LLC Keprait Partners), Ukrainian register code | live-source | legal-doc | OK |
The one non-OK row is adgrx.com: samsungads.ca serves an expired
certificate, so curl refuses it and the checker never sees a body. Fetched
by hand with –insecure: HTTP 200, 46,160 bytes, and the quoted row is
byte-verbatim in it. The checker was deliberately not given an
–insecure retry — an ownership citation whose host identity cannot be
validated should surface, not be swallowed.
L.5 Three defects the citation check had, found by these rows
owner_verify_sources.py was written for the 2026-09-05 rows and this pass
cited kinds of evidence it had never seen. Each defect would have reported a
correct citation as a fabrication, and each fix carries a control:
| Defect | Found on | Fix | Control |
|---|---|---|---|
No whois:// scheme, so every registry citation FETCHFAILed | 9 rows | a port-43 branch calling whois43.ask against the named server, not re-resolved through IANA, so the row is checked against the record it cites | the 9 rows now verify OK |
body.decode(“utf8”, “replace”) turns every byte of a windows-1251 page into U+FFFD, so a Cyrillic quote can never match | help.i.ua/agreement/ | a decode() that sniffs the charset the document declares, from the bytes rather than the header, because the cache holds the body alone | the quote is found; the same quote with one letter changed is not. The normaliser had already been fixed once for this class of bug, one layer later |
| A rate-limit stub cached as evidence, giving a permanent false NOTFOUND | sa-as.com, blogblog.com | refuse to return or cache a WHOIS answer under 400 bytes or one that does not name the queried domain; one retry 20 s later | both verify OK on the retry, and the guard was seen to fire before the retry was added |
Final tally over the 40 tail rows: FETCHFAIL=1, NOSOURCE=12, OK=27. The 12 NOSOURCE are the rows still unresolved, which have no citation by definition.
L.6 Hand corrections
Six rows a machine flagged and a human then re-checked. keep means the
citation holds and the checker or the transport was at fault; resource
means the claim holds and the row cited the wrong document for it;
rescore means the verdict was wrong on the brief's own rules.
| Domain | Action | Why |
|---|---|---|
app-us1.com | resource | The cited page is a 404 whose only evidence is a <link> tag, which the quote checker normalises to the empty string and reports as NOQUOTE – so the row's only citation was one the checker cannot see, on a page that does not exist. Re-issued the registrar WHOIS query live (whois.markmonitor.com, port 43): the registrant organisation is un-redacted and names ActiveCampaign, LLC. The verdict does not change; the evidence under it does, from a 404 page to the registry record. Rate effect: none. |
blogblog.com | resource | The row's source URL was prose – 'whois:blogblog.com (registrar WHOIS, whois.markmonitor.com)' – which no client can dereference, so it FETCHFAILed. Rewritten in the whois:<server>/<domain> form the checker parses and re-queried live: 'Registrant Organization: Google LLC'. Same record, same quote, a URL that resolves. Rate effect: none. |
ksearchnet.com | resource | The row cites an OV certificate's validated O= field – which the brief accepts as register-grade – but gave the https:// URL of the host, which now returns HTTP 404, so the checker fetched a 404 body and looked for a certificate subject in it. Repointed to the tls:// form the checker handles. Handshake redone 2026-09-11: subject=C=FI, ST=Uusimaa, O=Klevu Oy, CN=*.ksearchnet.com, issuer Sectigo Public Server Authentication CA OV R36. Rate effect: none. |
i.ua | rescore | Two things were wrong. (1) The row rested on the WHOIS registrant, 'Digital Ventures LLC'. The portal's own user agreement names a different company as its administration – ТОВ «КЕПРЕЙТ ПАРТНЕРС», register code 33500955 – so on this domain the registrant of record is NOT the operator. That is the single most important finding about the domain-register source kind this pass added, and it is kept in the residue notes rather than smoothed away. (2) The adjudicator scored Disconnect's entry 'I.UA' as error because the registrant's name differs from it. But the brief scores a trading brand for the same business as current, and the operator's own agreement calls the property 'порталу I.UA' – the portal I.UA. error means 'never the owner at any time', which this is not. Rescored current. Rate effect: moves Disconnect UP, which is the self-serving direction, so the alternative reading is published beside it and the sensitivity run reports what happens without it. |
adgrx.com | keep | FETCHFAIL 'HTTP 000' is the TLS layer, not the citation: samsungads.ca serves an expired certificate, so curl refuses it and the checker never sees a body. Re-fetched by hand with –insecure on 2026-09-11: HTTP 200, 46,160 bytes, and the quoted row 'adgrx.com</span></td><td>ADGRX_UID' is byte-verbatim in it. The checker is deliberately NOT given an –insecure retry: an ownership citation whose host identity cannot be validated should surface, not be swallowed. Same call the 2026-09-11 inter-rater pass made on its own expired-cert row. |
sa-as.com | keep | NOTFOUND was a cached rate-limit stub, not a bad citation. whois.markmonitor.com answers the fourth rapid query with a record that has no registrant block, and the checker cached it. Re-queried 25 seconds later: 3,357 bytes, 'Registrant Organization: FoundryCo, Inc.' present. owner_verify_sources.py now refuses to cache a whois answer under 400 bytes or one that does not name the queried domain, and blogblog.com hit the same stub. The row still rests on an UNCORROBORATED registrant organisation and is counted as such. |
Only one verdict changed, i.ua, and it moves Disconnect up, which is
the self-serving direction. The alternative reading is that “I.UA” names no
company at all and the row is an error; under it Disconnect's domain-level
current would be lower still than the 89.2% now published, so the published
figure is the more favourable of the two readings and not the less.
L.7 What is still unresolved, and every route tried
12 distinct domains of the 175, down from 40. These are not rows where the register was not tried. Each note below is the adjudicator's own account of what each of the three routes returned.
| Domain | In the residue of | What each list claims | Why it is still unresolved |
|---|---|---|---|
1rx.io | webXray | wx: Blinkx / TR: RhythmOne / DC: Nexxen | register-by-number: no company/VAT number available anywhere for 'Blinkx'/'RhythmOne'/'Nexxen' tied to this specific domain, so nothing to look up. rdap-history: RDAP has no data for the .io TLD via rdap.org; registry WHOIS (whois.nic.io) refers to GoDaddy, and the registrant is fully REDACTED (Domains By Proxy, LLC) with no organisation name, created 2015-05-28. archived-legal: root domain 1rx.io has zero Wayback ca |
agkn.com | webXray, Disconnect | wx: Neustar Marketing / TR: TransUnion LLC / DC: TransUnion | Register-by-number: SEC EDGAR full-text search for the exact phrase “agkn.com” across all filings returned 0 hits; “AdAdvisor” appears only in two old NEUSTAR INC (CIK 1265888) filings from before the 2021 TransUnion acquisition, nothing from TransUnion itself. RDAP/WHOIS history: registrant is fully privacy-proxied (Brandsight Privacy Customer, PO Box, Boise ID) with no organization field populated, at either GoDadd |
marphezis.com | Tracker Radar | TR: Online Media Solutions Ltd. dba Brightcom | register-by-number: no identifier available for 'Online Media Solutions Ltd. dba Brightcom' tied to this domain specifically. rdap-history: RDAP/registry WHOIS registrant is fully REDACTED (GoDaddy/Domains By Proxy pattern), registered 2015-07-14, last changed 2026-07-15 (routine renewal, not evidence of transfer). archived-legal: fetched the only 200-status legal-adjacent capture found – it is a GoDaddy domain-park |
cdnbasket.net | Tracker Radar, Disconnect | TR: Bounce Exchange / DC: cdnbasket.net | All three routes came back empty. Register-by-number: no identifying company/VAT number found anywhere for this domain or for Bounce Exchange/Wunderkind. RDAP/WHOIS history: registrant is privacy-proxied today (Contact Privacy Inc., Tucows), last changed 2026-08-15, with no un-redacted historical record available. Archived legal pages: the Wayback CDX index has zero captures of cdnbasket.net at any URL, ever. bouncee |
mapixl.com | Disconnect | DC: MarketingArchitects | All three routes tried and none produced an identifier for the actual operator, let alone one naming 'MarketingArchitects'. register-by-number: no imprint/legal page exists anywhere to pull a company number from. rdap-history: re-queried whois.godaddy.com by socket myself this session – registrant is 'Domains By Proxy, LLC', an explicit privacy proxy, which the brief excludes as a domain-register source; DNS is Clou |
htplayground.com | Disconnect | DC: htplayground.com | register-by-number: Disconnect's own entry is just the domain string, so there is no candidate company name to get an identifier for. rdap-history: rdap.org has no RDAP service for this registrar; registry WHOIS is thin, and the registrar-side WHOIS (whois.registrar.amazon) shows a proxy registrant ('c/o whoisproxy.com', 604 Cameron Street, Alexandria VA – a known privacy-proxy mail-drop address), no real organisati |
contentabc.com | Tracker Radar | TR: Aylo / DC: ContentABC | Register-by-number: no imprint, VAT, or company number for either candidate ('Aylo'/Ethical Capital Partners or a literal 'ContentABC') is attached to this domain anywhere I could find, so there is no number to look up. RDAP/WHOIS history: registrant is fully proxied via 'Whois Privacy (enumDNS dba)' in Luxembourg through EuroDNS S.A.; no organization ever disclosed; creation date 2009 does not narrow the candidates. |
stat-track.com | Disconnect | DC: StackTrack | All three routes were tried. Register-by-number: Disconnect's 'StackTrack' is not a findable real company, and a GitHub-issue claim of Moosend ownership could not be verified – Moosend's own privacy, cookie and terms pages (all fetched) never mention stat-track.com. RDAP/WHOIS history: rdap.verisign.com (queried live, quoted above) returns only a registrar (MarkMonitor) entity with no registrant object at all – ful |
mmstat.com | webXray | wx: Alimama / TR: Alibaba Group / DC: AlibabaGroup | Register-by-number: tried China's official ICP/beian registry at beian.miit.gov.cn directly – it returned HTTP 521 behind a JS/cookie challenge I could not solve with curl. A third-party ICP lookup aggregator (icplishi.com; not an accepted source, same category as a WHOIS aggregator) suggests the beian filer is an 'Alibaba (China) … Co., Ltd.' entity, but this is an unverified lead, not admissible evidence, so it |
collective-media.net | webXray | wx: Collective / TR: Collective, Inc. / DC: Collective | Tried all three routes. Register-by-number: SEC EDGAR full-text search for '“Collective, Inc.” advertising' returns 424 hits but none for the ad-network Collective, Inc. (it was VC-backed/private, never an SEC filer) – no identifier to look up. RDAP/WHOIS history: re-queried whois.godaddy.com this session; Registrant Organization is 'Domains By Proxy, LLC', a proxy, not acceptable per the domain-register rule. Archi |
hqseek.com | Tracker Radar | TR: Michal Vodrazka | All three routes tried. register-by-number: no imprint or identifiable number anywhere; the domain redirects to an adult site (hqporn.pics) which per content policy I did not fetch further. rdap-history: RDAP has no JSON endpoint for this domain; the two-hop WHOIS in the probe shows only registrar DNC Holdings/DirectNIC and registrant 'Jewella Privacy LLC' (a privacy service), no unredacted organization. archived-leg |
stripst.com | Disconnect | DC: Stripchat | Register-by-number: no imprint or company number is available anywhere for stripst.com, and nothing on the domain or Stripchat's reachable pages names an operating legal entity, so there is no number to look up. RDAP/WHOIS history: registrant is fully proxied ('Withheld For Privacy LLC', Delaware) via NameCheap; the mechanical probe's RDAP lookup also failed; TLS is a domain-validated Google Trust Services cert with |
Two obstacles are worth naming for anyone attempting this again:
mmstat.com would be settled by China's ICP registry at beian.miit.gov.cn,
which serves a JavaScript challenge no fetch here could pass; agkn.com by
TransUnion's own pages, which return 403 to everything that is not a browser
session. The remaining residue is a lower bound on what harder sourcing can
reach, not a claim that these domains have no owner.
L.8 What the pass could not establish
- Whether the 12 still-unresolved rows are worse than the rest is still unbounded, and this pass makes the bound narrower rather than tighter: the rows it recovered are by construction the easier end of the residue, so the test in section B of
report_tail_pass.pyis a test of the recovered rows, not of the residue. It can show the recovered rows are worse, or fail to; it cannot say anything about what is left.report_tail_pass.pysays so in its own docstring and the content page says so in prose. - The recovered/existing comparison has almost no power. 7, 9 and 12 recovered rows. A difference would have to be enormous to register, and none of the three p-values is interpretable as evidence of no difference.
- The adjudicators were again language models, one rater per domain, working to
out/tail_adj/INSTRUCTIONS_TAIL.md. No inter-rater agreement was measured for this pass — the +0.400 kappa on the content page is the 2026-09-05 rows. Measuring it would need a second model over the same 40 and was not attempted this sitting. - Three sub-agents could not re-issue a port-43 query themselves (this sandbox's
curlhas nowhois://support andbashhas no/dev/tcp) and relied on the probe's already-fetched WHOIS text, disclosing it per row. Every such citation was independently re-fetched byowner_verify_sources.pyafterwards, which is what closes that gap. cratecamera.comandsa-as.comrest on an uncorroborated registrant organisation and nothing further could be found for either: Admiral publishes no allowlist of its own domains, and no Foundry page namessa-as.com. They are in the sensitivity run for that reason.
J. scripts/report_ownership_resolution.mjs, in full
- report_ownership_resolution.mjs
// Every corpus figure on Design:Ownership resolution, with its denominator. // // node scripts/report_ownership_resolution.mjs > scripts/report_ownership_resolution-output.txt // node scripts/report_ownership_resolution.mjs --wiki # DokuWiki tables // // This page was carved out of Programming:Crawler:webXray on 2026-09-11. The // corpus-side figures it inherited were computed by scripts/report_webxray.mjs, // whose population is "papers that name webXray". That is the wrong population // for a page about ownership resolution in general, so this script re-derives // the landscape figures from the corpus independently and then CROSS-CHECKS the // overlapping rows against report_webxray.mjs's committed output. Two // implementations agreeing is a stronger audit trail than one copied number. // // Three things this script does that a naive sweep would not: // // 1. Every probe is run at TWO widths and the script asserts tight <= loose AND // tight subset-of loose. A narrowing probe that returns more hits, or hits // the loose one does not contain, is a broken probe, not a finding. The // Disconnect probe is the one that needs this: plain /disconnect/i is a // common English word and the CSP literature uses it as a technical term. // // 2. Every count is a PAPER count over data/fulltext/<year>/<venue>/<slug>/ // paper.cols.txt with whitespace collapsed, because a PDF line break inside // "Public Suffix List" silently undercounts it. // // 3. Papers with no full text are counted and printed, so the sweep's own // denominator is visible rather than assumed to be 5,859. // // Every row here is an UPPER BOUND on use unless the output says hand-verified: // a full-text match is a mention, not a use. The two rows that ARE hand // verified (webXray, Tracker Radar) carry their verdicts in report_webxray.mjs. import fs from 'node:fs'; import path from 'node:path'; import { dataRoot, loadExtractions, POPULATIONS, pct, table, wikiTable } from './lib.mjs'; const WIKI = process.argv.includes('--wiki'); const rows = loadExtractions(); const key = (p) => `${p.venue}/${p.year}/${p.slug}`; const FT = path.join(dataRoot(), 'fulltext'); const out = []; const P = (s = '') => out.push(s); const heading = (s) => { P(''); P(`=== ${s} ===`); P(''); }; const T = (headers, body) => P(WIKI ? wikiTable(headers, body) : table(headers, body)); // --- full text, whitespace collapsed ------------------------------------- // Streamed, not cached: the corpus is 5,859 papers and holding every // paper.cols.txt in a Map exhausts the default V8 heap. One pass, every // pattern tested against each paper's text, then the text is dropped. function readText(k) { const [venue, year, slug] = k.split('/'); const f = `${FT}/${year}/${venue}/${slug}/paper.cols.txt`; return fs.existsSync(f) ? fs.readFileSync(f, 'utf8').replace(/\s+/g, ' ') : null; } // Which papers have full text at all, and the hits for every pattern, in one // pass. PATTERNS is filled below before this runs. function sweepAll(patterns) { const hits = new Map([...patterns.keys()].map((n) => [n, []])); const haveText = new Set(); for (const p of rows) { const k = key(p); const t = readText(k); if (t === null) continue; haveText.add(k); for (const [name, re] of patterns) if (re.test(t)) hits.get(name).push(k); } for (const v of hits.values()) v.sort(); return { hits, haveText }; } // --- the probes, each at two widths -------------------------------------- // `loose` must be a strict superset of `tight` by construction; the script // checks that it is in fact one and throws if not. const PROBES = [ { name: 'webXray', tight: /webx[\s-]?ray/i, loose: /webx[\s-]?ray|libert/i, note: 'hand-verified: all 15, in report_webxray.mjs', }, { name: 'Tracker Radar', tight: /tracker[\s.-]?radar/i, loose: /tracker[\s.-]?radar|duckduckgo/i, note: 'hand-verified: all 32, in report_webxray.mjs', }, { name: 'WhoTracks.me', tight: /whotracks/i, loose: /whotracks|who ?tracks ?\.? ?me|ghostery/i, note: 'upper bound', }, { name: 'Crunchbase', tight: /crunchbase/i, loose: /crunchbase|crunch base/i, note: 'upper bound', }, { name: 'Disconnect (list sense)', tight: /disconnect'?s? (entit|list|block|black|tracking)|disconnect\.me|entities\.json/i, loose: /disconnect/i, note: 'upper bound; the loose width is the reason this one is tightened', }, { name: 'Public Suffix List', tight: /public suffix/i, loose: /public suffix|etld\+?1|effective top-?level domain/i, note: 'upper bound', }, ]; // One pass over the full text: every probe, both widths, plus the set of papers // that have a readable rendering at all. const PATTERNS = new Map(); for (const pr of PROBES) { PATTERNS.set(`${pr.name}|tight`, pr.tight); PATTERNS.set(`${pr.name}|loose`, pr.loose); } const { hits: SWEPT, haveText: HAVE_TEXT } = sweepAll(PATTERNS); const withText = rows.filter((p) => HAVE_TEXT.has(key(p))); const crawled = rows.filter(POPULATIONS.crawled); P('=============================================================================='); P('Design:Ownership resolution — every corpus figure, with its denominator'); P('=============================================================================='); P(''); P(`Corpus: ${rows.length} extracted papers, 7 venues (CCS, IMC, NDSS, PETS, USENIX`); P('Security, TheWebConf, IEEE S&P), 2010–2026. EuroS&P, ACSAC, RAID, AsiaCCS, CHI'); P('and SOUPS are absent, so every figure here is a claim about those seven venues.'); P(''); P(`Papers with a readable paper.cols.txt: ${withText.length} of ${rows.length} (${pct(withText.length, rows.length)}).`); P(`Papers that ran a crawl (crawlConfig present or studyTypes includes automated-web-crawl): ${crawled.length}.`); P('The sweep denominator is the full corpus; papers with no full text can only'); P('ever be misses, so every sweep count below is a floor on the true mention count.'); // --- A. probe widths ------------------------------------------------------ heading('A. Probe widths: tight must be contained in loose'); const results = new Map(); const widthRows = []; for (const pr of PROBES) { const tight = SWEPT.get(`${pr.name}|tight`); const loose = SWEPT.get(`${pr.name}|loose`); const tightSet = new Set(tight); const looseSet = new Set(loose); const escaped = tight.filter((k) => !looseSet.has(k)); if (tight.length > loose.length) { throw new Error(`probe ${pr.name}: tight (${tight.length}) > loose (${loose.length}) — the narrowing probe is broken`); } if (escaped.length > 0) { throw new Error(`probe ${pr.name}: ${escaped.length} tight hits are not in the loose set: ${escaped.slice(0, 5).join(', ')}`); } results.set(pr.name, tightSet); widthRows.push([pr.name, String(tight.length), String(loose.length), pr.note]); } T(['Resource', 'Tight probe', 'Loose probe', 'Status'], widthRows); P(''); P('Both widths are printed because the gap is the claim. Disconnect at the loose'); P('width matches 700 papers and almost none of them mean the list, which is why'); P('the published row is the tight one and why the page says so. A row whose two'); P('widths are close is a row whose count is not an artefact of the pattern.'); // --- B. the published landscape table ------------------------------------ heading('B. The ownership-resolution landscape (the page\'s main table)'); const landscape = PROBES.map((pr) => { const s = results.get(pr.name); const inCrawled = crawled.filter((p) => s.has(key(p))).length; return [pr.name, String(s.size), pct(s.size, rows.length), String(inCrawled), pct(inCrawled, crawled.length), pr.note.startsWith('hand-verified') ? 'yes' : 'no — upper bound']; }) .sort((a, b) => Number(b[1]) - Number(a[1])); T(['Resource named in full text', 'Papers', `Share of ${rows.length}`, `of which crawled`, `Share of ${crawled.length}`, 'Hand-verified?'], landscape); // --- C. the union --------------------------------------------------------- heading('C. The union: how widespread is ownership resolution at all?'); // The union is over the five OWNERSHIP resources. The Public Suffix List is // deliberately excluded: it answers "what is the registrable domain", not // "whose is it", and including it would nearly double the union on a resource // that every crawl paper touches for unrelated reasons. const UNION_NAMES = ['webXray', 'Tracker Radar', 'WhoTracks.me', 'Crunchbase', 'Disconnect (list sense)']; const union = new Set(); for (const n of UNION_NAMES) for (const k of results.get(n)) union.add(k); P(`Union of ${UNION_NAMES.join(' / ')}:`); P(` ${union.size} papers — ${pct(union.size, rows.length)} of the ${rows.length}-paper corpus,`); const unionCrawled = crawled.filter((p) => union.has(key(p))); P(` and ${pct(unionCrawled.length, crawled.length)} of the ${crawled.length} that ran a crawl (${unionCrawled.length} papers).`); P(''); P('The Public Suffix List is NOT in the union: it answers "what is the registrable'); P('domain", not "whose is it". Adding it would take the union to ' + (() => { const u2 = new Set(union); for (const k of results.get('Public Suffix List')) u2.add(k); return u2.size; })() + ' on a resource'); P('every crawling paper touches for unrelated reasons.'); P(''); P('Each member is an upper bound on USE, so the union is an upper bound too:'); P('136 papers mention one of these; far fewer resolved an owner with one.'); // --- D. per-year adoption ------------------------------------------------- heading('D. Per-year: is ownership resolution growing?'); const years = [...new Set(rows.map((p) => p.year))].sort(); const yearRows = years.map((y) => { const inYear = rows.filter((p) => p.year === y); const hit = inYear.filter((p) => union.has(key(p))); const star = y >= 2025 ? '*' : ''; return [`${y}${star}`, String(inYear.length), String(hit.length), pct(hit.length, inYear.length)]; }); T(['Year', 'Corpus papers', 'Naming an ownership resource', 'Share'], yearRows); P(''); P('* 2025 and 2026 are the provisional corpus edge: CCS 2026 and IMC 2026 have not'); P(' been held, and IEEE S&P / WWW 2026 abstracts are not in OpenAlex, so selection'); P(' under-samples them BY CONSTRUCTION. Do not read 2026 as a complete year.'); P(' The share, not the count, is the readable column for those two rows.'); // --- E. venue shape ------------------------------------------------------- heading('E. Venue shape of the union'); const venues = [...new Set(rows.map((p) => p.venue))].sort(); const venueRows = venues.map((v) => { const inVenue = rows.filter((p) => p.venue === v); const hit = inVenue.filter((p) => union.has(key(p))); return [v, String(inVenue.length), String(hit.length), pct(hit.length, inVenue.length)]; }).sort((a, b) => Number(b[2]) - Number(a[2])); T(['Venue', 'Corpus papers', 'Naming an ownership resource', 'Share of venue'], venueRows); P(''); P('Read the SHARE column, not the count: the venues differ in size by a factor of'); P('three. IMC and PETS are where this work lands; the shares are small everywhere,'); P('which is the honest headline — ownership resolution is a step inside a paper,'); P('not a paper topic, and most papers that do it do not say how.'); // --- F. the schema's own view -------------------------------------------- heading('F. What the extraction schema sees, against the sweep'); // classification[].groundTruthSource / resourceName naming an ownership list. const SCHEMA_RE = /webx[\s-]?ray|tracker[\s.-]?radar|whotracks|disconnect|crunchbase/i; const schemaPapers = new Set(); for (const p of rows) { const names = [ ...p.tools.map((t) => t.name), ...p.classification.map((c) => c.resourceName), ...p.classification.map((c) => c.groundTruthSource), ...p.otherToolsMentioned.map((t) => t.name), ].filter((s) => typeof s === 'string'); if (names.some((n) => SCHEMA_RE.test(n))) schemaPapers.add(key(p)); } const both = [...union].filter((k) => schemaPapers.has(k)); P(`Papers whose tools[] / classification[] / otherToolsMentioned[] name one of the`); P(`five resources: ${schemaPapers.size}`); P(`Papers whose FULL TEXT names one: ${union.size}`); P(`In both: ${both.length}`); P(`Full text only (schema misses): ${union.size - both.length}`); P(`Schema only (no full-text match): ${schemaPapers.size - both.length}`); P(''); P('The schema-only residue is printed rather than dropped:'); const schemaOnly = [...schemaPapers].filter((k) => !union.has(k)).sort(); for (const k of schemaOnly) { const p = rows.find((r) => key(r) === k); const hits = [ ...p.tools.map((t) => t.name), ...p.classification.map((c) => c.resourceName), ...p.classification.map((c) => c.groundTruthSource), ...p.otherToolsMentioned.map((t) => t.name), ].filter((s) => typeof s === 'string' && SCHEMA_RE.test(s)); P(` ${k}`); P(` named: ${[...new Set(hits)].join(' | ')}`); P(` full text present: ${HAVE_TEXT.has(k)}`); } P(''); P('This is why the page publishes the full-text sweep and not the schema field:'); P('the schema fires on a fraction of the papers that name these lists, because'); P('most mentions are in related work or a reference list rather than in a tools'); P('sentence the extractor reads as a tool. Neither signal is "the" population;'); P('both are reported.'); // --- G. cross-check against report_webxray.mjs --------------------------- heading('G. Cross-check against scripts/report_webxray-output.txt'); const committed = 'scripts/report_webxray-output.txt'; if (!fs.existsSync(committed)) { P(`${committed} not found — cross-check SKIPPED.`); } else { const txt = fs.readFileSync(committed, 'utf8'); const CHECKS = [ ['webXray', /^webXray\s+(\d+)\s/m], ['Tracker Radar', /^Tracker Radar\s+(\d+)\s/m], ['WhoTracks.me', /^WhoTracks\.me\s+(\d+)\s/m], ['Crunchbase', /^Crunchbase\s+(\d+)\s/m], ['Disconnect (list sense)', /^Disconnect \(list sense\)\s+(\d+)\s/m], ['Public Suffix List', /^Public Suffix List\s+(\d+)\s/m], ]; const checkRows = []; let bad = 0; for (const [name, re] of CHECKS) { const m = txt.match(re); if (m === null) throw new Error(`cross-check: could not find "${name}" in ${committed}`); const theirs = Number(m[1]); const mine = results.get(name).size; if (theirs !== mine) bad += 1; checkRows.push([name, String(mine), String(theirs), theirs === mine ? 'agree' : '*** DISAGREE ***']); } const um = txt.match(/^\s*(\d+) papers \((\d+\.\d)% of the corpus\)\. (\d+) of them are among the (\d+) that crawled -- (\d+\.\d)% of that population\./m); if (um === null) throw new Error(`cross-check: could not find the union line in ${committed}`); if (Number(um[1]) !== union.size) bad += 1; checkRows.push(['union of the five', String(union.size), um[1], Number(um[1]) === union.size ? 'agree' : '*** DISAGREE ***']); if (Number(um[3]) !== unionCrawled.length) bad += 1; checkRows.push(['union ∩ crawled', String(unionCrawled.length), um[3], Number(um[3]) === unionCrawled.length ? 'agree' : '*** DISAGREE ***']); if (Number(um[4]) !== crawled.length) bad += 1; checkRows.push(['crawled denominator', String(crawled.length), um[4], Number(um[4]) === crawled.length ? 'agree' : '*** DISAGREE ***']); T(['Figure', 'this script', 'report_webxray.mjs', 'verdict'], checkRows); P(''); if (bad > 0) { throw new Error(`${bad} figures disagree between the two independent implementations — do not publish either.`); } P('All rows agree. The two scripts share no code beyond lib.mjs and the'); P('whitespace-collapse convention; the patterns were written out separately.'); } // --- H. the page's non-corpus figures, listed so they are not mistaken ---- heading('H. Every number on the page that does NOT come from this corpus'); P('Listed so a reader auditing the page knows which script to re-run, and so'); P('that a refresh of this script is never mistaken for a refresh of the page.'); P(''); T(['Figure family', 'Script', 'Snapshot it is pinned to'], [ ['List shape (827 / 19,148 / 1,887 owners; 3,215 / 38,368 / 7,850 domains)', 'owner_dbs.py --cache out/webxray/cache', '2026-08-17'], ['Coverage census (669 / 5,581 / 2,268 of 32,369; the prevalence weights; the rank slices)', 'owner_dbs.py --cache out/webxray/cache', '2026-08-17'], ['Agreement between pairs (612 / 464 / 1,538 and the "neither" columns)', 'owner_dbs.py --cache out/webxray/cache', '2026-08-17'], ['The 28-row hand adjudication of high-prevalence disagreements', 'owner_adjudication.py --table', '2026-08-17'], ['Accuracy rates (69.2 / 73.9 / 89.2%, the encounter-weighted column, the CIs)', 'owner_random_sample.py --sample out/owner_sample.json --rows out/adj_rows_tail_final.json', '2026-09-05 draw, 2026-09-11 second pass'], ['What the second pass over the 40 unadjudicable rows changed', 'owner_tail_probe.sh | owner_tail_merge.py | owner_tail_corrections.py | report_tail_pass.py', '2026-09-11'], ['Coverage counts inside the random-sample section (664 / 5,566 / 2,264 of 32,337)', 'owner_sample.py --cache out/owner_cache', '2026-09-05'], ['Self-named entity shares (0.8% / 0.7% / 3.2%)', 'owner_selfname_census.py', '2026-09-05'], ['Inter-rater kappa (+0.400, +0.680)', 'owner_irr_kappa.py', '2026-09-11'], ['Repository state, licences, Wayback dates', 'hand checks + wayback_webxray.sh', '2026-08-17 / 2026-09-05'], ]); P(''); P('Two of the three lists change weekly. Re-run before citing.'); console.log(out.join('\n'));
K. Its unedited output
Run as node scripts/report_ownership_resolution.mjs.
- report_ownership_resolution-output.txt
============================================================================== Design:Ownership resolution — every corpus figure, with its denominator ============================================================================== Corpus: 5859 extracted papers, 7 venues (CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P), 2010–2026. EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS are absent, so every figure here is a claim about those seven venues. Papers with a readable paper.cols.txt: 5855 of 5859 (99.9%). Papers that ran a crawl (crawlConfig present or studyTypes includes automated-web-crawl): 1120. The sweep denominator is the full corpus; papers with no full text can only ever be misses, so every sweep count below is a floor on the true mention count. === A. Probe widths: tight must be contained in loose === Resource Tight probe Loose probe Status ----------------------- ----------- ----------- ---------------------------------------------------------------- webXray 15 233 hand-verified: all 15, in report_webxray.mjs Tracker Radar 32 108 hand-verified: all 32, in report_webxray.mjs WhoTracks.me 23 124 upper bound Crunchbase 23 23 upper bound Disconnect (list sense) 74 700 upper bound; the loose width is the reason this one is tightened Public Suffix List 101 157 upper bound Both widths are printed because the gap is the claim. Disconnect at the loose width matches 700 papers and almost none of them mean the list, which is why the published row is the tight one and why the page says so. A row whose two widths are close is a row whose count is not an artefact of the pattern. === B. The ownership-resolution landscape (the page's main table) === Resource named in full text Papers Share of 5859 of which crawled Share of 1120 Hand-verified? --------------------------- ------ ------------- ---------------- ------------- ---------------- Public Suffix List 101 1.7% 46 4.1% no — upper bound Disconnect (list sense) 74 1.3% 65 5.8% no — upper bound Tracker Radar 32 0.5% 28 2.5% yes WhoTracks.me 23 0.4% 19 1.7% no — upper bound Crunchbase 23 0.4% 11 1.0% no — upper bound webXray 15 0.3% 13 1.2% yes === C. The union: how widespread is ownership resolution at all? === Union of webXray / Tracker Radar / WhoTracks.me / Crunchbase / Disconnect (list sense): 136 papers — 2.3% of the 5859-paper corpus, and 9.6% of the 1120 that ran a crawl (107 papers). The Public Suffix List is NOT in the union: it answers "what is the registrable domain", not "whose is it". Adding it would take the union to 221 on a resource every crawling paper touches for unrelated reasons. Each member is an upper bound on USE, so the union is an upper bound too: 136 papers mention one of these; far fewer resolved an owner with one. === D. Per-year: is ownership resolution growing? === Year Corpus papers Naming an ownership resource Share ----- ------------- ---------------------------- ----- 2010 119 0 0.0% 2011 116 1 0.9% 2012 151 0 0.0% 2013 125 0 0.0% 2014 166 0 0.0% 2015 190 0 0.0% 2016 182 2 1.1% 2017 231 4 1.7% 2018 254 5 2.0% 2019 402 9 2.2% 2020 404 15 3.7% 2021 379 13 3.4% 2022 546 19 3.5% 2023 719 19 2.6% 2024 690 17 2.5% 2025* 770 21 2.7% 2026* 415 11 2.7% * 2025 and 2026 are the provisional corpus edge: CCS 2026 and IMC 2026 have not been held, and IEEE S&P / WWW 2026 abstracts are not in OpenAlex, so selection under-samples them BY CONSTRUCTION. Do not read 2026 as a complete year. The share, not the count, is the readable column for those two rows. === E. Venue shape of the union === Venue Corpus papers Naming an ownership resource Share of venue ------- ------------- ---------------------------- -------------- PETS 510 49 9.6% USENIX 1410 19 1.3% WWW 843 19 2.3% IMC 638 17 2.7% CCS 990 13 1.3% IEEE-SP 767 11 1.4% NDSS 701 8 1.1% Read the SHARE column, not the count: the venues differ in size by a factor of three. IMC and PETS are where this work lands; the shares are small everywhere, which is the honest headline — ownership resolution is a step inside a paper, not a paper topic, and most papers that do it do not say how. === F. What the extraction schema sees, against the sweep === Papers whose tools[] / classification[] / otherToolsMentioned[] name one of the five resources: 89 Papers whose FULL TEXT names one: 136 In both: 82 Full text only (schema misses): 54 Schema only (no full-text match): 7 The schema-only residue is printed rather than dropped: IEEE-SP/2024/to-auth-or-not-to-auth-a-comparative-analysis-of-the-pre-and-post-login-security named: Disconnect Tracker Protection List full text present: true PETS/2022/on-dark-patterns-and-manipulation-of-website-publishers-by-cmps named: Disconnect | Disconnect tracking filter list full text present: true PETS/2024/generalizable-active-privacy-choice-designing-a-graphical-user-interface-for-glo named: Disconnect Tracker Protection lists full text present: true PETS/2024/support-personas-a-concept-for-tailored-support-of-users-of-privacy-enhancing-te named: Disconnect full text present: true USENIX/2019/canvas-fast-and-inexpensive-automotive-network-mapping named: physical ECU access and disconnection full text present: true WWW/2018/the-cost-of-digital-advertisement-comparing-user-and-advertiser-views named: Disconnect | Disconnect blacklist full text present: true WWW/2023/who-funds-misinformation-a-systematic-analysis-of-the-ad-related-profit-routines named: Disconnect full text present: true This is why the page publishes the full-text sweep and not the schema field: the schema fires on a fraction of the papers that name these lists, because most mentions are in related work or a reference list rather than in a tools sentence the extractor reads as a tool. Neither signal is "the" population; both are reported. === G. Cross-check against scripts/report_webxray-output.txt === Figure this script report_webxray.mjs verdict ----------------------- ----------- ------------------ ------- webXray 15 15 agree Tracker Radar 32 32 agree WhoTracks.me 23 23 agree Crunchbase 23 23 agree Disconnect (list sense) 74 74 agree Public Suffix List 101 101 agree union of the five 136 136 agree union ∩ crawled 107 107 agree crawled denominator 1120 1120 agree All rows agree. The two scripts share no code beyond lib.mjs and the whitespace-collapse convention; the patterns were written out separately. === H. Every number on the page that does NOT come from this corpus === Listed so a reader auditing the page knows which script to re-run, and so that a refresh of this script is never mistaken for a refresh of the page. Figure family Script Snapshot it is pinned to ---------------------------------------------------------------------------------------- ------------------------------------------------------------------------------------------- --------------------------------------- List shape (827 / 19,148 / 1,887 owners; 3,215 / 38,368 / 7,850 domains) owner_dbs.py --cache out/webxray/cache 2026-08-17 Coverage census (669 / 5,581 / 2,268 of 32,369; the prevalence weights; the rank slices) owner_dbs.py --cache out/webxray/cache 2026-08-17 Agreement between pairs (612 / 464 / 1,538 and the "neither" columns) owner_dbs.py --cache out/webxray/cache 2026-08-17 The 28-row hand adjudication of high-prevalence disagreements owner_adjudication.py --table 2026-08-17 Accuracy rates (69.2 / 73.9 / 89.2%, the encounter-weighted column, the CIs) owner_random_sample.py --sample out/owner_sample.json --rows out/adj_rows_tail_final.json 2026-09-05 draw, 2026-09-11 second pass What the second pass over the 40 unadjudicable rows changed owner_tail_probe.sh | owner_tail_merge.py | owner_tail_corrections.py | report_tail_pass.py 2026-09-11 Coverage counts inside the random-sample section (664 / 5,566 / 2,264 of 32,337) owner_sample.py --cache out/owner_cache 2026-09-05 Self-named entity shares (0.8% / 0.7% / 3.2%) owner_selfname_census.py 2026-09-05 Inter-rater kappa (+0.400, +0.680) owner_irr_kappa.py 2026-09-11 Repository state, licences, Wayback dates hand checks + wayback_webxray.sh 2026-08-17 / 2026-09-05 Two of the three lists change weekly. Re-run before citing.
Related
- Ownership resolution — the page these notes are for.
- webxray — the live-list scripts, the fold residues, the Wayback log, the 2026-08-17 and 2026-09-05 run tables.
- random_sample — the 175-row adjudication, the estimator, the inter-rater draw and kappa.
- residue_pass — the 2026-09-11 second pass over the residue in full: its probe, its brief, its merge gate, its corrections, its report, its guard, its mutation harness and every unedited output, including the 40-domain probe log and the diff between the two estimates.
- Corpus — corpus-level selection and extraction caveats.
References
- [1]
- Matte, Célestin; Bielova, Nataliia; Santos, Cristiana Teixeira (2020): "Do Cookie Banners Respect my Choice? Measuring Legal Compliance of Banners from IAB Europe's Transparency and Consent Framework", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
- [2]
- Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)
- [3]
- Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
- [4]
- Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
- [5]
- Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)
- [6]
- Selmo, Carlos; Carisimo, Esteban; Bustamante, Fabián E.; Alvarez-Hamelin, J. Ignacio (2025): "Learning AS-to-Organization Mappings with Borges", in: Proceedings of the 2025 ACM Internet Measurement Conference, pp. 120-133. (DOI)
