User Tools

Site Tools


design:ownership_resolution

This is an old revision of the document!


Ownership resolution

Your crawl logged a request to adnxs.com. Whose is it? Every third-party measurement result of the form “N% of sites contact Google” rests on that step, and the step is almost never described: the paper names a list, sometimes, and moves on. This page is about what that choice costs — which sources exist, how much of the web each one can name, how often each is right, and what a reviewer needs to see before accepting the number that comes out.

The short answer, dated 2026-09-11.

  • Starting a new measurement: use Disconnect entities.json for the owner name and DuckDuckGo Tracker Radar for breadth, and say which file and which commit. On a stratified random sample of each list's own coverage, drawn 2026-09-05 and re-adjudicated over its residue on 2026-09-11, Disconnect names today's owner for 89.2% of the entries that could be settled against Tracker Radar's 73.9% — a gap that is significant if you count granularity as acceptable (p = 0.004) and is not if you require the exact legal entity (p = 0.083, against a Bonferroni threshold of 0.0167, and it never cleared it); Tracker Radar names an owner for 17.2% of the registrable third-party domains in the frame against Disconnect's 7.0%. Neither dominates, so most papers should merge them: Disconnect first for the owner name, Tracker Radar as the fallback when Disconnect has no entry — and record which list and which key matched each domain, so the mix is visible in your data rather than buried in the pipeline. Those two headline figures are the metrics on which this recommendation looks best, and the body of the page argues that two other metrics matter more: weighted by how often a crawl meets the domain, the live lists are near-tied on coverage (84.3% against 80.5%) and the frozen list tops accuracy (93.8%). The recommendation survives that because webXray's file has not moved since 2021 and its 93.8% is measured over the fifth of its entries a 2026 crawl still meets — but you should know the comparison is not one-sided before you cite it.
  • webXray's domain_owners.json is historical. It has not moved since 2021-03-04, covers 2.1% of the same frame, and the last corpus paper to use it was published in 2022. It is still the only one of the three with an ownership tree and with per-language policy URLs. See webXray for the file, the tool and its licence history.
  • No list removes the manual work, and the literature is consistent on this: [1Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] spent about eight person-hours hand-vetting same-entity pairs after using the lists, and [2Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] merged three of them and added a WHOIS-and-graph pass on top.
  • Three papers to read first, in this order: [1Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] for what the lists miss and what the manual work actually costs; [2Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] for how to merge several and state a priority order; [3Matte, Célestin; Bielova, Nataliia; Santos, Cristiana Teixeira (2020): "Do Cookie Banners Respect my Choice? Measuring Legal Compliance of Banners from IAB Europe's Transparency and Consent Framework", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] for the reproducibility table to copy.
  • WHOIS is not an ownership database. Used unfiltered it returns registrar-privacy strings, and one published top-ten table of “organizations” is a third registrar proxies — see the box in Choosing a resolution source now.

The question has two axes before it has an answer

Three lists look at doubleclick.net and give three answers. None of them is wrong:

doubleclick.net          webXray=['DoubleClick', 'Google', 'Alphabet']  TR='Google LLC'  Disconnect='Google'

That is not a disagreement, it is three different questions being answered. Before you compare lists, or blame one, fix your position on two independent axes — and then state both in the paper, because “sites contacting Google” is a different number under each:

  • Granularity. Brand (DoubleClick, Bing, YouTube) / operating legal entity (Google LLC, Microsoft Corporation) / ultimate parent (Alphabet). webXray can give you any of the three, because it carries a parent_id tree; Tracker Radar gives the legal entity; Disconnect gives a compact grouping name. The difference is not small: bing.com, linkedin.com and adnxs.com all become Microsoft at the entity level.
  • Vintage. Which snapshot, and when ownership was last checked. Ad tech consolidates continuously — the same problem [4Selmo, Carlos; Carisimo, Esteban; Bustamante, Fabián E.; Alvarez-Hamelin, J. Ignacio (2025): "Learning AS-to-Organization Mappings with Borges", in: Proceedings of the 2025 ACM Internet Measurement Conference, pp. 120-133. (DOI)] names at the network layer as “an Internet shaped by constant mergers, rebrandings, and regional variation”.1) adnxs.com is Xandr under AT&T in one list and Microsoft in another, and the difference is a 2022 acquisition, not an error.

A third axis appears only once you have chosen a list: what counts as the same company for a first-party check. Resolving the request domain is the easy half; the step almost always missing is resolving the visited site too, so that google.com embedding gstatic.com does not count as third-party tracking. See Assembling the pipeline.

The live sources, compared

The lists are not three attempts at the same artefact. They differ in size by more than an order of magnitude, in what a record means, and in what they are licensed for.

webXray domain_owners.json DuckDuckGo Tracker Radar entity_map.json Disconnect entities.json
Owners / entities 827 19,148 1,887
Domains covered 3,215 38,368 7,850
Ownership hierarchy yesparent_id, up to 6 levels no — flat, though per-entity files carry an Owner field no — flat
Extra per-owner data purpose, country, trade bodies, per-language policy URLs displayName, aliases; prevalence, categories, fingerprinting and cookie behaviour in sibling files properties vs resources split; category in services.json
How ownership is decided hand curation. Libert's own paper calls the database “the product of years of detective work” [5Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)] “automatically generated” by Tracker Radar Detector; “new development and bug fixes, other than broken sites, are handled internally”2) Disconnect's own process page describes step 4 as “Connect every tracker to its parent entity through DNS, WHOIS, and behavioral evidence”3)
Update cadence none — frozen; last public commit 2021-03-04 monthly regeneration; last commit on main 2026-08-28, releases tagged 2026.08.28, 2026.07.27, … the repository commits continuously (HEAD 2026-09-05), but entities.json itself last changed 2026-08-07 — read the file's history, not the repo's
Licence PolyForm Strict 1.0.0 (no redistribution) on the most-mirrored snapshot; the project ended MIT, the 2018 standalone list copy is GPL-3.0 — see webXray CC BY-NC-SA 4.0 CC BY-NC-SA 4.0 (GPL-3.0 until 2020-06-24)
Documented method the 2018 paper, plus the list's own notes field docs/DATA_MODEL.md, docs/FAQ.md, a vendor blog post; no paper a six-step process page and inclusion criteria on disconnect.me; entities.json's own schema is undocumented
How to get a correction in nowhere — the repository is gone file an issue; the repo says non-breakage development is handled internally not by pull request. The README says “Pull requests are not reviewed and will be closed”; corrections go to evaluations@disconnect.me and must argue from “publicly available materials and technical information”4)

A fourth live option, Ghostery's trackerdb, is not measured anywhere on this page: its ownership data is spread across per-company .eno files with a separate patterns layer rather than a single domain→owner map, so putting it in the same table would have meant writing a parser whose choices nobody could check against the other three. That is an omission, not a judgement — if you are choosing among the live lists, this page gives you numbers for two of the three and a description of the fourth in Choosing a resolution source now.

Both live lists are non-commercial. Tracker Radar and Disconnect's list are CC BY-NC-SA 4.0. That is fine for university research and share-alike publication of derived data; it is not fine for a spin-out, a consultancy deliverable, or an industry collaboration without a separate licence, and both vendors offer commercial terms on request. Check this before your data-availability statement, not after. Note also that Disconnect's licence was GPL-3.0 until 2020-06-24 — a paper that quotes the GPL terms is quoting a licence that no longer applies.5)

Coverage: how much of the third-party surface can each one name?

Size is the wrong comparison, because a domain you never meet costs you nothing. The right one is: of the third-party domains a crawl actually encounters, weighted by how often it encounters them, what share can this list attribute to a company?

The measurement below takes Tracker Radar's domain_summary.json as the universe of third-party domains and its prevalence field as the weight, folds each key to a registrable domain with the ICANN section of the current Public Suffix List, and asks each list for an owner. This universe is Tracker Radar's own view of the web, which flatters Tracker Radar and nobody else — read its column as “the list scored on its home ground” and the other two as measured against a denominator they had no part in choosing.

List Domains it can name an owner for Share of 32,369 Weighted by prevalence
Frame: the 2026-08-17 snapshot. The random-sample section below uses the 2026-09-05 snapshot, which is 32,337 domains — see the note there.
webXray 669 2.1% 58.6%
Tracker Radar 5,581 17.2% 84.3%
Disconnect (propertiesresources) 2,268 7.0% 80.5%

The 2.1% and the 58.6% in the same row are the whole story of webXray's file. It covers almost none of the web by domain count and well over half of it by weight, because it is a head list: the few hundred domains it knows are the ones on every page. Sliced by rank, the tail falls off a cliff — and it does so for all three:

Slice of the universe, by prevalence webXray Tracker Radar Disconnect
top 100 domains 71% 98% 94%
top 1,000 domains 26% 69% 68%
top 10,000 domains 5% 29% 17%

Libert said this about his own list in 2018, and the sentence is worth having to hand when a reviewer asks about coverage: “because webxray's database of domain ownership primarily contains major ad networks rather than small clients, and policyxray only searches for identified parties, variability in the long-tail of trackers may not have an outsized effect on overall findings related to disclosure. Nonetheless, it is important to point out that the number of parties being searched for is fewer than the total number of parties present.” [5Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)]6)

With the fold and the lookup done properly, 29,896 of the 32,369 domains (92.4%) have no owner in either webXray or Disconnect, but only 14.4% of the prevalence weight does. The most requested domain no list can name is tiktokw.us at prevalence 0.041 — Tracker Radar attributes it to ByteDance, the other two have nothing. That is what a real coverage hole looks like: recent, mid-tail, and concentrated in domains that appeared after a list was last curated.

Three traps, and the first will silently ruin any coverage figure

  • The Public Suffix List has two sections, and using the wrong one manufactures a coverage hole. googleapis.com is a PRIVATE PSL rule (6,941 ICANN rules against 3,290 private ones in the snapshot used here). Fold with the private section and fonts.googleapis.com, ajax.googleapis.com and maps.googleapis.com each survive as their own “registrable domain”; look each up by exact key and every one comes back unowned — even though all three lists name googleapis.com → Google. fonts.googleapis.com alone carries prevalence 0.369, so this single mistake moves webXray's weighted coverage from 58.6% to 54.5% and Tracker Radar's from 84.3% to 79.3%. Fold on the ICANN section, or walk parent labels at lookup time, or both. The four combinations are printed side by side by the script and the spread between them is larger than the difference between the lists.
  • A large fraction of domain_summary.json's 47,836 rows are keyed by hostname, and how large depends on the fold. An ICANN-section fold merges 15,651 of them; the ICANN+PRIVATE fold merges only 2,339, because googleapis.com, s3.amazonaws.com and cloudfront.net are themselves private-section suffixes. (A label-count heuristic — “three or more labels” — gives 16,396 and answers neither question; do not use it.) All three ownership lists key on the registrable domain, so comparing coverage without folding at all charges webXray and Disconnect for subdomains they were never meant to hold.
  • 19 rows are not hostnames at all — 18 bracketed IPv6 literals and the literal string “null”, the last with a prevalence of 0.024 and a full behaviour profile attached. Drop them explicitly and say you did; do not let a null domain become a data point.

Reading all three at once

The three files have three different shapes, and the shape is where the granularity choice becomes code. webXray is an array of owners each holding a domains list and a parent_id; Tracker Radar's domain_map.json is already keyed by domain; Disconnect's entities.json is keyed by entity, so it has to be inverted before you can look a domain up at all. Reading all three at once takes about twenty lines and immediately shows what you are choosing between:

owner_lookup.py
import json
 
WX = json.load(open('cache/webxray.json'))
TR = json.load(open('cache/tr_domain_map.json'))
DC = json.load(open('cache/disconnect_entities.json'))['entities']
 
# webXray: an array of owners, each with a `domains` list and a `parent_id`.
wx_by_id = {o['id']: o for o in WX}
wx = {d.lower(): o for o in WX for d in o['domains']}
 
def wx_chain(domain):
    o = wx.get(domain)
    if o is None:
        return None
    chain = [o['name']]
    seen = {o['id']}
    while o['parent_id'] is not None and o['parent_id'] not in seen:
        o = wx_by_id[o['parent_id']]
        seen.add(o['id'])
        chain.append(o['name'])
    return chain                      # brand first, ultimate parent last
 
# Tracker Radar: already keyed by domain.
tr = {d.lower(): v['entityName'] for d, v in TR.items()}
 
# Disconnect: keyed by ENTITY, so it has to be inverted. `properties` is the
# ownership claim; `resources` is what the tracker actually served from.
dc = {}
for name, v in DC.items():
    for d in v.get('resources', []):
        dc.setdefault(d.lower(), name)
for name, v in DC.items():
    for d in v.get('properties', []):
        dc[d.lower()] = name
 
for d in ['doubleclick.net', 'adnxs.com', 'facebook.net', 'fonts.googleapis.com']:
    print(f'{d:24} webXray={wx_chain(d)}  TR={tr.get(d)!r}  Disconnect={dc.get(d)!r}')

Its real output on the 2026-08-17 snapshots:

doubleclick.net          webXray=['DoubleClick', 'Google', 'Alphabet']  TR='Google LLC'  Disconnect='Google'
adnxs.com                webXray=['Xandr', 'AT&T']  TR='Microsoft Corporation'  Disconnect='Microsoft'
facebook.net             webXray=['Facebook']  TR='Facebook, Inc.'  Disconnect='Meta'
fonts.googleapis.com     webXray=None  TR=None  Disconnect=None

Four rows, four different lessons. doubleclick.net is the granularity axis, with nothing wrong anywhere. adnxs.com is the vintage axis, with webXray's whole chain superseded. facebook.net shows the two live lists disagreeing because one carries a five-year-old legal name. And fonts.googleapis.com comes back empty from all three — which is a bug in this snippet, not a coverage hole: every list names googleapis.com → Google, and the snippet looks up the exact key it was given. That is the PSL private-section trap above, reproduced in twenty lines. Fix it by trying each parent label:

def resolve(host, mapping):
    labels = host.split('.')
    for i in range(len(labels) - 1):          # stop at two labels
        candidate = '.'.join(labels[i:])
        if candidate in mapping:
            return mapping[candidate], candidate
    return None, None

with which fonts.googleapis.com resolves to Google in all three. Two further details worth copying: the seen set in wx_chain is not decoration — walk a hand-curated parent chain without a cycle guard and one bad edge hangs your pipeline; and Disconnect's resources are loaded before its properties so the ownership claim wins where the two disagree.

When two lists disagree

For every domain that two lists both cover, owner_dbs.py compares the owner strings after normalising punctuation and folding away legal-form suffixes (Inc, LLC, GmbH, S.A.S, …) but no synonyms — so the “agree” columns are a lower bound on real agreement and the last column an upper bound on real disagreement. The fold is deliberately timid: it leaves 782 of webXray's 827 owner names untouched, and of the 45 it does change only 16 lose a legal suffix — the other 29 change on punctuation alone (AT&T, JD.com, “Here, There & Everywhere”). If you want the agreement figures to go up, a synonym table is what you would have to add, and there is no principled one.

Pair Domains both cover Same after suffix fold One name contains the other Neither
webXray vs Tracker Radar 612 305 (49.8%) 109 (17.8%) 198 (32.4%)
webXray vs Disconnect 464 213 (45.9%) 34 (7.3%) 217 (46.8%)
Tracker Radar vs Disconnect 1,538 634 (41.2%) 320 (20.8%) 584 (38.0%)

Between a third and a half of jointly covered domains get different company names — including between the two live lists, which is the row most papers need and the one least often noticed.

Before concluding that some list is broken, note what happens if you resolve webXray up its ownership tree first — the obvious fix, since webXray says “DoubleClick” where the others say “Google”:

Pair Neither, using webXray's immediate owner Neither, using the root of webXray's tree
vs Tracker Radar 198 (32.4%) 239 (39.1%)
vs Disconnect 217 (46.8%) 230 (49.6%)

Resolving to the root makes agreement worse, and the reason is instructive: webXray's roots are holding companies and, worse, historical ones. google.com resolves to “Alphabet” (against “Google LLC” and “Google”), adnxs.com to “AT&T”, yahoo.com to “Verizon”, tapad.com to “Telenor”, turn.com to “Singtel”. Every one of those was true when the file was last curated and none is true now. Climbing a hierarchy trades a granularity error for a vintage error; it does not remove either.

Which list is right, when they disagree?

30 of the highest-prevalence disagreements were adjudicated one by one against primary sources — company newsrooms, SEC filings, or the domain's own legal documents — on 2026-08-17, by Sonnet sub-agents working to a published brief, single-rated, not by a human expert. Read every verdict below with that in mind; the inter-rater figure further down is measured on the 2026-09-05 random sample, not on this table. Two could not be settled from a primary source and are excluded from the table, leaving 28.

List Names today's owner Stale (a real former owner) Granularity only Outright error No entry
webXray 1 18 4 1 4
Tracker Radar 8 12 6 2 0
Disconnect 26 0 2 0 0

This is not an accuracy ranking, and it cannot be turned into one. The 28 rows were chosen because the lists disagreed on them, ranked by prevalence; a random sample would be dominated by domains all three get right, so no percentage taken from this table means anything about how often a list is correct in general. Tracker Radar's zero in “No entry” is an artefact of how the sample was drawn rather than a result: the rows were ranked by Tracker Radar's own prevalence field.7)

What the table does establish is the shape of the disagreement. It is concentrated in acquisitions and renames rather than spread across the list; it points overwhelmingly one way; and the list regenerated most conservatively is not the one that is most current. If you need an accuracy rate, you have to draw a stratified random sample and adjudicate it — the next section.

Named examples, useful as regression tests for your own pipeline: adnxs.com (Xandr, AT&T → Microsoft, closed 2022-06-06); facebook.com (Facebook, Inc. → Meta Platforms, 2021-10-28 — Tracker Radar still says “Facebook, Inc.”); outbrain.com and teads.tv (Outbrain acquired Teads 2025-02-03 and then took its name, inverting the pair); postrelease.com (Nativo → Life360, completed 2026-01-05, which only Disconnect has); crwdcntrl.net (Lotame → Publicis, announced 2025-03-06). Three rows are errors rather than staleness: Tracker Radar attributes jsdelivr.net to “Prospect One”, jsDelivr's infrastructure contractor, where jsDelivr's own data-processing agreement names Volentio JSD Limited; Tracker Radar attributes stackadapt.com to “Collective Roll”, which is StackAdapt's own pre-2014 founding name and not a separate owner; and webXray's root for 1rx.io is “Marimedia”, which no primary source corroborates.

How often is each list right? A random sample

The table above is what you see when you look where the lists disagree. It is not an accuracy rate. This section is the measurement that was missing until 2026-09-05: 60 registrable domains drawn at random from each list's own coverage, stratified by prevalence quartile, each adjudicated against a primary source. The adjudicators were language models — Sonnet sub-agents working to a published brief, one rater per domain, with a second model re-rating a random fifth so the agreement could be measured rather than assumed (below). No human expert adjudicated these rows. 180 draws over 175 distinct domains (five were drawn for more than one list); each domain was adjudicated once and scored against all three lists, giving 525 verdicts. Only a list's own draws enter that list's rate, and of each list's 60 the ones that could be settled and make an ownership claim are 56, 56 and 54 — 166 rows carry every accuracy figure below. Those counts are after a second pass on 2026-09-11 that re-adjudicated the 40 domains the first pass could not settle, using sourcing the first pass did not attempt: national company registers searched by company number rather than by name, the registrar's own port-43 WHOIS rather than RDAP, and archived captures of the domains' own legal pages. It settled 28 of the 40, cutting the unadjudicable share from 23% of the draws to 7.8% — see the section on it for what that did to each rate.

Two numbers per verdict, because there are two different questions:

  • domain-level — of the domains this list names an owner for, what share does it name correctly? Not a raw proportion: the sample is stratified 24/12/12/12 over prevalence quartiles, so each stratum is weighted back to its share of the list's coverage. The raw counts are 48 of 54, 42 of 56 and 40 of 56 — divide those and you get 88.9 / 75.0 / 71.4, which is why they do not equal the 89.2 / 73.9 / 69.2 in the table.
  • encounter-weighted — when a crawl resolves a third-party request through this list, what share of resolutions are right? Weighted by Tracker Radar prevalence — its own per-domain figure for the share of the sites it crawls that request the domain, so gstatic.com at 0.40 is met on 40% of them. This is the number a paper's attribution error depends on. It is dominated by a handful of very common domains — gstatic.com alone carries 38.8% of webXray's — which is why its intervals are three times wider than the domain-level ones.
List Entries settled, of 60 drawn Names today's owner (domain-level) 95% CI Encounter-weighted 95% CI
webXray 56 69.2% 56.1–81.4% 93.8% 80.9–98.6%
Tracker Radar 56 73.9% 61.6–85.3% 74.3% 41.1–97.8%
Disconnect 54 89.2% 80.5–96.8% 82.8% 53.8–100.0%

30 of the 525 verdicts were changed after the adjudicators returned them, and the direction is overwhelmingly favourable to the lists. 29 on 2026-09-05: 14 landed on current, 2 on granularity, 1 on stale and 9 on unknown (which removes the row from every rate); one more on 2026-09-11, i.ua from error to current, which also favours a list. Most of that is retiring a verdict the brief should never have offered, but the net effect is to move rates up, and every figure in the table above should be discounted accordingly. The full list with reasons is on the provenance pages.

Counting granularity (right company, wrong level of the corporate tree) as acceptable rather than wrong — which is the right choice if your unit of analysis is “which company”, not “which legal entity” — the domain-level figures become webXray 78.0% (65.8–88.8), Tracker Radar 76.4% (64.6–87.5), Disconnect 96.8% (91.5–100.0).

Two frames, 32 domains apart — read the dates, not just the numbers. The coverage table earlier on this page counts 669 / 5,581 / 2,268 of 32,369 domains; this section counts 664 / 5,566 / 2,264 of 32,337. Nothing is inconsistent: the earlier figures are the 2026-08-17 snapshot that the 28-row disagreement adjudication was done against, and this section's are the 2026-09-05 snapshot the random sample was drawn from. Tracker Radar's domain_summary.json and the Public Suffix List both changed in the nineteen days between. The coverage shares are identical to one decimal (2.1% / 17.2% / 7.0%), which is what makes this easy to miss. The earlier figures are deliberately not re-derived, because re-deriving them would orphan the adjudication pinned to them.

What the sample does and does not license

Four readings, each of which needs the random sample and could not come from the disagreement table:

  • error — a name that was never the owner — is uncommon, and it is no longer absent from two of the three lists. Zero of webXray's 56 settled entries, and the rule-of-three bound on that zero is below 5.4% at 95% confidence. Tracker Radar has 3 of 56: elfsight.com attributed to “Vladimir Fedotov”, Elfsight's co-founder rather than the owning entity (elfsight.com/terms-of-service/ names Elfsight, SL); cnevids.com attributed to the law firm “Sabin, Bermant & Gould LLP”, which appears in the record as the registrant's counsel and never operated the domain (player.cnevids.com is Condé Nast Entertainment's own player); and cedscdn.it attributed to “The Trustico Group Ltd”, a UK certificate reseller, where the .it registry names CED DIGITAL & SERVIZI Srl. Disconnect has 1 of 54: km0trk.com, whose entry is the bare domain string, where Good On You Pty Ltd's own privacy policy names www.km0trk.com as its own. That one is a scoring decision, not an attribution blunder like the other three, and the rule separating it from i.ua — where Disconnect's bare-looking “I.UA” was scored current — is whether the label is a brand the operator's own documents use for the property. “I.UA” is (the user agreement says “порталу I.UA”); “km0trk.com” is not. Reasonable people would score these the other way, and the entry-shape census below counts the whole class without needing a verdict at all. Three of the four are new: elfsight.com was already there, and cnevids.com, cedscdn.it and km0trk.com came out of the 2026-09-11 second pass, because an unsettled row is scored against nobody. Staleness is still the dominant failure, but “error is rare” was partly an artefact of what could not be checked.
  • Tracker Radar's encounter-weighted accuracy (74.3%) is the lowest of the three, despite its being the broadest — and partly because the three rates are not strictly on one axis. A brand name survives an acquisition and a legal-entity name does not: for crwdcntrl.net, “Lotame” scores current and “Lotame Solutions, Inc.” scores stale, on one fact about one company. A list that carries legal names is therefore penalised by this measure for being more precise. Tracker Radar regenerates monthly but carries legal-entity names, and legal names lag renames: zoominfo.com is still “Zoom Information, Inc.”, mountain.com “Mountain Digital, Inc.” for MNTN, 20min.ch “Tamedia AG” where the group renamed to TX Group, tvtime.com “Whipclip”.
  • The encounter-weighted column is concentrated enough that one domain moves it. gstatic.com alone carries 38.8% of webXray's encounter-weighted estimate, and webXray's top three drawn domains carry 52.2%; had gstatic.com been stale, webXray's 93.8% would read 55.0%. The equivalent heaviest rows are omtrdc.net for Tracker Radar (20.0%, and 74.3% would read 54.3%) and id5-sync.com for Disconnect (16.6%, 82.8% → 66.2%). That is not a defect in the estimator — it is what “weighted by how often you meet the domain” means when prevalence spans five orders of magnitude — but it is why the encounter-weighted column supports neither a number nor an ordering: all three intervals overlap every other, and no pairwise test is run on them. It is here because it is the quantity a paper's attribution error actually depends on, not because this sample can pin it down.
  • What could not be settled has shrunk from a quarter of the sample to a fourteenth, and settling it moved the answers. After the second pass, 14 of the 180 draws are unresolved (7.8%), over 12 distinct domains of the 175 — down from 42 draws and 40 domains. Per list it is 4, 4 and 6 of each 60, i.e. 6.7% / 6.7% / 10.0%, where it was 18% / 22% / 30%. Every rate above is still conditional on the entry being adjudicable, but that condition now binds on a seventh of what it bound on. The worst case, every remaining unresolved entry being stale, gives 66.7% / 70.0% / 80.0% where it gave 58.3% / 58.3% / 65.0%. What is left is genuinely hard: dead domains with a privacy-proxied registrant, no archived legal page and, in four cases, nothing but tracking endpoints ever captured.

Only one of the three pairwise differences survives a correction for having tested all three. Fisher's exact on the eligible raw counts, two-sided:

Pair Counting current only Counting current+granularity
Eligible rows p p before Eligible rows p p before
Disconnect vs webXray 48 of 54 against 40 of 56 0.031 0.014 52 of 54 against 44 of 56 0.008 0.010
Disconnect vs Tracker Radar 48 of 54 against 42 of 56 0.083 0.025 52 of 54 against 43 of 56 0.004 0.004
Tracker Radar vs webXray 42 of 56 against 40 of 56 0.83 0.82 43 of 56 against 44 of 56 1.00 0.81

Three tests on one sample, so the threshold is 0.0167 rather than 0.05. Which comparisons survive it depends entirely on whether granularity counts as right, and that is a choice about your unit of analysis, not about the data:

  • If you care which legal entity owns the domain (current only), none of the three survives. Disconnect is still the highest and by the widest margin, but on 54 and 56 adjudicated rows that ordering is not established at this threshold and should be described rather than asserted.
  • If you care which company owns it — the unit most papers use, and the one this page recommends two paragraphs above — Disconnect beats both of the others, at p = 0.008 and p = 0.004, and it did before the second pass too.

Quote the p-values rather than pointing at the gap between the intervals: a percentile bootstrap on 54 rows undercovers at the top, which is why one endpoint reads 100.0.

What the second pass actually overturned is narrower than it first looks, and the narrowing matters. Before 2026-09-11, Disconnect-versus-webXray cleared the corrected threshold on the current-only metric at p = 0.014; it is now 0.031 and does not. On current+granularity nothing changed: Disconnect cleared the bar against both lists before (0.010, 0.004) and clears it after (0.008, 0.004). So the second pass cost Disconnect a legal-entity-level lead over a list this page calls historical, and cost it nothing at the company level. Settling 28 of the 40 previously unadjudicable rows added 7 eligible rows to webXray, 9 to Tracker Radar and 12 to Disconnect — and on the current-only count, 9 of Disconnect's 12 were current, against 39 of its previous 42; its point estimate fell from 94.4% to 89.2% there, and from 98.8% to 96.8% on the looser metric. An earlier version of this page asserted that the Disconnect-versus-Tracker-Radar difference could not be established without having computed it; on current only it was 0.025 then and 0.083 now, and the correction is recorded on the provenance page.

Two further limits on the 69.2%, both of which cut against reading it as a rehabilitation of webXray. First, those 56 settled entries come from the 664 webXray domains that are still in a 2026 crawl frame at all — a fifth of its 3,215 domain→owner pairs. A domain still requested in 2026 is a survivor, and survivors are where “current” concentrates, so the rate is over the surviving fifth and not over the file. Second, seven domains were adjudicated in both the 2026-08-17 and the 2026-09-05 sittings against a bit-identical copy of the file; of the five settled in both, webXray's verdict differs in four, every one towards current, and the difference is a change of scoring convention (which level of the tree, and whether the acquiring group or the operating subsidiary is “the owner”) rather than a change of sample.

If you cite one figure from this page, cite a coverage figure, not an accuracy figure. Coverage is a census over the whole frame with no sampling error; accuracy is 54 to 56 adjudicated rows per list, with intervals from ±8 (Disconnect) to ±13 (webXray) percentage points. The strongest defensible claims are: no list names an owner for more than 17.2% of the registrable third-party domains a Tracker Radar crawl sees; and among entries that exist and can be checked, staleness rather than error is what goes wrong.

What the second pass over the residue changed

The 2026-09-05 sample left 40 of its 175 domains unadjudicable, and every rate it produced was conditional on that. Because the share rose toward the prevalence tail and was worst for the list this page recommends, the obvious worry was that the unsettled rows were systematically worse and all three rates were flattered — which nothing inside that sample could test. On 2026-09-11 those 40 rows, and only those 40, were re-adjudicated — again by Sonnet sub-agents working to a published brief, one rater per domain, with no inter-rater agreement measured for this pass — against harder sourcing: national company registers searched by company number, the registrar's own port-43 WHOIS rather than RDAP (which returns a thin registry record with no registrant for .com and .net), and archived captures of each domain's own legal pages. 28 of the 40 settled.

Before After
Distinct domains unresolved, of 175 40 (22.9%) 12 (6.9%)
Draws unresolved, of 180 42 (23.3%) 14 (7.8%)
Eligible rows carrying the rates 49 / 47 / 42 56 / 56 / 54
Disconnect, domain-level current 94.4% 89.2%
Tracker Radar, domain-level current 73.7% 73.9%
webXray, domain-level current 69.9% 69.2%
Tracker Radar, encounter-weighted current 68.6% 74.3%
error verdicts found 1 (Tracker Radar) 4 (3 Tracker Radar, 1 Disconnect)
Pairwise comparisons surviving Bonferroni 1 of 3 0 of 3

The largest single movement of the pass is not in the domain-level column at all: Tracker Radar's encounter-weighted rate rose 68.6% → 74.3%, almost all of it one recovered row, spot.im at prevalence 0.011. That is the column the page calls “the number a paper's attribution error depends on”, and it moved five times as far as any domain-level rate.

The formal test does not confirm the worry, and the point estimates do not refute it. Comparing the recovered rows against the rows already in each rate — current+granularity against stale+error, Fisher's exact two-sided — gives p = 0.64 for webXray (5 of 7 recovered good, against 39 of 49), p = 1.00 for Tracker Radar (7 of 9, against 36 of 47) and p = 0.40 for Disconnect (11 of 12, against 41 of 42). None is close to significant. But 7, 9 and 12 rows have almost no power to detect anything, and Disconnect's estimate still fell 5.2 points, which is the direction the worry predicted for the list the worry named. And the answer depends on the metric: counting granularity as acceptable, Disconnect's recovered rows are 11 of 12 against 41 of 42 and its estimate falls 2.0 points rather than 5.2. The honest summary is that nothing here establishes that the tail was worse for any list, that Disconnect's point estimate moved in the direction the worry predicted on the stricter metric and barely at all on the looser one, and that none of the three tests has the power to tell those apart.

The second pass used a bar the first pass did not, and you should know where it is weaker. Three source kinds are new: domain-register (an un-redacted Registrant Organization in the registry or registrar WHOIS), archived-legal-doc (a Wayback capture of the domain's own legal page), and tls-san (another organisation's certificate on this host — corroboration only, never used alone). Of the 28 recovered rows, 9 rest on domain-register and 2 on archived-legal-doc; the other 17 clear the original bar.

domain-register is genuinely weaker: the field is asserted by the registrant and validated by nobody. This pass produced its own counter-example. i.ua's registry record names Digital Ventures LLC; the portal's own user agreement names a different company as its administration, ТОВ «КЕПРЕЙТ ПАРТНЕРС», register code 33500955. On that domain the registrant of record is not the operator, and the row was re-sourced to the legal document. Dropping every remaining recovered row that rests on an uncorroborated registrant organisation — cratecamera.com and sa-as.com — moves nothing: webXray +0.0, Tracker Radar −0.6, Disconnect +0.0 percentage points.

The comparison that actually answers the objection is the same-bar one, and it is the one to quote if you distrust the new sources: drop all 11 rows resting on a source kind the first pass would not have accepted, so that what is left is only what the 2026-09-05 bar could itself have reached. The rates barely move — webXray 69.2% (+0.0), Tracker Radar 72.1% (−1.8), Disconnect 89.0% (−0.2) — and the ordering, the significance verdicts and the residue conclusion are all unchanged. So the second pass's result does not depend on the bar it widened. Under the other reading of i.ua — that “I.UA” names no company and the row is an error rather than current, the reading this pass argued against — Disconnect is 88.6% rather than 89.2%. All three runs are in section E3 of the report script's output.

What is left is a different kind of hard. The 12 domains still unresolved — 1rx.io, agkn.com, cdnbasket.net, collective-media.net, contentabc.com, hqseek.com, htplayground.com, mapixl.com, marphezis.com, mmstat.com, stat-track.com, stripst.com — are not domains where the register was not tried. Ten of the twelve are domains with a privacy-proxied registrant, no live site, no archived legal page, and in several cases nothing in the Archive but tracking endpoints, so there is no document to read at any date. Two have a further obstacle worth naming for anyone attempting this: mmstat.com would be settled by China's ICP registry, which serves a JavaScript challenge no fetch here could pass, and agkn.com by TransUnion's own pages, which return 403 to everything that is not a browser session.

The other two are limits of this pass rather than of the evidence, and a person could still close them. hqseek.com redirects to a live adult site and the adjudicator declined to follow it on content policy — so the row is unresolved because a rater stopped, not because nothing is there; its note also records that Tracker Radar's claim, a personal name, matches nothing the domain's own pages ever surfaced, which would be an error rather than an unknown if it were run down. stripst.com's operator pages returned HTTP 406 to every non-browser fetch — a transport failure, not an absence. Both would be worth a browser session; neither was given one here.

The residue is also not a random slice of the sample, and the way it is skewed matters for the bound below. Three of Disconnect's six remaining rows are entries where Disconnect's own value carries no company name a register could be searched for — cdnbasket.net and htplayground.com are the bare domain string, and stat-track.com's “StackTrack” is not a findable company. That is the same class as km0trk.com, the one entry Disconnect was scored error on, so the worst case is more likely to bind on Disconnect than the even split of the residue across the three lists would suggest. The remaining residue is a lower bound on what harder sourcing can reach, not a claim that these domains have no owner. Every attempted route per row is printed on ownership_resolution.

Inter-rater agreement, measured rather than assumed

A random 20% of the 2026-09-05 sample (35 of 175 domains) was re-adjudicated on 2026-09-11 by a different model against the identical brief and the identical evidence blocks. On the rows the brief asks about, the two raters agree on 58.5% of 65 verdicts, Cohen kappa +0.400 (95% CI +0.241 to +0.551); per list, +0.027 for webXray on 12 rows — chance agreement — +0.478 for Tracker Radar on 30 and +0.460 for Disconnect on 23. So for webXray specifically the two raters agreed no better than chance, and its 69.2% carries rater uncertainty on top of the sampling interval already quoted. That kappa was measured on the 2026-09-05 rows; no inter-rater agreement was measured for the second pass over the residue, so the 28 rows it recovered carry no agreement figure at all.

About two-thirds of the disagreement is one rater finding a primary source where the other did not, rather than two readings of the same evidence — restricted to rows both settled, the pooled kappa is +0.680. If you are planning your own adjudication, that is the transferable result: specify how hard the adjudicator must look, not just what the labels mean. Rater 2 was more favourable to webXray than rater 1 (70.0% of its adjudicable rows current against 50.0%), so the disagreement does not point at the published rates being too kind. See also Interrater agreement for what a kappa of that size means and what it does not.

An entry-shape figure that needs no sample

Some entities' names are just a domain rather than a company — Disconnect's owner for cdnbasket.net is the string “cdnbasket.net”. That can be counted over every pair in every file rather than estimated: it is 0.8% of the 3,215 domains webXray names an owner for, 0.7% of Tracker Radar's 38,368 and 3.2% of Disconnect's 7,850. An entry like that resolves to itself and tells a measurement nothing it did not already have; if your pipeline counts “domains with a known owner”, these should not be in it.

Choosing a resolution source now

Dated, because a list of what the literature did is not advice about what to do:

Source Status — repository state re-checked 2026-09-11, accuracy figures drawn 2026-09-05 and re-adjudicated over the residue 2026-09-11 Use it when
Disconnect entities.json current, but less current than the repository looks: the repo's HEAD is 2026-09-05 and entities.json has not changed since 2026-08-07, five weeks earlier — the 2026-09-05 commit touched services.json. The file is byte-identical (412,191 bytes) at 2026-08-07, at 2026-09-05 and on master today.8) CC BY-NC-SA 4.0; feeds Firefox's tracking protection via Mozilla's shavar-prod-lists9) and, per Disconnect, Microsoft Edge as well you want the most current owner name, and you can live with a sparser list and no hierarchy. Note that properties (5,843 domains) and resources (4,148) are different claims and 2,007 domains appear only in the second; the schema is not documented, so state which key you read. Do not quote Disconnect's own headline “14,332 verified domains and entity mappings” as the size of entities.json — that file holds 7,850 domains; count the file you read.10)
DuckDuckGo Tracker Radar entity_map.json / domain_map.json current; regenerated monthly, last commit on main 2026-08-28 — the commit the newest release tag 2026.08.28 points at. The repository's pushed_at field reads later than that (2026-09-02 on 2026-09-11) because unmerged automation branches count towards it, so read the branch, not the field; CC BY-NC-SA 4.0 you want the broadest coverage, prevalence weights, or per-domain categories and fingerprinting scores in the same dataset. Expect legal-entity names, and expect renames to lag.
Ghostery trackerdb / WhoTracks.me current; trackerdb last commit 2026-09-01, CC BY-NC-SA 4.0 (its package.json says so explicitly; GitHub reports no SPDX id, so do not trust the API field); WhoTracks.me data repo last commit 2026-09-02 (“August update”); the site now redirects to ghostery.com/whotracksme you want an organizations + patterns model where one company can carry several independently categorised behaviours (Google Analytics separate from Google Tag Manager), or the [6Karaj, Arjaldo; Macbeth, Sam; Berson, Rémi; Pujol, Josep M. (2018): "WhoTracks.Me: Shedding light on the opaque world of online tracking". arXiv:1804.08959v2, revised 2019-04-25 (Link)] longitudinal data
webXray domain_owners.json historical; the file is frozen at 2021-03-04 with no update path, and the tool has had no commit since 2023. Measured accuracy on a random sample of its own coverage: 69.2% current (56.1–81.4), 22.0% stale — but it covers only 2.1% of the frame reproducing or extending a pre-2022 result that used it, or you specifically need the parent_id tree or the per-language policy URLs and will re-verify each owner you rely on. Not the list to start a new measurement with: Disconnect is 89.2% current on the same test and Tracker Radar covers eight times as many domains
WHOIS / RDAP current, but see the box below as a fallback for domains no list covers, and only with privacy-proxy filtering
TLS certificates, DNS/SOA, CNAME chains current first-party CDN and sibling-domain detection, which the ownership lists are worst at: [1Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] needed it because those lists “frequently miss connections among two hostnames, e.g., twitch.tv and twitchcdn.net
Crunchbase live but paid: the v4 API returns HTTP 401 without a key you have institutional access; used this way by [7Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)] and [8Cui, Hao; Trimananda, Rahmadi; Markopoulou, Athina; Jordan, Scott (2023): "PoliGraph: Automated Privacy Policy Analysis using Knowledge Graphs", in: Proceedings of the USENIX Security Symposium. (Link)]
LLM-based entity resolution emerging, and only just reaching the domain layer. In this corpus the 2025 work is at the network layer: [4Selmo, Carlos; Carisimo, Esteban; Bustamante, Fabián E.; Alvarez-Hamelin, J. Ignacio (2025): "Learning AS-to-Organization Mappings with Borges", in: Proceedings of the 2025 ACM Internet Measurement Conference, pp. 120-133. (DOI)] maps AS numbers to organisations with an LLM and releases prompts and code; [9Gouda, Deepak; Dainotti, Alberto; Testart, Cecilia (2025): "Prefix2Org: Mapping BGP Prefixes to Organizations", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] maps BGP prefixes. No corpus paper applies this to third-party domain ownership — but outside the seven venues it has started, and the first result is a warning rather than an endorsement: a 2026-06 preprint evaluates four models on domain-to-brand attribution over 36 heavily-phished brands and finds they enumerate a brand's domains at up to 82% precision from memory alone yet fail at ownership verification without external tools, macro F1 at most 0.37, rising by up to 0.65 with WHOIS lookup.11) you are willing to build and validate it yourself. The problem statement transfers directly; the validation burden does too — and on the one published attempt, an unaugmented model is worse than the frozen list this page calls historical

Do not use WHOIS as an ownership database without filtering privacy proxies. [7Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)] tried four sources in order — CrunchBase, webXray's list, TLS certificates, then WHOIS — resolving 411 of its 762 mobile-specific trackers from the first three and a further 251 from WHOIS. Its top-ten table of “organizations” then contains Redacted For Privacy (34), Domains By Proxy (25), Whois Guard (14), Global Domain Privacy Services (8) and Whois Privacy (7). Those five registrar-privacy strings account for 88 trackers, more than the largest real company in the table (Adobe, 48). The paper says so itself: “these five organizations cannot represent the real organizations of those trackers”. [2Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] built a deliberately conservative WHOIS-plus-graph pipeline instead, and still had to note that the registration ecosystem “does not aim at providing full transparency on the organizations behind each domain”.

“Tracker Radar” names three different artefacts, and citing the name tells a reader nothing. Of the 32 corpus papers that mention it, 11 used the ownership dataset, 9 used it as a tracker or category database, and 9 used Tracker Radar Collector — a Puppeteer crawler that carries no ownership data at all and has its own page, Tracker Radar Collector. A third artefact, Tracker Radar Detector, is the build pipeline. Name the artefact, the file and the commit.

Assembling the pipeline

The step most often missing from a paper is not the lookup, it is what surrounds it. To turn a request log into “N% of sites contact Google”:

  1. Extract the request host, and keep the visited site's host beside it.
  2. Fold both to a registrable domain using the ICANN section of a dated PSL. Not the private section — see the traps above. Do not ship a frozen copy of the list [10McQuistin, Stephen; Snyder, Peter; Perkins, Colin; Haddadi, Hamed; Tyson, Gareth (2023): "A First Look at the Privacy Harms of the Public Suffix List", in: Proceedings of the ACM Internet Measurement Conference. (DOI)].
  3. Resolve the request domain to an owner, walking parent labels rather than looking up the exact key, and record which key matched.
  4. Resolve the visited site to an owner too, and drop the request if they are the same owner. This is the step that is almost always silently skipped, and it is what the ownership list is for: [11Wu, Xiaoyuan; Hu, Lydia; Zeng, Eric; Habib, Hana; Bauer, Lujo (2025): "Transparency or Information Overload? Evaluating Users’ Comprehension and Perceptions of the iOS App Privacy Report", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] describes webXray's list as “allowing distinction between first- and third-party domains”. Without it, google.com embedding gstatic.com counts as third-party tracking, and any site whose CDN is on a sibling domain is over-counted. eTLD+1 comparison alone does not do it — that is exactly why [1Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] needed eight person-hours of hand-vetting.
  5. Aggregate per site, not per request, and state the unit: “sites with at least one request to an owner” is not “requests”, is not “domains”, and is not “owners”.
  6. Report the unattributed remainder twice: as a share of domains and as a share of sites or requests. Those differ by more than an order of magnitude here — a list can cover 2.1% of domains and 58.6% of prevalence weight.

Every one of those six steps is a place where two papers measuring “the same thing” diverge, and only the third is about which list you picked.

What to report in a paper

If a result depends on attributing domains to companies, report all of:

  • which list, which file, and which commit or content hash. “We used Tracker Radar” is not reproducible: name build-data/generated/entity_map.json at a commit, or publish the hash. For Disconnect, say whether you read properties, resources or the union.
  • the date the snapshot was taken, separately from the crawl date. Ownership changes between the two.
  • the granularity you resolved to — brand, operating legal entity, or ultimate parent — and, if you walked a hierarchy, how deep and how you handled cycles and missing parents.
  • the fallback order, if you merged sources, and the conflict rule. [2Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] and [7Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)] both used strict priority orders and both said so; that is the standard to meet.
  • how many domains you could not attribute, as a count and as a share of requests or sites, not just of domains. Say which denominator your unattributed share uses.
  • your privacy-proxy filter, if WHOIS was involved, and the list of proxy strings you removed.
  • the eTLD+1 rule and the PSL version, since every one of these lists keys on the registrable domain and the PSL changes.
  • the manual verification you did, its sample size, and the disagreements you found. A prevalence-by-company figure with no manual check is a figure about a list, not about the web. If you adjudicate, say how hard the adjudicator was told to look — that, not the label definitions, is where two raters diverge.

If the claim is “N% of sites contact Google”, the reader cannot evaluate it without the granularity and the snapshot: it silently includes or excludes YouTube, DoubleClick, Bing-adjacent Microsoft properties and whatever Alphabet has bought or sold since the file was written.

Use in publications

Figures below are from the seven venues on Corpus (CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P, 2010–2026, 5,859 extracted papers), not from the web-measurement literature as a whole. 5,855 of the 5,859 have a readable full-text rendering, so every sweep count is a floor.

How widespread is ownership resolution at all? A full-text sweep for the named resources gives 136 papers2.3% of the corpus. 107 of those 136 are among the 1,120 papers that ran a crawl, which is 9.6% of that population. Each row below is an upper bound — a match, not a verified use — except the two that were hand-verified in full:

Resource named in full text Papers Share of 5,859 of which crawled Share of 1,120 Hand-verified?
Public Suffix List 101 1.7% 46 4.1% no — upper bound
Disconnect, in a list/entity sense 74 1.3% 65 5.8% no — upper bound
Tracker Radar (all three artefacts) 32 0.5% 28 2.5% yes, all 32
WhoTracks.me 23 0.4% 19 1.7% no — upper bound
Crunchbase 23 0.4% 11 1.0% no — upper bound
webXray 15 0.3% 13 1.2% yes, all 15

The Public Suffix List is not in the 136: it answers “what is the registrable domain”, not “whose is it”, and adding it takes the union to 221 on a resource crawling papers touch for unrelated reasons. The Disconnect row is the one to be careful with — plain /disconnect/i matches 700 papers and almost none mean the list, because the CSP literature uses the word as a technical term, so the published 74 comes from a tightened pattern. Every probe was run at two widths and both are printed on the provenance page, with what the tightening costs.

It is not growing. The share of each year's papers naming an ownership resource has moved between 1% and 4% since 2016 with no trend this corpus can establish — 2022 against 2023 is 19 of 546 against 19 of 719 (Fisher's exact, p = 0.41), and pooling 2020–2022 against 2023–2026 gives p = 0.11. Read the table as a level, not a curve:

Year Corpus papers Naming an ownership resource Share
2016 182 2 1.1%
2017 231 4 1.7%
2018 254 5 2.0%
2019 402 9 2.2%
2020 404 15 3.7%
2021 379 13 3.4%
2022 546 19 3.5%
2023 719 19 2.6%
2024 690 17 2.5%
2025* 770 21 2.7%
2026* 415 11 2.7%

2010–2015 contributes one paper in total and is omitted from the table rather than padded with zeros. The starred years are the provisional corpus edge — CCS 2026 and IMC 2026 have not been held, and IEEE S&P and TheWebConf 2026 are under-selected by construction — so read the share column for those two rows and not the count.

It is a PETS topic. Read the share, not the count: the venues differ in size by a factor of three.

Venue Corpus papers Naming an ownership resource Share of venue
PETS 510 49 9.6%
IMC 638 17 2.7%
TheWebConf 843 19 2.3%
IEEE S&P 767 11 1.4%
USENIX Security 1,410 19 1.3%
CCS 990 13 1.3%
NDSS 701 8 1.1%

PETS is where this work lands, at more than eight times NDSS's rate (9.6% of 510 against 1.1% of 701). Everywhere else the shares are small, which is the honest headline: ownership resolution is a step inside a paper, not a paper topic, and that is precisely why most papers that do it do not say how.

Which source the literature used, and when

webXray's list was superseded, and the corpus can date it. The last corpus paper to use webXray's crawler or list was published in 2022. Ownership use of Tracker Radar begins in 2021 and rises through the corpus edge — [12Jannett, Louis; Mayer, Andreas; Westers, Maximilian; Mladenov, Vladislav; Mainka, Christian; Schwenk, Jörg (2026): "The State of Passkeys: Studying the Adoption and Security of Passkeys on the Web", in: Proceedings of the USENIX Security Symposium. (Link)] is a 2026 example, using the Entity Map to avoid treating gmail.com and google.com as separate authentication systems:

Year Tracker Radar used for ownership webXray crawler or ownership list
2018 0 1
2019 0 0
2020 0 2
2021 1 1
2022 0 4
2023 2 0
2024 1 0
2025* 4 0
2026* 3 0

The counts are small enough that the crossover is a signal about direction, not a market share, and 2025–2026 are the provisional edge. Combined with the repository and licence findings on webXray, the corpus supports “webXray's ownership list is historical and Tracker Radar's is current practice”; it does not support any claim about which is more accurate, which is what the random sample above is for.

Why these are full-text counts and not schema counts. The extraction's own tool and classification fields name one of the five resources in 89 papers against the full text's 136, because most mentions sit in related work or a reference list rather than in a sentence an extractor reads as a tool. Both signals are reported; neither is “the” population. The overlap, the 54 the schema misses, and the 7 it finds that the sweep does not are itemised on the provenance page.

Three findings that matter more than the counts

Nobody says which snapshot they used. Across the papers that used webXray's crawler or list, 2 of 8 say anything at all about which version, and 1 names a commit. That one, [3Matte, Célestin; Bielova, Nataliia; Santos, Cristiana Teixeira (2020): "Do Cookie Banners Respect my Choice? Measuring Legal Compliance of Banners from IAB Europe's Transparency and Consent Framework", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)], pins both lists it uses in a reproducibility table — “Disconnect list commit eb817fb1 (2019-12-10)” and “WebXRay commit 04c3c8e8 (2019-06-18)” — beside the Chromium build, the kernel, the user agent, the vantage point and the Tranco list id. Copy that table. None of these files has a version field or release tags for the data, so a commit or an archive URL is the only thing that makes the figure reproducible.

No list removes the manual work. [1Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] is the most honest account of the real cost. They needed same-entity relations for the Tranco top 10,000, found that the curated lists “frequently miss connections among two hostnames, e.g., twitch.tv and twitchcdn.net”, mined their own crawl for candidates, and hand-vetted 2,175 candidate site pairs down to 1,146 confirmed same-entity pairs in about eight person-hours. webXray's list then contributed 133 further relations their own method had missed, and their conclusion was that “it alone does not suffice for our purposes”. Two lessons: budget the person-hours, and treat the lists as additive rather than as alternatives. [2Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] took the same route from the other end — it merged Disconnect, WhoTracks.me and webXray by priority (“Disconnect first, WhoTracks.me second, and webxray third”, explicitly “preferring those that have been updated most recently”), added an automated WHOIS-and-graph pass “to increase the coverage”, found exactly one conflict against the 3,913 domains in any of the three lists, and filed a bug report that fixed an error in the Disconnect list. [13Dambra, Savino; Sanchez-Rola, Iskander; Bilge, Leyla; Balzarotti, Davide (2022): "When Sally Met Trackers: Web Tracking From the Users' Perspective", in: Proceedings of the USENIX Security Symposium. (Link)] merged the same three.

Categorisation is a different problem from ownership, and the two get conflated. [14Utz, Christine; Amft, Sabrina; Degeling, Martin; Holz, Thorsten; Fahl, Sascha; Schaub, Florian (2023): "Privacy Rarely Considered: Exploring Considerations in the Adoption of Third-Party Services by Websites", Proceedings on Privacy Enhancing Technologies 2023(1):5-28. (DOI)] is worth reading before you pick a source: it compared five third-party categorisations, including WhoTracks.me and Tracker Radar, and reports that “categorizations differ in granularity and focus” while overlapping substantially — the granularity axis again, stated from inside the literature. If your question is “is this request a tracker?” rather than “whose is it?”, the instrument is a filter list, not an owner database: see Requests.

Other corpus uses worth knowing: [15Musa, Maaz Bin; Nithyanand, Rishab (2022): "ATOM: Ad-network Tomography", in: Proceedings on Privacy Enhancing Technologies. (DOI)] combined webXray's list with WHOIS records and TLS certificates; [16Cassel, Darion; Lin, Su-Chin; Buraggina, Alessio; Wang, William; Zhang, Andrew; Bauer, Lujo; Hsiao, Hsu-Chun; Jia, Limin; Libert, Timothy (2022): "OmniCrawl: Comprehensive Measurement of Web Tracking With Real Desktop and Mobile Browsers", in: Proceedings on Privacy Enhancing Technologies. (DOI)] used it “to determine the provenance of the requests”; [17Kashaf, Aqsa; Sekar, Vyas; Agarwal, Yuvraj (2020): "Analyzing Third Party Service Dependencies in Modern Web Services: Have We Learned from the Mirai-Dyn Incident?", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] cited it but built its own TLD + certificate-SAN + SOA heuristic instead; and [18Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)]'s critique of webXray was of the crawler, not the list.

  • Nothing on this wiki or in this corpus measures Ghostery trackerdb against the other three. It is live, it is the fourth real option, and it is the one gap in the comparison above.
  • webXray's parent_id tree is the only hierarchy any of these files carries. Whether a hierarchy built from a current list — Tracker Radar's entities plus a corporate-structure source — would beat both flat lookup and webXray's frozen tree is unmeasured.
  • The accuracy rates here are conditional on a domain being adjudicable. That was 18–30% of each list's draws until 2026-09-11, when a second pass with harder sourcing settled 28 of the 40 and brought it to 6.7% / 6.7% / 10.0%; it is not zero and the 12 domains still unresolved cannot be bounded from inside the sample. Whether those differ systematically between lists still needs a different instrument.

Methodology and limitations of these figures

  • Every query, every unedited script output, the fold residues, the sources rejected and the figures deliberately not published are on ownership_resolution. Corpus-wide selection and extraction caveats are on Corpus. The full 175-row adjudication table and the scripts behind the accuracy figures are on random_sample, which was written before this page was split out and keeps its id so the links in the published record still resolve.
  • The corpus figures come from scripts/report_ownership_resolution.mjs, which re-derives them independently of scripts/report_webxray.mjs and cross-checks every overlapping row against it; both scripts and their unedited output are on the provenance page. The live-database comparison is scripts/owner_dbs.py; the 28-row hand adjudication is scripts/owner_adjudication.py; the random sample is scripts/owner_sample.py (draw) and scripts/owner_random_sample.py (estimate); the inter-rater re-adjudication is scripts/owner_irr_kappa.py; the second pass over the residue is scripts/owner_tail_probe.sh (evidence), scripts/owner_tail_merge.py (gate) and scripts/report_tail_pass.py (what it changed).
  • The coverage and agreement figures use a single snapshot per list, taken 2026-08-17; the random sample uses a second set taken 2026-09-05, and two of the six files had moved between them. Both live lists move between months — Tracker Radar regenerates monthly, Disconnect commits continuously. Re-run the scripts rather than quoting these numbers in 2027.
  • The coverage denominator is Tracker Radar's own crawl output. There is no neutral census of third-party domains, so this favours Tracker Radar; the page says so wherever a coverage figure appears rather than pretending otherwise.
  • The 28 adjudicated rows are the highest-prevalence disagreements, deliberately not a random sample, so they characterise the shape of disagreement and not any list's accuracy.
  • The random sample is 60 domains per list from that list's own coverage, stratified into prevalence quartiles with allocation 24/12/12/12, seed 20260905. Its limits, in the order they matter: every rate is conditional on the entry being adjudicable, and after the 2026-09-11 second pass 4–6 of each 60 still are not (11–18 before it); the frame is Tracker Radar's domain_summary.json, so no list is scored on domains that crawl never saw; and the encounter-weighted column is a ratio estimator concentrated in a handful of head domains, which is why its intervals reach 68 percentage points wide. 29 of the 525 verdicts were changed after the adjudicators returned them, and the direction is overwhelmingly favourable to the lists — discount accordingly. Every change is listed with its reason on the provenance pages.
  • The corpus is seven venues and omits EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS, so “136 papers” is a statement about those seven. Ownership resolution is also published outside them: webXray's own author placed most of his results in The BMJ, JAMA, Communications of the ACM and New Media and Society, none of which this corpus can see.
  • Full-text counts are paper counts over paper.cols.txt with whitespace collapsed, so a term broken across a PDF column boundary still matches. They are mentions, not verified uses, except where the table says hand-verified.
  • webXray — the tool the frozen list came from: its architecture, its licence history, and why you cannot install it.
  • Tracker Radar Collector — the Puppeteer crawler that shares Tracker Radar's name and carries none of its ownership data.
  • Requests — filter lists, which answer “is this a tracker?” rather than “whose is it?”, and how the two get conflated.
  • Cookies — where cookie attribution needs an owner name, and what the same granularity choice does to it.
  • IP classification — the network-layer version of this problem, including AS-to-organisation mapping, where the LLM work above actually lives.
  • Website classification — topic and industry labels for a site, which is a different question from who owns it.
  • Traffic files — capturing the requests you are about to attribute.
  • Interrater agreement — what the kappa above means, and what it does not.

References

[1]
Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[2]
Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[3]
Matte, Célestin; Bielova, Nataliia; Santos, Cristiana Teixeira (2020): "Do Cookie Banners Respect my Choice? Measuring Legal Compliance of Banners from IAB Europe's Transparency and Consent Framework", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[4]
Selmo, Carlos; Carisimo, Esteban; Bustamante, Fabián E.; Alvarez-Hamelin, J. Ignacio (2025): "Learning AS-to-Organization Mappings with Borges", in: Proceedings of the 2025 ACM Internet Measurement Conference, pp. 120-133. (DOI)
[5]
Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)
[6]
Karaj, Arjaldo; Macbeth, Sam; Berson, Rémi; Pujol, Josep M. (2018): "WhoTracks.Me: Shedding light on the opaque world of online tracking". arXiv:1804.08959v2, revised 2019-04-25 (Link)
[7]
Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[8]
Cui, Hao; Trimananda, Rahmadi; Markopoulou, Athina; Jordan, Scott (2023): "PoliGraph: Automated Privacy Policy Analysis using Knowledge Graphs", in: Proceedings of the USENIX Security Symposium. (Link)
[9]
Gouda, Deepak; Dainotti, Alberto; Testart, Cecilia (2025): "Prefix2Org: Mapping BGP Prefixes to Organizations", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[10]
McQuistin, Stephen; Snyder, Peter; Perkins, Colin; Haddadi, Hamed; Tyson, Gareth (2023): "A First Look at the Privacy Harms of the Public Suffix List", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[11]
Wu, Xiaoyuan; Hu, Lydia; Zeng, Eric; Habib, Hana; Bauer, Lujo (2025): "Transparency or Information Overload? Evaluating Users’ Comprehension and Perceptions of the iOS App Privacy Report", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[12]
Jannett, Louis; Mayer, Andreas; Westers, Maximilian; Mladenov, Vladislav; Mainka, Christian; Schwenk, Jörg (2026): "The State of Passkeys: Studying the Adoption and Security of Passkeys on the Web", in: Proceedings of the USENIX Security Symposium. (Link)
[13]
Dambra, Savino; Sanchez-Rola, Iskander; Bilge, Leyla; Balzarotti, Davide (2022): "When Sally Met Trackers: Web Tracking From the Users' Perspective", in: Proceedings of the USENIX Security Symposium. (Link)
[14]
Utz, Christine; Amft, Sabrina; Degeling, Martin; Holz, Thorsten; Fahl, Sascha; Schaub, Florian (2023): "Privacy Rarely Considered: Exploring Considerations in the Adoption of Third-Party Services by Websites", Proceedings on Privacy Enhancing Technologies 2023(1):5-28. (DOI)
[15]
Musa, Maaz Bin; Nithyanand, Rishab (2022): "ATOM: Ad-network Tomography", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[16]
Cassel, Darion; Lin, Su-Chin; Buraggina, Alessio; Wang, William; Zhang, Andrew; Bauer, Lujo; Hsiao, Hsu-Chun; Jia, Limin; Libert, Timothy (2022): "OmniCrawl: Comprehensive Measurement of Web Tracking With Real Desktop and Mobile Browsers", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[17]
Kashaf, Aqsa; Sekar, Vyas; Agarwal, Yuvraj (2020): "Analyzing Third Party Service Dependencies in Modern Web Services: Have We Learned from the Mirai-Dyn Incident?", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[18]
Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
1)
That sentence is contiguous in the published PDF — checked with pypdf on 2026-09-11 — but the string does not appear in this corpus's column-repaired paper.cols.txt rendering of the paper, which reorders the surrounding text. If you are quote-checking two-column ACM PDFs against a single rendering, this is the failure mode: the two renderings fail on different sentences, and a miss in one is not evidence of a miss in the source.
2)
docs/DATA_MODEL.md in duckduckgo/tracker-radar, checked 2026-08-17.
5)
Verified from the commit history of LICENSE in disconnectme/disconnect-tracking-protection: GPLv3 in the initial commit 89d421e8 (2015-10-13), rewritten to CC BY-NC-SA 4.0 by two commits on 2020-06-24. Checked 2026-08-17.
6)
That sentence is interleaved with a table caption in the paper's two-column layout, so it does not appear as one contiguous string in any text extraction of the PDF. Re-verified 2026-09-11 against paper.cols.txt, paper.norm.txt and paper.txt; every clause is verbatim and in this order, with “Table 1: Third-Party Prevalence, SSL Use, and First-Party Disclosure” and “†Denotes Company has Consumer Services” cutting across it.
7)
This table was corrected on 2026-09-05. Its “No entry” column was wrong in 9 of its 90 verdicts — Disconnect has an entry for all eight rows it was scored absent on, five of them under resources rather than properties, and webXray has “PulsePoint” for contextweb.com. absent is a fact about a file, not a judgement, and the 2026-08-17 script took the adjudicator's word for it; scripts/owner_adjudication_absent_audit.py now re-derives every one from the files and the script refuses to print until it passes.
8)
api.github.com/repos/disconnectme/disconnect-tracking-protection/commits?path=entities.jsonab6ff5a3a8, 2026-08-07T17:07:16Z, “Latest list updates.”; the repo's HEAD 4b592c288a, 2026-09-05T20:21:14Z, “Recent list updates.”, does not touch it. The three raw files fetched and hashed, 2026-09-11. Query the path, not the repo — a repository that commits every few days can carry a file that has not moved in over a month.
9)
mozilla-services/shavar-prod-lists README, checked 2026-08-17: “Firefox's Enhanced Tracking Protection features rely on lists of trackers maintained by Disconnect… Mozilla does not maintain these lists”, and disconnect-blacklist.json is “a version controlled copy of Disconnect's list of trackers”.
10)
https://disconnect.me/trackerprotection, checked 2026-08-17, states “14,332 Verified domains and entity mappings” and that the lists “power Microsoft Edge, Mozilla Firefox, other partners”. The 7,850 figure is measured from entities.json by owner_dbs.py.
11)
Can LLMs Reason About Brand Ownership? An Empirical Study of Domain Attribution Intelligence, Mashood and Nabeel, arXiv:2606.20868v1, submitted 2026-06-18, abstract fetched 2026-09-11. Its task is phishing and squatting defence — is this domain the brand's own? — not third-party tracker attribution, and 36 brands is not a coverage claim. It is cited here as evidence that the domain layer is no longer untouched, not as a method to adopt.
You could leave a comment if you were logged in.
design/ownership_resolution.1789166421.txt.gz · Last modified: by karel.kubicek.claude