User Tools

Site Tools


programming:crawler:webxray

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
programming:crawler:webxray [2026/09/05 19:35] – Per-list ownership accuracy from a stratified random sample (2026-09-05) with intervals and its confounds stated; Wayback dating of webXray's disappearance; four-state licence history; 30-row table corrected (9 wrong absent verdicts); repository state ref karel.kubicek.claudeprogramming:crawler:webxray [2026/09/11 19:20] (current) – webXray's measured rate refreshed after the 2026-09-11 second pass over the residue: 69.9% -> 69.2% (56.1-81.4) on 56 settled entries. Authored by Claude karel.kubicek.claude
Line 3: Line 3:
 webXray {[libert2015_invisible]} is a third-party request and privacy-policy measurement tool by Timothy Libert, and it was one of the three hand-curated lists the pre-2022 literature used to put a **company name** next to a third-party domain — the other two being Disconnect and WhoTracks.me, which two corpus papers merge with it explicitly {[sanchezrola2021_journey]} {[dambra2022_sally]}. Its distinguishing part was never the crawler: it was the **domain-ownership database** — 827 corporate owners, arranged in a parent/child tree, each carrying a purpose, a country, trade-body memberships and privacy-policy URLs in dozens of languages. webXray {[libert2015_invisible]} is a third-party request and privacy-policy measurement tool by Timothy Libert, and it was one of the three hand-curated lists the pre-2022 literature used to put a **company name** next to a third-party domain — the other two being Disconnect and WhoTracks.me, which two corpus papers merge with it explicitly {[sanchezrola2021_journey]} {[dambra2022_sally]}. Its distinguishing part was never the crawler: it was the **domain-ownership database** — 827 corporate owners, arranged in a parent/child tree, each carrying a purpose, a country, trade-body memberships and privacy-policy URLs in dozens of languages.
  
-Two things a new measurement needs to know before citing it:+<WRAP important> 
 +**If your question is "how do I attribute a third-party domain to a company?", this is not the page you want.** That question has its own page — [[Design:Ownership resolution]] — which compares webXray's file against the two live alternatives, censuses what share of the third-party surface each one can name, gives measured accuracy rates from a random sample, and says which source to start a new measurement with. It is dated and it is maintained; this page is about the tool.
  
-  * **The tool is gone from where every paper points, but it is not lost and it did not end up proprietary.** ''github.com/timlib/webXray'' returns HTTP 404 and there has never been a PyPI package — but Libert relicensed webXray twice after its most widely mirrored snapshot, **back** to GPLv3 in 2021 — the licence its own website claimed in 2015 — and then to **MIT in 2023**, and that history survives in a fork. The author now runs a commercial product under the same name at ''webxray.ai''. +This page answers the narrower questions: what webXray was, why you cannot install it, what licence the copy you find is under, and what the corpus actually did with it. 
-  * **The database survives and it is frozen at 2021-03-04. How stale that makes it has now been measured, and the honest answer is "less than this page used to imply, and by an amount this page cannot fully separate from how it was measured".** Drawn at random from its own coverage — 60 domains, adjudicated against primary sources on 2026-09-05 — webXray names today's owner for **69.9%** of the entries that could be settled (95% CI 55.5–83.1%), and for **93.9%** of the third-party surface weighted by how often a crawl meets the domain, though that second figure is concentrated enough that ''gstatic.com'' alone carries 39% of it. The older table of 28 hand-picked //disagreements// on this page reads far worse — 1 of the 24 domains webXray covers there — but **the two numbers are not comparable**, for two reasons at once, and the section below spells both out. What is not in doubt is **coverage**: webXray names an owner for 2.1% of the registrable third-party domains in the frame, against Tracker Radar's 17.2% and Disconnect's 7.0%. Tracker Radar and Disconnect are the live alternatives, and they are not equivalent to each other either.+</WRAP>
  
-This page is therefore two pages in one: what webXray was and why you cannot install it, and — the part you actually need — **how domain-to-company ownership resolution works now**, measured. If your question is "which crawler should I run", go to [[Programming:Crawler]]; if it is "how do I decide which company a request went to", you are in the right place, and the short answer is in [[#Choosing a resolution source now|Choosing a resolution source now]] with the pipeline around it in [[#Assembling the pipeline|Assembling the pipeline]]. Everything between those and here is the evidence for them.+Three things a new measurement needs to know before citing webXray: 
 + 
 +  * **The tool is gone from where every paper points, but it is not lost and it did not end up proprietary.** ''github.com/timlib/webXray'' returns HTTP 404 and there has never been a PyPI package — but Libert relicensed webXray twice after its most widely mirrored snapshot, **back** to GPLv3 in 2021 — the licence its own website claimed in 2015 — and then to **MIT in 2023**, and that history survives in a fork. The author now runs a commercial product under the same name at ''webxray.ai''. 
 +  * **The database survives and it is frozen at 2021-03-04.** How stale that makes it has been measured: on a stratified random sample of its own coverage, drawn 2026-09-05 and re-adjudicated on 2026-09-11 where the first pass could not settle a row, it names today's owner for **69.2%** of the entries that could be settled (95% CI 56.1–81.4%) — a rate over the fifth of its entries a 2026 crawl still meets, not over the file, and one that sits beside a **coverage** figure of **2.1%** of the frame. The adjudicators were language models, the comparison has two snapshot frames, and there are several reasons not to read 69.2% as a rehabilitation. All of that is on [[Design:Ownership resolution]], and this page deliberately does not restate it: take the two figures from there, with their caveats, rather than from this bullet. 
 +  * **The list and the tool came apart in the literature years ago.** Of the 15 corpus papers that name webXray, **7 used only the ownership list** and exactly **one ran the crawler** — and that one is Libert's own paper.
  
 <WRAP important> <WRAP important>
-Do not cite webXray as the tool you used unless you really ran it. Cite {[libert2015_invisible]} or {[libert2018_automated]} for the **method** — third-party request measurement with corporate attribution, and automated privacy-policy auditing — and cite the ownership list separately from the crawler, because in the literature they came apart years ago: of 15 corpus papers that name webXray, **7 used only the ownership list** and exactly **one ran the crawler**, and that one is Libert's own paper.+Do not cite webXray as the tool you used unless you really ran it. Cite {[libert2015_invisible]} or {[libert2018_automated]} for the **method** — third-party request measurement with corporate attribution, and automated privacy-policy auditing — and cite the ownership list separately from the crawler.
 </WRAP> </WRAP>
  
Line 31: Line 36:
 | ''github.com/timlib/webXray'' | **HTTP 404**. The ''timlib'' account itself still exists (HTTP 200) with 0 public repositories. | | ''github.com/timlib/webXray'' | **HTTP 404**. The ''timlib'' account itself still exists (HTTP 200) with 0 public repositories. |
 | ''github.com/timlib/webXray_Domain_Owner_List'' | **HTTP 404**. This is the URL cited by {[kashaf2020_dependencies]}, among others. | | ''github.com/timlib/webXray_Domain_Owner_List'' | **HTTP 404**. This is the URL cited by {[kashaf2020_dependencies]}, among others. |
-| ''webxray.org'' | HTTP 200, but a placeholder: a heading and the line "Public interest projects for the interested public." No source link, no version, no download. Until **2024-03-28** it hosted a live demo search engine over "scans of 100,000 sites", with a company drop-down drawn from the ownership list; by 2024-05-24 that returned ''401 Authorization Required'', and from 2024-07-24 to at least 2026-02-05 the domain 301-redirected to ''webxray.ai''. See [[#Methodology and limitations of these figures|the Wayback timeline below]]. |+| ''webxray.org'' | HTTP 200, but a placeholder: a heading and the line "Public interest projects for the interested public." No source link, no version, no download. Until **2024-03-28** it hosted a live demo search engine over "scans of 100,000 sites", with a company drop-down drawn from the ownership list; by 2024-05-24 that returned ''401 Authorization Required'', and from 2024-07-24 to at least 2026-02-05 the domain 301-redirected to ''webxray.ai''. See [[#What happened to webxray.org, and when|the Wayback timeline below]]. |
 | PyPI ''webxray'' / ''web-xray'' / ''policyxray'' | 404 each. webXray was never packaged on PyPI. | | PyPI ''webxray'' / ''web-xray'' / ''policyxray'' | 404 each. webXray was never packaged on PyPI. |
 | ''github.com/thezedwards/webXray'' | The most widely mirrored copy of webXray 3.x, last commit **2021-03-04**. Its README still instructs ''git clone https://github.com/timlib/webXray.git''. Every figure on this page comes from here — and it is **not** the most complete copy. | | ''github.com/thezedwards/webXray'' | The most widely mirrored copy of webXray 3.x, last commit **2021-03-04**. Its README still instructs ''git clone https://github.com/timlib/webXray.git''. Every figure on this page comes from here — and it is **not** the most complete copy. |
Line 75: Line 80:
 Two features of this schema have no equivalent in Tracker Radar or Disconnect, and are the only reasons to still reach for this file: the **per-language policy URLs** (useful if your study needs a company's German privacy statement, and directly connected to policyXray), and the **''parent_id'' tree**, which lets you resolve at the level your research question wants instead of the level the list happens to record. Two features of this schema have no equivalent in Tracker Radar or Disconnect, and are the only reasons to still reach for this file: the **per-language policy URLs** (useful if your study needs a company's German privacy statement, and directly connected to policyXray), and the **''parent_id'' tree**, which lets you resolve at the level your research question wants instead of the level the list happens to record.
  
-That tree is not cosmetic. It runs up to **six levels deep** (508 owners at the root, 211 at depth 2, 80 at depth 3, 24 at depth 4, 2 each at depths 5 and 6), and for **2,175 of the 3,215 domains (67.7%)** the immediate owner differs from the root of its tree. ''doubleclick.net'' is owned by "DoubleClick", whose parent chain ends at "Alphabet". Tracker Radar says "Google LLC" and Disconnect says "Google". None of the three is wrong; they are answers to three different questions, and §"Two lists disagree" below measures what happens when you forget that.+That tree is not cosmetic. It runs up to **six levels deep** (508 owners at the root, 211 at depth 2, 80 at depth 3, 24 at depth 4, 2 each at depths 5 and 6), and for **2,175 of the 3,215 domains (67.7%)** the immediate owner differs from the root of its tree. ''doubleclick.net'' is owned by "DoubleClick", whose parent chain ends at "Alphabet". Tracker Radar says "Google LLC" and Disconnect says "Google". None of the three is wrong; they are answers to three different questions, and [[Design:Ownership resolution#When two lists disagree|Design:Ownership resolution]] measures what happens when you forget that — resolving webXray to the root of its tree makes agreement with both live lists **worse**, not better, because its roots are historical holding companies.
  
 <WRAP important> <WRAP important>
Line 81: Line 86:
 </WRAP> </WRAP>
  
-===== How it compares to Tracker Radar and Disconnect =====+===== What happened to webxray.org, and when =====
  
-The three lists are not three attempts at the same artefact. They differ in size by more than an order of magnitude, in what a record means, and in what they are licensed for. A fourth live option, Ghostery's ''trackerdb'', is **not** measured here: its ownership data is spread across per-company ''.eno'' files with a separate ''patterns'' layer rather than a single domain→owner map, so putting it in the same table would have meant writing a parser whose choices nobody could check against the other three. That is an omission, not a judgement — if you are choosing among the live lists, this page gives you numbers for two of the three.+The Wayback Machine can date webXray's disappearance, and the dates are **bounds, not events**, because the Archive has no capture inside either window. Every check below was redone on 2026-09-05 after ''web.archive.org'' returned 502/503 all day on 2026-08-17; the CDX queries and the raw ''id_'' captures are on [[provenance:programming:crawler:webxray|the provenance page]].
  
-^ ^ webXray ''domain_owners.json'' ^ DuckDuckGo Tracker Radar ''entity_map.json'' ^ Disconnect ''entities.json'' ^ +^ Artefact ^ Last archived HTTP 200 ^ First archived 404 ^ What that bounds ^ 
-| Owners / entities | 827 | 19,148 | 1,887 | +| ''github.com/timlib/webXray'' | **2023-03-31** | **2023-11-15** | consistent with ''peterjoles/webXray'''s last push on 2023-03-12 — the fork was taken, and then upstream went | 
-| Domains covered | 3,215 | 38,368 | 7,850 | +| ''github.com/timlib/webXray_Domain_Owner_List'' | **2022-10-06** | **2025-01-18** | a much wider window; this is the URL {[kashaf2020_dependencies]} and others cite |
-| Ownership hierarchy | **yes** — ''parent_id'', up to 6 levels | no — flat, though per-entity files carry an ''Owner'' field | no — flat | +
-| Extra per-owner data | purpose, country, trade bodies, per-language policy URLs | ''displayName'', ''aliases''; prevalence, categories, fingerprinting and cookie behaviour in sibling files | ''properties'' vs ''resources'' split; category in ''services.json'' | +
-| How ownership is decided | hand curation. Libert's own paper calls the database "the product of years of detective work" {[libert2018_automated]} | "automatically generated" by Tracker Radar Detector; "new development and bug fixes, other than broken sites, are handled internally"((''docs/DATA_MODEL.md'' in ''duckduckgo/tracker-radar'', checked 2026-08-17.)) | Disconnect's own process page describes step 4 as "Connect every tracker to its parent entity through DNS, WHOIS, and behavioral evidence"((''https://disconnect.me/trackerprotection'', checked 2026-08-17.)) | +
-| Update cadence | **none** — frozen; last public commit 2021-03-04 | monthly regeneration; last commit on ''main'' 2026-08-28, releases tagged ''2026.08.28'', ''2026.07.27'', … | continuous; last commit 2026-08-28 | +
-| Licence | PolyForm Strict 1.0.0 (no redistribution). The 2018 standalone copy is GPL-3.0 | **CC BY-NC-SA 4.0** | **CC BY-NC-SA 4.0** (GPL-3.0 until 2020-06-24) | +
-| Documented method | the 2018 paper, plus the list's own ''notes'' field | ''docs/DATA_MODEL.md'', ''docs/FAQ.md'', a vendor blog post; no paper | a six-step process page and inclusion criteria on ''disconnect.me''; ''entities.json'''s own schema is undocumented | +
-| How to get a correction in | nowhere — the repository is gone | file an issue; the repo says non-breakage development is handled internally | **not by pull request.** The README says "Pull requests are not reviewed and will be closed"; corrections go to ''evaluations@disconnect.me'' and must argue from "publicly available materials and technical information"((''https://disconnect.me/domain_evaluations'', checked 2026-08-17.)) |+
  
-<WRAP important> +''webxray.org'' itself went through four states:
-**Both live lists are non-commercial.** Tracker Radar and Disconnect's list are CC BY-NC-SA 4.0. That is fine for university research and share-alike publication of derived data; it is not fine for a spin-out, a consultancy deliverable, or an industry collaboration without a separate licence, and both vendors offer commercial terms on request. Check this before your data-availability statement, not after. Note also that Disconnect's licence was **GPL-3.0 until 2020-06-24** — a paper that quotes the GPL terms is quoting a licence that no longer applies.((Verified from the commit history of ''LICENSE'' in ''disconnectme/disconnect-tracking-protection'': GPLv3 in the initial commit ''89d421e8'' (2015-10-13), rewritten to CC BY-NC-SA 4.0 by two commits on 2020-06-24. Checked 2026-08-17.)) +
-</WRAP>+
  
-==== Coverage: how much of the third-party surface can each one name? ====+  * **2010-11-08 to at least 2011-09-27** — something else entirely, before Libert: a page titled "Web Xray - Free Website Data Mashup Tool" carrying a "beta 0.5" badge, a public directory of scraped sites with Google AdSense units. (Two separate elements, the ''%%<title>%%'' and a badge beside the logo, not one sentence. The domain's //first// archived capture, 2010-10-08, is a DreamHost parking page.) 
 +  * **2022-01-23 to 2024-03-28** — a **live demo search engine** over "scans of 100,000 sites", with a company drop-down drawn from ''domain_owners.json''. This is the closest thing to a public interface the ownership list ever had, and it is gone. 
 +  * **2024-05-24 to 2026-02-05** — ''401 Authorization Required'', then from 2024-07-24 a 301 redirect to ''webxray.ai''. 
 +  * **by 2026-08-17** — the present placeholder. The Archive has **no capture** of it, so when the domain came back off ''webxray.ai'' can only be bounded to between 2026-02-05 and 2026-08-17.
  
-Size is the wrong comparison, because a domain you never meet costs you nothing. The right one is: **of the third-party domains a crawl actually encounters, weighted by how often it encounters them, what share can this list attribute to a company?** 
- 
-The measurement below takes Tracker Radar's ''domain_summary.json'' as the universe of third-party domains and its ''prevalence'' field as the weight, folds each key to a registrable domain with the **ICANN section** of the current Public Suffix List, and asks each list for an owner. This universe is Tracker Radar's own view of the web, which flatters Tracker Radar and nobody else — read its column as "the list scored on its home ground" and the other two as measured against a denominator they had no part in choosing. 
- 
-^ List ^ Domains it can name an owner for ^ Share of 32,369 ^ Weighted by prevalence ^ 
-^ //Frame: the 2026-08-17 snapshot. The random-sample section below uses the 2026-09-05 snapshot, which is 32,337 domains — see the note there.// ^^^^ 
-| webXray | 669 | 2.1% | **58.6%** | 
-| Tracker Radar | 5,581 | 17.2% | **84.3%** | 
-| Disconnect (''properties'' ∪ ''resources'') | 2,268 | 7.0% | **80.5%** | 
- 
-The 2.1% and the 58.6% in the same row are the whole story of this file. webXray covers almost none of the web by domain count and well over half of it by weight, because it is a **head list**: the few hundred domains it knows are the ones on every page. Sliced by rank, the tail falls off a cliff: 
- 
-^ Slice of the universe, by prevalence ^ webXray ^ Tracker Radar ^ Disconnect ^ 
-| top 100 domains | 71% | 98% | 94% | 
-| top 1,000 domains | 26% | 69% | 68% | 
-| top 10,000 domains | 5% | 29% | 17% | 
- 
-Libert said this himself in 2018, and the sentence is worth having to hand when a reviewer asks about coverage: "because webxray's database of domain ownership primarily contains major ad networks rather than small clients, and policyxray only searches for identified parties, variability in the long-tail of trackers may not have an outsized effect on overall findings related to disclosure. Nonetheless, it is important to point out that the number of parties being searched for is fewer than the total number of parties present." {[libert2018_automated]} 
- 
-Three traps here, and the first one will silently ruin any coverage figure you compute: 
- 
-  * **The Public Suffix List has two sections, and using the wrong one manufactures a coverage hole.** ''googleapis.com'' is a **PRIVATE** PSL rule (6,941 ICANN rules against 3,290 private ones in the snapshot used here). Fold with the private section and ''fonts.googleapis.com'', ''ajax.googleapis.com'' and ''maps.googleapis.com'' each survive as their own "registrable domain"; look each up by exact key and every one comes back unowned — even though **all three lists name ''googleapis.com'' → Google**. ''fonts.googleapis.com'' alone carries prevalence 0.369, so this single mistake moves webXray's weighted coverage from 58.6% to 54.5% and Tracker Radar's from 84.3% to 79.3%. Fold on the **ICANN section**, or walk parent labels at lookup time, or both. The four combinations are printed side by side by the script and the spread between them is larger than the difference between the lists. 
-  * **A large fraction of ''domain_summary.json'''s 47,836 rows are keyed by hostname, and //how// large depends on the fold.** An ICANN-section fold merges **15,651** of them; the ICANN+PRIVATE fold merges only 2,339, because ''googleapis.com'', ''s3.amazonaws.com'' and ''cloudfront.net'' are themselves private-section suffixes. (A label-count heuristic — "three or more labels" — gives 16,396 and answers neither question; do not use it.) All three ownership lists key on the registrable domain, so comparing coverage without folding at all charges webXray and Disconnect for subdomains they were never meant to hold. 
-  * **19 rows are not hostnames at all** — 18 bracketed IPv6 literals and the literal string ''"null"'', the last with a prevalence of 0.024 and a full behaviour profile attached. Drop them explicitly and say you did; do not let a ''null'' domain become a data point. 
- 
-With the fold and the lookup done properly, **29,896 of the 32,369 domains (92.4%) have no owner in either webXray or Disconnect, but only 14.4% of the prevalence weight does**. The most requested domain no list can name is ''tiktokw.us'' at prevalence 0.041 — Tracker Radar attributes it to ByteDance, the other two have nothing. That is what a real coverage hole looks like: recent, mid-tail, and concentrated in domains that appeared after the list was last curated. 
- 
-==== Two lists disagree: is that an error, or a different question? ==== 
- 
-For every domain that two lists both cover, ''owner_dbs.py'' compares the owner strings after normalising punctuation and folding away legal-form suffixes (''Inc'', ''LLC'', ''GmbH'', ''S.A.S'', …) but **no** synonyms — so the "agree" columns are a lower bound on real agreement and the last column an upper bound on real disagreement. The fold is deliberately timid: it leaves **782 of webXray's 827** owner names untouched, and of the 45 it does change only **16** lose a legal suffix — the other 29 change on punctuation alone (''AT&T'', ''JD.com'', "Here, There & Everywhere"). If you want the agreement figures to go up, a synonym table is what you would have to add, and there is no principled one. 
- 
-^ Pair ^ Domains both cover ^ Same after suffix fold ^ One name contains the other ^ Neither ^ 
-| webXray vs Tracker Radar | 612 | 305 (49.8%) | 109 (17.8%) | **198 (32.4%)** | 
-| webXray vs Disconnect | 464 | 213 (45.9%) | 34 (7.3%) | **217 (46.8%)** | 
-| Tracker Radar vs Disconnect | 1,538 | 634 (41.2%) | 320 (20.8%) | **584 (38.0%)** | 
- 
-Between a third and a half of jointly covered domains get different company names. Before concluding that some list is broken, note what happens if you resolve webXray up its ownership tree first — the obvious fix, since webXray says "DoubleClick" where the others say "Google": 
- 
-^ Pair ^ Neither, using webXray's immediate owner ^ Neither, using the root of webXray's tree ^ 
-| vs Tracker Radar | 198 (32.4%) | **239 (39.1%)** | 
-| vs Disconnect | 217 (46.8%) | **230 (49.6%)** | 
- 
-Resolving to the root makes agreement **worse**, and the reason is instructive. webXray's roots are holding companies and, worse, //historical// ones: ''google.com'' resolves to "Alphabet" (against "Google LLC" and "Google"), ''adnxs.com'' to "AT&T", ''yahoo.com'' to "Verizon", ''tapad.com'' to "Telenor", ''turn.com'' to "Singtel". Every one of those was true when the file was last curated and none is true now. So the disagreements decompose into **two independent axes**, and a paper has to state its position on both: 
- 
-  * **Granularity.** Brand (''DoubleClick'', ''Bing'', ''YouTube'') / operating legal entity (''Google LLC'', ''Microsoft Corporation'') / ultimate parent (''Alphabet''). webXray can give you any of the three; Tracker Radar gives the legal entity; Disconnect gives a compact grouping name. "Sites contacting Google" is a different number under each, and the difference is not small: ''bing.com'', ''linkedin.com'' and ''adnxs.com'' all become Microsoft at the entity level. 
-  * **Vintage.** Which snapshot, and when was ownership last checked. Ad tech consolidates continuously — the same problem {[selmo2025_borges]} names at the network layer as "an Internet shaped by constant mergers, rebrandings, and regional variation". 
- 
-==== Reading all three at once ==== 
- 
-The three files have three different shapes, and the shape is where the granularity choice becomes code. webXray is an array of owners each holding a ''domains'' list and a ''parent_id''; Tracker Radar's ''domain_map.json'' is already keyed by domain; Disconnect's ''entities.json'' is keyed by **entity**, so it has to be inverted before you can look a domain up at all. Reading all three at once takes about twenty lines and immediately shows what you are choosing between: 
- 
-<file python owner_lookup.py> 
-import json 
- 
-WX = json.load(open('cache/webxray.json')) 
-TR = json.load(open('cache/tr_domain_map.json')) 
-DC = json.load(open('cache/disconnect_entities.json'))['entities'] 
- 
-# webXray: an array of owners, each with a `domains` list and a `parent_id`. 
-wx_by_id = {o['id']: o for o in WX} 
-wx = {d.lower(): o for o in WX for d in o['domains']} 
- 
-def wx_chain(domain): 
-    o = wx.get(domain) 
-    if o is None: 
-        return None 
-    chain = [o['name']] 
-    seen = {o['id']} 
-    while o['parent_id'] is not None and o['parent_id'] not in seen: 
-        o = wx_by_id[o['parent_id']] 
-        seen.add(o['id']) 
-        chain.append(o['name']) 
-    return chain                      # brand first, ultimate parent last 
- 
-# Tracker Radar: already keyed by domain. 
-tr = {d.lower(): v['entityName'] for d, v in TR.items()} 
- 
-# Disconnect: keyed by ENTITY, so it has to be inverted. `properties` is the 
-# ownership claim; `resources` is what the tracker actually served from. 
-dc = {} 
-for name, v in DC.items(): 
-    for d in v.get('resources', []): 
-        dc.setdefault(d.lower(), name) 
-for name, v in DC.items(): 
-    for d in v.get('properties', []): 
-        dc[d.lower()] = name 
- 
-for d in ['doubleclick.net', 'adnxs.com', 'facebook.net', 'fonts.googleapis.com']: 
-    print(f'{d:24} webXray={wx_chain(d)}  TR={tr.get(d)!r}  Disconnect={dc.get(d)!r}') 
-</file> 
- 
-Its real output on the 2026-08-17 snapshots: 
- 
-<code> 
-doubleclick.net          webXray=['DoubleClick', 'Google', 'Alphabet']  TR='Google LLC'  Disconnect='Google' 
-adnxs.com                webXray=['Xandr', 'AT&T']  TR='Microsoft Corporation'  Disconnect='Microsoft' 
-facebook.net             webXray=['Facebook']  TR='Facebook, Inc.'  Disconnect='Meta' 
-fonts.googleapis.com     webXray=None  TR=None  Disconnect=None 
-</code> 
- 
-Four rows, four different lessons. ''doubleclick.net'' is the granularity axis, with nothing wrong anywhere. ''adnxs.com'' is the vintage axis, with webXray's whole chain superseded. ''facebook.net'' shows the two live lists disagreeing because one carries a five-year-old legal name. And ''fonts.googleapis.com'' comes back empty from all three — **which is a bug in this snippet, not a coverage hole**: every list names ''googleapis.com'' → Google, and the snippet looks up the exact key it was given. That is the PSL private-section trap above, reproduced in twenty lines. Fix it by trying each parent label: 
- 
-<code python> 
-def resolve(host, mapping): 
-    labels = host.split('.') 
-    for i in range(len(labels) - 1):          # stop at two labels 
-        candidate = '.'.join(labels[i:]) 
-        if candidate in mapping: 
-            return mapping[candidate], candidate 
-    return None, None 
-</code> 
- 
-with which ''fonts.googleapis.com'' resolves to Google in all three. Two further details worth copying: the ''seen'' set in ''wx_chain'' is not decoration — walk a hand-curated parent chain without a cycle guard and one bad edge hangs your pipeline; and Disconnect's ''resources'' are loaded before its ''properties'' so the ownership claim wins where the two disagree. 
- 
-==== Which list is right, when they disagree? ==== 
- 
-30 of the highest-prevalence disagreements were adjudicated one by one against primary sources (by Sonnet sub-agents working to a published brief, single-rated — see the methodology section) against primary sources — company newsrooms, SEC filings, or the domain's own legal documents — on 2026-08-17. Every row and source is in ''scripts/owner_adjudication.py'' and on the provenance page. 
- 
-Two of the 30 could not be settled from a primary source and are **excluded from the table**, leaving 28. 
- 
-^ List ^ Names today's owner ^ Stale (a real former owner) ^ Granularity only ^ Outright error ^ No entry ^ 
-| webXray | 1 | **18** | 4 | 1 | 4 | 
-| Tracker Radar | 8 | 12 | 6 | 2 | 0 | 
-| Disconnect | **26** | 0 | 2 | 0 | 0 | 
- 
-<WRAP important> 
-**This is not an accuracy ranking, and it cannot be turned into one.** The 28 rows were chosen //because// the lists disagreed on them, ranked by prevalence; a random sample would be dominated by domains all three get right, so no percentage taken from this table means anything about how often a list is correct in general. Tracker Radar's zero in "No entry" is an artefact of how the sample was drawn rather than a result: the rows were ranked by Tracker Radar's own prevalence field.((This table was **corrected on 2026-09-05**. Its "No entry" column was wrong in 9 of its 90 verdicts — Disconnect has an entry for all eight rows it was scored absent on, five of them under ''resources'' rather than ''properties'', and webXray has "PulsePoint" for ''contextweb.com''. ''absent'' is a fact about a file, not a judgement, and the 2026-08-17 script took the adjudicator's word for it; ''scripts/owner_adjudication_absent_audit.py'' now re-derives every one from the files and the script refuses to print until it passes. The 2026-09-05 pipeline always did this. Every guard on this page had passed, because the page matched the script and the script's data was wrong.)) 
- 
-What the table //does// establish, and what a random sample would show less sharply, is the **shape** of the disagreement. It is concentrated in acquisitions and renames rather than spread across the list; it points overwhelmingly one way; and the list regenerated most conservatively is not the one that is most current. If you need an accuracy rate, you have to draw a stratified random sample and adjudicate it. That is now done, on 2026-09-05, and it is the next section. It gives a very different number for webXray — but **not only because it is a random sample**, and the next section is explicit about the part that is a change of scoring convention rather than a change of sample. 
-</WRAP> 
- 
-Named examples, useful as regression tests for your own pipeline: ''adnxs.com'' (Xandr, AT&T → Microsoft, closed 2022-06-06); ''facebook.com'' (Facebook, Inc. → Meta Platforms, 2021-10-28 — Tracker Radar still says "Facebook, Inc."); ''outbrain.com'' and ''teads.tv'' (Outbrain acquired Teads 2025-02-03 and then took its name, inverting the pair); ''postrelease.com'' (Nativo → Life360, completed 2026-01-05, which only Disconnect has); ''crwdcntrl.net'' (Lotame → Publicis, announced 2025-03-06). Three rows are errors rather than staleness: Tracker Radar attributes ''jsdelivr.net'' to "Prospect One", jsDelivr's infrastructure contractor, where jsDelivr's own data-processing agreement names **Volentio JSD Limited**; Tracker Radar attributes ''stackadapt.com'' to "Collective Roll", which is StackAdapt's own pre-2014 founding name and not a separate owner; and webXray's root for ''1rx.io'' is "Marimedia", which no primary source corroborates. 
- 
-==== How often is each list right? A random sample ==== 
- 
-The table above is what you see when you look **where the lists disagree**. It is 
-not an accuracy rate, and every version of this page has said so. This section is 
-the measurement that was missing until 2026-09-05: **60 registrable domains drawn 
-at random from each list's own coverage**, stratified by prevalence quartile, 
-each adjudicated against a primary source under the same **sourcing** bar as the 30 above — though, as it turns out, not under the same **scoring** convention; see the box below. 
-180 draws over **175 distinct domains** (five were drawn for more than one 
-list); each domain was adjudicated **once** and scored against all three lists, 
-giving 525 verdicts. Only a list's own draws enter that list's rate, and of each 
-list's 60 the ones that could be settled //and// make an ownership claim are 
-**49, 47 and 42** — 138 rows carry every accuracy figure below. 
- 
-The sampler is ''scripts/owner_sample.py'' (seed 20260905, allocation 24/12/12/12 
-over quartiles, snapshots hashed), the estimator ''scripts/owner_random_sample.py'' 
-(stratified, bootstrap intervals). The design, the denominators and the limits 
-are on [[provenance:programming:crawler:webxray|the provenance page]]; both 
-scripts and every unedited output, including the full 175-row adjudication 
-table, are on its [[provenance:programming:crawler:webxray:random_sample|code 
-appendix]]. 
- 
-Two numbers per verdict, because there are two different questions: 
- 
-  * **domain-level** — of the domains this list names an owner for, what share does it name correctly? Every entry weighs the same. 
-  * **encounter-weighted** — when a crawl resolves a third-party request through this list, what share of //resolutions// are right? Weighted by Tracker Radar prevalence — Tracker Radar's own per-domain figure for the share of the sites it crawls that request the domain, so ''gstatic.com'' at 0.40 is met on 40% of them. This is the number a paper's attribution error depends on. It is dominated by a handful of very common domains — ''gstatic.com'' alone carries 39.4% of webXray's — which is why its intervals are three times wider than the domain-level ones. 
- 
-^ List ^ Entries settled, of 60 drawn ^ Names today's owner (domain-level) ^ 95% CI ^ Encounter-weighted ^ 95% CI ^ 
-| webXray | 49 | 69.9% | 55.5–83.1% | **93.9%** | 80.9–98.8% | 
-| Tracker Radar | 47 | 73.7% | 60.8–85.8% | 68.6% | 29.3–97.0% | 
-| Disconnect | 42 | **94.4%** | 86.9–100.0% | 83.1% | 53.7–100.0% | 
- 
-<WRAP tip> 
-**Two frames, 32 domains apart — read the dates, not just the numbers.** The 
-coverage table earlier on this page counts 669 / 5,581 / 2,268 of **32,369** 
-domains; this section counts 664 / 5,566 / 2,264 of **32,337**. Nothing is 
-inconsistent: the earlier figures are the **2026-08-17** snapshot that the 30-row 
-disagreement adjudication was done against, and this section's are the 
-**2026-09-05** snapshot the random sample was drawn from. Tracker Radar's 
-''domain_summary.json'' and the Public Suffix List both changed in the nineteen 
-days between. The coverage //shares// are identical to one decimal (2.1% / 
-17.2% / 7.0%), which is what makes this easy to miss. The earlier figures are 
-deliberately **not** re-derived, because re-deriving them would orphan the 
-adjudication pinned to them. 
-</WRAP> 
- 
-Counting //granularity// (right company, wrong level of the corporate tree) as 
-acceptable rather than wrong — which is the right choice if your unit of analysis 
-is "which company", not "which legal entity" — the domain-level figures become 
-webXray 80.2% (67.7–90.8), Tracker Radar 76.5% (63.5–88.0), Disconnect 98.8% 
-(96.2–100.0). Full per-verdict tables, including //stale// and //error// 
-separately, are in section B of the estimator's output on the appendix page. 
- 
-<WRAP important> 
-**The two figures on this page for webXray — 1 of 24, and 69.9% — differ for two 
-reasons at once, and this page can only separate one of them.** Say so when you 
-cite either. 
- 
-  * **Selection.** The 28 rows are hand-picked //disagreements// ranked by prevalence. A random sample is dominated by domains all three lists get right, so the two sets were never going to agree. This is the reason the page anticipated. 
-  * **Scoring convention, which it did not anticipate.** Seven domains were adjudicated in both sittings against a **bit-identical** copy of the file (''e53760188e6dc9aa''). Of the five settled in both, webXray's verdict differs in **four**, every one of them towards //current//. Two are a tree-level difference — ''lijit.com'' is "Sovrn" at the entry and "Federated Media" at webXray's root, ''simpli.fi'' is "simpli.fi" and "GTCR" — and 2026-08-17 scored the root where 2026-09-05 scored the entry. The other two, ''bidr.io'' and ''crwdcntrl.net'', are a difference in what counts as "the owner": 2026-08-17 named the acquiring group (Comcast, Publicis) and marked the operating brand stale; 2026-09-05 named the operating subsidiary (Beeswax Inc., Epsilon under the Lotame name) and marked the same string current. Only ''smartadserver.com'' agrees. 
- 
-So **the honest reading is not "a random sample rehabilitates webXray"** — it is 
-that the two measurements are on different instruments and the gap between them 
-is not a finding. What **is** a finding is the 69.9% / 93.9% pair itself, which is 
-internally consistent, computed under one convention over one sample, and stated 
-with its intervals. Treat the older table as what it always claimed to be: a 
-description of the //shape// of disagreement, not a rate. 
- 
-The one number here that is a measurement of agreement rather than of a list is 
-uncomfortable and belongs on the page: **webXray verdicts agreed on 1 of the 5 
-domains adjudicated in both sittings.** Both sittings were single-rated. That is 
-the inter-rater evidence this page has, and it is thin. 
- 
-That does **not** make webXray the list to use. Three things it does not fix: 
- 
-  * **Coverage, which is counted rather than estimated.** webXray names an owner for 664 of the 32,337 registrable domains in this section's frame (2.1%), against Tracker Radar's 5,566 (17.2%) and Disconnect's 2,264 (7.0%). A list that is right about the 2% it knows still leaves 98% unattributed. Note also what "its own coverage" means here: those 664 are **a fifth of webXray's 3,215 domain→owner pairs** — the rest name domains a 2026 crawl never met. A domain still requested in 2026 is a survivor, and survivors are where "current" concentrates, so the 69.9% is a rate over the surviving fifth and not over the file. 
-  * **The gap is one-directional and it is growing.** webXray's file has not moved since 2021-03-04. Every rate here degrades from this date forward and from no other cause. 
-  * **Disconnect is better on the measured axis anyway** — 94.4% domain-level against webXray's 69.9%, and on the eligible rows Fisher's exact test gives **p = 0.014** two-sided (39 of 42 against 35 of 49). That is the claim to make, rather than pointing at the gap between the intervals: a percentile bootstrap on 42 rows undercovers at the top, which is why one endpoint reads 100.0. webXray against //Tracker Radar// domain-level, 69.9 against 73.7, is by contrast **not** a difference this sample can establish (p = 0.82). 
- 
-The earlier wording "badly stale wherever it has been checked" was a statement 
-about **where, and how,** it had been checked. Drawing the sample is what made 
-both halves of that visible. 
-</WRAP> 
- 
-Four further readings, each of which needs the sample and could not come from 
-the disagreement table: 
- 
-  * **//error// — a name that was never the owner — is rare in all three lists.** Zero of webXray's 49 settled entries and zero of Disconnect's 42; the only one in Tracker Radar's 47 is ''elfsight.com'', attributed to "Vladimir Fedotov", who is Elfsight's co-founder and not the owning entity (''elfsight.com/terms-of-service/'' names **Elfsight, SL**). A zero cell's bootstrap interval is [0,0] and asserts nothing; the usable statement is the rule-of-three bound — below **6.1%** for webXray and **7.1%** for Disconnect at 95% confidence. Almost all list error is **staleness**, not invention. 
-  * **Tracker Radar is the list whose encounter-weighted accuracy is worst (68.6%), despite being the broadest — and the reason is partly that the three rates are not strictly on one axis.** A brand name survives an acquisition and a legal-entity name does not: for ''crwdcntrl.net'', "Lotame" scores current and "Lotame Solutions, Inc." scores stale, on one fact about one company. A list that carries legal names is therefore penalised by this measure for being more precise. It regenerates monthly but carries legal-entity names, and legal names lag renames: ''zoominfo.com'' is still "Zoom Information, Inc.", ''mountain.com'' "Mountain Digital, Inc." for MNTN, ''20min.ch'' "Tamedia AG" where the group renamed to TX Group, ''tvtime.com'' "Whipclip". The interval is very wide (29.3–97.0%) because a handful of head domains carry the weight — but the point estimate being the lowest of the three, on the list most papers use, is worth knowing. 
-  * **The encounter-weighted column is concentrated enough that one domain moves it.** ''gstatic.com'' alone carries **39.4%** of webXray's encounter-weighted estimate, and webXray's top three drawn domains carry 53.1%; had ''gstatic.com'' been stale, webXray's 93.9% would read 54.4%. The equivalent heaviest rows are ''omtrdc.net'' for Tracker Radar (24.9%, and 68.6% would read 43.7%) and ''id5-sync.com'' for Disconnect (16.7%, 83.1% → 66.4%). That is not a defect in the estimator — it is what "weighted by how often you meet the domain" means when prevalence spans five orders of magnitude — but it is why the encounter-weighted intervals are three times the domain-level ones, and why the encounter-weighted column should be read as an ordering rather than as a number. The concentration figures and the flip test are printed by the estimator itself, in section C2 of its output. 
-  * **A third of the sample could not be settled at all**, and the share rises toward the tail: 11, 13 and 18 of each 60 drawn rows. Those are domains that serve nothing, have no legal document and no newsroom, and only third-party aggregators to go on. Every rate above is therefore **conditional on the entry being adjudicable**, which is a real limit, not a footnote. And the shares are **not** equal — 18%, 22% and **30%** of each list's draws — so if unadjudicable entries are worse than adjudicable ones, the list flattered most is **Disconnect**, the one this page recommends. The worst case, every unresolved entry being stale, gives 58.3% / 58.3% / 65.0%: it narrows the ordering's margin without reversing it. 
- 
-<WRAP tip> 
-**If you cite one figure from this page, cite a coverage figure, not an accuracy 
-figure.** Coverage is a census over the whole frame with no sampling error; 
-accuracy is 42 to 49 adjudicated rows per list and the intervals are ±13 
-percentage points at best. The strongest defensible claims are: no list names an 
-owner for more than **17.2%** of the registrable third-party domains a Tracker 
-Radar crawl sees; and among entries that exist and can be checked, **staleness 
-rather than error** is what goes wrong. 
-</WRAP> 
- 
-One entry-shape figure comes out of the same work and needs no sample at all, 
-because it can be counted over every pair in every file 
-(''scripts/owner_selfname_census.py''): entities whose //name// is just a domain 
-rather than a company — Disconnect's owner for ''cdnbasket.net'' is the string 
-"cdnbasket.net". That is **0.8%** of the 3,215 domains webXray names an owner for, **0.7%** 
-of Tracker Radar's 38,368 and **3.2%** of Disconnect's 7,850. An entry like that 
-resolves to itself and tells a measurement nothing it did not already have; if 
-your pipeline counts "domains with a known owner", these should not be in it. 
- 
-==== Where these figures come from, and how to redo them ==== 
- 
-Every **coverage and agreement** figure on this page — everything above the random-sample section — comes from ''owner_dbs.py''; every **accuracy** figure comes from ''owner_sample.py'' and ''owner_random_sample.py''. All three are **published in full, with their unedited output, on [[provenance:programming:crawler:webxray|the provenance page]]** and its appendix — 1,049 lines between them, and reading them is an audit task rather than a way to learn the topic. Each of the three prints the SHA-256 prefix of the exact bytes it computed from, so a figure quoted from any of them is pinned to a snapshot rather than to "the list"; ''owner_dbs.py'' fetches the lists live and prints the residue of every fold as well. **Re-run it before citing any number here**: two of the three lists move weekly. 
- 
-Two rules in it are easy to get wrong and both change the answer. Do **not** use Tracker Radar's own ''entity_map.json'' or ''domain_map.json'' as the universe — that scores Tracker Radar at 100% by construction, where ''domain_summary.json'' is a crawl //result// and is defensible. And weight a registrable domain by the **largest** prevalence among its hostnames, never the sum, because one site can request both ''fonts.googleapis.com'' and ''ajax.googleapis.com'' and summing double-counts it. 
- 
-===== Choosing a resolution source now ===== 
- 
-Dated, because a list of what the literature //did// is not advice about what to do: 
- 
-^ Source ^ Status — repository state re-checked 2026-09-05, accuracy figures measured 2026-09-05 ^ Use it when ^ 
-| **Disconnect ''entities.json''** | current; last commit 2026-08-28; CC BY-NC-SA 4.0; feeds Firefox's tracking protection via Mozilla's ''shavar-prod-lists''((''mozilla-services/shavar-prod-lists'' README, checked 2026-08-17: "Firefox's Enhanced Tracking Protection features rely on lists of trackers maintained by Disconnect… Mozilla does not maintain these lists", and ''disconnect-blacklist.json'' is "a version controlled copy of Disconnect's list of trackers".)) and, per Disconnect, Microsoft Edge as well | you want the **most current** owner name, and you can live with a sparser list and no hierarchy. Note that ''properties'' (5,843 domains) and ''resources'' (4,148) are different claims and 2,007 domains appear only in the second; the schema is not documented, so state which key you read. Do not quote Disconnect's own headline "14,332 verified domains and entity mappings" as the size of ''entities.json'' — that file holds 7,850 domains; count the file you read.((''https://disconnect.me/trackerprotection'', checked 2026-08-17, states "14,332 Verified domains and entity mappings" and that the lists "power Microsoft Edge, Mozilla Firefox, other partners". The 7,850 figure is measured from ''entities.json'' by ''owner_dbs.py''.)) | 
-| **DuckDuckGo Tracker Radar** ''entity_map.json'' / ''domain_map.json'' | current; regenerated monthly, last commit on ''main'' 2026-08-28 — the commit the newest release tag ''2026.08.28'' points at. The repository's ''pushed_at'' field reads later than that because unmerged automation branches count towards it, so read the branch, not the field; CC BY-NC-SA 4.0 | you want the **broadest** coverage, prevalence weights, or per-domain categories and fingerprinting scores in the same dataset. Expect legal-entity names, and expect renames to lag. | 
-| **Ghostery ''trackerdb''** / WhoTracks.me | current; ''trackerdb'' last commit 2026-09-01, **CC BY-NC-SA 4.0** (its ''package.json'' says so explicitly; GitHub reports no SPDX id, so do not trust the API field); WhoTracks.me data repo last commit 2026-09-02 ("August update"); the site now redirects to ''ghostery.com/whotracksme'' | you want an ''organizations'' + ''patterns'' model where one company can carry several independently categorised behaviours (Google Analytics separate from Google Tag Manager), or the {[karaj2018whotracksme]} longitudinal data | 
-| **webXray ''domain_owners.json''** | **historical**; the file is frozen at 2021-03-04 with no update path, and the tool has had no commit since 2023. Measured accuracy on a random sample of its own coverage: 69.9% current (55.5–83.1), 19.8% stale — but it covers only 2.1% of the frame | reproducing or extending a pre-2022 result that used it, or you specifically need the ''parent_id'' tree or the per-language policy URLs and will re-verify each owner you rely on. Not the list to start a new measurement with: Disconnect is 94.4% current on the same test and Tracker Radar covers eight times as many domains | 
-| **WHOIS / RDAP** | current, but see the trap below | as a **fallback** for domains no list covers, and only with privacy-proxy filtering | 
-| **TLS certificates, DNS/SOA, CNAME chains** | current | first-party CDN and sibling-domain detection, which the ownership lists are worst at: {[steffens2021_blockparty]} needed it because "those lists frequently miss connections among two hostnames, e.g., ''twitch.tv'' and ''twitchcdn.net''" | 
-| **Crunchbase** | live but **paid**: the v4 API returns HTTP 401 without a key | you have institutional access; used this way by {[yang2020_comparative]} and {[cui2023_poligraph]} | 
-| **LLM-based entity resolution** | **emerging, and not yet at the domain layer.** In this corpus the 2025 work is at the network layer: {[selmo2025_borges]} maps AS numbers to organisations with an LLM and releases prompts and code; {[gouda2025_prefix2org]} maps BGP prefixes. No corpus paper applies this to third-party //domain// ownership | you are willing to build and validate it yourself. The problem statement transfers directly; the validation burden does too | 
- 
-<WRAP important> 
-**Do not use WHOIS as an ownership database without filtering privacy proxies.** {[yang2020_comparative]} resolved organisations for 411 of 762 mobile-specific trackers using CrunchBase, webXray's list, TLS certificates and WHOIS in that order — and its top-ten table of "organizations" contains ''Redacted For Privacy'' (34), ''Domains By Proxy'' (25), ''Whois Guard'' (14), ''Global Domain Privacy Services'' (8) and ''Whois Privacy'' (7). Those five registrar-privacy strings account for 88 trackers, more than the largest real company in the table. {[sanchezrola2021_journey]} built a deliberately conservative WHOIS-plus-graph pipeline instead, and still had to note that the registration ecosystem "does not aim at providing full transparency on the organizations behind each domain". 
-</WRAP> 
- 
-<WRAP important> 
-**"Tracker Radar" names three different artefacts**, and citing the name tells a reader nothing. Of the 32 corpus papers that mention it, **11** used the ownership dataset, **9** used it as a tracker or category database, and **9** used **Tracker Radar Collector** — a Puppeteer crawler that carries no ownership data at all and has its own page, [[Programming:Crawler:Tracker Radar Collector]]. A third artefact, Tracker Radar Detector, is the build pipeline. Name the artefact, the file and the commit. 
-</WRAP> 
  
 ===== Use in publications ===== ===== Use in publications =====
Line 377: Line 106:
 Figures below are from the seven venues on [[literature:corpus]] (CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P, 2010–2026), not from the web-measurement literature as a whole. Figures below are from the seven venues on [[literature:corpus]] (CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P, 2010–2026), not from the web-measurement literature as a whole.
  
-**15 papers** name webXray anywhere in their full text — 0.3% of the 5,859-paper corpus, and 1.3% of the 1,120 papers that ran a crawl. The extraction's tool field fires on only 7 of them, because most of the rest cite it in a related-work sentence or a reference list; both signals are reported rather than merged. Every verdict below was read off a quoted sentence.+**15 papers** name webXray anywhere in their full text — 0.3% of the 5,859-paper corpus. **13 of those 15** are among the 1,120 papers that ran a crawl, which is **1.2%** of that population.((Published as "1.3%" until 2026-09-11: the script was dividing the corpus-wide count of 15 by 1,120, a numerator and a denominator from two different populations. See the methodology section.)) The extraction's tool field fires on only 7 of them, because most of the rest cite it in a related-work sentence or a reference list; both signals are reported rather than merged. Every verdict below was read off a quoted sentence.
  
 ^ Role webXray plays ^ Papers ^ Share of 15 ^ ^ Role webXray plays ^ Papers ^ Share of 15 ^
Line 395: Line 124:
 | Yang & Yue, PoPETs 2020, //Web Tracking on Mobile and Desktop// {[yang2020_comparative]} | ownership list (cited as "Tim Libert's library"), **second** of four sources tried in order: CrunchBase, then webXray, then TLS certificates, then WHOIS | | Yang & Yue, PoPETs 2020, //Web Tracking on Mobile and Desktop// {[yang2020_comparative]} | ownership list (cited as "Tim Libert's library"), **second** of four sources tried in order: CrunchBase, then webXray, then TLS certificates, then WHOIS |
 | Steffens et al., NDSS 2021, //Who's Hosting the Block Party?// {[steffens2021_blockparty]} | ownership list, **retrieved from the Internet Archive** because upstream was already hard to obtain | | Steffens et al., NDSS 2021, //Who's Hosting the Block Party?// {[steffens2021_blockparty]} | ownership list, **retrieved from the Internet Archive** because upstream was already hard to obtain |
-| Sánchez-Rola et al., IEEE S&P 2021, //Journey to the Center of the Cookie Ecosystem// {[sanchezrola2021_journey]} | ownership list, ranked **third** of three merged lists. //The corpus files this paper under 2022 and the year table below follows the corpus; Crossref and the paper's own header put it at IEEE S&P 2021.// |+| Sánchez-Rola et al., IEEE S&P 2021, //Journey to the Center of the Cookie Ecosystem// {[sanchezrola2021_journey]} | ownership list, ranked **third** of three merged lists. //The corpus files this paper under 2022; Crossref and the paper's own header put it at IEEE S&P 2021. Year tables derived from this corpus, here and on [[Design:Ownership resolution]], follow the corpus.// |
 | Musa & Nithyanand, PoPETs 2022, //ATOM// {[musa2022_atom]} | ownership list, alongside WHOIS records and TLS certificates | | Musa & Nithyanand, PoPETs 2022, //ATOM// {[musa2022_atom]} | ownership list, alongside WHOIS records and TLS certificates |
 | Cassel et al., PoPETs 2022, //OmniCrawl// {[cassel2022_omnicrawl]} | ownership list, "to determine the provenance of the requests" | | Cassel et al., PoPETs 2022, //OmniCrawl// {[cassel2022_omnicrawl]} | ownership list, "to determine the provenance of the requests" |
Line 406: Line 135:
 Three findings from that table matter more than the counts. Three findings from that table matter more than the counts.
  
-**The tool and the database came apart, and then the database was superseded.** The last corpus paper to use either was published in **2022**. Ownership use of Tracker Radar begins in 2021 and rises through the corpus edge — {[jannett2026_passkeys]} is a 2026 example, using the Entity Map to avoid treating ''gmail.com'' and ''google.com'' as separate authentication systems: +**The tool and the database came apart, and then the database was superseded.** The last corpus paper to use either was published in **2022**. Ownership use of Tracker Radar begins in 2021 and rises through the corpus edge; the year-by-year crossover, and what it does and does not license, is on [[Design:Ownership resolution#Which source the literature used, and when|Design:Ownership resolution]].
- +
-^ Year ^ Tracker Radar used for ownership ^ webXray crawler or ownership list ^ +
-| 2018 | 0 | 1 | +
-| 2019 | 0 | 0 | +
-| 2020 | 0 | 2 | +
-| 2021 | 1 | 1 | +
-| 2022 | 0 | 4 | +
-| 2023 | 2 | 0 | +
-| 2024 | 1 | 0 | +
-| 2025* | 4 | 0 | +
-| 2026* | 3 | 0 | +
- +
-The starred years are the provisional corpus edge — CCS 2026 and IMC 2026 have not been held, and IEEE S&P and TheWebConf 2026 are under-selected by construction — so do not read 2026 as a complete year. The counts are small enough that the crossover is a signal about direction, not a market share. Combined with the licence and 404 findings above, the corpus supports "webXray's ownership list is historical and Tracker Radar's is current practice"; it does not support any claim about which is more accurate, which is what the random-sample section is for. +
- +
-**Nobody says which snapshot they used.** Of the 8 papers that used the crawler or the list, **2** say anything at all about which version, and **1** names a commit. That one, {[matte2020_cookie]}, pins **both** lists it uses in a reproducibility table — "Disconnect list commit eb817fb1 (2019-12-10)" and "WebXRay commit 04c3c8e8 (2019-06-18)" — beside the Chromium build, the kernel, the user agent, the vantage point and the Tranco list id. Copy that table. The file has no version field and no release tags, so a commit or an archive URL is the only thing that makes the figure reproducible; ''owner_dbs.py'' prints a content hash for the same reason. +
- +
-**No list removes the manual work.** {[steffens2021_blockparty]} is the most honest account of the real cost. They needed same-entity relations for the Tranco top 10,000, found that the curated lists "frequently miss connections among two hostnames, e.g., ''twitch.tv'' and ''twitchcdn.net''", mined their own crawl for candidates, and hand-vetted **2,175 candidate site pairs down to 1,146 confirmed same-entity pairs in about eight person-hours**. webXray's list then contributed **133 further relations** their own method had missed, and their conclusion was that "it alone does not suffice for our purposes". Two lessons: budget the person-hours, and treat the lists as //additive// rather than as alternatives. {[sanchezrola2021_journey]} took the same route from the other end — it merged Disconnect, WhoTracks.me and webXray by priority ("Disconnect first, WhoTracks.me second, and webxray third", explicitly "preferring those that have been updated most recently"), added an automated WHOIS-and-graph pass "to increase the coverage", found exactly one conflict against the **3,913** domains in any of the three lists, and filed a bug report that fixed an error in the Disconnect list. +
- +
-**How widespread is ownership resolution at all?** A full-text sweep for the named resources gives **136 papers (2.3% of the corpus; 12.1% of those that crawled)** that mention at least one. Each row is an upper bound — a match, not a verified use — except the two that were hand-verified in full: +
- +
-^ Resource named in full text ^ Papers ^ Hand-verified? ^ +
-| Public Suffix List | 101 | no — upper bound | +
-| Disconnect, in a list/entity sense | 74 | no — upper bound | +
-| Tracker Radar (all three artefacts) | 32 | **yes, all 32** | +
-| WhoTracks.me | 23 | no — upper bound | +
-| Crunchbase | 23 | no — upper bound | +
-| webXray | 15 | **yes, all 15** | +
- +
-{[utz2023_rarely]} is worth reading before you pick one: it compared five third-party categorisations, including WhoTracks.me and Tracker Radar, and reports that "categorizations differ in granularity and focus" while overlapping substantially — the granularity axis again, stated from inside the literature. +
- +
-===== Assembling the pipeline ===== +
- +
-The step most often missing from a paper is not the lookup, it is what surrounds it. To turn a request log into "//N//% of sites contact Google": +
- +
-  - **Extract the request host**, and keep the visited site's host beside it. +
-  - **Fold both to a registrable domain** using the **ICANN section** of a dated PSL. Not the private section — see the trap above. +
-  - **Resolve the request domain to an owner**, walking parent labels rather than looking up the exact key, and record which key matched. +
-  - **Resolve the visited site to an owner too, and drop the request if they are the same owner.** This is the step that is almost always silently skipped, and it is what the ownership list is //for//: {[wu2025_appprivacyreport]} describes webXray's list as "allowing distinction between first- and third-party domains". Without it, ''google.com'' embedding ''gstatic.com'' counts as third-party tracking, and any site whose CDN is on a sibling domain is over-counted. eTLD+1 comparison alone does not do it — that is exactly why {[steffens2021_blockparty]} needed eight person-hours of hand-vetting. +
-  - **Aggregate per site, not per request**, and state the unit: "sites with at least one request to an owner" is not "requests", is not "domains", and is not "owners". +
-  - **Report the unattributed remainder** twice: as a share of domains and as a share of sites or requests. Those differ by more than an order of magnitude here. +
- +
-Every one of those six steps is a place where two papers measuring "the same thing" diverge, and only the third is about which list you picked.+
  
-===== What to report in a paper =====+**Nobody says which snapshot they used.** Of the 8 papers that used the crawler or the list, **2** say anything at all about which version, and **1** names a commit — {[matte2020_cookie]}, which pins "WebXRay commit 04c3c8e8 (2019-06-18)" beside the Disconnect list, the Chromium build, the kernel, the user agent, the vantage point and the Tranco list id. The file has no version field and no release tags, so a commit or an archive URL is the only thing that makes a figure derived from it reproducible. [[Design:Ownership resolution#What to report in a paper|Design:Ownership resolution]] treats that reproducibility table as the standard to meet and says what else belongs in it.
  
-If a result depends on attributing domains to companies, report all of:+**The corpus cannot see most of webXray's impact.** The seven venues omit EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS, and webXray's own author published most of its results outside those seven — his homepage lists The BMJ, JAMA, //Communications of the ACM// and //New Media and Society//((''https://timlibert.me/'', checked 2026-08-17.)). "15 papers" is a statement about seven security and networking venues, not a measure of how much webXray was used.
  
-  * **which list, which file, and which commit or content hash.** "We used Tracker Radar" is not reproducible: name ''build-data/generated/entity_map.json'' at a commit, or publish the hash. For Disconnect, say whether you read ''properties'', ''resources'' or the union. +The wider question — how many corpus papers do domain-to-company attribution **at all**, with what, and whether it is growing — is on [[Design:Ownership resolution#Use in publications|Design:Ownership resolution]]: 136 papers name an ownership resource, 107 of them among the 1,120 that ran a crawl.
-  * **the date the snapshot was taken**, separately from the crawl date. Ownership changes between the two. +
-  * **the granularity you resolved to** — brand, operating legal entity, or ultimate parent — and, if you walked a hierarchy, how deep and how you handled cycles and missing parents. +
-  * **the fallback order**, if you merged sources, and the conflict rule. {[sanchezrola2021_journey]} and {[yang2020_comparative]} both used strict priority orders and both said so; that is the standard to meet. +
-  * **how many domains you could not attribute**, as a count and as a share of //requests or sites//, not just of domains. The two differ enormously: a list can cover 2.1% of domains and 58.6% of prevalence weight. Say which denominator your unattributed share uses. +
-  * **your privacy-proxy filter**, if WHOIS was involved, and the list of proxy strings you removed. +
-  * **the eTLD+1 rule and the PSL version**, since every one of these lists keys on the registrable domain and the PSL changes. Do not ship a frozen copy {[mcquistin2023_psl]}. +
-  * **the manual verification you did**, its sample size, and the disagreements you found. A prevalence-by-company figure with no manual check is a figure about a list, not about the web.+
  
-If the claim is "//N//% of sites contact Google", the reader cannot evaluate it without the granularity and the snapshot: it silently includes or excludes YouTube, DoubleClick, Bing-adjacent Microsoft properties and whatever Alphabet has bought or sold since the file was written. 
  
 ===== Methodology and limitations of these figures ===== ===== Methodology and limitations of these figures =====
  
-  * The corpus audit script is ''scripts/report_webxray.mjs''; the live-database comparison is ''scripts/owner_dbs.py''; the hand adjudication of the 30 disagreements is ''scripts/owner_adjudication.py''; the random sample is ''scripts/owner_sample.py'' (draw) and ''scripts/owner_random_sample.py'' (estimate), with ''owner_merge_rows.py'', ''owner_verify_sources.py'', ''owner_verify_rendered.mjs'', ''owner_corrections.py'', ''owner_verdict_fixes.py'', ''owner_citation_audit.py'' and ''owner_selfname_census.py'' between them. Every query, every unedited output, the fold residues, the sources rejected, the corrections applied and the figures deliberately **not** published are on [[provenance:programming:crawler:webxray|the provenance page]] and its [[provenance:programming:crawler:webxray:random_sample|code appendix]]. Corpus-wide selection and extraction caveats are on [[literature:corpus]].+  * The corpus audit script is ''scripts/report_webxray.mjs'', published in full with its unedited output on [[provenance:programming:crawler:webxray|the provenance page]] along with every query, every fold residue, the sources rejected and the figures deliberately **not** published. Corpus-wide selection and extraction caveats are on [[literature:corpus]]. **The ownership-database comparison, the coverage census and the accuracy rates moved to [[Design:Ownership resolution]] on 2026-09-11**; their scripts (''owner_dbs.py'', ''owner_adjudication.py'', ''owner_sample.py'', ''owner_random_sample.py'', ''owner_irr_kappa.py'') and outputs stay on this page's provenance page and its [[provenance:programming:crawler:webxray:random_sample|code appendix]], whose ids are unchanged so the published record still resolves. [[provenance:design:ownership_resolution]] records the split and carries the new page's own queries.
   * The webXray population is a **full-text sweep**, not the extraction's tool field, because the tool field finds 7 papers where the sweep finds 15 and the difference is exactly the citation-only cases this page needs to separate. All 15 carry a hand verdict with a quoted sentence; the residue between sweep and hand map is zero, and the script fails loudly if that changes.   * The webXray population is a **full-text sweep**, not the extraction's tool field, because the tool field finds 7 papers where the sweep finds 15 and the difference is exactly the citation-only cases this page needs to separate. All 15 carry a hand verdict with a quoted sentence; the residue between sweep and hand map is zero, and the script fails loudly if that changes.
-  * The 26 literal per-paper figures and quotes on this page were checked against each paper's ''paper.cols.txt'' rendering: 26 of 26 located. The 16 extraction evidence quotes behind the tool and classification tuples were checked the same way: 4 exact, 11 partial, 1 below threshold; the below-threshold quote is present in the source and mangled by a column splice, not unsupported. +  * ''report_webxray.mjs'' section G checks **26** literal per-paper figures and quotes, 26 of 26 located against each paper's ''paper.cols.txt'' rendering. Since the 2026-09-11 split those 26 are divided between this page and [[Design:Ownership resolution]] — the script still checks both sets, and the count has not been re-partitioned. The 16 extraction evidence quotes behind the tool and classification tuples were checked the same way: 4 exact, 11 partial, 1 below threshold; the below-threshold quote is present in the source and mangled by a column splice, not unsupported. 
-  * One figure was **not published**: {[steffens2021_blockparty]} contains the sentence "webXray's list does account for 1,096 of our 1,146 found connections meaning that it alone does not suffice for our purposes", whose two halves contradict each other. The published PDF reads the same way, so no coverage percentage is taken from it; the unambiguous parts of that paragraph are quoted above instead. +  * One figure was **not published**: {[steffens2021_blockparty]} contains the sentence "webXray's list does account for 1,096 of our 1,146 found connections meaning that it alone does not suffice for our purposes", whose two halves contradict each other. The published PDF reads the same way, so no coverage percentage is taken from it. 
-  * The coverage and agreement comparison uses a **single snapshot per list, taken on 2026-08-17**; the random sample uses a second set taken on **2026-09-05**, and two of the six files had moved between them. Both sets of SHA-256 prefixes, and what the difference does to the frame, are on the provenance page. Two of the three lists change weekly. Re-run the script rather than quoting these numbers in 2027. +  * **Every repository, licence and URL claim on this page was checked on 2026-08-17 and the Wayback bounds on 2026-09-05.** Repository state changes; re-check before citing. The licence finding in particular was **wrong in this page's first version**, which asserted from one snapshot that webXray "is not open source and redistributing it is prohibited" — true of that snapshot and false of the project's final state. 
-  * The coverage denominator is Tracker Radar's own crawl output. There is no neutral census of third-party domains, so this favours Tracker Radar; the page says so wherever a coverage figure appears rather than pretending otherwise. +  * The **1.2% of the 1,120 crawling papers** above was published as 1.3% until 2026-09-11, when the split found that ''report_webxray.mjs'' was dividing a corpus-wide numerator (15) by the crawled denominator (1,120) instead of counting the 13 sweep hits that are actually in that population. The same bug on a larger figure read 12.1% where the comparable number is 9.6%; both are corrected and the script now computes the intersection.
-  * The 30 adjudicated rows are the highest-prevalence **disagreements**, deliberately not a random sample, so they characterise the shape of disagreement and not any list's accuracy. Two rows (the acquisition date behind ''360yield.com'', and ''fwmrm.net'''s post-2026-spinoff status) could not be settled from a primary source and are recorded as unresolved rather than guessed. +
-  * The **random sample** that does carry accuracy rates is a separate draw made on 2026-09-05: 60 domains per list from that list's own coverage, stratified into prevalence quartiles with allocation 24/12/12/12, seed 20260905. Its limits, in the order they matter: every rate is **conditional on the entry being adjudicable**, and 11–18 of each 60 were not; the frame is Tracker Radar's ''domain_summary.json'', so it is Tracker Radar's view of the third-party surface and no list is scored on domains that crawl never saw; the encounter-weighted column is a ratio estimator concentrated in a handful of head domains — ''gstatic.com'' alone carries 39.4% of webXray's estimate, and its top three carry 53.1% — which is why its intervals reach 68 percentage points wide; and the brief asked for prevalence //deciles// where quartiles were used, because 60 draws over ten strata support no stratum-level statement. **29 of the 525 verdicts were changed after the adjudicators returned them** — 26 by retiring a verdict the brief should never have offered (see the next bullet), 2 because they scored webXray's ownership-tree //root// rather than its entry, and 1 demotion to ''unresolved''. **The direction is overwhelmingly favourable +
-to the lists and you should discount it accordingly**: of the 28 re-scores, 14 +
-landed on ''current'', 2 on ''granularity'', 1 on ''stale'' and 9 on +
-''unknown'' (which removes the row from every rate), and both of the +
-root-scoring fixes made the verdict //less// severe. That is what retiring a +
-verdict which had been firing on correct company names does. Separately, **10 of the 175 rows had their cited source or quote replaced** after a machine re-fetch could not find the quoted sentence on the cited page; 9 of those 10 left the verdicts untouched. Every change is listed with its reason on the provenance page, and after all of them **0 rows are left citing a source a machine could not confirm**. +
-  * A ''self-named'' verdict ("the list's owner is just the domain") was in the adjudication vocabulary and was **retired mid-run**: it fired on "TrustArc" and "Klaviyo", which are the companies. Whether an entity name is a domain is a property of the file, counted over every pair rather than sampled, and it is reported that way above. +
-  * The Wayback checks that failed on 2026-08-17 (''web.archive.org'' was returning 502/503 all day) were redone on 2026-09-05. Both answers are **bounds, not dates**, because the Archive has no capture inside either window: ''github.com/timlib/webXray'' has its last archived HTTP 200 on **2023-03-31** and its first archived 404 on **2023-11-15**, with nothing in between — consistent with ''peterjoles/webXray'''s last push on 2023-03-12, i.e. the fork was taken and then upstream went. ''github.com/timlib/webXray_Domain_Owner_List'' has a much wider window: last 200 **2022-10-06**, first 404 **2025-01-18**. ''webxray.org'' ran a **live demo search engine** over "scans of 100,000 sites" — with a company drop-down drawn from ''domain_owners.json'' — from 2022-01-23 to its last capture on **2024-03-28**; on 2024-05-24 the site returned ''401 Authorization Required'', and from 2024-07-24 to at least 2026-02-05 every capture is a 301 to ''webxray.ai''. The Archive has **no capture of the present placeholder**, so when the domain came back off ''webxray.ai'' can only be bounded to between 2026-02-05 and 2026-08-17. Before webXray, the same domain was something else entirely: from 2010-11-08 to at least 2011-09-27 it served a page titled "Web Xray - Free Website Data Mashup Tool" carrying a "beta 0.5" badge — a public directory of scraped sites with Google AdSense units. (Two separate elements, the ''%%<title>%%'' and a badge beside the logo, not one sentence. The domain's //first// archived capture, 2010-10-08, is a DreamHost parking page.)((CDX queries and the raw ''id_'' captures behind each of these, 2026-09-05, are on the provenance page.)) +
-  * The seven-venue corpus omits EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS, so "15 papers name webXray" is a statement about those seven venues. webXray's own author published most of its results outside those seven venues — his homepage lists The BMJ, JAMA, Communications of the ACM and //New Media and Society//((''https://timlibert.me/'', checked 2026-08-17.)) — so the corpus cannot see the work the tool's reputation mostly rests on, and "15 papers" is not a measure of how much webXray was used.+
  
 ===== Related pages ===== ===== Related pages =====
  
 +  * [[Design:Ownership resolution]] — **the question this tool's database answered**: which source to use now, how much of the web each one covers, and how often each is right.
   * [[Programming:Crawler]] — generic automation libraries and the other specialised crawlers.   * [[Programming:Crawler]] — generic automation libraries and the other specialised crawlers.
   * [[Programming:Crawler:Tracker Radar Collector]] — the Puppeteer crawler that shares Tracker Radar's name and carries none of its ownership data.   * [[Programming:Crawler:Tracker Radar Collector]] — the Puppeteer crawler that shares Tracker Radar's name and carries none of its ownership data.
   * [[Programming:Crawler:OpenWPM]] — the instrument webXray was compared against in 2016, and the one still maintained.   * [[Programming:Crawler:OpenWPM]] — the instrument webXray was compared against in 2016, and the one still maintained.
-  * [[Privacy:Requests]] — filter lists, which answer "is this a tracker?" rather than "whose is it?", and how the two get conflated. +  * [[Privacy:Policies]] — what policyXray did, and what has replaced it. 
-  * [[Privacy:Cookies]] — where cookie attribution needs an owner name, and what the same granularity choice does to it. +  * [[Programming:Traffic files]] — capturing the requests webXray captured.
-  * [[Design:IP classification]] — the network-layer version of this problem, including AS-to-organisation mapping. +
-  * [[Programming:Traffic files]] — capturing the requests you are about to attribute.+
  
 ====== References ====== ====== References ======
programming/crawler/webxray.1788636928.txt.gz · Last modified: by karel.kubicek.claude