| Next revision | Previous revision |
| programming:crawler:webxray [2026/08/17 07:54] – New page: webXray, its domain-ownership database, and a measured three-way comparison with DuckDuckGo Tracker Radar and Disconnect entities.json (coverage, prevalence-weighted coverage, pairwise disagreement, and 30 disagreements adjudicated against prima karel.kubicek.claude | programming:crawler:webxray [2026/09/11 19:20] (current) – webXray's measured rate refreshed after the 2026-09-11 second pass over the residue: 69.9% -> 69.2% (56.1-81.4) on 56 settled entries. Authored by Claude karel.kubicek.claude |
|---|
| webXray {[libert2015_invisible]} is a third-party request and privacy-policy measurement tool by Timothy Libert, and it was one of the three hand-curated lists the pre-2022 literature used to put a **company name** next to a third-party domain — the other two being Disconnect and WhoTracks.me, which two corpus papers merge with it explicitly {[sanchezrola2021_journey]} {[dambra2022_sally]}. Its distinguishing part was never the crawler: it was the **domain-ownership database** — 827 corporate owners, arranged in a parent/child tree, each carrying a purpose, a country, trade-body memberships and privacy-policy URLs in dozens of languages. | webXray {[libert2015_invisible]} is a third-party request and privacy-policy measurement tool by Timothy Libert, and it was one of the three hand-curated lists the pre-2022 literature used to put a **company name** next to a third-party domain — the other two being Disconnect and WhoTracks.me, which two corpus papers merge with it explicitly {[sanchezrola2021_journey]} {[dambra2022_sally]}. Its distinguishing part was never the crawler: it was the **domain-ownership database** — 827 corporate owners, arranged in a parent/child tree, each carrying a purpose, a country, trade-body memberships and privacy-policy URLs in dozens of languages. |
| |
| Two things a new measurement needs to know before citing it: | <WRAP important> |
| | **If your question is "how do I attribute a third-party domain to a company?", this is not the page you want.** That question has its own page — [[Design:Ownership resolution]] — which compares webXray's file against the two live alternatives, censuses what share of the third-party surface each one can name, gives measured accuracy rates from a random sample, and says which source to start a new measurement with. It is dated and it is maintained; this page is about the tool. |
| | |
| | This page answers the narrower questions: what webXray was, why you cannot install it, what licence the copy you find is under, and what the corpus actually did with it. |
| | </WRAP> |
| |
| * **The tool is gone.** ''github.com/timlib/webXray'' returns HTTP 404, there has never been a PyPI package, the surviving copies are dormant, and the licence forbids redistributing it. The author now runs a commercial product under the same name. | Three things a new measurement needs to know before citing webXray: |
| * **The database survives, but it is frozen at 2021 and badly stale wherever it has been checked.** Of 30 high-prevalence ownership disagreements adjudicated against primary sources on 2026-08-17, webXray names today's owner for **3** of the 25 domains it covers at all. Those 30 were selected //because// the lists disagreed on them, so that is not an error rate — but the direction is not in doubt. Tracker Radar and Disconnect are the live alternatives, and they are not equivalent to each other either. | |
| |
| This page is therefore two pages in one: what webXray was and why you cannot install it, and — the part you actually need — **how domain-to-company ownership resolution works now**, measured. If your question is "which crawler should I run", go to [[Programming:Crawler]]; if it is "how do I decide which company a request went to", you are in the right place. | * **The tool is gone from where every paper points, but it is not lost and it did not end up proprietary.** ''github.com/timlib/webXray'' returns HTTP 404 and there has never been a PyPI package — but Libert relicensed webXray twice after its most widely mirrored snapshot, **back** to GPLv3 in 2021 — the licence its own website claimed in 2015 — and then to **MIT in 2023**, and that history survives in a fork. The author now runs a commercial product under the same name at ''webxray.ai''. |
| | * **The database survives and it is frozen at 2021-03-04.** How stale that makes it has been measured: on a stratified random sample of its own coverage, drawn 2026-09-05 and re-adjudicated on 2026-09-11 where the first pass could not settle a row, it names today's owner for **69.2%** of the entries that could be settled (95% CI 56.1–81.4%) — a rate over the fifth of its entries a 2026 crawl still meets, not over the file, and one that sits beside a **coverage** figure of **2.1%** of the frame. The adjudicators were language models, the comparison has two snapshot frames, and there are several reasons not to read 69.2% as a rehabilitation. All of that is on [[Design:Ownership resolution]], and this page deliberately does not restate it: take the two figures from there, with their caveats, rather than from this bullet. |
| | * **The list and the tool came apart in the literature years ago.** Of the 15 corpus papers that name webXray, **7 used only the ownership list** and exactly **one ran the crawler** — and that one is Libert's own paper. |
| |
| <WRAP important> | <WRAP important> |
| Do not cite webXray as the tool you used unless you really ran it. Cite {[libert2015_invisible]} or {[libert2018_automated]} for the **method** — third-party request measurement with corporate attribution, and automated privacy-policy auditing — and cite the ownership list separately from the crawler, because in the literature they came apart years ago: of 15 corpus papers that name webXray, **7 used only the ownership list** and exactly **one ran the crawler**, and that one is Libert's own paper. | Do not cite webXray as the tool you used unless you really ran it. Cite {[libert2015_invisible]} or {[libert2018_automated]} for the **method** — third-party request measurement with corporate attribution, and automated privacy-policy auditing — and cite the ownership list separately from the crawler. |
| </WRAP> | </WRAP> |
| |
| | ''github.com/timlib/webXray'' | **HTTP 404**. The ''timlib'' account itself still exists (HTTP 200) with 0 public repositories. | | | ''github.com/timlib/webXray'' | **HTTP 404**. The ''timlib'' account itself still exists (HTTP 200) with 0 public repositories. | |
| | ''github.com/timlib/webXray_Domain_Owner_List'' | **HTTP 404**. This is the URL cited by {[kashaf2020_dependencies]}, among others. | | | ''github.com/timlib/webXray_Domain_Owner_List'' | **HTTP 404**. This is the URL cited by {[kashaf2020_dependencies]}, among others. | |
| | ''webxray.org'' | HTTP 200, but a placeholder: a heading and the line "Public interest projects for the interested public." No source link, no version, no download. | | | ''webxray.org'' | HTTP 200, but a placeholder: a heading and the line "Public interest projects for the interested public." No source link, no version, no download. Until **2024-03-28** it hosted a live demo search engine over "scans of 100,000 sites", with a company drop-down drawn from the ownership list; by 2024-05-24 that returned ''401 Authorization Required'', and from 2024-07-24 to at least 2026-02-05 the domain 301-redirected to ''webxray.ai''. See [[#What happened to webxray.org, and when|the Wayback timeline below]]. | |
| | PyPI ''webxray'' / ''web-xray'' / ''policyxray'' | 404 each. webXray was never packaged on PyPI. | | | PyPI ''webxray'' / ''web-xray'' / ''policyxray'' | 404 each. webXray was never packaged on PyPI. | |
| | ''github.com/thezedwards/webXray'' | The most complete surviving copy of webXray 3.x, last commit **2021-03-04**. Its README still instructs ''git clone https://github.com/timlib/webXray.git''. | | | ''github.com/thezedwards/webXray'' | The most widely mirrored copy of webXray 3.x, last commit **2021-03-04**. Its README still instructs ''git clone https://github.com/timlib/webXray.git''. Every figure on this page comes from here — and it is **not** the most complete copy. | |
| | the 19 forks of that copy | all dormant; the newest activity anywhere in the network is 2023-03-12, three small commits in a personal working copy.((''api.github.com/repos/thezedwards/webXray/forks?per_page=100'', checked 2026-08-17. Most recent by ''pushed_at'': ''peterjoles/webXray'' 2023-03-12.)) | | | ''github.com/peterjoles/webXray'' | The **most complete** surviving copy: 36 commits ahead of the above, preserving upstream history to 2023-02-01 including Libert's own post-2021 feature work and both relicensing commits. Last push 2023-03-12. This is the copy to take. | |
| | | the fork network | dormant; nothing pushed anywhere since 2023-03-12. ''forks_count'' reports 19 while the forks endpoint returns 20 objects.((''api.github.com/repos/thezedwards/webXray'' and ''.../forks?per_page=100'', checked 2026-08-17.)) | |
| | ''github.com/RDBinns/webXray_Domain_Owner_List'' | The ownership list split out as a standalone, **GPL-3.0** repository, created and last pushed on 2018-04-05. Dormant since, and an older schema than the in-tool copy (''owner_name'' rather than ''name'', and no ''uses'', ''platforms'' or ''trade_groups''). | | | ''github.com/RDBinns/webXray_Domain_Owner_List'' | The ownership list split out as a standalone, **GPL-3.0** repository, created and last pushed on 2018-04-05. Dormant since, and an older schema than the in-tool copy (''owner_name'' rather than ''name'', and no ''uses'', ''platforms'' or ''trade_groups''). | |
| | ''webxray.ai'' | HTTP 200: a **commercial** litigation-support product ("Top US class action law firms and Fortune 100 in-house compliance teams use webXray to find actionable privacy violations first"). Libert's own homepage states "(Dr.) Timothy Libert is founder and CEO of webXray LLC."((''https://webxray.ai/'' and ''https://timlibert.me/'', both fetched 2026-08-17. The ''timlib'' GitHub profile lists ''company: webXray.ai''.)) | | | ''webxray.ai'' | HTTP 200: a **commercial** litigation-support product ("Top US class action law firms and Fortune 100 in-house compliance teams use webXray to find actionable privacy violations first"). Libert's own homepage states "(Dr.) Timothy Libert is founder and CEO of webXray LLC."((''https://webxray.ai/'' and ''https://timlibert.me/'', both fetched 2026-08-17. The ''timlib'' GitHub profile lists ''company: webXray.ai''.)) | |
| |
| <WRAP important> | <WRAP important> |
| **webXray is not open source, and redistributing it is prohibited.** The surviving copy's ''LICENSE.md'' is the **PolyForm Strict License 1.0.0**, which grants use for any noncommercial purpose — explicitly including "public research organization" and "educational institution" — but grants **no** right to distribute the software or to make "changes or new works based on the software". PolyForm's own summary is blunter: Strict "removes permission to distribute copies and make changes, leaving only permission to use for noncommercial purposes". 1.0.0 is still the only version, and note that PolyForm Strict has **no SPDX identifier** — if your artefact metadata expects one, there is none to give.((''https://polyformproject.org/licenses'' and ''.../licenses/strict/1.0.0'', checked 2026-08-17: ''strict/1.0.0'' is the only Strict version listed. SPDX's licence list carries ''PolyForm-Noncommercial-1.0.0'' and ''PolyForm-Small-Business-1.0.0'' but no Strict entry, checked against ''spdx/license-list-data'' on the same date.)) Its README says the same in plainer words: "This software is //not// open source, it is //source available// and licensed for non-commercial use only. You may not distribute webXray in whole or in part or sell data generated by webXray without prior written permission." | **Check the licence of the copy you actually take: webXray has been under three licences in four states, and the surviving copies disagree.** The first state is not in the table below because no repository snapshot carries it: ''webxray.org'' in 2015 described webXray 1.0 as "Free!*" and "(*Subject to terms of the GNU Public License.)", linking the phrase to ''gnu.org/copyleft/gpl.html'' — which served **GPLv3** both then and now.((Wayback capture ''web.archive.org/web/20150530080456/http://webxray.org/'', read 2026-09-05; the quote joins two adjacent ''%%<p>%%'' elements. The linked ''gnu.org/copyleft/gpl.html'' is archived as "Version 3, 29 June 2007" at ''web/20150531125231'' and 302s to ''gnu.org/licenses/gpl-3.0.html'' today. The same page's install instructions end ''git clone https://github.com/timlib/webXray.git''.)) So the sequence is GPLv3 (2015) → PolyForm Strict (by 2021-03) → **back to** GPLv3 (2021-06-14) → MIT (2023-02-01): the 2021 commit titled "Now open-source" was a return, not a first opening, and the restrictive snapshot every mirror carries is the interlude between two free-software states. This is the most misleading thing about webXray's remains, and it caught this page — the first version asserted from one snapshot that webXray "is not open source and redistributing it is prohibited". That is true of that snapshot and false of the project's final state. |
| |
| The practical consequence is not academic. Upstream is a 404, so the only remaining route to the code is a redistribution the licence forbids. Treat webXray as **unobtainable** and do not plan a study around it. Even if you obtain it, it no longer installs cleanly: ''requirements.txt'' pins ''lxml==4.6.2'' and ''psycopg2-binary==2.8.6'', and neither has a PyPI wheel beyond CPython 3.9 — which has been end-of-life since October 2025 — so ''pip install -r requirements.txt'' on a current interpreter falls back to source builds that need matching ''libxml2''/''libxslt'' and PostgreSQL headers.((Checked on PyPI, 2026-08-17: the newest wheels for ''lxml'' 4.6.2 and ''psycopg2-binary'' 2.8.6 are ''cp39''. ''websocket-client'' 0.57.0 and ''textstat'' 0.7.0 ship universal/py3 wheels and are fine.)) The 2018 standalone ownership list at ''RDBinns/webXray_Domain_Owner_List'' is a separate matter: its own README licenses it under **GPLv3**, so that snapshot can be used and redistributed — but it is the 2018 schema, not the 2021 file every figure on this page is measured from. Verify the licence of whatever file you actually download rather than assuming one licence covers the project. | ^ Copy ^ Licence ^ What it permits ^ |
| | | ''thezedwards/webXray'', last push 2021-03-04 — the most widely mirrored snapshot, and the one every figure on this page was computed from | ''LICENSE.md'' is **PolyForm Strict 1.0.0**; its README says "not open source… source available… non-commercial use only" | use for any noncommercial purpose, explicitly including a "public research organization" and an "educational institution" — but **no** right to distribute, or to make "changes or new works based on the software". GitHub reports ''NOASSERTION'', and PolyForm Strict has **no SPDX identifier**, so artefact metadata expecting one has nothing to give | |
| | | the same tree at 2021-06-14, commit ''245ec5d7'' "Update LICENSE.md — Now open-source" | **GPLv3** | free software; several surviving forks still report ''GPL-3.0'' | |
| | | the last state, commit ''73fe0fc9'' of 2023-02-01, authored by Tim Libert and visible today at ''peterjoles/webXray'' | **MIT**, "Copyright (c) 2023 Tim Libert" | anything, including redistribution. Its README still says "GPLv3, open source" — the README lagged the file | |
| | |
| | So webXray ended **MIT-licensed**, and the fork carrying that history is 36 commits ahead of the copy measured here.((''api.github.com/repos/peterjoles/webXray'' reports ''spdx_id: MIT'' and ''pushed_at: 2023-03-12''; ''compare/master...peterjoles:master'' reports ''ahead_by: 36''; the ''LICENSE'' commit ''73fe0fc9'' is authored by "Tim Libert". Checked 2026-08-17. Also checked: ''polyformproject.org/licenses'' lists ''strict/1.0.0'' as the only Strict version, and ''spdx/license-list-data'' carries no Strict entry.)) That does not make webXray live — upstream is still a 404, no fork has been pushed since 2023-03-12, and it still will not install: ''requirements.txt'' pins ''lxml==4.6.2'' and ''psycopg2-binary==2.8.6'', neither of which has a PyPI wheel beyond CPython 3.9, end-of-life since October 2025.((Checked on PyPI, 2026-08-17. ''websocket-client'' 0.57.0 and ''textstat'' 0.7.0 ship universal/py3 wheels and are fine.)) But it changes what you may lawfully do with it, and it means **every figure here is measured from the most restrictively licensed copy in existence**. If you reuse webXray, take the 2023 MIT tree and name the commit. |
| | |
| | The ownership list has its own history: ''RDBinns/webXray_Domain_Owner_List'' is **GPLv3** by its own README, but it is the 2018 schema (''owner_name'', and no ''uses''/''platforms''/''trade_groups''), not the 2021 file measured here. |
| </WRAP> | </WRAP> |
| |
| Two features of this schema have no equivalent in Tracker Radar or Disconnect, and are the only reasons to still reach for this file: the **per-language policy URLs** (useful if your study needs a company's German privacy statement, and directly connected to policyXray), and the **''parent_id'' tree**, which lets you resolve at the level your research question wants instead of the level the list happens to record. | Two features of this schema have no equivalent in Tracker Radar or Disconnect, and are the only reasons to still reach for this file: the **per-language policy URLs** (useful if your study needs a company's German privacy statement, and directly connected to policyXray), and the **''parent_id'' tree**, which lets you resolve at the level your research question wants instead of the level the list happens to record. |
| |
| That tree is not cosmetic. It runs up to **six levels deep** (508 owners at the root, 211 at depth 2, 80 at depth 3, 24 at depth 4, 2 each at depths 5 and 6), and for **2,175 of the 3,215 domains (67.7%)** the immediate owner differs from the root of its tree. ''doubleclick.net'' is owned by "DoubleClick", whose parent chain ends at "Alphabet". Tracker Radar says "Google LLC" and Disconnect says "Google". None of the three is wrong; they are answers to three different questions, and §"Two lists disagree" below measures what happens when you forget that. | That tree is not cosmetic. It runs up to **six levels deep** (508 owners at the root, 211 at depth 2, 80 at depth 3, 24 at depth 4, 2 each at depths 5 and 6), and for **2,175 of the 3,215 domains (67.7%)** the immediate owner differs from the root of its tree. ''doubleclick.net'' is owned by "DoubleClick", whose parent chain ends at "Alphabet". Tracker Radar says "Google LLC" and Disconnect says "Google". None of the three is wrong; they are answers to three different questions, and [[Design:Ownership resolution#When two lists disagree|Design:Ownership resolution]] measures what happens when you forget that — resolving webXray to the root of its tree makes agreement with both live lists **worse**, not better, because its roots are historical holding companies. |
| |
| <WRAP important> | <WRAP important> |
| </WRAP> | </WRAP> |
| |
| ===== How it compares to Tracker Radar and Disconnect ===== | ===== What happened to webxray.org, and when ===== |
| |
| The three lists are not three attempts at the same artefact. They differ in size by more than an order of magnitude, in what a record means, and in what they are licensed for. | The Wayback Machine can date webXray's disappearance, and the dates are **bounds, not events**, because the Archive has no capture inside either window. Every check below was redone on 2026-09-05 after ''web.archive.org'' returned 502/503 all day on 2026-08-17; the CDX queries and the raw ''id_'' captures are on [[provenance:programming:crawler:webxray|the provenance page]]. |
| |
| ^ ^ webXray ''domain_owners.json'' ^ DuckDuckGo Tracker Radar ''entity_map.json'' ^ Disconnect ''entities.json'' ^ | ^ Artefact ^ Last archived HTTP 200 ^ First archived 404 ^ What that bounds ^ |
| | Owners / entities | 827 | 19,148 | 1,887 | | | ''github.com/timlib/webXray'' | **2023-03-31** | **2023-11-15** | consistent with ''peterjoles/webXray'''s last push on 2023-03-12 — the fork was taken, and then upstream went | |
| | Domains covered | 3,215 | 38,368 | 7,850 | | | ''github.com/timlib/webXray_Domain_Owner_List'' | **2022-10-06** | **2025-01-18** | a much wider window; this is the URL {[kashaf2020_dependencies]} and others cite | |
| | Ownership hierarchy | **yes** — ''parent_id'', up to 6 levels | no — flat, though per-entity files carry an ''Owner'' field | no — flat | | |
| | Extra per-owner data | purpose, country, trade bodies, per-language policy URLs | ''displayName'', ''aliases''; prevalence, categories, fingerprinting and cookie behaviour in sibling files | ''properties'' vs ''resources'' split; category in ''services.json'' | | |
| | How ownership is decided | hand curation. Libert's own paper calls the database "the product of years of detective work" {[libert2018_automated]} | "automatically generated" by Tracker Radar Detector; "new development and bug fixes, other than broken sites, are handled internally"((''docs/DATA_MODEL.md'' in ''duckduckgo/tracker-radar'', checked 2026-08-17.)) | Disconnect's own process page describes step 4 as "Connect every tracker to its parent entity through DNS, WHOIS, and behavioral evidence"((''https://disconnect.me/trackerprotection'', checked 2026-08-17.)) | | |
| | Update cadence | **none** — frozen; last public commit 2021-03-04 | monthly regeneration; last commit 2026-08-12, releases tagged ''2026.07.27'', ''2026.06.08'', … | continuous; last commit 2026-08-07 | | |
| | Licence | PolyForm Strict 1.0.0 (no redistribution). The 2018 standalone copy is GPL-3.0 | **CC BY-NC-SA 4.0** | **CC BY-NC-SA 4.0** (GPL-3.0 until 2020-06-24) | | |
| | Documented method | the 2018 paper, plus the list's own ''notes'' field | ''docs/DATA_MODEL.md'', ''docs/FAQ.md'', a vendor blog post; no paper | a six-step process page and inclusion criteria on ''disconnect.me''; ''entities.json'''s own schema is undocumented | | |
| | How to get a correction in | nowhere — the repository is gone | file an issue; the repo says non-breakage development is handled internally | **not by pull request.** The README says "Pull requests are not reviewed and will be closed"; corrections go to ''evaluations@disconnect.me'' and must argue from "publicly available materials and technical information"((''https://disconnect.me/domain_evaluations'', checked 2026-08-17.)) | | |
| |
| <WRAP important> | ''webxray.org'' itself went through four states: |
| **Both live lists are non-commercial.** Tracker Radar and Disconnect's list are CC BY-NC-SA 4.0. That is fine for university research and share-alike publication of derived data; it is not fine for a spin-out, a consultancy deliverable, or an industry collaboration without a separate licence, and both vendors offer commercial terms on request. Check this before your data-availability statement, not after. Note also that Disconnect's licence was **GPL-3.0 until 2020-06-24** — a paper that quotes the GPL terms is quoting a licence that no longer applies.((Verified from the commit history of ''LICENSE'' in ''disconnectme/disconnect-tracking-protection'': GPLv3 in the initial commit ''89d421e8'' (2015-10-13), rewritten to CC BY-NC-SA 4.0 by two commits on 2020-06-24. Checked 2026-08-17.)) | |
| </WRAP> | |
| |
| ==== Coverage: how much of the third-party surface can each one name? ==== | * **2010-11-08 to at least 2011-09-27** — something else entirely, before Libert: a page titled "Web Xray - Free Website Data Mashup Tool" carrying a "beta 0.5" badge, a public directory of scraped sites with Google AdSense units. (Two separate elements, the ''%%<title>%%'' and a badge beside the logo, not one sentence. The domain's //first// archived capture, 2010-10-08, is a DreamHost parking page.) |
| | * **2022-01-23 to 2024-03-28** — a **live demo search engine** over "scans of 100,000 sites", with a company drop-down drawn from ''domain_owners.json''. This is the closest thing to a public interface the ownership list ever had, and it is gone. |
| | * **2024-05-24 to 2026-02-05** — ''401 Authorization Required'', then from 2024-07-24 a 301 redirect to ''webxray.ai''. |
| | * **by 2026-08-17** — the present placeholder. The Archive has **no capture** of it, so when the domain came back off ''webxray.ai'' can only be bounded to between 2026-02-05 and 2026-08-17. |
| |
| Size is the wrong comparison, because a domain you never meet costs you nothing. The right one is: **of the third-party domains a crawl actually encounters, weighted by how often it encounters them, what share can this list attribute to a company?** | |
| |
| The measurement below takes Tracker Radar's ''domain_summary.json'' as the universe of third-party domains and its ''prevalence'' field as the weight, folds each key to a registrable domain with the current Public Suffix List, and asks each list for an owner. This universe is Tracker Radar's own view of the web, which flatters Tracker Radar and nobody else — read its column as "the list scored on its home ground" and the other two as measured against a denominator they had no part in choosing. | |
| |
| ^ List ^ Domains it can name an owner for ^ Share of 45,525 ^ Weighted by prevalence ^ | |
| | webXray | 657 | 1.4% | **54.5%** | | |
| | Tracker Radar | 5,539 | 12.2% | **79.3%** | | |
| | Disconnect (''properties'' ∪ ''resources'') | 2,277 | 5.0% | **75.8%** | | |
| |
| The 1.4% and the 54.5% in the same row are the whole story of this file. webXray covers almost none of the web by domain count and more than half of it by weight, because it is a **head list**: the few hundred domains it knows are the ones on every page. Sliced by rank, the tail falls off a cliff: | |
| |
| ^ Slice of the universe, by prevalence ^ webXray ^ Tracker Radar ^ Disconnect ^ | |
| | top 100 domains | 70% | 95% | 91% | | |
| | top 1,000 domains | 26% | 67% | 67% | | |
| | top 10,000 domains | 5% | 27% | 16% | | |
| |
| Libert said this himself in 2018, and the sentence is worth having to hand when a reviewer asks about coverage: "because webxray's database of domain ownership primarily contains major ad networks rather than small clients, and policyxray only searches for identified parties, variability in the long-tail of trackers may not have an outsized effect on overall findings related to disclosure. Nonetheless, it is important to point out that the number of parties being searched for is fewer than the total number of parties present." {[libert2018_automated]} | |
| |
| Two traps in the denominator, both of which will bite anyone who repeats this measurement: | |
| |
| * **Tracker Radar's ''domain_summary.json'' is keyed by //hostname// for 16,396 of its 47,836 rows**, so ''fonts.googleapis.com'', ''ajax.googleapis.com'' and ''maps.googleapis.com'' are separate rows while ''googleapis.com'' is not one at all. All three ownership lists key on the registrable domain, so comparing coverage without folding to eTLD+1 charges webXray and Disconnect for subdomains they were never meant to hold. After folding, ''fonts.googleapis.com'' remains the single most prevalent third-party name in the data (0.369) with **no** owner in any of the three lists. | |
| * **19 rows are not hostnames at all** — 18 bracketed IPv6 literals and the literal string ''"null"'', the last with a prevalence of 0.024 and a full behaviour profile attached. Drop them explicitly and say you did; do not let a ''null'' domain become a data point. | |
| |
| ==== Two lists disagree: is that an error, or a different question? ==== | |
| |
| For every domain that two lists both cover, ''owner_dbs.py'' compares the owner strings after folding away legal-form suffixes (''Inc'', ''LLC'', ''GmbH'', ''S.A.S'', …) but **no** synonyms — so the "agree" columns are a lower bound on real agreement and the last column an upper bound on real disagreement. | |
| |
| ^ Pair ^ Domains both cover ^ Same after suffix fold ^ One name contains the other ^ Neither ^ | |
| | webXray vs Tracker Radar | 601 | 304 (50.6%) | 100 (16.6%) | **197 (32.8%)** | | |
| | webXray vs Disconnect | 461 | 212 (46.0%) | 33 (7.2%) | **216 (46.9%)** | | |
| | Tracker Radar vs Disconnect | 1,532 | 630 (41.1%) | 318 (20.8%) | **584 (38.1%)** | | |
| |
| Between a third and a half of jointly covered domains get different company names. Before concluding that some list is broken, note what happens if you resolve webXray up its ownership tree first — the obvious fix, since webXray says "DoubleClick" where the others say "Google": | |
| |
| ^ Pair ^ Neither, using webXray's immediate owner ^ Neither, using the root of webXray's tree ^ | |
| | vs Tracker Radar | 197 (32.8%) | **236 (39.3%)** | | |
| | vs Disconnect | 216 (46.9%) | **228 (49.5%)** | | |
| |
| Resolving to the root makes agreement **worse**, and the reason is instructive. webXray's roots are holding companies and, worse, //historical// ones: ''google.com'' resolves to "Alphabet" (against "Google LLC" and "Google"), ''adnxs.com'' to "AT&T", ''yahoo.com'' to "Verizon", ''tapad.com'' to "Telenor", ''turn.com'' to "Singtel". Every one of those was true when the file was last curated and none is true now. So the disagreements decompose into **two independent axes**, and a paper has to state its position on both: | |
| |
| * **Granularity.** Brand (''DoubleClick'', ''Bing'', ''YouTube'') / operating legal entity (''Google LLC'', ''Microsoft Corporation'') / ultimate parent (''Alphabet''). webXray can give you any of the three; Tracker Radar gives the legal entity; Disconnect gives a compact grouping name. "Sites contacting Google" is a different number under each, and the difference is not small: ''bing.com'', ''linkedin.com'' and ''adnxs.com'' all become Microsoft at the entity level. | |
| * **Vintage.** Which snapshot, and when was ownership last checked. Ad tech consolidates continuously — the same problem {[selmo2025_borges]} names at the network layer as "an Internet shaped by constant mergers, rebrandings, and regional variation". | |
| |
| ==== Reading all three at once ==== | |
| |
| The three files have three different shapes, and the shape is where the granularity choice becomes code. webXray is an array of owners each holding a ''domains'' list and a ''parent_id''; Tracker Radar's ''domain_map.json'' is already keyed by domain; Disconnect's ''entities.json'' is keyed by **entity**, so it has to be inverted before you can look a domain up at all. Reading all three at once takes about twenty lines and immediately shows what you are choosing between: | |
| |
| <file python owner_lookup.py> | |
| import json | |
| |
| WX = json.load(open('cache/webxray.json')) | |
| TR = json.load(open('cache/tr_domain_map.json')) | |
| DC = json.load(open('cache/disconnect_entities.json'))['entities'] | |
| |
| # webXray: an array of owners, each with a `domains` list and a `parent_id`. | |
| wx_by_id = {o['id']: o for o in WX} | |
| wx = {d.lower(): o for o in WX for d in o['domains']} | |
| |
| def wx_chain(domain): | |
| o = wx.get(domain) | |
| if o is None: | |
| return None | |
| chain = [o['name']] | |
| seen = {o['id']} | |
| while o['parent_id'] is not None and o['parent_id'] not in seen: | |
| o = wx_by_id[o['parent_id']] | |
| seen.add(o['id']) | |
| chain.append(o['name']) | |
| return chain # brand first, ultimate parent last | |
| |
| # Tracker Radar: already keyed by domain. | |
| tr = {d.lower(): v['entityName'] for d, v in TR.items()} | |
| |
| # Disconnect: keyed by ENTITY, so it has to be inverted. `properties` is the | |
| # ownership claim; `resources` is what the tracker actually served from. | |
| dc = {} | |
| for name, v in DC.items(): | |
| for d in v.get('resources', []): | |
| dc.setdefault(d.lower(), name) | |
| for name, v in DC.items(): | |
| for d in v.get('properties', []): | |
| dc[d.lower()] = name | |
| |
| for d in ['doubleclick.net', 'adnxs.com', 'facebook.net', 'fonts.googleapis.com']: | |
| print(f'{d:24} webXray={wx_chain(d)} TR={tr.get(d)!r} Disconnect={dc.get(d)!r}') | |
| </file> | |
| |
| Its real output on the 2026-08-17 snapshots: | |
| |
| <code> | |
| doubleclick.net webXray=['DoubleClick', 'Google', 'Alphabet'] TR='Google LLC' Disconnect='Google' | |
| adnxs.com webXray=['Xandr', 'AT&T'] TR='Microsoft Corporation' Disconnect='Microsoft' | |
| facebook.net webXray=['Facebook'] TR='Facebook, Inc.' Disconnect='Meta' | |
| fonts.googleapis.com webXray=None TR=None Disconnect=None | |
| </code> | |
| |
| Four rows, four different lessons: ''doubleclick.net'' is the granularity axis with nothing wrong anywhere; ''adnxs.com'' is the vintage axis, with webXray's whole chain superseded; ''facebook.net'' shows the two live lists disagreeing because one carries a five-year-old legal name; and ''fonts.googleapis.com'' is the single most requested third-party name in the data with **no owner in any list**, which is what a coverage figure looks like from the inside. The ''seen'' set in ''wx_chain'' is not defensive decoration — walk a hand-curated parent chain without a cycle guard and a single bad edge hangs your pipeline. | |
| |
| ==== Which list is right, when they disagree? ==== | |
| |
| 30 of the highest-prevalence disagreements were adjudicated by hand against primary sources — company newsrooms, SEC filings, or the domain's own legal documents — on 2026-08-17. Every row and source is in ''scripts/owner_adjudication.py'' and on the provenance page. | |
| |
| ^ List ^ Names today's owner ^ Stale (a real former owner) ^ Granularity only ^ Outright error ^ No entry ^ | |
| | webXray | 3 | **17** | 4 | 1 | 5 | | |
| | Tracker Radar | 9 | 12 | 7 | 2 | 0 | | |
| | Disconnect | **22** | 0 | 0 | 0 | 8 | | |
| |
| Read as a share of each list's own entries: webXray is current on **3 of 25 (12%)**, Tracker Radar on **9 of 30 (30%)**, Disconnect on **22 of 22 (100%)**. | |
| |
| <WRAP important> | |
| These 30 rows were chosen **because** the lists disagreed on them, ranked by prevalence. They are not a random sample, so the table above is not an error rate for any list — a random sample would be dominated by domains all three get right. What it does establish, and what a random sample would not show as sharply, is the //shape// of the disagreement: it is concentrated in acquisitions and renames, it points overwhelmingly one way, and the list that is regenerated most conservatively is not the one that is most current. | |
| </WRAP> | |
| |
| Named examples, useful as regression tests for your own pipeline: ''adnxs.com'' (Xandr, AT&T → Microsoft, closed 2022-06-06); ''facebook.com'' (Facebook, Inc. → Meta Platforms, 2021-10-28 — Tracker Radar still says "Facebook, Inc."); ''outbrain.com'' and ''teads.tv'' (Outbrain acquired Teads 2025-02-03 and then took its name, inverting the pair); ''postrelease.com'' (Nativo → Life360, completed 2026-01-05, which only Disconnect has); ''crwdcntrl.net'' (Lotame → Publicis, announced 2025-03-06). Three rows are errors rather than staleness: Tracker Radar attributes ''jsdelivr.net'' to "Prospect One", jsDelivr's infrastructure contractor, where jsDelivr's own data-processing agreement names **Volentio JSD Limited**; Tracker Radar attributes ''stackadapt.com'' to "Collective Roll", which is StackAdapt's own pre-2014 founding name and not a separate owner; and webXray's root for ''1rx.io'' is "Marimedia", which no primary source corroborates. | |
| |
| ==== Where these figures come from, and how to redo them ==== | |
| |
| Every figure in the two sections above comes from ''owner_dbs.py'', which is **published in full, with its unedited output, on [[provenance:programming:crawler:webxray|the provenance page]]** — 396 lines, and reading it is an audit task rather than a way to learn the topic. It fetches each list live, prints the SHA-256 prefix of the exact bytes it computed from, and prints the residue of every fold, so a figure quoted from it is pinned to a snapshot rather than to "the list". **Re-run it before citing any number here**: two of the three lists move weekly. | |
| |
| Two rules in it are easy to get wrong and both change the answer. Do **not** use Tracker Radar's own ''entity_map.json'' or ''domain_map.json'' as the universe — that scores Tracker Radar at 100% by construction, where ''domain_summary.json'' is a crawl //result// and is defensible. And weight a registrable domain by the **largest** prevalence among its hostnames, never the sum, because one site can request both ''fonts.googleapis.com'' and ''ajax.googleapis.com'' and summing double-counts it. | |
| |
| ===== Choosing a resolution source now ===== | |
| |
| Dated, because a list of what the literature //did// is not advice about what to do: | |
| |
| ^ Source ^ Status, 2026-08-17 ^ Use it when ^ | |
| | **Disconnect ''entities.json''** | current; last commit 2026-08-07; CC BY-NC-SA 4.0; feeds Firefox's tracking protection via Mozilla's ''shavar-prod-lists''((''mozilla-services/shavar-prod-lists'' README, checked 2026-08-17: "Firefox's Enhanced Tracking Protection features rely on lists of trackers maintained by Disconnect… Mozilla does not maintain these lists", and ''disconnect-blacklist.json'' is "a version controlled copy of Disconnect's list of trackers".)) and, per Disconnect, Microsoft Edge as well | you want the **most current** owner name, and you can live with a sparser list and no hierarchy. Note that ''properties'' (5,843 domains) and ''resources'' (4,148) are different claims and 2,007 domains appear only in the second; the schema is not documented, so state which key you read. Do not quote Disconnect's own headline "14,332 verified domains and entity mappings" as the size of ''entities.json'' — that file holds 7,850 domains; count the file you read.((''https://disconnect.me/trackerprotection'', checked 2026-08-17, states "14,332 Verified domains and entity mappings" and that the lists "power Microsoft Edge, Mozilla Firefox, other partners". The 7,850 figure is measured from ''entities.json'' by ''owner_dbs.py''.)) | | |
| | **DuckDuckGo Tracker Radar** ''entity_map.json'' / ''domain_map.json'' | current; regenerated monthly, last commit 2026-08-12; CC BY-NC-SA 4.0 | you want the **broadest** coverage, prevalence weights, or per-domain categories and fingerprinting scores in the same dataset. Expect legal-entity names, and expect renames to lag. | | |
| | **Ghostery ''trackerdb''** / WhoTracks.me | current; ''trackerdb'' last commit 2026-08-06, **CC BY-NC-SA 4.0** (its ''package.json'' says so explicitly; GitHub reports no SPDX id, so do not trust the API field); WhoTracks.me data repo last commit 2026-08-04 ("July update"); the site now redirects to ''ghostery.com/whotracksme'' | you want an ''organizations'' + ''patterns'' model where one company can carry several independently categorised behaviours (Google Analytics separate from Google Tag Manager), or the {[karaj2018whotracksme]} longitudinal data | | |
| | **webXray ''domain_owners.json''** | **historical**; frozen at 2021-03-04, no update path, tool unobtainable | reproducing or extending a pre-2022 result that used it, or you specifically need the ''parent_id'' tree or the per-language policy URLs and will re-verify each owner you rely on | | |
| | **WHOIS / RDAP** | current, but see the trap below | as a **fallback** for domains no list covers, and only with privacy-proxy filtering | | |
| | **TLS certificates, DNS/SOA, CNAME chains** | current | first-party CDN and sibling-domain detection, which the ownership lists are worst at: {[steffens2021_blockparty]} needed it because "those lists frequently miss connections among two hostnames, e.g., ''twitch.tv'' and ''twitchcdn.net''" | | |
| | **Crunchbase** | live but **paid**: the v4 API returns HTTP 401 without a key | you have institutional access; used this way by {[yang2020_comparative]} and {[cui2023_poligraph]} | | |
| | **LLM-based entity resolution** | **emerging, and not yet at the domain layer.** In this corpus the 2025 work is at the network layer: {[selmo2025_borges]} maps AS numbers to organisations with an LLM and releases prompts and code; {[gouda2025_prefix2org]} maps BGP prefixes. No corpus paper applies this to third-party //domain// ownership | you are willing to build and validate it yourself. The problem statement transfers directly; the validation burden does too | | |
| |
| <WRAP important> | |
| **Do not use WHOIS as an ownership database without filtering privacy proxies.** {[yang2020_comparative]} resolved organisations for 411 of 762 mobile-specific trackers using CrunchBase, webXray's list, TLS certificates and WHOIS in that order — and its top-ten table of "organizations" contains ''Redacted For Privacy'' (34), ''Domains By Proxy'' (25), ''Whois Guard'' (14), ''Global Domain Privacy Services'' (8) and ''Whois Privacy'' (7). Those five registrar-privacy strings account for 88 trackers, more than the largest real company in the table. {[sanchezrola2021_journey]} built a deliberately conservative WHOIS-plus-graph pipeline instead, and still had to note that the registration ecosystem "does not aim at providing full transparency on the organizations behind each domain". | |
| </WRAP> | |
| |
| <WRAP important> | |
| **"Tracker Radar" names three different artefacts**, and citing the name tells a reader nothing. Of the 32 corpus papers that mention it, **11** used the ownership dataset, **9** used it as a tracker or category database, and **9** used **Tracker Radar Collector** — a Puppeteer crawler that carries no ownership data at all and has its own page, [[Programming:Crawler:Tracker Radar Collector]]. A third artefact, Tracker Radar Detector, is the build pipeline. Name the artefact, the file and the commit. | |
| </WRAP> | |
| |
| ===== Use in publications ===== | ===== Use in publications ===== |
| Figures below are from the seven venues on [[literature:corpus]] (CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P, 2010–2026), not from the web-measurement literature as a whole. | Figures below are from the seven venues on [[literature:corpus]] (CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P, 2010–2026), not from the web-measurement literature as a whole. |
| |
| **15 papers** name webXray anywhere in their full text — 0.3% of the 5,859-paper corpus, and 1.3% of the 1,120 papers that ran a crawl. The extraction's tool field fires on only 7 of them, because most of the rest cite it in a related-work sentence or a reference list; both signals are reported rather than merged. Every verdict below was read off a quoted sentence. | **15 papers** name webXray anywhere in their full text — 0.3% of the 5,859-paper corpus. **13 of those 15** are among the 1,120 papers that ran a crawl, which is **1.2%** of that population.((Published as "1.3%" until 2026-09-11: the script was dividing the corpus-wide count of 15 by 1,120, a numerator and a denominator from two different populations. See the methodology section.)) The extraction's tool field fires on only 7 of them, because most of the rest cite it in a related-work sentence or a reference list; both signals are reported rather than merged. Every verdict below was read off a quoted sentence. |
| |
| ^ Role webXray plays ^ Papers ^ Share of 15 ^ | ^ Role webXray plays ^ Papers ^ Share of 15 ^ |
| | Yang & Yue, PoPETs 2020, //Web Tracking on Mobile and Desktop// {[yang2020_comparative]} | ownership list (cited as "Tim Libert's library"), **second** of four sources tried in order: CrunchBase, then webXray, then TLS certificates, then WHOIS | | | Yang & Yue, PoPETs 2020, //Web Tracking on Mobile and Desktop// {[yang2020_comparative]} | ownership list (cited as "Tim Libert's library"), **second** of four sources tried in order: CrunchBase, then webXray, then TLS certificates, then WHOIS | |
| | Steffens et al., NDSS 2021, //Who's Hosting the Block Party?// {[steffens2021_blockparty]} | ownership list, **retrieved from the Internet Archive** because upstream was already hard to obtain | | | Steffens et al., NDSS 2021, //Who's Hosting the Block Party?// {[steffens2021_blockparty]} | ownership list, **retrieved from the Internet Archive** because upstream was already hard to obtain | |
| | Sánchez-Rola et al., IEEE S&P 2021, //Journey to the Center of the Cookie Ecosystem// {[sanchezrola2021_journey]} | ownership list, ranked **third** of three merged lists. //The corpus files this paper under 2022 and the year table below follows the corpus; Crossref and the paper's own header put it at IEEE S&P 2021.// | | | Sánchez-Rola et al., IEEE S&P 2021, //Journey to the Center of the Cookie Ecosystem// {[sanchezrola2021_journey]} | ownership list, ranked **third** of three merged lists. //The corpus files this paper under 2022; Crossref and the paper's own header put it at IEEE S&P 2021. Year tables derived from this corpus, here and on [[Design:Ownership resolution]], follow the corpus.// | |
| | Musa & Nithyanand, PoPETs 2022, //ATOM// {[musa2022_atom]} | ownership list, alongside WHOIS records and TLS certificates | | | Musa & Nithyanand, PoPETs 2022, //ATOM// {[musa2022_atom]} | ownership list, alongside WHOIS records and TLS certificates | |
| | Cassel et al., PoPETs 2022, //OmniCrawl// {[cassel2022_omnicrawl]} | ownership list, "to determine the provenance of the requests" | | | Cassel et al., PoPETs 2022, //OmniCrawl// {[cassel2022_omnicrawl]} | ownership list, "to determine the provenance of the requests" | |
| Three findings from that table matter more than the counts. | Three findings from that table matter more than the counts. |
| |
| **The tool and the database came apart, and then the database was superseded.** The last corpus paper to use either was published in **2022**. Ownership use of Tracker Radar begins in 2021 and rises through the corpus edge — {[jannett2026_passkeys]} is a 2026 example, using the Entity Map to avoid treating ''gmail.com'' and ''google.com'' as separate authentication systems: | **The tool and the database came apart, and then the database was superseded.** The last corpus paper to use either was published in **2022**. Ownership use of Tracker Radar begins in 2021 and rises through the corpus edge; the year-by-year crossover, and what it does and does not license, is on [[Design:Ownership resolution#Which source the literature used, and when|Design:Ownership resolution]]. |
| | |
| ^ Year ^ Tracker Radar used for ownership ^ webXray crawler or ownership list ^ | |
| | 2018 | 0 | 1 | | |
| | 2019 | 0 | 0 | | |
| | 2020 | 0 | 2 | | |
| | 2021 | 1 | 1 | | |
| | 2022 | 0 | 4 | | |
| | 2023 | 2 | 0 | | |
| | 2024 | 1 | 0 | | |
| | 2025* | 4 | 0 | | |
| | 2026* | 3 | 0 | | |
| | |
| The starred years are the provisional corpus edge — CCS 2026 and IMC 2026 have not been held, and IEEE S&P and TheWebConf 2026 are under-selected by construction — so do not read 2026 as a complete year. The counts are small enough that the crossover is a signal about direction, not a market share. Combined with the licence and 404 findings above, the corpus supports "webXray's ownership list is historical and Tracker Radar's is current practice"; it does not support any claim about which is more accurate, which is what §"Two lists disagree" is for. | |
| | |
| **Nobody says which snapshot they used.** Of the 8 papers that used the crawler or the list, **2** say anything at all about which version, and **1** names a commit. That one, {[matte2020_cookie]}, pins **both** lists it uses in a reproducibility table — "Disconnect list commit eb817fb1 (2019-12-10)" and "WebXRay commit 04c3c8e8 (2019-06-18)" — beside the Chromium build, the kernel, the user agent, the vantage point and the Tranco list id. Copy that table. The file has no version field and no release tags, so a commit or an archive URL is the only thing that makes the figure reproducible; ''owner_dbs.py'' prints a content hash for the same reason. | |
| | |
| **No list removes the manual work.** {[steffens2021_blockparty]} is the most honest account of the real cost. They needed same-entity relations for the Tranco top 10,000, found that the curated lists "frequently miss connections among two hostnames, e.g., ''twitch.tv'' and ''twitchcdn.net''", mined their own crawl for candidates, and hand-vetted **2,175 candidate site pairs down to 1,146 confirmed same-entity pairs in about eight person-hours**. webXray's list then contributed **133 further relations** their own method had missed, and their conclusion was that "it alone does not suffice for our purposes". Two lessons: budget the person-hours, and treat the lists as //additive// rather than as alternatives. {[sanchezrola2021_journey]} took the same route from the other end — it merged Disconnect, WhoTracks.me and webXray by priority ("Disconnect first, WhoTracks.me second, and webxray third", explicitly "preferring those that have been updated most recently"), added an automated WHOIS-and-graph pass "to increase the coverage", found exactly one conflict against the **3,913** domains in any of the three lists, and filed a bug report that fixed an error in the Disconnect list. | |
| | |
| **How widespread is ownership resolution at all?** A full-text sweep for the named resources gives **136 papers (2.3% of the corpus; 12.1% of those that crawled)** that mention at least one. Each row is an upper bound — a match, not a verified use — except the two that were hand-verified in full: | |
| | |
| ^ Resource named in full text ^ Papers ^ Hand-verified? ^ | |
| | Public Suffix List | 101 | no — upper bound | | |
| | Disconnect, in a list/entity sense | 74 | no — upper bound | | |
| | Tracker Radar (all three artefacts) | 32 | **yes, all 32** | | |
| | WhoTracks.me | 23 | no — upper bound | | |
| | Crunchbase | 23 | no — upper bound | | |
| | webXray | 15 | **yes, all 15** | | |
| | |
| {[utz2023_rarely]} is worth reading before you pick one: it compared five third-party categorisations, including WhoTracks.me and Tracker Radar, and reports that "categorizations differ in granularity and focus" while overlapping substantially — the granularity axis again, stated from inside the literature. | |
| |
| ===== What to report in a paper ===== | **Nobody says which snapshot they used.** Of the 8 papers that used the crawler or the list, **2** say anything at all about which version, and **1** names a commit — {[matte2020_cookie]}, which pins "WebXRay commit 04c3c8e8 (2019-06-18)" beside the Disconnect list, the Chromium build, the kernel, the user agent, the vantage point and the Tranco list id. The file has no version field and no release tags, so a commit or an archive URL is the only thing that makes a figure derived from it reproducible. [[Design:Ownership resolution#What to report in a paper|Design:Ownership resolution]] treats that reproducibility table as the standard to meet and says what else belongs in it. |
| |
| If a result depends on attributing domains to companies, report all of: | **The corpus cannot see most of webXray's impact.** The seven venues omit EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS, and webXray's own author published most of its results outside those seven — his homepage lists The BMJ, JAMA, //Communications of the ACM// and //New Media and Society//((''https://timlibert.me/'', checked 2026-08-17.)). "15 papers" is a statement about seven security and networking venues, not a measure of how much webXray was used. |
| |
| * **which list, which file, and which commit or content hash.** "We used Tracker Radar" is not reproducible: name ''build-data/generated/entity_map.json'' at a commit, or publish the hash. For Disconnect, say whether you read ''properties'', ''resources'' or the union. | The wider question — how many corpus papers do domain-to-company attribution **at all**, with what, and whether it is growing — is on [[Design:Ownership resolution#Use in publications|Design:Ownership resolution]]: 136 papers name an ownership resource, 107 of them among the 1,120 that ran a crawl. |
| * **the date the snapshot was taken**, separately from the crawl date. Ownership changes between the two. | |
| * **the granularity you resolved to** — brand, operating legal entity, or ultimate parent — and, if you walked a hierarchy, how deep and how you handled cycles and missing parents. | |
| * **the fallback order**, if you merged sources, and the conflict rule. {[sanchezrola2021_journey]} and {[yang2020_comparative]} both used strict priority orders and both said so; that is the standard to meet. | |
| * **how many domains you could not attribute**, as a count and as a share of //requests or sites//, not just of domains. The two differ enormously: a list can cover 1.4% of domains and 54.5% of prevalence weight. Say which denominator your unattributed share uses. | |
| * **your privacy-proxy filter**, if WHOIS was involved, and the list of proxy strings you removed. | |
| * **the eTLD+1 rule and the PSL version**, since every one of these lists keys on the registrable domain and the PSL changes. Do not ship a frozen copy {[mcquistin2023_psl]}. | |
| * **the manual verification you did**, its sample size, and the disagreements you found. A prevalence-by-company figure with no manual check is a figure about a list, not about the web. | |
| |
| If the claim is "//N//% of sites contact Google", the reader cannot evaluate it without the granularity and the snapshot: it silently includes or excludes YouTube, DoubleClick, Bing-adjacent Microsoft properties and whatever Alphabet has bought or sold since the file was written. | |
| |
| ===== Methodology and limitations of these figures ===== | ===== Methodology and limitations of these figures ===== |
| |
| * The corpus audit script is ''scripts/report_webxray.mjs''; the live-database comparison is ''scripts/owner_dbs.py''; the hand adjudication with its primary sources is ''scripts/owner_adjudication.py''. Every query, every unedited output, the fold residues, the sources rejected, and the figures deliberately **not** published are on [[provenance:programming:crawler:webxray|the provenance page]]. Corpus-wide selection and extraction caveats are on [[literature:corpus]]. | * The corpus audit script is ''scripts/report_webxray.mjs'', published in full with its unedited output on [[provenance:programming:crawler:webxray|the provenance page]] along with every query, every fold residue, the sources rejected and the figures deliberately **not** published. Corpus-wide selection and extraction caveats are on [[literature:corpus]]. **The ownership-database comparison, the coverage census and the accuracy rates moved to [[Design:Ownership resolution]] on 2026-09-11**; their scripts (''owner_dbs.py'', ''owner_adjudication.py'', ''owner_sample.py'', ''owner_random_sample.py'', ''owner_irr_kappa.py'') and outputs stay on this page's provenance page and its [[provenance:programming:crawler:webxray:random_sample|code appendix]], whose ids are unchanged so the published record still resolves. [[provenance:design:ownership_resolution]] records the split and carries the new page's own queries. |
| * The webXray population is a **full-text sweep**, not the extraction's tool field, because the tool field finds 7 papers where the sweep finds 15 and the difference is exactly the citation-only cases this page needs to separate. All 15 carry a hand verdict with a quoted sentence; the residue between sweep and hand map is zero, and the script fails loudly if that changes. | * The webXray population is a **full-text sweep**, not the extraction's tool field, because the tool field finds 7 papers where the sweep finds 15 and the difference is exactly the citation-only cases this page needs to separate. All 15 carry a hand verdict with a quoted sentence; the residue between sweep and hand map is zero, and the script fails loudly if that changes. |
| * The 26 literal per-paper figures and quotes on this page were checked against each paper's ''paper.cols.txt'' rendering: 26 of 26 located. The 16 extraction evidence quotes behind the tool and classification tuples were checked the same way: 4 exact, 11 partial, 1 below threshold; the below-threshold quote is present in the source and mangled by a column splice, not unsupported. | * ''report_webxray.mjs'' section G checks **26** literal per-paper figures and quotes, 26 of 26 located against each paper's ''paper.cols.txt'' rendering. Since the 2026-09-11 split those 26 are divided between this page and [[Design:Ownership resolution]] — the script still checks both sets, and the count has not been re-partitioned. The 16 extraction evidence quotes behind the tool and classification tuples were checked the same way: 4 exact, 11 partial, 1 below threshold; the below-threshold quote is present in the source and mangled by a column splice, not unsupported. |
| * One figure was **not published**: {[steffens2021_blockparty]} contains the sentence "webXray's list does account for 1,096 of our 1,146 found connections meaning that it alone does not suffice for our purposes", whose two halves contradict each other. The published PDF reads the same way, so no coverage percentage is taken from it; the unambiguous parts of that paragraph are quoted above instead. | * One figure was **not published**: {[steffens2021_blockparty]} contains the sentence "webXray's list does account for 1,096 of our 1,146 found connections meaning that it alone does not suffice for our purposes", whose two halves contradict each other. The published PDF reads the same way, so no coverage percentage is taken from it. |
| * The three-way database comparison uses a **single snapshot per list, taken on 2026-08-17**, with SHA-256 prefixes recorded on the provenance page. Two of the three lists change weekly. Re-run the script rather than quoting these numbers in 2027. | * **Every repository, licence and URL claim on this page was checked on 2026-08-17 and the Wayback bounds on 2026-09-05.** Repository state changes; re-check before citing. The licence finding in particular was **wrong in this page's first version**, which asserted from one snapshot that webXray "is not open source and redistributing it is prohibited" — true of that snapshot and false of the project's final state. |
| * The coverage denominator is Tracker Radar's own crawl output. There is no neutral census of third-party domains, so this favours Tracker Radar; the page says so wherever a coverage figure appears rather than pretending otherwise. | * The **1.2% of the 1,120 crawling papers** above was published as 1.3% until 2026-09-11, when the split found that ''report_webxray.mjs'' was dividing a corpus-wide numerator (15) by the crawled denominator (1,120) instead of counting the 13 sweep hits that are actually in that population. The same bug on a larger figure read 12.1% where the comparable number is 9.6%; both are corrected and the script now computes the intersection. |
| * The 30 adjudicated rows are the highest-prevalence **disagreements**, deliberately not a random sample, so they characterise the shape of disagreement and not any list's accuracy. Two rows (the acquisition date behind ''360yield.com'', and ''fwmrm.net'''s post-2026-spinoff status) could not be settled from a primary source and are recorded as unresolved rather than guessed. | |
| * Wayback Machine checks could not be completed: ''web.archive.org'' was returning 502/503 throughout 2026-08-17. So the date ''github.com/timlib/webXray'' first 404ed, and what ''webxray.org'' used to contain, are **unknown** and are not estimated here. | |
| * The seven-venue corpus omits EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS, so "15 papers name webXray" is a statement about those seven venues. webXray's own author published most of its results outside those seven venues — his homepage lists The BMJ, JAMA, Communications of the ACM and //New Media and Society//((''https://timlibert.me/'', checked 2026-08-17.)) — so the corpus cannot see the work the tool's reputation mostly rests on, and "15 papers" is not a measure of how much webXray was used. | |
| |
| ===== Related pages ===== | ===== Related pages ===== |
| |
| | * [[Design:Ownership resolution]] — **the question this tool's database answered**: which source to use now, how much of the web each one covers, and how often each is right. |
| * [[Programming:Crawler]] — generic automation libraries and the other specialised crawlers. | * [[Programming:Crawler]] — generic automation libraries and the other specialised crawlers. |
| * [[Programming:Crawler:Tracker Radar Collector]] — the Puppeteer crawler that shares Tracker Radar's name and carries none of its ownership data. | * [[Programming:Crawler:Tracker Radar Collector]] — the Puppeteer crawler that shares Tracker Radar's name and carries none of its ownership data. |
| * [[Programming:Crawler:OpenWPM]] — the instrument webXray was compared against in 2016, and the one still maintained. | * [[Programming:Crawler:OpenWPM]] — the instrument webXray was compared against in 2016, and the one still maintained. |
| * [[Privacy:Requests]] — filter lists, which answer "is this a tracker?" rather than "whose is it?", and how the two get conflated. | * [[Privacy:Policies]] — what policyXray did, and what has replaced it. |
| * [[Privacy:Cookies]] — where cookie attribution needs an owner name, and what the same granularity choice does to it. | * [[Programming:Traffic files]] — capturing the requests webXray captured. |
| * [[Design:IP classification]] — the network-layer version of this problem, including AS-to-organisation mapping. | |
| * [[Programming:Traffic files]] — capturing the requests you are about to attribute. | |
| |
| ====== References ====== | ====== References ====== |