| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| programming:crawler:webxray [2026/09/11 17:50] – Carve out the ownership-resolution material to the new Design:Ownership resolution and leave this as a page about webXray itself: architecture, availability, the three-licence history, the domain_owners.json schema and its frozen PSL, a new Wayback timeli karel.kubicek.claude | programming:crawler:webxray [2026/09/11 19:20] (current) – webXray's measured rate refreshed after the 2026-09-11 second pass over the residue: 69.9% -> 69.2% (56.1-81.4) on 56 settled entries. Authored by Claude karel.kubicek.claude |
|---|
| |
| * **The tool is gone from where every paper points, but it is not lost and it did not end up proprietary.** ''github.com/timlib/webXray'' returns HTTP 404 and there has never been a PyPI package — but Libert relicensed webXray twice after its most widely mirrored snapshot, **back** to GPLv3 in 2021 — the licence its own website claimed in 2015 — and then to **MIT in 2023**, and that history survives in a fork. The author now runs a commercial product under the same name at ''webxray.ai''. | * **The tool is gone from where every paper points, but it is not lost and it did not end up proprietary.** ''github.com/timlib/webXray'' returns HTTP 404 and there has never been a PyPI package — but Libert relicensed webXray twice after its most widely mirrored snapshot, **back** to GPLv3 in 2021 — the licence its own website claimed in 2015 — and then to **MIT in 2023**, and that history survives in a fork. The author now runs a commercial product under the same name at ''webxray.ai''. |
| * **The database survives and it is frozen at 2021-03-04.** How stale that makes it has been measured: on a stratified random sample of its own coverage, drawn and adjudicated on 2026-09-05, it names today's owner for **69.9%** of the entries that could be settled (95% CI 55.5–83.1%). That figure is a rate over the fifth of its entries a 2026 crawl still meets, not over the file, and it sits beside a **coverage** figure of **2.1%** of the registrable third-party domains in the frame against Tracker Radar's 17.2% and Disconnect's 7.0%. The measurement, its two frames and every caveat on reading it are on [[Design:Ownership resolution]]; nothing on this page restates them. | * **The database survives and it is frozen at 2021-03-04.** How stale that makes it has been measured: on a stratified random sample of its own coverage, drawn 2026-09-05 and re-adjudicated on 2026-09-11 where the first pass could not settle a row, it names today's owner for **69.2%** of the entries that could be settled (95% CI 56.1–81.4%) — a rate over the fifth of its entries a 2026 crawl still meets, not over the file, and one that sits beside a **coverage** figure of **2.1%** of the frame. The adjudicators were language models, the comparison has two snapshot frames, and there are several reasons not to read 69.2% as a rehabilitation. All of that is on [[Design:Ownership resolution]], and this page deliberately does not restate it: take the two figures from there, with their caveats, rather than from this bullet. |
| * **The list and the tool came apart in the literature years ago.** Of the 15 corpus papers that name webXray, **7 used only the ownership list** and exactly **one ran the crawler** — and that one is Libert's own paper. | * **The list and the tool came apart in the literature years ago.** Of the 15 corpus papers that name webXray, **7 used only the ownership list** and exactly **one ran the crawler** — and that one is Libert's own paper. |
| |
| | ''github.com/timlib/webXray'' | **HTTP 404**. The ''timlib'' account itself still exists (HTTP 200) with 0 public repositories. | | | ''github.com/timlib/webXray'' | **HTTP 404**. The ''timlib'' account itself still exists (HTTP 200) with 0 public repositories. | |
| | ''github.com/timlib/webXray_Domain_Owner_List'' | **HTTP 404**. This is the URL cited by {[kashaf2020_dependencies]}, among others. | | | ''github.com/timlib/webXray_Domain_Owner_List'' | **HTTP 404**. This is the URL cited by {[kashaf2020_dependencies]}, among others. | |
| | ''webxray.org'' | HTTP 200, but a placeholder: a heading and the line "Public interest projects for the interested public." No source link, no version, no download. Until **2024-03-28** it hosted a live demo search engine over "scans of 100,000 sites", with a company drop-down drawn from the ownership list; by 2024-05-24 that returned ''401 Authorization Required'', and from 2024-07-24 to at least 2026-02-05 the domain 301-redirected to ''webxray.ai''. See [[#Methodology and limitations of these figures|the Wayback timeline below]]. | | | ''webxray.org'' | HTTP 200, but a placeholder: a heading and the line "Public interest projects for the interested public." No source link, no version, no download. Until **2024-03-28** it hosted a live demo search engine over "scans of 100,000 sites", with a company drop-down drawn from the ownership list; by 2024-05-24 that returned ''401 Authorization Required'', and from 2024-07-24 to at least 2026-02-05 the domain 301-redirected to ''webxray.ai''. See [[#What happened to webxray.org, and when|the Wayback timeline below]]. | |
| | PyPI ''webxray'' / ''web-xray'' / ''policyxray'' | 404 each. webXray was never packaged on PyPI. | | | PyPI ''webxray'' / ''web-xray'' / ''policyxray'' | 404 each. webXray was never packaged on PyPI. | |
| | ''github.com/thezedwards/webXray'' | The most widely mirrored copy of webXray 3.x, last commit **2021-03-04**. Its README still instructs ''git clone https://github.com/timlib/webXray.git''. Every figure on this page comes from here — and it is **not** the most complete copy. | | | ''github.com/thezedwards/webXray'' | The most widely mirrored copy of webXray 3.x, last commit **2021-03-04**. Its README still instructs ''git clone https://github.com/timlib/webXray.git''. Every figure on this page comes from here — and it is **not** the most complete copy. | |
| | Yang & Yue, PoPETs 2020, //Web Tracking on Mobile and Desktop// {[yang2020_comparative]} | ownership list (cited as "Tim Libert's library"), **second** of four sources tried in order: CrunchBase, then webXray, then TLS certificates, then WHOIS | | | Yang & Yue, PoPETs 2020, //Web Tracking on Mobile and Desktop// {[yang2020_comparative]} | ownership list (cited as "Tim Libert's library"), **second** of four sources tried in order: CrunchBase, then webXray, then TLS certificates, then WHOIS | |
| | Steffens et al., NDSS 2021, //Who's Hosting the Block Party?// {[steffens2021_blockparty]} | ownership list, **retrieved from the Internet Archive** because upstream was already hard to obtain | | | Steffens et al., NDSS 2021, //Who's Hosting the Block Party?// {[steffens2021_blockparty]} | ownership list, **retrieved from the Internet Archive** because upstream was already hard to obtain | |
| | Sánchez-Rola et al., IEEE S&P 2021, //Journey to the Center of the Cookie Ecosystem// {[sanchezrola2021_journey]} | ownership list, ranked **third** of three merged lists. //The corpus files this paper under 2022 and the year table below follows the corpus; Crossref and the paper's own header put it at IEEE S&P 2021.// | | | Sánchez-Rola et al., IEEE S&P 2021, //Journey to the Center of the Cookie Ecosystem// {[sanchezrola2021_journey]} | ownership list, ranked **third** of three merged lists. //The corpus files this paper under 2022; Crossref and the paper's own header put it at IEEE S&P 2021. Year tables derived from this corpus, here and on [[Design:Ownership resolution]], follow the corpus.// | |
| | Musa & Nithyanand, PoPETs 2022, //ATOM// {[musa2022_atom]} | ownership list, alongside WHOIS records and TLS certificates | | | Musa & Nithyanand, PoPETs 2022, //ATOM// {[musa2022_atom]} | ownership list, alongside WHOIS records and TLS certificates | |
| | Cassel et al., PoPETs 2022, //OmniCrawl// {[cassel2022_omnicrawl]} | ownership list, "to determine the provenance of the requests" | | | Cassel et al., PoPETs 2022, //OmniCrawl// {[cassel2022_omnicrawl]} | ownership list, "to determine the provenance of the requests" | |
| **The tool and the database came apart, and then the database was superseded.** The last corpus paper to use either was published in **2022**. Ownership use of Tracker Radar begins in 2021 and rises through the corpus edge; the year-by-year crossover, and what it does and does not license, is on [[Design:Ownership resolution#Which source the literature used, and when|Design:Ownership resolution]]. | **The tool and the database came apart, and then the database was superseded.** The last corpus paper to use either was published in **2022**. Ownership use of Tracker Radar begins in 2021 and rises through the corpus edge; the year-by-year crossover, and what it does and does not license, is on [[Design:Ownership resolution#Which source the literature used, and when|Design:Ownership resolution]]. |
| |
| **Nobody says which snapshot they used.** Of the 8 papers that used the crawler or the list, **2** say anything at all about which version, and **1** names a commit. That one, {[matte2020_cookie]}, pins **both** lists it uses in a reproducibility table — "Disconnect list commit eb817fb1 (2019-12-10)" and "WebXRay commit 04c3c8e8 (2019-06-18)" — beside the Chromium build, the kernel, the user agent, the vantage point and the Tranco list id. Copy that table. The file has no version field and no release tags, so a commit or an archive URL is the only thing that makes a figure derived from it reproducible. | **Nobody says which snapshot they used.** Of the 8 papers that used the crawler or the list, **2** say anything at all about which version, and **1** names a commit — {[matte2020_cookie]}, which pins "WebXRay commit 04c3c8e8 (2019-06-18)" beside the Disconnect list, the Chromium build, the kernel, the user agent, the vantage point and the Tranco list id. The file has no version field and no release tags, so a commit or an archive URL is the only thing that makes a figure derived from it reproducible. [[Design:Ownership resolution#What to report in a paper|Design:Ownership resolution]] treats that reproducibility table as the standard to meet and says what else belongs in it. |
| |
| **The corpus cannot see most of webXray's impact.** The seven venues omit EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS, and webXray's own author published most of its results outside those seven — his homepage lists The BMJ, JAMA, //Communications of the ACM// and //New Media and Society//((''https://timlibert.me/'', checked 2026-08-17.)). "15 papers" is a statement about seven security and networking venues, not a measure of how much webXray was used. | **The corpus cannot see most of webXray's impact.** The seven venues omit EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS, and webXray's own author published most of its results outside those seven — his homepage lists The BMJ, JAMA, //Communications of the ACM// and //New Media and Society//((''https://timlibert.me/'', checked 2026-08-17.)). "15 papers" is a statement about seven security and networking venues, not a measure of how much webXray was used. |
| * The corpus audit script is ''scripts/report_webxray.mjs'', published in full with its unedited output on [[provenance:programming:crawler:webxray|the provenance page]] along with every query, every fold residue, the sources rejected and the figures deliberately **not** published. Corpus-wide selection and extraction caveats are on [[literature:corpus]]. **The ownership-database comparison, the coverage census and the accuracy rates moved to [[Design:Ownership resolution]] on 2026-09-11**; their scripts (''owner_dbs.py'', ''owner_adjudication.py'', ''owner_sample.py'', ''owner_random_sample.py'', ''owner_irr_kappa.py'') and outputs stay on this page's provenance page and its [[provenance:programming:crawler:webxray:random_sample|code appendix]], whose ids are unchanged so the published record still resolves. [[provenance:design:ownership_resolution]] records the split and carries the new page's own queries. | * The corpus audit script is ''scripts/report_webxray.mjs'', published in full with its unedited output on [[provenance:programming:crawler:webxray|the provenance page]] along with every query, every fold residue, the sources rejected and the figures deliberately **not** published. Corpus-wide selection and extraction caveats are on [[literature:corpus]]. **The ownership-database comparison, the coverage census and the accuracy rates moved to [[Design:Ownership resolution]] on 2026-09-11**; their scripts (''owner_dbs.py'', ''owner_adjudication.py'', ''owner_sample.py'', ''owner_random_sample.py'', ''owner_irr_kappa.py'') and outputs stay on this page's provenance page and its [[provenance:programming:crawler:webxray:random_sample|code appendix]], whose ids are unchanged so the published record still resolves. [[provenance:design:ownership_resolution]] records the split and carries the new page's own queries. |
| * The webXray population is a **full-text sweep**, not the extraction's tool field, because the tool field finds 7 papers where the sweep finds 15 and the difference is exactly the citation-only cases this page needs to separate. All 15 carry a hand verdict with a quoted sentence; the residue between sweep and hand map is zero, and the script fails loudly if that changes. | * The webXray population is a **full-text sweep**, not the extraction's tool field, because the tool field finds 7 papers where the sweep finds 15 and the difference is exactly the citation-only cases this page needs to separate. All 15 carry a hand verdict with a quoted sentence; the residue between sweep and hand map is zero, and the script fails loudly if that changes. |
| * The 26 literal per-paper figures and quotes on this page were checked against each paper's ''paper.cols.txt'' rendering: 26 of 26 located. The 16 extraction evidence quotes behind the tool and classification tuples were checked the same way: 4 exact, 11 partial, 1 below threshold; the below-threshold quote is present in the source and mangled by a column splice, not unsupported. | * ''report_webxray.mjs'' section G checks **26** literal per-paper figures and quotes, 26 of 26 located against each paper's ''paper.cols.txt'' rendering. Since the 2026-09-11 split those 26 are divided between this page and [[Design:Ownership resolution]] — the script still checks both sets, and the count has not been re-partitioned. The 16 extraction evidence quotes behind the tool and classification tuples were checked the same way: 4 exact, 11 partial, 1 below threshold; the below-threshold quote is present in the source and mangled by a column splice, not unsupported. |
| * One figure was **not published**: {[steffens2021_blockparty]} contains the sentence "webXray's list does account for 1,096 of our 1,146 found connections meaning that it alone does not suffice for our purposes", whose two halves contradict each other. The published PDF reads the same way, so no coverage percentage is taken from it. | * One figure was **not published**: {[steffens2021_blockparty]} contains the sentence "webXray's list does account for 1,096 of our 1,146 found connections meaning that it alone does not suffice for our purposes", whose two halves contradict each other. The published PDF reads the same way, so no coverage percentage is taken from it. |
| * **Every repository, licence and URL claim on this page was checked on 2026-08-17 and the Wayback bounds on 2026-09-05.** Repository state changes; re-check before citing. The licence finding in particular was **wrong in this page's first version**, which asserted from one snapshot that webXray "is not open source and redistributing it is prohibited" — true of that snapshot and false of the project's final state. | * **Every repository, licence and URL claim on this page was checked on 2026-08-17 and the Wayback bounds on 2026-09-05.** Repository state changes; re-check before citing. The licence finding in particular was **wrong in this page's first version**, which asserted from one snapshot that webXray "is not open source and redistributing it is prohibited" — true of that snapshot and false of the project's final state. |