User Tools

Site Tools


programming:crawler

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
programming:crawler [2026/08/29 01:41] – Shrink 'Being Detected' to a pointer at the new programming:crawler_detection, which now covers it in full; keep the two points specific to choosing a library. Also correct the anti-detection patches row from 12/2 to 11/1: tool_fold.mjs matched 'Stealth A karel.kubicek.claudeprogramming:crawler [2026/09/17 10:45] (current) – Clarify consentAction schema statistic; Authored by Claude karel.kubicek.claude
Line 3: Line 3:
 Every automated web measurement makes two separate choices that papers routinely report as one: **which browser** renders the page, and **which control channel** drives it. A "Selenium crawl" says nothing about the first; a "Chrome crawl" says nothing about the second. They fail differently, they are detected differently, and they give you access to different data. Every automated web measurement makes two separate choices that papers routinely report as one: **which browser** renders the page, and **which control channel** drives it. A "Selenium crawl" says nothing about the first; a "Chrome crawl" says nothing about the second. They fail differently, they are detected differently, and they give you access to different data.
  
-This page compares the options. It covers the generic automation libraries ([[#Selenium|Selenium]], [[#Puppeteer|Puppeteer]], [[#Playwright|Playwright]], [[#Plain CDP|plain CDP]]) and then the specialised privacy and security crawlers built on top of them, each of which has — or will have — its own page: [[Programming:Crawler:OpenWPM]], [[Programming:Crawler:webXray]], [[Programming:Crawler:Tracker Radar Collector]], [[Programming:Crawler:PageGraph]], [[Programming:Crawler:Foxhound]], [[Programming:Crawler:PanoptiChrome]].+This page compares the options. It covers the generic automation libraries ([[#Selenium|Selenium]], [[#Puppeteer|Puppeteer]], [[#Playwright|Playwright]], [[#Plain CDP|plain CDP]]) and then the specialised privacy and security crawlers built on top of them, each of which has — or will have — its own page: [[Programming:Crawler:OpenWPM]], [[Programming:Crawler:webXray]], [[Programming:Crawler:Tracker Radar Collector]], [[Programming:Crawler:PageGraph]], [[Programming:Crawler:Foxhound]], [[Programming:Crawler:PanoptiChrome]]. A newer child, [[Programming:Crawler:LLM Agents]], covers the one instrument on this page that is not a control channel at all: an agent that decides the next click with a model.
  
 It pairs with [[Design:Automated measurements]] (whether to crawl at all), [[Design:Crawling location]] (where from), [[Programming:Stateful stateless]] (with or without a profile), [[Programming:Interaction]] (what to do on the page), and [[Privacy:Consent]] (what to do with the banner). It pairs with [[Design:Automated measurements]] (whether to crawl at all), [[Design:Crawling location]] (where from), [[Programming:Stateful stateless]] (with or without a profile), [[Programming:Interaction]] (what to do on the page), and [[Privacy:Consent]] (what to do with the banner).
Line 241: Line 241:
 ==== Plain CDP ==== ==== Plain CDP ====
  
-You can skip the libraries entirely: launch Chromium with ''--remote-debugging-port'', open a WebSocket, enable the domains you want, and read the events. **42 papers (3.8%) did**, most of them between 2018 and 2021 — a window that closes as Puppeteer matures.+You can skip the libraries entirely: launch Chromium with ''%%--remote-debugging-port%%'', open a WebSocket, enable the domains you want, and read the events. **42 papers (3.8%) did**, most of them between 2018 and 2021 — a window that closes as Puppeteer matures.
  
 Reach for it when your analysis lives in a language with no good automation binding, when you need a CDP domain the wrapper does not expose, or when you want to know exactly what your instrument is doing. The cost, visible in the code below, is that everything a library gives you is now yours to write: waiting for the port, waiting for the load, waiting for the network to settle, and reconnecting when a target dies. Reach for it when your analysis lives in a language with no good automation binding, when you need a CDP domain the wrapper does not expose, or when you want to know exactly what your instrument is doing. The cost, visible in the code below, is that everything a library gives you is now yours to write: waiting for the port, waiting for the load, waiting for the network to settle, and reconnecting when a target dies.
Line 314: Line 314:
  
 <WRAP info> <WRAP info>
-**Sandbox gotchas we hit, since they cost half a day.** On this aarch64 container, //every full Chrome/Chromium binary// — Debian's ''/usr/bin/chromium'', Chrome for Testing, and Playwright's own ''chromium-*'' build — died at startup with ''chrome_crashpad_handler: --database is required'' and never opened the DevTools port. Only Playwright's ''chromium_headless_shell'' build started. Chrome for Testing ships no aarch64 build (the downloaded ''linux_arm'' binary is x86-64 and fails under emulation), and neither does ''chromedriver''; a working aarch64 ''chromedriver'' came from Debian's ''chromium-driver'' package, extracted with ''dpkg -x'' without root. Firefox 153 + geckodriver 0.37.1 never handed over its Marionette port in this container, with or without Xvfb, so the Selenium example is the Chromium one. None of this is a property of the libraries — it is a property of running x86-64-first browser tooling on ARM in a container, which is increasingly where CI lives.+**Sandbox gotchas we hit, since they cost half a day.** On this aarch64 container, //every full Chrome/Chromium binary// — Debian's ''/usr/bin/chromium'', Chrome for Testing, and Playwright's own ''chromium-*'' build — died at startup with ''chrome_crashpad_handler: %%--database%% is required'' and never opened the DevTools port. Only Playwright's ''chromium_headless_shell'' build started. Chrome for Testing ships no aarch64 build (the downloaded ''linux_arm'' binary is x86-64 and fails under emulation), and neither does ''chromedriver''; a working aarch64 ''chromedriver'' came from Debian's ''chromium-driver'' package, extracted with ''dpkg -x'' without root. Firefox 153 + geckodriver 0.37.1 never handed over its Marionette port in this container, with or without Xvfb, so the Selenium example is the Chromium one. None of this is a property of the libraries — it is a property of running x86-64-first browser tooling on ARM in a container, which is increasingly where CI lives.
 </WRAP> </WRAP>
  
Line 334: Line 334:
   * **OpenWPM is still Firefox and still Selenium.** Its README's first paragraph says so, and it pins an unbranded Firefox build.(([[https://github.com/openwpm/OpenWPM|github.com/openwpm/OpenWPM]], README and ''scripts/install-firefox.sh'', checked 2026-08-06.)) It is actively maintained. If your study is about Chrome's behaviour, OpenWPM is measuring a different browser — a point Krumnow et al. {[krumnow2022_gullible]} press further by analysing how detectable OpenWPM is, how resilient its recording is, and how widespread OpenWPM-specific detection is in the wild.   * **OpenWPM is still Firefox and still Selenium.** Its README's first paragraph says so, and it pins an unbranded Firefox build.(([[https://github.com/openwpm/OpenWPM|github.com/openwpm/OpenWPM]], README and ''scripts/install-firefox.sh'', checked 2026-08-06.)) It is actively maintained. If your study is about Chrome's behaviour, OpenWPM is measuring a different browser — a point Krumnow et al. {[krumnow2022_gullible]} press further by analysing how detectable OpenWPM is, how resilient its recording is, and how widespread OpenWPM-specific detection is in the wild.
   * **PageGraph does not track your automation.** Its own documentation warns that "PageGraph currently does not track puppeteer / automation scripts, and so modifying or interacting with the document through devtools/puppeteer while recording a PageGraph file will likely fail".((''pagegraph-crawl'' README, [[https://github.com/brave/pagegraph-crawl|github.com/brave/pagegraph-crawl]], checked 2026-08-06.)) It is a recorder of what the page did, not a driver of interaction. Its own limitations list is unusually honest and worth reading before you rely on it: no WebSocket tracking, no worker tracking, request headers not recorded (responses only), no ''style='' mutation tracking.((PageGraph wiki, [[https://github.com/brave/brave-browser/wiki/PageGraph|"Would Be Nice / Someday / Known Limitations"]], checked 2026-08-06. Query tooling: ''pagegraph-query'' (Python) under ''brave-experiments'', ''pagegraph-rust'' under ''brave''.))   * **PageGraph does not track your automation.** Its own documentation warns that "PageGraph currently does not track puppeteer / automation scripts, and so modifying or interacting with the document through devtools/puppeteer while recording a PageGraph file will likely fail".((''pagegraph-crawl'' README, [[https://github.com/brave/pagegraph-crawl|github.com/brave/pagegraph-crawl]], checked 2026-08-06.)) It is a recorder of what the page did, not a driver of interaction. Its own limitations list is unusually honest and worth reading before you rely on it: no WebSocket tracking, no worker tracking, request headers not recorded (responses only), no ''style='' mutation tracking.((PageGraph wiki, [[https://github.com/brave/brave-browser/wiki/PageGraph|"Would Be Nice / Someday / Known Limitations"]], checked 2026-08-06. Query tooling: ''pagegraph-query'' (Python) under ''brave-experiments'', ''pagegraph-rust'' under ''brave''.))
-  * **Tracker Radar Collector caps concurrency at 38** crawlers, injects a simple anti-bot-detection script into every frame by defaultand can be pointed at a Selenium Hub or a specific Chromium version.((''tracker-radar-collector'' README, [[https://github.com/duckduckgo/tracker-radar-collector|github.com/duckduckgo/tracker-radar-collector]], checked 2026-08-06.)) It has no canonical academic paper; cite the repository and the Tracker Radar dataset. +  * **Tracker Radar Collector's README is stale, and one of its stale numbers was on this page.** It still claims "a hard limit of 38" concurrent crawlers; ''38'' appears nowhere in the code and the cap was removed in commit ''963f9ec'' (2025-04-16). The default is ''Math.floor(cores * 0.8)''then capped at the number of URLs. The README's ''maxLoadTimeMs'' default of 30 s is stale the same way — the code fallback is 60 s. TRC does inject an anti-bot script into every frame by default (''-a'' disables it) and can be pointed at a Selenium Hub or a specific Chromium version.((Checked against a clone at commit ''8b64006691a1ce3929cfdfeeb425e7cc64be6543'' on 2026-09-03: ''crawlerConductor.js'' concurrency''crawler.js'' ''maxLoadTimeMs || 60000'', and ''git log -S38 -- crawlerConductor.js''[[https://github.com/duckduckgo/tracker-radar-collector|github.com/duckduckgo/tracker-radar-collector]].)) It has no canonical academic paper; cite the repository and the Tracker Radar dataset — and see [[Programming:Crawler:Tracker Radar Collector]] for the rest of the README-versus-code divergences
-  * **webXray cannot be installed, and the reason is now known.** The repository its own documentation points at, ''github.com/timlib/webXray'', **no longer exists**, the ''timlib'' account has no public repositories, and there is no ''webxray'' package on PyPI. webXray is now a **commercial** product, ''webxray.ai'', run by webXray LLC with Libert as founder and CEO; ''webxray.org'' has been reduced to a placeholder.((Checked 2026-08-17: ''api.github.com/repos/timlib/webXray'' → HTTP 404, ''users/timlib'' → HTTP 200 with ''public_repos: 0'' and ''company: webXray.ai''; ''pypi.org/pypi/webxray/json'' → 404; ''https://timlibert.me/'' states "(Dr.) Timothy Libert is founder and CEO of webXray LLC".)) The most complete surviving copy is ''thezedwards/webXray'', last pushed **2021-03-04** — webXray 3.x, which drives **consumer Chrome over the DevTools protocol**, not PhantomJSits licence is PolyForm Strict 1.0.0, which permits noncommercial use but **forbids redistribution**, so there is no lawful route to the code now that upstream is gone. Treat {[libert2015_invisible]} and {[libert2018_automated]} as citations for the //method// — third-party request measurement with company attribution — and see [[Programming:Crawler:webXray]] for what survives of itthe ownership database, and how it compares to Tracker Radar and Disconnect today.+  * **webXray cannot be installed, and the reason is now known.** The repository its own documentation points at, ''github.com/timlib/webXray'', **no longer exists**, the ''timlib'' account has no public repositories, and there is no ''webxray'' package on PyPI. webXray is now a **commercial** product, ''webxray.ai'', run by webXray LLC with Libert as founder and CEO; ''webxray.org'' has been reduced to a placeholder.((Checked 2026-08-17: ''api.github.com/repos/timlib/webXray'' → HTTP 404, ''users/timlib'' → HTTP 200 with ''public_repos: 0'' and ''company: webXray.ai''; ''pypi.org/pypi/webxray/json'' → 404; ''https://timlibert.me/'' states "(Dr.) Timothy Libert is founder and CEO of webXray LLC".)) The most widely mirrored copy is ''thezedwards/webXray'', last pushed **2021-03-04** — webXray 3.x, which drives **consumer Chrome over the DevTools protocol**, not PhantomJS. It is **not** the copy to take, and its licence is the reason to be careful rather than the reason to give up: that snapshot is PolyForm Strict 1.0.0, which permits noncommercial use but forbids redistribution — but Libert relicensed the project **back to GPLv3 in 2021 and then to MIT in 2023**, and that history survives in ''peterjoles/webXray'', 36 commits ahead. There **is** a lawful route to the code; it is just not the copy most mirrors carry. Treat {[libert2015_invisible]} and {[libert2018_automated]} as citations for the //method// — third-party request measurement with company attribution — and see [[Programming:Crawler:webXray]] for the tool's remains. The part that outlived itthe **domain-ownership database**, and how it compares to the two live alternatives today, moved on 2026-09-11 to [[Design:Ownership resolution]]: that is the page to read if your question is "whose domain is this?" rather than "which crawler should I run?".
   * **Taint tracking is a different question.** Foxhound and PanoptiChrome answer "did this value reach that sink?", not "what did the page load". If your question is about tracking prevalence, they are the wrong instrument and cost you a browser build; if it is about how data escapes — client-side XSS, DOM-based leaks, fingerprinting inputs — nothing else answers it. **Between the two, start with Foxhound**: the one head-to-head evaluation measured PanoptiChrome at 50% compatibility and 36.7× overhead against Foxhound's 95% and 1.4×, and no paper in this corpus has used PanoptiChrome as an instrument {[calzavara2025_dynamic]} — see [[Programming:Crawler:PanoptiChrome]] for when it is nevertheless the only option. Foxhound tells you what to cite: its README's "Cite us!" section asks for the EuroS&P paper in which the browser is described, Klein et al. {[klein2022_handsanitizers]}, and its wiki separately lists the papers that have used it.   * **Taint tracking is a different question.** Foxhound and PanoptiChrome answer "did this value reach that sink?", not "what did the page load". If your question is about tracking prevalence, they are the wrong instrument and cost you a browser build; if it is about how data escapes — client-side XSS, DOM-based leaks, fingerprinting inputs — nothing else answers it. **Between the two, start with Foxhound**: the one head-to-head evaluation measured PanoptiChrome at 50% compatibility and 36.7× overhead against Foxhound's 95% and 1.4×, and no paper in this corpus has used PanoptiChrome as an instrument {[calzavara2025_dynamic]} — see [[Programming:Crawler:PanoptiChrome]] for when it is nevertheless the only option. Foxhound tells you what to cite: its README's "Cite us!" section asks for the EuroS&P paper in which the browser is described, Klein et al. {[klein2022_handsanitizers]}, and its wiki separately lists the papers that have used it.
 +
 +===== Agent-Driven Crawling =====
 +
 +Since 2026 there is an option that does not fit the two-layer model above: a
 +**crawler whose next action is chosen by a model** rather than scripted. The open-source
 +ones — Browser Use, BrowserGym/AgentLab, Skyvern — still drive Chromium through CDP or
 +Playwright, so they add no new control channel; the vendor computer-use APIs act on
 +screenshots and coordinates and leave the channel to you. What all of them change is that
 +the sequence of actions is non-deterministic, the completion rate becomes an empirical
 +quantity, and the cost per site is measured in cents rather than fractions of one.
 +
 +**Five of the 1,120 crawling papers here drove a crawl this way, all of them in 2026.**((Four
 +in the table below, which counts only ''tools[]'' entries in an automation category, as
 +every other row of that table does. A fifth paper drove its crawl with Claude's Computer
 +Use API, which the extraction filed under category ''llm''; it is counted on the child
 +page and explained there.)) That is not enough to call it practice, and this page does not
 +recommend it as a default. It is enough for a page of its own, because those five — and
 +four neighbours that measure agents rather than crawl with them — have already published
 +completion rates, per-site costs and (on benchmark tasks) run-to-run variance that nobody
 +starting an agent crawl should have to rediscover.
 +**[[Programming:Crawler:LLM Agents]]** has those numbers, the tool currency, and what to
 +report.
  
 ===== Being Detected ===== ===== Being Detected =====
Line 345: Line 367:
 gives you away with, why cloaking is the shape that ends up in your abstract, how to gives you away with, why cloaking is the shape that ends up in your abstract, how to
 measure the obstruction instead of fighting it, and what the corpus says about how rarely measure the obstruction instead of fighting it, and what the corpus says about how rarely
-anyone reports it. Two points from there bear on the choice made //on this page//:+anyone reports it. Three points bear on the choice made //on this page//:
  
-  * **The framework is detectable, not just automation in general.** Vastel et al. +  * **The framework is detectable, not just automation in general.** Vastel et al. {[vastel2018_scanner]} show fingerprint inconsistencies distinguish instrumented browsers from ordinary ones, and Krumnow et al. {[krumnow2022_gullible]} show OpenWPM specifically has detectors deployed against it in the wild. 
-    {[vastel2018_scanner]} show fingerprint inconsistencies distinguish instrumented +  * **The anti-detection patches in the tables below are an arms race you lose quietly**, and most of those packages are now unmaintained or have moved to a successor under a different name; [[Programming:Crawler Detection]] dates each of them. Using them is also an ethics decision — see [[Practices:Ethics]] — because you are overriding a site operator's expressed access preference
-    browsers from ordinary ones, and Krumnow et al. {[krumnow2022_gullible]} show OpenWPM +  * **If you take the agent option above, you are on a different detection surface, not just a more visible one.** Bot management now scores an //agent acting on a user's behalf// separately from both a human and a crawler collecting training data; [[Programming:Crawler Detection]] names the vendor categories and dates their defaults, which were still moving through 2026. The rules you are judged by are those, and they are not the two above: [[Programming:Crawler:LLM Agents#being_detected_and_the_ethics_of_not_being|the child page's detection section]] has them, including three 2026 papers that took three different positions on whether to evade at all.
-    specifically has detectors deployed against it in the wild. +
-  * **The anti-detection patches in the tables below are an arms race you lose quietly**, +
-    and most of those packages are now unmaintained or have moved to a successor under a +
-    different name; [[Programming:Crawler Detection]] dates each of them. Using them is +
-    also an ethics decision — see [[Practices:Ethics]] — because you are overriding a +
-    site operator's expressed access preference.+
  
 ===== Use in Publications ===== ===== Use in Publications =====
Line 377: Line 393:
 ^ Family ^ Papers ^ Share of 1,120 ^ ^ Family ^ Papers ^ Share of 1,120 ^
 | Selenium | 242 | 21.6% | | Selenium | 242 | 21.6% |
-| //Bespoke crawler, given its own name// (''SSOScan'', ''AdFisher'', ''PhishPrint'', ''CryptoScamTracker''…) | 184 | 16.4% |+| //Bespoke crawler, given its own name// (''SSOScan'', ''AdFisher'', ''PhishPrint'', ''CryptoScamTracker''…) | 181 | 16.2% |
 | //Bespoke crawler, described generically// ("our crawler", "a custom Python crawler", "a crawling extension") | 147 | 13.1% | | //Bespoke crawler, described generically// ("our crawler", "a custom Python crawler", "a crawling extension") | 147 | 13.1% |
 | Puppeteer | 76 | 6.8% | | Puppeteer | 76 | 6.8% |
Line 398: Line 414:
 | Anti-detection patches (''stealth'', ''undetected-chromedriver'') | 9 | 0.8% | | Anti-detection patches (''stealth'', ''undetected-chromedriver'') | 9 | 0.8% |
 | Fuzzers and monkey testers | 5 | 0.4% | | Fuzzers and monkey testers | 5 | 0.4% |
 +| [[Programming:Crawler:LLM Agents|LLM browser agents]] (Browser Use, BrowserGym/AgentLab)((Added as a family on 2026-08-29. All four papers are 2026, i.e. entirely inside the provisional slice. The family is tested //before// Playwright and Puppeteer because several of these agents are built on those libraries, so a name matched in the wrong order would land in the library's row instead. A fifth crawling paper drove its crawl with an agent the extraction filed under category ''llm'', outside this table's population; [[Programming:Crawler:LLM Agents]] counts it and explains the boundary.)) | 4 | 0.4% |
 | webXray | 1 | 0.1% | | webXray | 1 | 0.1% |
  
 <WRAP important> <WRAP important>
-Put the two bespoke rows together — 10 papers are in both — and **321 of 1,120 crawling papers (28.7%) crawled with something home-grown or too obscure to have a family here.** That is more than used Selenium (242), and more than the number naming any of Puppeteer, Playwright, direct CDP, Scrapy or OpenWPM (215). Of the 184 that gave their crawler a proper name, **75 named nothing else at all** — the name is the only identification of the instrument in the paper, and it means nothing to a reader who does not have the code.+Put the two bespoke rows together — 10 papers are in both — and **318 of 1,120 crawling papers (28.4%) crawled with something home-grown or too obscure to have a family here.** That is more than used Selenium (242), and more than the number naming any of Puppeteer, Playwright, direct CDP, Scrapy or OpenWPM (215). Of the 181 that gave their crawler a proper name, **74 named nothing else at all** — the name is the only identification of the instrument in the paper, and it means nothing to a reader who does not have the code.
 </WRAP> </WRAP>
  
Line 457: Line 474:
 | Headless or headful | 140 | 12.5% | | Headless or headful | 140 | 12.5% |
 | Stateful or stateless | 219 | 19.6% | | Stateful or stateless | 219 | 19.6% |
-| Consent action | 349 | 31.2% |+| Consent action field populated (schema; not an audited paper claim) | 349 | 31.2% |
 | Interaction depth | 841 | 75.1% | | Interaction depth | 841 | 75.1% |
 | Authentication | 779 | 69.6% | | Authentication | 779 | 69.6% |
  
-^ Family ^ N ^ States headless ^ States statefulness ^ States consent action ^ Public artifact ^+The consent-action count above is field population, not proof that 349 papers stated a choice. The 2026-09-05 audit found **55 of 1,120 crawling papers (4.9%)** with an audited consent-action claim and **279 of 313 (89.1%)** unsupported ''no-interaction'' labels. The family percentages below retain the schema-field denominator and should be read the same way. 
 + 
 +^ Family ^ N ^ States headless ^ States statefulness ^ Consent field populated (schema) ^ Public artifact ^
 | Selenium | 242 | 21.1% | 27.7% | 36.0% | 57.0% | | Selenium | 242 | 21.1% | 27.7% | 36.0% | 57.0% |
 | Puppeteer | 76 | 35.5% | 35.5% | 57.9% | 67.1% | | Puppeteer | 76 | 35.5% | 35.5% | 57.9% | 67.1% |
Line 481: Line 500:
 | OpenWPM | 60 | 3 | 2015–2026 | | OpenWPM | 60 | 3 | 2015–2026 |
 | ''tbselenium'' / tor-browser-crawler | 21 | 0 | 2017–2026 | | ''tbselenium'' / tor-browser-crawler | 21 | 0 | 2017–2026 |
-| Tracker Radar Collector | 21 | 1 | 2021–2026 |+| Tracker Radar Collector((This row is the whole ''tracker.?radar'' fold, and the family is named after the crawler. Only **10** of the 21 ran the Collector; the other **11** used the //Tracker Radar dataset//, the entity list, the entity map or the wiki, filed by the extraction as a ''classification-service'' (9 papers), a ''cookie-database'' (1) or ''other'' (1). The folded-framework row above (10 papers, 0.9%) is the crawler-only count and agrees with [[Programming:Crawler:Tracker Radar Collector]]. Split by ''trc_index_reconcile.mjs''; the 11 papers are listed on [[provenance:programming:crawler:tracker_radar_collector]].)) | 21 | 1 | 2021–2026 |
 | Anti-detection patches | 11 | 1 | 2021–2026 | | Anti-detection patches | 11 | 1 | 2021–2026 |
 | VisibleV8 | 9 | 0 | 2019–2026 | | VisibleV8 | 9 | 0 | 2019–2026 |
Line 489: Line 508:
 | PanoptiChrome | 2 | 0 | 2024–2025 | | PanoptiChrome | 2 | 0 | 2024–2025 |
  
-**OpenWPM is still the field's only broadly shared instrument**, but by less than it was: extending the corpus to 2026 roughly **tripled Tracker Radar Collector (8 → 21) and quadrupled Foxhound (2 → 8)** while OpenWPM grew 53 → 60. The 2024 cut-off in the earlier version of this table was doing real work — it caught the newer instruments mid-growth and made them look more marginal than they are. Everything below OpenWPM is still either niche, new, or a tool used mainly by the group that built it.+**OpenWPM is still the field's only broadly shared instrument**, but by less than it was: extending the corpus to 2026 roughly **tripled the Tracker Radar family (8 → 21) and quadrupled Foxhound (2 → 8)** while OpenWPM grew 53 → 60. Read the Tracker Radar growth with its footnote: most of this family is people //using the dataset//, and the crawler itself is the 10-paper row in the folded table above. The 2024 cut-off in the earlier version of this table was doing real work — it caught the newer instruments mid-growth and made them look more marginal than they are. Everything below OpenWPM is still either niche, new, or a tool used mainly by the group that built it.
  
 ==== Methodology and limitations of these figures ==== ==== Methodology and limitations of these figures ====
  
   * **How they were produced.** One structured record per paper was extracted from full text, each tuple carrying a verbatim evidence quote and its section, so any figure here can be traced to the sentence that supports it. The script that produces every number on this page, with its denominators, is ''report_crawler.mjs''; the folding rules are in ''tool_fold.mjs''. Every query, the script's unedited output and the full residue are on [[provenance:programming:crawler]]; corpus-level caveats are on [[literature:corpus]].   * **How they were produced.** One structured record per paper was extracted from full text, each tuple carrying a verbatim evidence quote and its section, so any figure here can be traced to the sentence that supports it. The script that produces every number on this page, with its denominators, is ''report_crawler.mjs''; the folding rules are in ''tool_fold.mjs''. Every query, the script's unedited output and the full residue are on [[provenance:programming:crawler]]; corpus-level caveats are on [[literature:corpus]].
-  * **How names were folded.** Tool names are free text and agree run-to-run on only about a fifth of exact strings, so nothing here is counted by exact string. Names were folded into the families shown by an explicit, ordered list of regular expressions — specific tools before the generic libraries they are built on, so ''puppeteer-extra-plugin-stealth'' lands in //Anti-detection patches// and not in //Puppeteer//. Folding matters: the exact string ''Selenium'' appears in 194 papers, while the folded family covers 242, so counting exact strings would undercount Selenium by 19.8% — the difference is ''Selenium WebDriver'', ''Selenium Webdriver'', ''Python Selenium WebDriver'', ''selenium'', ''ChromeDriver'' and ''Selenium's ChromeDriver''. Of **1,075 tool mentions across 501 distinct strings** in the crawling population, **204 distinct strings across 210 mentions match no family**. They are not discarded — they are the //Bespoke crawler, given its own name// row, because almost all of them are one paper's own tool (''SSOScan'', ''AdFisher'', ''Formlock'', ''CryptoScamTracker'', ''Spider-Scents'': 184 papers, essentially one name each). A handful are third-party tools we chose not to give a family of their own (''OmniCrawl'', ''JAW'', ''BrowserStack'', ''MetaMask automator'', ''Headless Chromium'', the ''measurement framework of Demir et al.''), so read that row as "home-grown or obscure" rather than strictly "home-grown". ''report_crawler.mjs'' prints the full list, so nothing vanishes. Browser names folded to 1 unclassified string out of 529 papers (''Ghostery'', which is an extension, not a browser).+  * **How names were folded.** Tool names are free text and agree run-to-run on only about a fifth of exact strings, so nothing here is counted by exact string. Names were folded into the families shown by an explicit, ordered list of regular expressions — specific tools before the generic libraries they are built on, so ''puppeteer-extra-plugin-stealth'' lands in //Anti-detection patches// and not in //Puppeteer//. Folding matters: the exact string ''Selenium'' appears in 194 papers, while the folded family covers 242, so counting exact strings would undercount Selenium by 19.8% — the difference is ''Selenium WebDriver'', ''Selenium Webdriver'', ''Python Selenium WebDriver'', ''selenium'', ''ChromeDriver'' and ''Selenium's ChromeDriver''. Of **1,075 tool mentions across 501 distinct strings** in the crawling population, **199 distinct strings across 205 mentions match no family**. They are not discarded — they are the //Bespoke crawler, given its own name// row, because almost all of them are one paper's own tool (''SSOScan'', ''AdFisher'', ''Formlock'', ''CryptoScamTracker'', ''Spider-Scents'': 181 papers, essentially one name each). A handful are third-party tools we chose not to give a family of their own (''OmniCrawl'', ''JAW'', ''BrowserStack'', ''MetaMask automator'', ''Headless Chromium'', the ''measurement framework of Demir et al.''), so read that row as "home-grown or obscure" rather than strictly "home-grown". ''report_crawler.mjs'' prints the full list, so nothing vanishes. Browser names folded to 1 unclassified string out of 529 papers (''Ghostery'', which is an extension, not a browser).
   * **A paper counts once per family**, never once per mention, and shares do not sum to 100% because a paper can name several tools. A family's share is of the population named in its heading.   * **A paper counts once per family**, never once per mention, and shares do not sum to 100% because a paper can name several tools. A family's share is of the population named in its heading.
   * **Silence is not absence.** "Does not name a framework" means the paper did not say, not that the authors used none. These are reporting figures.   * **Silence is not absence.** "Does not name a framework" means the paper did not say, not that the authors used none. These are reporting figures.
Line 512: Line 531:
     * Multi-browser or multi-language, or a lab already fluent in it → Selenium.     * Multi-browser or multi-language, or a lab already fluent in it → Selenium.
     * The lowest-effort modern tracking crawl → Tracker Radar Collector.     * The lowest-effort modern tracking crawl → Tracker Radar Collector.
-  - **Do not write a new crawler for a solved problem.** Nearly three in ten crawling papers used a home-grown one, and 75 of them identify it by nothing but a name they invented. If you must build one, say what it is built on, and publish it.+  - **Do not write a new crawler for a solved problem.** Nearly three in ten crawling papers used a home-grown one, and 74 of them identify it by nothing but a name they invented. If you must build one, say what it is built on, and publish it.
   - **State headless or headful and justify it.** Headless is the most detectable configuration you can choose {[vastel2018_scanner]} and only 12.5% of papers say which they used.   - **State headless or headful and justify it.** Headless is the most detectable configuration you can choose {[vastel2018_scanner]} and only 12.5% of papers say which they used.
-  - **Pin the browser binary, not just the library.** Puppeteer and Playwright do this for you; Selenium does not. Record the exact build and archive it with your research artefact (see [[:Artifacts]], a page this wiki still owes you).+  - **Pin the browser binary, not just the library.** Puppeteer and Playwright do this for you; Selenium does not. Record the exact build and archive it with your research artefact (see [[:Artifacts]]).
   - **Validate against something that is not your crawler.** A manual visit to a sample, a second browser engine, or a second vantage point {[jueckstock2021_realistic]}. Crawls diverge from human browsing in ways your crawl cannot see {[zeber2020representativeness]}, and repeated crawls diverge from each other {[demir2022_reproducibility]}.   - **Validate against something that is not your crawler.** A manual visit to a sample, a second browser engine, or a second vantage point {[jueckstock2021_realistic]}. Crawls diverge from human browsing in ways your crawl cannot see {[zeber2020representativeness]}, and repeated crawls diverge from each other {[demir2022_reproducibility]}.
   - **Run more than once.** A single crawl is a point estimate. Report the spread.   - **Run more than once.** A single crawl is a point estimate. Report the spread.
programming/crawler.1787967699.txt.gz · Last modified: by karel.kubicek.claude