User Tools

Site Tools


programming:crawler

This is an old revision of the document!


Comparison of Crawling Libraries

Every automated web measurement makes two separate choices that papers routinely report as one: which browser renders the page, and which control channel drives it. A “Selenium crawl” says nothing about the first; a “Chrome crawl” says nothing about the second. They fail differently, they are detected differently, and they give you access to different data.

This page compares the options. It covers the generic automation libraries (Selenium, Puppeteer, Playwright, plain CDP) and then the specialised privacy and security crawlers built on top of them, each of which has — or will have — its own page: OpenWPM, webXray, Tracker Radar Collector, PageGraph, Foxhound, PanoptiChrome.

It pairs with Automated measurements (whether to crawl at all), Crawling location (where from), Stateful stateless (with or without a profile), Interaction (what to do on the page), and Consent (what to do with the banner).

The single most consequential thing on this page is not which library is best. It is that only 12.0% of the papers in our corpus that name an automation tool also state its version (see Use in Publications). “We used Selenium” spans fifteen years of incompatible releases and two different wire protocols. Report the version.

The Two Layers

  • The browser is what the website sees: a rendering engine, a JavaScript engine, a TLS stack, a fingerprint. Chrome and Firefox disagree about cookie partitioning, about tracking protection defaults, and about which APIs exist at all, so they see different tracking. OmniCrawl [1Cassel, Darion; Lin, Su-Chin; Buraggina, Alessio; Wang, William; Zhang, Andrew; Bauer, Lujo; Hsiao, Hsu-Chun; Jia, Limin; Libert, Timothy (2022): "OmniCrawl: Comprehensive Measurement of Web Tracking With Real Desktop and Mobile Browsers", in: Proceedings on Privacy Enhancing Technologies. (DOI)] drove 42 non-emulated desktop and mobile browsers simultaneously precisely because a single-browser crawl is a single-browser result.
  • The control channel is how your code talks to it. There are exactly three in practice:
    • W3C WebDriver — an HTTP protocol spoken by a per-browser driver process (chromedriver, geckodriver). Portable across browsers; classic WebDriver deliberately models a user, not a debugger, and has no network commands at all. Its successor WebDriver BiDi adds them, over a WebSocket, standardised and cross-browser — see Selenium and WebDriver BiDi.
    • The Chrome DevTools Protocol (CDP) — a WebSocket protocol into Chromium's internals. Gives you the network layer, the JS engine and the cookie jar. Chromium only.1)
    • Nothing — a patched browser build that instruments itself and writes its own logs. This is what the specialised crawlers do.

For most of the period this page's corpus covers, that produced a hard boundary: classic WebDriver cannot capture network requests. There is no such command in the specification. If your research question involves third-party requests, cookies set by responses, or JavaScript API calls — which is to say, most privacy measurement — classic WebDriver will not answer it, and you end up adding a proxy, a browser extension, or a CDP side-channel. That gap is a large part of why the specialised crawlers exist, and it is why a paper that says only “we used Selenium” leaves its instrument unidentified.

WebDriver BiDi closes the boundary, and it is new enough that essentially none of the corpus could have used it. If you are starting a crawler today it changes the recommendation; it does not change how to read the literature.

The Generic Libraries

Selenium 4 Puppeteer Playwright Plain CDP
Control channel classic WebDriver, WebDriver BiDi, or a CDP bridge on Chromium CDP CDP (Chromium), own protocol for its Firefox/WebKit builds CDP
Browsers Chrome, Firefox, Edge, Safari, and anything with a driver Chrome, Chromium, and Firefox — production-ready since v23, over WebDriver BiDi Chromium, Firefox, WebKit — Playwright's own builds; real Chrome/Edge via channel Chromium only
Languages Java, Python, JS, C#, Ruby, Rust… JS/TS (pyppeteer unmaintained) JS/TS, Python, Java, .NET any language with a WebSocket client
Network interception ✗ classic WebDriver; ✓ via BiDi or the CDP bridge
Response bodies ✗ classic; via BiDi/CDP otherwise
All cookies incl. HttpOnly ✓ — the driver returns them, unlike in-page JS ✓ (Network.getAllCookies)
Waits for the network to settle ✗ you sleep networkidle0 networkidle ✗ you sleep
Multiple isolated profiles per browser ✗ one session per browser process contexts contexts, cheap and first-class Target.createBrowserContext
Ships its own browser ✗ — Selenium Manager fetches a matching driver, and will fetch a browser if it finds none ✓ pinned Chrome for Testing build ✓ pinned, patched builds
Best for multi-language work; an existing Selenium codebase Chromium measurement; the base of most modern research tooling reproducibility, parallelism, cross-engine work minimal dependencies, exotic CDP domains, other languages

Selenium

Selenium is the field's default: 21.6% of all crawling papers in the corpus name it, and it has been the most-used off-the-shelf library in every four-year window since 2014. Its virtues are real — it drives Firefox properly, it has first-class bindings in Python and Java, and code written against it in 2016 mostly still runs.

Its limitation was architectural. Classic WebDriver was designed to test websites, so it exposes what a user can do and hides what a debugger can see. The older escape hatch is a CDP session tunnelled over the WebDriver connection, which works but is Chromium-only and gives up the portability you chose Selenium for; the modern answer is WebDriver BiDi. This is worth stating plainly because it is the most common silent methodological compromise in the corpus: a paper says “Selenium” and reports third-party requests, and the reader cannot tell whether those came from a proxy, an extension, performance.getEntriesByType(“resource”), or CDP — four instruments with four different blind spots.

selenium.mjs
// Selenium 4 (W3C WebDriver) driving Chromium through chromedriver.
//
// The point of this script is what is missing from it. Classic WebDriver has no
// network-interception command, so there is no `page.on('request')` to write:
// cookies and the DOM you can have, the request log you cannot. What follows is
// the two escape hatches researchers actually used before WebDriver BiDi.
import { Builder, Browser } from 'selenium-webdriver';
import chrome from 'selenium-webdriver/chrome.js';
 
const URL = process.argv[2] ?? 'http://127.0.0.1:8099/';
const t0 = Date.now();
 
const options = new chrome.Options();
options.setChromeBinaryPath('/path/to/chrome');
options.addArguments('--headless=new', '--no-sandbox', '--disable-dev-shm-usage');
const service = new chrome.ServiceBuilder('/path/to/chromedriver');
 
const driver = await new Builder()
  .forBrowser(Browser.CHROME)
  .setChromeOptions(options)
  .setChromeService(service)
  .build();
 
const requests = [];
try {
  // Escape hatch 1: a raw CDP session tunnelled over the WebDriver connection.
  // Chromium only -- this call does not exist for Firefox or Safari.
  const cdp = await driver.createCDPConnection('page');
  await cdp.execute('Network.enable', {});
  cdp._wsConnection.on('message', (raw) => {
    const msg = JSON.parse(raw.toString());
    if (msg.method === 'Network.requestWillBeSent')
      requests.push({ method: msg.params.request.method, url: msg.params.request.url, type: msg.params.type });
  });
 
  await driver.get(URL);
  await driver.sleep(1500);
 
  const title = await driver.getTitle();
  const cookies = await driver.manage().getCookies();
  // Escape hatch 2: whatever the page itself exposes. Portable across browsers,
  // but the page can erase it with performance.clearResourceTimings().
  const perf = await driver.executeScript(
    'return performance.getEntriesByType("resource").map(e => ({name: e.name, type: e.initiatorType}))'
  );
 
  console.log(JSON.stringify({ title, requestsViaCDP: requests, cookies, resourceTimingEntries: perf, ms: Date.now() - t0 }, null, 1));
} finally {
  await driver.quit();
}

That CDP bridge is racy. In five runs of the script above against an identical local page, four recorded all four requests and one recorded only the main document. We have not isolated the mechanism — the listener is attached before the navigation, so this is not simply “subscribed too late” — but the outcome reproduced on a second machine, and it is a silent 75% data loss with no error.2) If you build on this, assert on the number of requests you expected rather than trusting the log.

Selenium and WebDriver BiDi

The paragraph above is the historical picture, and it is the one the literature was written under. It is no longer the whole story: WebDriver BiDi is a bidirectional WebSocket protocol standardised alongside WebDriver, and Selenium exposes it directly. Turn it on with options.enableBidi() and you subscribe to network events without CDP and without giving up cross-browser portability, because BiDi is a W3C protocol that Firefox implements too.

selenium_bidi.mjs
// Selenium 4 using WebDriver BiDi rather than classic WebDriver.
//
// MODE = 'req' | 'resp' | 'both'. With 'both', every callback receives both event
// types in selenium-webdriver 4.46 -- see the table below the listing. Nothing
// here is asserted: the response data is read off the events.
import { Builder, Browser } from 'selenium-webdriver';
import chrome from 'selenium-webdriver/chrome.js';
import { Network } from 'selenium-webdriver/bidi/network.js';
 
const MODE = process.argv[2] ?? 'both';
 
const options = new chrome.Options();
options.setChromeBinaryPath('/path/to/chrome');
options.addArguments('--headless', '--no-sandbox', '--disable-dev-shm-usage');
options.enableBidi(); // sets the webSocketUrl capability
const service = new chrome.ServiceBuilder('/path/to/chromedriver');
 
const driver = await new Builder()
  .forBrowser(Browser.CHROME)
  .setChromeOptions(options)
  .setChromeService(service)
  .build();
 
const requestEvents = [];
const responseEvents = [];
try {
  const network = await Network(driver);
  if (MODE !== 'resp')
    await network.beforeRequestSent((e) => requestEvents.push({ method: e.request.method, url: e.request.url }));
  if (MODE !== 'req')
    await network.responseCompleted((e) =>
      // e.response is undefined on the cross-delivered request events, so read
      // what is actually there rather than what should be.
      responseEvents.push({ url: e.request.url, status: e.response === undefined ? null : e.response.status })
    );
 
  await driver.get('http://127.0.0.1:8099/');
  await driver.sleep(1500);
 
  console.log({
    requestEvents: requestEvents.length,
    responseEvents: responseEvents.length,
    responseEventsWithData: responseEvents.filter((r) => r.status !== null).length,
    statuses: responseEvents.map((r) => r.status),
    uniqueUrls: [...new Set([...requestEvents, ...responseEvents].map((r) => r.url))],
  });
} finally {
  await driver.quit();
}

Subscribe to one network event and it is exactly right. Subscribe to two and, in selenium-webdriver 4.46, every callback starts receiving both event types. Measured on the fixture page, whose four requests we know:

Handlers registered beforeRequestSent fired responseCompleted fired …of which carried response data Statuses seen
beforeRequestSent only 4
responseCompleted only 4 4 200, 200, 200, 204
both 8 8 4 null, 200, null, null, 200, 200, null, 204

Both counts double, and half the “responses” are request events wearing a response event's clothing: e.response is undefined for them, so the obvious e.response.status throws on roughly every other event. Nothing warns you.

Two consequences, in order of how badly they bite:

  1. Do not paper over this by deduplicating on URL. Of the two copies of each event, one carries the response data and one does not, and a URL-keyed dedupe has no way to prefer the right one — you can keep the empty copy and conclude the status was unavailable. We did exactly that, and published the wrong claim before catching it. Filter on the field you need (e.response !== undefined) instead.
  2. Register one handler, or verify the counts. The straightforward workaround is to subscribe only to responseCompleted, which carries the request and the response. If you need both events, assert your event count against a page whose requests you already know before you point the crawler at a hundred thousand sites.

So BiDi is the right direction and, in this binding, not yet a drop-in replacement. The protocol is not at fault: the W3C specification defines a response field of type network.ResponseData on network.responseCompleted, and Chromium delivers it. Check what your binding does with it.

Puppeteer

Puppeteer is a thin, well-documented wrapper over CDP maintained by the Chrome team. It appears in 6.8% of crawling papers, first in 2018, and it is what most of the modern research tooling is built on — DuckDuckGo's Tracker Radar Collector and Brave's pagegraph-crawl are both Puppeteer programs. Recent versions download and pin a Chrome for Testing build, which fixes the “which Chrome was that?” reproducibility problem for free, provided you record the version.

Puppeteer used to be Chromium-and-CDP all the way down. It is not any more: it supports Firefox over WebDriver BiDi, and when you launch Firefox, BiDi is the default. Chrome still defaults to CDP “since not all CDP features are supported by WebDriver BiDi yet”, and a Puppeteer feature that BiDi cannot express throws UnsupportedOperation.3) So Puppeteer is now a plausible cross-engine choice, with the caveat that the two engines do not give you the same API surface — check that the calls your measurement depends on work on both before you report a cross-browser comparison.

puppeteer.mjs
import puppeteer from 'puppeteer-core';
 
const URL = process.argv[2] ?? 'http://127.0.0.1:8099/';
const t0 = Date.now();
const browser = await puppeteer.launch({
  executablePath: '/path/to/chrome',
  headless: true,
  args: ['--no-sandbox'],
});
const page = await browser.newPage();
 
const requests = [];
page.on('request', (r) => requests.push({ method: r.method(), url: r.url(), type: r.resourceType() }));
const responses = [];
page.on('response', (r) => responses.push({ status: r.status(), url: r.url() }));
 
await page.goto(URL, { waitUntil: 'networkidle0' });
const cookies = await browser.defaultBrowserContext().cookies();
const title = await page.title();
await browser.close();
 
console.log(JSON.stringify({ title, requests, responses, cookies, ms: Date.now() - t0 }, null, 1));

Playwright

Playwright is the newest of the three and the only one that drives three different engines — Chromium, Firefox and WebKit — through browser builds it ships itself, pinned per Playwright version. For a measurement study that is the interesting property: the browser binary is part of the dependency, so naming the Playwright version pins the engine too, and a cross-engine comparison stops requiring three separate harnesses.

It is the youngest tool on this page — 34 papers (3.0% of crawling papers), all of them 2022 or later — and the only one whose adoption curve is still going up steeply: in the provisional 2025–2026 window it overtakes Puppeteer, 9.6% against 8.1%. Its users also report its version better than most, though the margin is narrow: 26.5%, just ahead of OpenWPM's 25.9% and more than double everyone else (see Use in Publications).

Three caveats for measurement work.

  1. What Playwright ships is not what your subjects run. By default it uses open-source Chromium builds, which are ahead of released Chrome: “when the world is on Google Chrome N, Playwright already supports Chromium N+1”.4) Firefox and WebKit ship as Playwright's own builds rather than Mozilla's or Apple's. If the browser itself is the object of study — its fingerprint, its tracking protection, its defaults — use channel: “chrome” / “msedge” to drive the real installed browser, and say which you did.
  2. Headed and headless are different binaries. Playwright ships a full Chromium for headed runs and a separate chromium_headless_shell for headless ones. The headless mode is therefore not simply “the same browser without a window”, which matters if you are asking whether headless is detectable.
  3. The default context is empty. A fresh isolated profile per run is excellent hygiene and quietly makes your crawl stateless unless you say otherwise.
playwright.mjs
import { chromium } from 'playwright';
 
const URL = process.argv[2] ?? 'http://127.0.0.1:8099/';
const t0 = Date.now();
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext();
const page = await context.newPage();
 
const requests = [];
page.on('request', (r) => requests.push({ method: r.method(), url: r.url(), type: r.resourceType() }));
const responses = [];
page.on('response', (r) => responses.push({ status: r.status(), url: r.url() }));
 
await page.goto(URL, { waitUntil: 'networkidle' });
const cookies = await context.cookies();
const title = await page.title();
await browser.close();
 
console.log(JSON.stringify({ title, requests, responses, cookies, ms: Date.now() - t0 }, null, 1));

Plain CDP

You can skip the libraries entirely: launch Chromium with –remote-debugging-port, open a WebSocket, enable the domains you want, and read the events. 42 papers (3.8%) did, most of them between 2018 and 2021 — a window that closes as Puppeteer matures.

Reach for it when your analysis lives in a language with no good automation binding, when you need a CDP domain the wrapper does not expose, or when you want to know exactly what your instrument is doing. The cost, visible in the code below, is that everything a library gives you is now yours to write: waiting for the port, waiting for the load, waiting for the network to settle, and reconnecting when a target dies.

cdp.mjs
// Plain Chrome DevTools Protocol: no automation library at all.
import { spawn } from 'node:child_process';
import CDP from 'chrome-remote-interface';
 
const URL = process.argv[2] ?? 'http://127.0.0.1:8099/';
const PORT = 9222;
const t0 = Date.now();
 
const proc = spawn('/path/to/chrome', [
  '--headless',
  `--remote-debugging-port=${PORT}`,
  '--no-sandbox',
  '--user-data-dir=/tmp/cdp-profile',
  'about:blank',
]);
 
// No launcher means you wait for the port yourself.
async function waitForPort(retries = 60) {
  for (let i = 0; i < retries; i++) {
    try {
      return await CDP.Version({ port: PORT });
    } catch {
      await new Promise((r) => setTimeout(r, 250));
    }
  }
  throw new Error('devtools port never opened');
}
const version = await waitForPort();
 
const client = await CDP({ port: PORT });
const { Network, Page, Runtime } = client;
const requests = [];
const responses = [];
Network.requestWillBeSent(({ request, type }) => requests.push({ method: request.method, url: request.url, type }));
Network.responseReceived(({ response }) => responses.push({ status: response.status, url: response.url }));
 
await Promise.all([Network.enable(), Page.enable(), Runtime.enable()]);
const loaded = new Promise((r) => Page.loadEventFired(r));
await Page.navigate({ url: URL });
await loaded;
await new Promise((r) => setTimeout(r, 1000)); // no networkidle helper: you sleep
const { cookies } = await Network.getAllCookies();
const { result } = await Runtime.evaluate({ expression: 'document.title' });
 
await client.close();
proc.kill();
 
console.log(JSON.stringify({ browser: version['Browser'], title: result.value, requests, responses, cookies, ms: Date.now() - t0 }, null, 1));

What We Ran

The four scripts above were run against an identical local fixture page — one HTML document, one image, one script that sets a cookie and fires a fetch(), one Set-Cookie response header on each — so that the differences in the table below are differences between instruments and not between page loads. Chromium 151.0.7922.34 (Playwright's build), Node 22.23, aarch64 Linux, 2026-08-06.5)

Library Lines6) Requests seen Response statuses Cookies Saw the HttpOnly cookie Wall clock
Playwright 15 4 4 3 yes 779 ms
Puppeteer 18 4 4 3 yes 877 ms
Selenium 4, WebDriver BiDi 36 4 unique, from 8 events 4 of 8 response events carried one 3 yes 2038 ms
Selenium 4, CDP bridge 33 4 (1 in one run of five) 0 (not exposed) 3 yes 1838 ms
Plain CDP 39 4 4 3 yes 1329 ms

Three things in that table are worth more than the timings:

  1. Both Selenium paths cost more code and delivered less. Neither returned response status codes: the CDP bridge does not surface them without extra handlers, and the BiDi response event does not carry them at all. Selenium's portable fallback, the Resource Timing API, returned 3 entries rather than 4 — it does not report the main document, and a page can erase it with performance.clearResourceTimings().
  2. All four saw the HttpOnly cookie, which is the one thing a JavaScript-based instrument (document.cookie, an in-page script) cannot see. If a paper reports cookie counts collected in-page, its numbers exclude session cookies by construction.
  3. The wall-clock ordering is mostly setup cost, not throughput. Do not choose a library on these numbers; choose it on the columns of the comparison table.

Sandbox gotchas we hit, since they cost half a day. On this aarch64 container, every full Chrome/Chromium binary — Debian's /usr/bin/chromium, Chrome for Testing, and Playwright's own chromium-* build — died at startup with chrome_crashpad_handler: –database is required and never opened the DevTools port. Only Playwright's chromium_headless_shell build started. Chrome for Testing ships no aarch64 build (the downloaded linux_arm binary is x86-64 and fails under emulation), and neither does chromedriver; a working aarch64 chromedriver came from Debian's chromium-driver package, extracted with dpkg -x without root. Firefox 153 + geckodriver 0.37.1 never handed over its Marionette port in this container, with or without Xvfb, so the Selenium example is the Chromium one. None of this is a property of the libraries — it is a property of running x86-64-first browser tooling on ARM in a container, which is increasingly where CI lives.

Specialised Measurement Crawlers

The generic libraries give you the network layer and the DOM. They do not tell you which script set a cookie, which fingerprinting API was called with what arguments, or how a value flowed from document.cookie into a request URL. Getting that means instrumenting the browser itself, and that is the entire business of the tools below.

Tool Base Automation What it adds over a generic library Output Setup Papers7)
OpenWPM [2Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] unbranded Firefox Selenium + geckodriver JS API call monitoring, HTTP request/response, navigation, DNS and cookie instruments, configurable per browser8); profile management SQLite or Parquet for structured data, LevelDB or gzip for response bodies, S3/GCS remotes Python/conda, one install script; no Windows9) 60
Tracker Radar Collector Chromium Puppeteer modular collectors (requests, cookies, API calls, screenshots, cookie popups); built-in autoconsent opt-in/opt-out; mobile emulation; SOCKS proxy one JSON per site + metadata.json npm i; easiest on this list 21
Brave PageGraph patched Brave (Blink + V8) pagegraph-crawl (Node) a causal graph: every DOM mutation, script execution and request attributed to its cause, plus Shields filter-rule effects; the production successor to AdGraph [3Iqbal, Umar; Snyder, Peter; Zhu, Shitong; Livshits, Benjamin; Qian, Zhiyun; Shafiq, Zubair (2020): "AdGraph: A Graph-Based Approach to Ad and Tracker Blocking", in: 2020 IEEE Symposium on Security and Privacy (SP), pp. 763-776. (DOI)] GraphML in Brave ≥ 1.46, but JS-API recording needs Brave Nightly 8
VisibleV8 [4Jueckstock, Jordan; Kapravelos, Alexandros (2019): "VisibleV8: In-browser Monitoring of JavaScript in the Wild", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] patched Chromium (V8) any CDP library native-code JS API tracing inside V8, below the reach of page JavaScript append-only trace logs custom Chromium build 9
SAP Project Foxhound patched Firefox (Gecko + SpiderMonkey) Playwright10) dynamic taint tracking: string-level flows from sources to sinks __taintreport DOM events full Firefox build toolchain, or prebuilt binaries 8
PanoptiChrome [5Kanyal, Rahul; Sarangi, Smruti R. (2024): "PanoptiChrome: A Modern In-browser Taint Analysis Framework", in: Proceedings of the ACM Web Conference. (DOI)] patched Chromium/V8, pinned to Chrome 116 dynamic taint tracking in Chromium, arbitrary sources and sinks see the paper build a pinned Chrome 116 fork from patches 2
webXray [6Libert, Timothy (2015): "Exposing the Invisible Web: An Analysis of Third-Party HTTP Requests on 1 Million Websites", International Journal of Communication 9. (Link)] PhantomJS (historically) third-party request detection plus domain-to-company attribution MySQL, CSV reports see below 7

Notes that matter more than the table:

  • OpenWPM is still Firefox and still Selenium. Its README's first paragraph says so, and it pins an unbranded Firefox build.11) It is actively maintained. If your study is about Chrome's behaviour, OpenWPM is measuring a different browser — a point Krumnow et al. [7Krumnow, Benjamin; Jonker, Hugo; Karsch, Stefan (2022): "How gullible are web measurement tools?", in: Proceedings of the 18th International Conference on Emerging Networking EXperiments and Technologies, pp. 171-186. (DOI)] press further by analysing how detectable OpenWPM is, how resilient its recording is, and how widespread OpenWPM-specific detection is in the wild.
  • PageGraph does not track your automation. Its own documentation warns that “PageGraph currently does not track puppeteer / automation scripts, and so modifying or interacting with the document through devtools/puppeteer while recording a PageGraph file will likely fail”.12) It is a recorder of what the page did, not a driver of interaction. Its own limitations list is unusually honest and worth reading before you rely on it: no WebSocket tracking, no worker tracking, request headers not recorded (responses only), no style= mutation tracking.13)
  • Tracker Radar Collector caps concurrency at 38 crawlers, injects a simple anti-bot-detection script into every frame by default, and can be pointed at a Selenium Hub or a specific Chromium version.14) It has no canonical academic paper; cite the repository and the Tracker Radar dataset.
  • webXray needs care. The repository its own documentation points at, github.com/timlib/webXray, no longer exists — the account has no public repositories at all, and there is no webxray package on PyPI. What survives on GitHub are stale third-party mirrors, the newest of which was last pushed in 2015 and targets PhantomJS, whose own maintainer suspended development in 2018.15) webxray.org is a live landing page with no source link.16) Treat the 2015 paper [6Libert, Timothy (2015): "Exposing the Invisible Web: An Analysis of Third-Party HTTP Requests on 1 Million Websites", International Journal of Communication 9. (Link)] as the citation for the method — third-party request measurement with company attribution — and not as a tool you can install today. If you know where webXray is currently developed, please correct this.
  • Taint tracking is a different question. Foxhound and PanoptiChrome answer “did this value reach that sink?”, not “what did the page load”. If your question is about tracking prevalence, they are the wrong instrument and cost you a browser build; if it is about how data escapes — client-side XSS, DOM-based leaks, fingerprinting inputs — nothing else answers it. Foxhound tells you what to cite: its README's “Cite us!” section asks for the EuroS&P paper in which the browser is described, Klein et al. [8Klein, David; Barber, Thomas; Bensalim, Souphiane; Stock, Ben; Johns, Martin (2022): "Hand Sanitizers in the Wild: A Large-scale Study of Custom JavaScript Sanitizer Functions", in: 2022 IEEE 7th European Symposium on Security and Privacy (EuroS&P), pp. 236-250. (DOI)], and its wiki separately lists the papers that have used it.

Being Detected

Whatever you drive, the website may notice. This is a measurement-validity problem, not just an engineering nuisance: a bot-managed site serves your crawler a different page, and that difference is silently attributed to whatever you were studying.

  • Automation is detectable in the browser. Vastel et al. [9Vastel, Antoine; Laperdrix, Pierre; Rudametkin, Walter; Rouvoy, Romain (2018): "Fp-Scanner: The Privacy Implications of Browser Fingerprint Inconsistencies", in: Proceedings of the USENIX Security Symposium. (Link)] show that fingerprint inconsistencies distinguish instrumented and spoofed browsers from ordinary ones; headless Chrome and driver-injected properties are among the easiest signals to read.
  • Specific frameworks are detectable specifically. Krumnow et al. [7Krumnow, Benjamin; Jonker, Hugo; Karsch, Stefan (2022): "How gullible are web measurement tools?", in: Proceedings of the 18th International Conference on Emerging Networking EXperiments and Technologies, pp. 171-186. (DOI)] study OpenWPM detection in the wild. A tool that 58 papers in this corpus share is a tool worth writing a detector for.
  • Crawls differ from humans even when nobody is detecting you. Zeber et al. [10Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] quantified how far automated crawls diverge from real browsing on common tracking and fingerprinting metrics, and found crawls fail to capture the diversity of user environments.
  • The vantage point compounds this, since datacenter IP ranges are themselves a low-trust signal [11Jueckstock, Jordan; Sarker, Shaown; Snyder, Peter; Beggs, Aidan; Papadopoulos, Panagiotis; Varvello, Matteo; Livshits, Benjamin; Kapravelos, Alexandros (2021): "Towards Realistic and Reproducible Web Crawl Measurements", in: Proceedings of the ACM Web Conference. (DOI)] — see Crawling location.
  • Repeating a crawl is not free either. Demir et al. [12Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)] found substantial variation between repetitions of the same web measurement, so a single crawl of a single configuration is a point estimate with unstated error bars.

Nine papers in the corpus reach for explicit anti-detection patches (puppeteer-extra-plugin-stealth, undetected-chromedriver), all of them 2021 or later. Two warnings if you follow them. They are an arms race you will lose quietly and without notice, so any result that depends on them needs a validity check that does not. And evading bot management is a decision with an ethics dimension — see Ethics — because you are deliberately overriding a site operator's expressed access preference.

Use in Publications

The figures below come from a structured extraction over 5,859 full-text papers from CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf and IEEE S&P, 2010–2026. Unless stated otherwise the population is the 1,120 papers that ran a crawl, and sentinel values (not-stated) are counted as what they are rather than as answers. The 2025 and 2026 venue-years are provisional — CCS and IMC 2026 have not been held, and IEEE S&P and WWW 2026 abstracts are not yet in the selection source — so any per-year row that reaches them is under-represented by construction. Methodology and limitations are at the end of this section.

Most papers do not say what drove the browser

What the paper states Papers Share of 1,120
Names an automation framework or library 723 64.6%
Names a browser 529 47.2%
Names both 443 39.6%
Names neither 311 27.8%

More than a quarter of crawling papers describe neither layer. Naming the browser is rarer than naming the library, which is the wrong way round: the browser is what the website reacts to.

Which framework

Folded into families (see the note on folding) and counted by paper:

Family Papers Share of 1,120
Selenium 242 21.6%
Bespoke crawler, given its own name (SSOScan, AdFisher, PhishPrint, CryptoScamTracker…) 184 16.4%
Bespoke crawler, described generically (“our crawler”, “a custom Python crawler”, “a crawling extension”) 147 13.1%
Puppeteer 76 6.8%
OpenWPM 58 5.2%
Chrome DevTools Protocol, used directly 42 3.8%
Playwright 34 3.0%
PhantomJS and other legacy scriptable browsers 32 2.9%
Vulnerability and state-space crawlers (Crawljax, Black Widow, ZAP…) 29 2.6%
Scrapy 21 1.9%
Tor Browser Crawler / tbselenium 20 1.8%
App-store and platform scrapers 18 1.6%
HTTP-level scrapers, no browser (BeautifulSoup, requests) 17 1.5%
Web-performance harnesses (WebPageTest, Browsertime, Lighthouse) 16 1.4%
Third-party browser extension used as the instrument (Ghostery, uBlock Origin, NoScript…)17) 14 1.3%
OS-level GUI automation (pyautogui, xdotool) 14 1.3%
Archival crawlers (Heritrix, HTTrack, wget -r) 12 1.1%
Tracker Radar Collector 10 0.9%
Consent-interaction crawlers (BannerClick, Consent-O-Matic, Priv-Accept) 10 0.9%
Ad and tracker collection crawlers 9 0.8%
Anti-detection patches (stealth, undetected-chromedriver) 9 0.8%
Fuzzers and monkey testers 5 0.4%
webXray 1 0.1%

Put the two bespoke rows together — 10 papers are in both — and 321 of 1,120 crawling papers (28.7%) crawled with something home-grown or too obscure to have a family here. That is more than used Selenium (242), and more than the number naming any of Puppeteer, Playwright, direct CDP, Scrapy or OpenWPM (215). Of the 184 that gave their crawler a proper name, 75 named nothing else at all — the name is the only identification of the instrument in the paper, and it means nothing to a reader who does not have the code.

And when

Framework 2010–2013 (n=102) 2014–2017 (n=167) 2018–2021 (n=308) 2022–2024 (n=345) 2025–2026 (provisional) (n=198)
Selenium 4 (3.9%) 33 (19.8%) 76 (24.7%) 81 (23.5%) 48 (24.2%)
Puppeteer 0 0 21 (6.8%) 39 (11.3%) 16 (8.1%)
Playwright 0 0 0 15 (4.3%) 19 (9.6%)
CDP used directly 0 1 (0.6%) 23 (7.5%) 15 (4.3%) 3 (1.5%)
PhantomJS and other legacy scriptable browsers 3 (2.9%) 16 (9.6%) 10 (3.2%) 1 (0.3%) 2 (1.0%)
OpenWPM 0 7 (4.2%) 18 (5.8%) 26 (7.5%) 7 (3.5%)

The shape of the field, in one table. Selenium rose to a fifth of crawling papers by 2014–2017 and has sat at a quarter ever since — it is not being displaced, it is being supplemented. PhantomJS collapsed from 9.6% to 0.3% once headless Chrome shipped in 2017, which is the cleanest example available of a measurement tool becoming a liability: papers using it were measuring a WebKit engine that no user ran. Direct CDP use peaked in 2018–2021 and has fallen away since, absorbed by Puppeteer. Playwright's first appearance in this corpus is 2022, and in 2025–2026 it overtakes Puppeteer (9.6% against 8.1%) — the one clear reordering the extended corpus produced. Read that last column with the provisional caveat above: it rests on two venue-years that are incomplete by construction.

Which browser

529 of the 1,120 crawling papers (47.2%) name at least one browser. (1,080 of the 1,120 have a crawl-configuration record at all, so the missing 40 could not have named one; the denominator below stays 1,120 either way, because a paper that says nothing counts as saying nothing.)

Browser family Papers Share of 529 naming a browser
Chrome / Chromium 304 57.5%
Firefox 155 29.3%
Names a library, not a browser (“Selenium”, “Puppeteer”, “Playwright-controlled browser”) 49 9.3%
Tor Browser 38 7.2%
PhantomJS and other headless shells 24 4.5%
Internet Explorer 16 3.0%
Unnamed or custom browser (“a real browser”, “instrumented Chromium”) 16 3.0%
Safari / WebKit 16 3.0%
Edge 15 2.8%
Brave 10 1.9%
Opera 9 1.7%
Mobile device or WebView 8 1.5%
Other named browser (Whale, Kiwi, QQ, UC…) 5 0.9%
Messenger web client (WhatsApp, Telegram, WeChat) 2 0.4%

Two rows are findings rather than data. 49 papers answered “which browser?” with the name of a library, which is exactly the conflation this page opened with — and the row grew faster than the corpus did, because the 2025–2026 papers write “a Playwright-controlled browser” where older ones wrote “Chrome”. And Safari/WebKit is 3.0% of the papers that name a browser, far below any published estimate of its real-world usage share — the engine that behaves least like the other two, and whose tracking protection is on by default, is the one the field almost never measures. OmniCrawl [1Cassel, Darion; Lin, Su-Chin; Buraggina, Alessio; Wang, William; Zhang, Andrew; Bauer, Lujo; Hsiao, Hsu-Chun; Jia, Limin; Libert, Timothy (2022): "OmniCrawl: Comprehensive Measurement of Web Tracking With Real Desktop and Mobile Browsers", in: Proceedings on Privacy Enhancing Technologies. (DOI)] is the notable exception.

Almost nobody reports the version

Family Papers naming it of which state a version Share
Selenium 242 27 11.2%
Puppeteer 76 10 13.2%
Playwright 34 9 26.5%
CDP used directly 42 2 4.8%
PhantomJS and other legacy scriptable browsers 32 5 15.6%
OpenWPM 58 15 25.9%
Tor Browser Crawler / tbselenium 20 3 15.0%
Any automation tool 723 87 12.0%

This is the largest reporting gap on the page and the cheapest to fix. Playwright's cohort is best at it (26.5%) and OpenWPM's is a fraction behind (25.9%) — both more than double the 12.0% average, and in OpenWPM's case plausibly because its own releases are numbered and cited. On the previous corpus Playwright led by a wide margin on a base of ten papers; with 34 the two are effectively tied. Everyone else writes “we used Selenium”.

Configuration reporting, and whether the tool predicts it

Configuration detail States it Share of 1,120
Headless or headful 140 12.5%
Stateful or stateless 219 19.6%
Consent action 349 31.2%
Interaction depth 841 75.1%
Authentication 779 69.6%
Family N States headless States statefulness States consent action Public artifact
Selenium 242 21.1% 27.7% 36.0% 57.0%
Puppeteer 76 35.5% 35.5% 57.9% 67.1%
Playwright 34 26.5% 35.3% 47.1% 70.6%
CDP used directly 42 21.4% 33.3% 59.5% 69.0%
PhantomJS and other legacy scriptable browsers 32 40.6% 21.9% 37.5% 50.0%
OpenWPM 58 17.2% 55.2% 51.7% 55.2%
Tor Browser Crawler / tbselenium 20 10.0% 20.0% 30.0% 70.0%

Read the rows against each other, not against 100%. OpenWPM papers report statefulness at nearly twice the rate of everyone else (55.2%), which is what you would expect from a tool whose configuration file makes the stateful/stateless choice explicit and whose documentation names it — the instrument shapes what gets written down. Conversely OpenWPM papers report headless least often (17.2%), and its configuration suggests why: OpenWPM's display_mode takes three values — native, headless and xvfb18) — so a paper reporting “we ran under Xvfb” has stated something the headless/headful dichotomy does not have a slot for, and the extractor scores it as unstated.

The Public artifact column is confounded by year and should not be read as an effect of the tool: Playwright and Puppeteer papers are recent, and among crawling papers the share releasing a public artifact rose from 21.6% (2010–2013) to 61.2% (2022–2024), and to 71.7% in the provisional 2025–2026 window, for reasons that have nothing to do with crawler choice.

Which specialised crawlers actually get used

Matched by name across every tool mention in all 5,859 papers, so these counts include uses outside a formal crawl configuration and are not restricted to the 1,120. The two rows that also appear in the folded framework table use the same patterns there, so the two tables agree:

Tool Papers using or producing it Papers only citing it Years used
OpenWPM 60 3 2015–2026
tbselenium / tor-browser-crawler 21 0 2017–2026
Tracker Radar Collector 21 1 2021–2026
Anti-detection patches 12 2 2021–2026
VisibleV8 9 0 2019–2026
Brave PageGraph 8 0 2020–2025
SAP Project Foxhound 8 1 2024–2026
webXray 7 0 2018–2022
PanoptiChrome 2 0 2024–2025

OpenWPM is still the field's only broadly shared instrument, but by less than it was: extending the corpus to 2026 roughly tripled Tracker Radar Collector (8 → 21) and quadrupled Foxhound (2 → 8) while OpenWPM grew 53 → 60. The 2024 cut-off in the earlier version of this table was doing real work — it caught the newer instruments mid-growth and made them look more marginal than they are. Everything below OpenWPM is still either niche, new, or a tool used mainly by the group that built it.

Methodology and limitations of these figures

  • How they were produced. One structured record per paper was extracted from full text, each tuple carrying a verbatim evidence quote and its section, so any figure here can be traced to the sentence that supports it. The script that produces every number on this page, with its denominators, is report_crawler.mjs; the folding rules are in tool_fold.mjs. Every query, the script's unedited output and the full residue are on crawler; corpus-level caveats are on corpus.
  • How names were folded. Tool names are free text and agree run-to-run on only about a fifth of exact strings, so nothing here is counted by exact string. Names were folded into the families shown by an explicit, ordered list of regular expressions — specific tools before the generic libraries they are built on, so puppeteer-extra-plugin-stealth lands in Anti-detection patches and not in Puppeteer. Folding matters: the exact string Selenium appears in 194 papers, while the folded family covers 242, so counting exact strings would undercount Selenium by 19.8% — the difference is Selenium WebDriver, Selenium Webdriver, Python Selenium WebDriver, selenium, ChromeDriver and Selenium's ChromeDriver. Of 1,075 tool mentions across 501 distinct strings in the crawling population, 204 distinct strings across 210 mentions match no family. They are not discarded — they are the Bespoke crawler, given its own name row, because almost all of them are one paper's own tool (SSOScan, AdFisher, Formlock, CryptoScamTracker, Spider-Scents: 184 papers, essentially one name each). A handful are third-party tools we chose not to give a family of their own (OmniCrawl, JAW, BrowserStack, MetaMask automator, Headless Chromium, the measurement framework of Demir et al.), so read that row as “home-grown or obscure” rather than strictly “home-grown”. report_crawler.mjs prints the full list, so nothing vanishes. Browser names folded to 1 unclassified string out of 529 papers (Ghostery, which is an extension, not a browser).
  • A paper counts once per family, never once per mention, and shares do not sum to 100% because a paper can name several tools. A family's share is of the population named in its heading.
  • Silence is not absence. “Does not name a framework” means the paper did not say, not that the authors used none. These are reporting figures.
  • used and produced both count as driving a crawl — a paper that built its own crawler crawled with it. Tools only compared or cited do not, which is why the last table separates the two.
  • Quotes were spot-checked, by scripts/quote_check.mjs rather than by hand. Of the 106 evidence quotes attached to an OpenWPM or Playwright tool mention anywhere in the corpus, 58 match the source text exactly once whitespace and line-break hyphenation are normalised, and 32 more match on at least 60% of their five-word windows. The remaining 16 were then read by hand against the full text: all sixteen are present in the paper and none was unsupported. The mismatch is always either a bracketed citation marker the extraction dropped (the Open-WPM platform [52] on the Firefox browser) or the two-column reading order still interleaving mid-sentence (we created a Playwrightconstrains the valid child elements, and everything else is moved based [18] crawler).
  • Venue coverage. EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS are absent entirely; any claim here is a claim about seven venues. Notably, several of the tool papers this page recommends were published outside them — Krumnow et al. at CoNEXT, Klein et al. at EuroS&P, Libert in a communication journal — so a tool's paper count in this table is a lower bound on its standing.
  • crawled is defined as a paper whose crawl configuration was recorded or whose study types include an automated web crawl (1,120 papers, 19.1% of the corpus). This corpus is seven broad security venues, not a web-measurement corpus, so shares of all 5,859 papers would be meaningless here.

Recommendations

  1. Report both layers and both versions. “Chromium 121 driven by Playwright 1.41” is a complete answer; “Selenium” is not. This costs one sentence and is the difference between a replicable and an unreplicable crawl. It is also the field's biggest gap: 12.0%.
  2. Pick the instrument from the question, not the habit.
    • Third-party requests, cookies, response bodies → Puppeteer, Playwright or CDP. With Selenium, use WebDriver BiDi and verify what your binding actually delivers; never classic WebDriver alone.
    • Which script did it, and why → PageGraph or VisibleV8.
    • Did this value reach that sink → Foxhound or PanoptiChrome.
    • Firefox, a stable published baseline, and instrumentation that other papers share → OpenWPM.
    • Multi-browser or multi-language, or a lab already fluent in it → Selenium.
    • The lowest-effort modern tracking crawl → Tracker Radar Collector.
  3. Do not write a new crawler for a solved problem. Nearly three in ten crawling papers used a home-grown one, and 75 of them identify it by nothing but a name they invented. If you must build one, say what it is built on, and publish it.
  4. State headless or headful and justify it. Headless is the most detectable configuration you can choose [9Vastel, Antoine; Laperdrix, Pierre; Rudametkin, Walter; Rouvoy, Romain (2018): "Fp-Scanner: The Privacy Implications of Browser Fingerprint Inconsistencies", in: Proceedings of the USENIX Security Symposium. (Link)] and only 12.5% of papers say which they used.
  5. Pin the browser binary, not just the library. Puppeteer and Playwright do this for you; Selenium does not. Record the exact build and archive it with your research artefact (see Artifacts, a page this wiki still owes you).
  6. Validate against something that is not your crawler. A manual visit to a sample, a second browser engine, or a second vantage point [11Jueckstock, Jordan; Sarker, Shaown; Snyder, Peter; Beggs, Aidan; Papadopoulos, Panagiotis; Varvello, Matteo; Livshits, Benjamin; Kapravelos, Alexandros (2021): "Towards Realistic and Reproducible Web Crawl Measurements", in: Proceedings of the ACM Web Conference. (DOI)]. Crawls diverge from human browsing in ways your crawl cannot see [10Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)], and repeated crawls diverge from each other [12Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)].
  7. Run more than once. A single crawl is a point estimate. Report the spread.

Open Questions

* A head-to-head comparison of what OpenWPM, Tracker Radar Collector and a plain Playwright crawl each detect on the same sample, with the same vantage point and the same date, does not exist in the literature we found. It would settle a design question the whole field guesses at. * Playwright's patched Firefox and WebKit builds are used as stand-ins for the real browsers. How far the patching moves the fingerprint, and whether it changes what trackers do, is unmeasured. * Where webXray is developed today (see above). * Whether Selenium's steady quarter-share reflects a considered choice or institutional inertia — a question for a survey of authors, not for this corpus.

References

[1]
Cassel, Darion; Lin, Su-Chin; Buraggina, Alessio; Wang, William; Zhang, Andrew; Bauer, Lujo; Hsiao, Hsu-Chun; Jia, Limin; Libert, Timothy (2022): "OmniCrawl: Comprehensive Measurement of Web Tracking With Real Desktop and Mobile Browsers", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[2]
Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
[3]
Iqbal, Umar; Snyder, Peter; Zhu, Shitong; Livshits, Benjamin; Qian, Zhiyun; Shafiq, Zubair (2020): "AdGraph: A Graph-Based Approach to Ad and Tracker Blocking", in: 2020 IEEE Symposium on Security and Privacy (SP), pp. 763-776. (DOI)
[4]
Jueckstock, Jordan; Kapravelos, Alexandros (2019): "VisibleV8: In-browser Monitoring of JavaScript in the Wild", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[5]
Kanyal, Rahul; Sarangi, Smruti R. (2024): "PanoptiChrome: A Modern In-browser Taint Analysis Framework", in: Proceedings of the ACM Web Conference. (DOI)
[6]
Libert, Timothy (2015): "Exposing the Invisible Web: An Analysis of Third-Party HTTP Requests on 1 Million Websites", International Journal of Communication 9. (Link)
[7]
Krumnow, Benjamin; Jonker, Hugo; Karsch, Stefan (2022): "How gullible are web measurement tools?", in: Proceedings of the 18th International Conference on Emerging Networking EXperiments and Technologies, pp. 171-186. (DOI)
[8]
Klein, David; Barber, Thomas; Bensalim, Souphiane; Stock, Ben; Johns, Martin (2022): "Hand Sanitizers in the Wild: A Large-scale Study of Custom JavaScript Sanitizer Functions", in: 2022 IEEE 7th European Symposium on Security and Privacy (EuroS&P), pp. 236-250. (DOI)
[9]
Vastel, Antoine; Laperdrix, Pierre; Rudametkin, Walter; Rouvoy, Romain (2018): "Fp-Scanner: The Privacy Implications of Browser Fingerprint Inconsistencies", in: Proceedings of the USENIX Security Symposium. (Link)
[10]
Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
[11]
Jueckstock, Jordan; Sarker, Shaown; Snyder, Peter; Beggs, Aidan; Papadopoulos, Panagiotis; Varvello, Matteo; Livshits, Benjamin; Kapravelos, Alexandros (2021): "Towards Realistic and Reproducible Web Crawl Measurements", in: Proceedings of the ACM Web Conference. (DOI)
[12]
Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)
1)
Firefox implements a partial CDP shim and, more actively, WebDriver BiDi. Playwright drives Firefox and WebKit through its own patched builds rather than through CDP.
2)
Measured in the sandbox described in What We Ran, five consecutive runs, 2026-08-06; independently reproduced (4, 4, 1, 4, 4) in a separate five-run sequence during review of this page.
3)
Puppeteer documentation, WebDriver BiDi support, checked 2026-08-06.
4)
Playwright documentation, Browsers, checked 2026-08-06.
5)
Our own runs for this page, in a Linux container, 2026-08-06. The listings above are the programs, with the sandbox's absolute paths replaced by placeholders and a little JSON-shaping trimmed; the fixture is a 40-line Node HTTP server. Numbers are one representative run of five, except where a row says otherwise. See the page history for who ran them.
6)
Non-comment, non-blank lines of the listing above, which is the whole program.
7)
All from the same source: the mention-matching table in Which specialised crawlers actually get used, over all 5,859 papers. Do not compare these against the folded framework table, whose population is the 1,120 crawling papers.
8)
The callstack_instrument, which recorded JS call stacks for HTTP requests, is currently broken and raises a ConfigError if enabled — docs/Configuration.md and issue #557.
9)
OpenWPM does not support windows — README, and issue #503.
10)
The only automation layer the repository documents; it carries its own .PLAYWRIGHT_VERSION and a CI workflow for it.
11)
github.com/openwpm/OpenWPM, README and scripts/install-firefox.sh, checked 2026-08-06.
12)
pagegraph-crawl README, github.com/brave/pagegraph-crawl, checked 2026-08-06.
13)
PageGraph wiki, "Would Be Nice / Someday / Known Limitations", checked 2026-08-06. Query tooling: pagegraph-query (Python) under brave-experiments, pagegraph-rust under brave.
14)
tracker-radar-collector README, github.com/duckduckgo/tracker-radar-collector, checked 2026-08-06.
15)
ariya/phantomjs#15344, “Archiving the project: suspending the development”, opened 2018-03-03.
16)
Checked 2026-08-06: api.github.com/repos/timlib/webXray → HTTP 404; api.github.com/users/timlib/repos → empty; pypi.org/pypi/webxray/json → not found.
17)
An extension the authors wrote themselves is a bespoke crawler and is counted in the two bespoke rows instead, because that is the distinction that decides whether a reader can identify the instrument.
18)
OpenWPM README.md and docs/Configuration.md; the xvfb mode runs a real, non-headless Firefox inside a virtual X display.
You could leave a comment if you were logged in.
programming/crawler.1786529129.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki