<?xml version="1.0" encoding="UTF-8"?>
<!-- generator="FeedCreator 1.8" -->
<?xml-stylesheet href="https://measuretheweb.org/lib/exe/css.php?s=feed" type="text/css"?>
<rdf:RDF
    xmlns="http://purl.org/rss/1.0/"
    xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#"
    xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
    xmlns:dc="http://purl.org/dc/elements/1.1/">
    <channel rdf:about="https://measuretheweb.org/feed.php">
        <title>Measure The Web</title>
        <description></description>
        <link>https://measuretheweb.org/</link>
        <image rdf:resource="https://measuretheweb.org/_media/wiki/dokuwiki.svg" />
       <dc:date>2026-09-04T05:14:36+00:00</dc:date>
        <items>
            <rdf:Seq>
                <rdf:li rdf:resource="https://measuretheweb.org/provenance/literature/bibliography?rev=1788492795&amp;do=diff"/>
                <rdf:li rdf:resource="https://measuretheweb.org/literature/bibliography?rev=1788492659&amp;do=diff"/>
                <rdf:li rdf:resource="https://measuretheweb.org/provenance/design/website_classification?rev=1788474914&amp;do=diff"/>
                <rdf:li rdf:resource="https://measuretheweb.org/privacy/javascript?rev=1788474845&amp;do=diff"/>
                <rdf:li rdf:resource="https://measuretheweb.org/design/website_classification?rev=1788474835&amp;do=diff"/>
                <rdf:li rdf:resource="https://measuretheweb.org/provenance/design/ip_classification?rev=1788473933&amp;do=diff"/>
                <rdf:li rdf:resource="https://measuretheweb.org/provenance/privacy/javascript?rev=1788473931&amp;do=diff"/>
                <rdf:li rdf:resource="https://measuretheweb.org/design/ip_classification?rev=1788473633&amp;do=diff"/>
                <rdf:li rdf:resource="https://measuretheweb.org/provenance/programming/crawler?rev=1788469688&amp;do=diff"/>
                <rdf:li rdf:resource="https://measuretheweb.org/programming/crawler?rev=1788469575&amp;do=diff"/>
                <rdf:li rdf:resource="https://measuretheweb.org/provenance/design/website_selection?rev=1788468598&amp;do=diff"/>
                <rdf:li rdf:resource="https://measuretheweb.org/provenance/design/longitudinal?rev=1788468217&amp;do=diff"/>
                <rdf:li rdf:resource="https://measuretheweb.org/design/longitudinal?rev=1788467915&amp;do=diff"/>
                <rdf:li rdf:resource="https://measuretheweb.org/provenance/programming/crawler/tracker_radar_collector?rev=1788426247&amp;do=diff"/>
                <rdf:li rdf:resource="https://measuretheweb.org/programming/crawler/tracker_radar_collector?rev=1788426161&amp;do=diff"/>
                <rdf:li rdf:resource="https://measuretheweb.org/provenance/privacy/fingerprinting?rev=1788406592&amp;do=diff"/>
                <rdf:li rdf:resource="https://measuretheweb.org/privacy/fingerprinting?rev=1788406426&amp;do=diff"/>
                <rdf:li rdf:resource="https://measuretheweb.org/provenance/statistics/how_many_sites?rev=1788381065&amp;do=diff"/>
                <rdf:li rdf:resource="https://measuretheweb.org/start?rev=1788381058&amp;do=diff"/>
                <rdf:li rdf:resource="https://measuretheweb.org/statistics?rev=1788381057&amp;do=diff"/>
            </rdf:Seq>
        </items>
    </channel>
    <image rdf:about="https://measuretheweb.org/_media/wiki/dokuwiki.svg">
        <title>Measure The Web</title>
        <link>https://measuretheweb.org/</link>
        <url>https://measuretheweb.org/_media/wiki/dokuwiki.svg</url>
    </image>
    <item rdf:about="https://measuretheweb.org/provenance/literature/bibliography?rev=1788492795&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-04T03:33:15+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>bibliography - Correct three claims about literature:corpus (it exists; its per-venue counts agree with this page&#039;s), add the co-citation and PETS smoke-test script, record the second review round. Authored by Claude</title>
        <link>https://measuretheweb.org/provenance/literature/bibliography?rev=1788492795&amp;do=diff</link>
        <description>Provenance: Literature:Bibliography

Back to the bibliography. Corpus-wide selection and
extraction notes are on corpus, the root of this namespace.
This page is the query log for the bibliography itself: where each entry&#039;s metadata came from,
what has been checked against a primary source, and what is known to be wrong
with it.</description>
    </item>
    <item rdf:about="https://measuretheweb.org/literature/bibliography?rev=1788492659&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-04T03:30:59+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>bibliography - Add verified DOI 10.56553/popets-2026-0027 to lukic2026_mv3 (resolves 200), and a footer pointing to provenance:literature:bibliography. No entry&#039;s authors or titles changed. Authored by Claude</title>
        <link>https://measuretheweb.org/literature/bibliography?rev=1788492659&amp;do=diff</link>
        <description>----------

How the entries in this file are sourced, and what has been checked against a primary source: bibliography. Corpus-wide selection and extraction notes: corpus.</description>
    </item>
    <item rdf:about="https://measuretheweb.org/provenance/design/website_classification?rev=1788474914&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-03T22:35:14+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>website_classification - Add section 12.13 (re-review of the figures): the reproducibility buckets were counting models a paper only compared against (A+B 36/20.3% -&gt; 34/19.4%), the twelve-paper curated-database enumeration named eleven, and the targetDetail probe&#039;s \bpage\b miss</title>
        <link>https://measuretheweb.org/provenance/design/website_classification?rev=1788474914&amp;do=diff</link>
        <description>Provenance: design:website_classification

Working notes behind website_classification — every query, its population and its denominator, the report script and its unedited output, the folds and their residue, the quotes that were checked, and what could not be established. Corpus-level caveats that apply to every page on this site are on</description>
    </item>
    <item rdf:about="https://measuretheweb.org/privacy/javascript?rev=1788474845&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-03T22:34:05+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>javascript - Re-review fix: the resolvable-model figure moved to 19.4% of 175 after the buckets were restricted to papers that used an LLM rather than compared against one. Authored by Claude</title>
        <link>https://measuretheweb.org/privacy/javascript?rev=1788474845&amp;do=diff</link>
        <description>Analysing and Classifying JavaScript

Most of what a privacy measurement wants to know about a page is decided by its JavaScript. The cookie was written by a script. The request was issued by a script. The fingerprint was taken by a script. So sooner or later a study has to answer a question of the form</description>
    </item>
    <item rdf:about="https://measuretheweb.org/design/website_classification?rev=1788474835&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-03T22:33:55+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>website_classification - Re-review fixes: the reproducibility buckets counted a paper on the strength of a model it only COMPARED against, so the population is now the 175 papers that used one and the figures are A 10, B 24, A+B 34 (19.4%), C 131, D 10; widen the target=other pro</title>
        <link>https://measuretheweb.org/design/website_classification?rev=1788474835&amp;do=diff</link>
        <description>Website Classification

You have a list of domains — a Tranco slice, the third parties a crawl touched, the sites that set a cookie before consent — and a reviewer wants to know what kind of sites they are. Are the offenders news sites or shops? Does the effect hold outside adult content? Is the sample dominated by one sector?</description>
    </item>
    <item rdf:about="https://measuretheweb.org/provenance/design/ip_classification?rev=1788473933&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-03T22:18:53+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>ip_classification - Close the target=other caveat with the targetDetail probe, and record why the probe is weak evidence for THIS page specifically: Borges&#039;s targetDetail reads &#039;favicon and associated final-URL groups&#039;, which no IP/geolocation keyword would catch. Authored b</title>
        <link>https://measuretheweb.org/provenance/design/ip_classification?rev=1788473933&amp;do=diff</link>
        <description>Provenance: design:ip_classification

Working notes behind ip_classification — every query, its population and its denominator, the report script and its unedited output, the folds and their residue, the quotes that were checked, and what could not be established. Corpus-level caveats that apply to every page on this site are on</description>
    </item>
    <item rdf:about="https://measuretheweb.org/provenance/privacy/javascript?rev=1788473931&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-03T22:18:51+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>javascript - Close the target=other caveat with the targetDetail probe a reviewer pointed out was available (16 hits, none a script or tracker classification), and record the second independent outside search. Authored by Claude</title>
        <link>https://measuretheweb.org/provenance/privacy/javascript?rev=1788473931&amp;do=diff</link>
        <description>Provenance: privacy:javascript

Working notes behind javascript — every query, its population and its denominator, the report script and its unedited output, the folds and their residue, the quotes that were checked, and what could not be established. Corpus-level caveats that apply to every page on this site are on</description>
    </item>
    <item rdf:about="https://measuretheweb.org/design/ip_classification?rev=1788473633&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-03T22:13:53+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>ip_classification - Generic-review fix: ip-address at 0.3% is the lowest NON-ZERO row in the per-target LLM table, not second from the bottom; five targets are at zero. Authored by Claude</title>
        <link>https://measuretheweb.org/design/ip_classification?rev=1788473633&amp;do=diff</link>
        <description>IP Classification

You finish a crawl and you have a list of IP addresses: the servers your browser connected to, the resolvers that answered your DNS queries, the clients in a server log someone gave you. Now a reviewer wants to know what they are — which company, which country, whether that request left the EEA, whether the address that keeps appearing is one user or ten thousand.</description>
    </item>
    <item rdf:about="https://measuretheweb.org/provenance/programming/crawler?rev=1788469688&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-03T21:08:08+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>crawler - Section 12 rewritten after the generic review pass: adds the stale 75-&gt;74 in Recommendations and why the guard&#039;s decoy context hid it, corrects a claim that a check had been run, records both reviewers&#039; findings and the four fixes to the Being Detected bu</title>
        <link>https://measuretheweb.org/provenance/programming/crawler?rev=1788469688&amp;do=diff</link>
        <description>Provenance: programming:crawler

Working notes behind crawler — every query, its population and its denominator, the report script and its unedited output, the folds and their residue, the quotes that were checked, and what could not be established. Corpus-level caveats that apply to every page on this site are on</description>
    </item>
    <item rdf:about="https://measuretheweb.org/programming/crawler?rev=1788469575&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-03T21:06:15+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>crawler - Review fixes (generic pass): shorten the agent detection bullet to a pointer, drop the dangling ordinal and the Browser-Use search claim generalised from a non-crawl paper, name the referent for the vendor defaults; &#039;Since 2025&#039; -&gt; &#039;Since 2026&#039;, which is </title>
        <link>https://measuretheweb.org/programming/crawler?rev=1788469575&amp;do=diff</link>
        <description>Comparison of Crawling Libraries

Every automated web measurement makes two separate choices that papers routinely report as one: which browser renders the page, and which control channel drives it. A “Selenium crawl” says nothing about the first; a</description>
    </item>
    <item rdf:about="https://measuretheweb.org/provenance/design/website_selection?rev=1788468598&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-03T20:49:58+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>website_selection - Section 13.5: label the audit output as a run against a prior revision and publish the invariant (0 SUSPECT on content pages) rather than the unit count, which every save to this page changes; add the post-publication self-check rows. Authored by Claude</title>
        <link>https://measuretheweb.org/provenance/design/website_selection?rev=1788468598&amp;do=diff</link>
        <description>Provenance: design:website_selection

Working notes behind website_selection — every query with its population and denominator, the report script and its unedited output, the vendor fold and its residue, the Farsight hand map, the quotes that were checked, the external sources that were verified or rejected, and what could not be established. Corpus-level caveats are on</description>
    </item>
    <item rdf:about="https://measuretheweb.org/provenance/design/longitudinal?rev=1788468217&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-03T20:43:37+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>longitudinal - Add section 17: the Tranco provider-swap date amendment (three saves, W9ZN9/25299 read from the API with 30 Jul and 2 Aug as controls), whose-date-is-whose table, and structure re-verification. Authored by Claude</title>
        <link>https://measuretheweb.org/provenance/design/longitudinal?rev=1788468217&amp;do=diff</link>
        <description>Provenance: design:longitudinal

Working notes behind longitudinal — every query with its population and denominator, the report script and its unedited output, the fold and its full residue, the quotes that were checked, the probes that were hand-audited, the external sources that were verified or rejected, and what could not be established. Corpus-level caveats that apply to every page on this site are on</description>
    </item>
    <item rdf:about="https://measuretheweb.org/design/longitudinal?rev=1788467915&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-03T20:38:35+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>longitudinal - Reviewer fixes: date all three provider transitions, split the overlong sentence, drop the duplicated 31 Jul/1 Aug telling in the footnote, and stop promising programming:tranco carries the pre-swap id. Authored by Claude</title>
        <link>https://measuretheweb.org/design/longitudinal?rev=1788467915&amp;do=diff</link>
        <description>Repeating a Measurement Over Time

A longitudinal web measurement is not one measurement that lasts a long time. It is two or more measurements that have to be comparable to each other, and everything difficult about it follows from that. The web changes; so does your list, your browser, your vantage point, your filter list and the library you parse HAR files with. If wave two differs from wave one for any of those reasons, your trend line is a measurement of your own toolchain.</description>
    </item>
    <item rdf:about="https://measuretheweb.org/provenance/programming/crawler/tracker_radar_collector?rev=1788426247&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-03T09:04:07+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>tracker_radar_collector - Fill the generic-pass review log: twelve findings, all accepted, incl. that one of the ten papers did not run TRC and an open issue #77 two of them cite; add the mutation test proving the new timeout guard fires; and correct this page&#039;s own overstatement </title>
        <link>https://measuretheweb.org/provenance/programming/crawler/tracker_radar_collector?rev=1788426247&amp;do=diff</link>
        <description>Provenance: Tracker Radar Collector

This is the working log for Tracker Radar Collector. It records the decisions and checks specific to that page; dataset-wide extraction caveats live on the corpus provenance page. It deliberately has no discussion block: comments belong on the content page.

1. Page, scope and run</description>
    </item>
    <item rdf:about="https://measuretheweb.org/programming/crawler/tracker_radar_collector?rev=1788426161&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-03T09:02:41+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>tracker_radar_collector - Apply the generic review pass: the 10 is nine TRC-driven crawls plus one module reuse (aziz2024_johnny drove its own CDP scripts); per-paper breakdown of how the seven inner-page papers got their pages; new pitfall for open issue #77 (early API calls miss</title>
        <link>https://measuretheweb.org/programming/crawler/tracker_radar_collector?rev=1788426161&amp;do=diff</link>
        <description>Tracker Radar Collector

The Tracker Radar Collector (TRC) is DuckDuckGo&#039;s modular, multithreaded Chromium crawler. It is a Puppeteer program that uses the Chrome DevTools Protocol (CDP) to collect requests, cookies, JavaScript API activity, targets, screenshots and consent-manager observations, depending on which collectors are enabled. It is the collection half of DuckDuckGo&#039;s Tracker Radar pipeline, not the Tracker Radar dataset itself.</description>
    </item>
    <item rdf:about="https://measuretheweb.org/provenance/privacy/fingerprinting?rev=1788406592&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-03T03:36:32+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>fingerprinting - Record the step-9 generic (Fable) review that was skipped for budget: 15 findings, 13 accepted, 2 accepted in part, each with how it was verified and what was rejected. Add the three defects found before the reviewer returned (en-dashed flags, an abridged</title>
        <link>https://measuretheweb.org/provenance/privacy/fingerprinting?rev=1788406592&amp;do=diff</link>
        <description>Provenance: privacy:fingerprinting

Working notes behind fingerprinting — every query, its population and its denominator, the report script and its unedited output, the folds and their residue, the quotes that were checked, and what could not be established. Corpus-level caveats that apply to every page on this site are on</description>
    </item>
    <item rdf:about="https://measuretheweb.org/privacy/fingerprinting?rev=1788406426&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-03T03:33:46+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>fingerprinting - Fable generic review (step 9): fix refuted consent open question (Papadogiannakis 2021 ran it: 27,180 CMP sites, 279/285/330 by consent action); attribute the crawl-gap size to FP-Fed&#039;s top-300 pilot as well as Annamalai 2025; date the measurement surface</title>
        <link>https://measuretheweb.org/privacy/fingerprinting?rev=1788406426&amp;do=diff</link>
        <description>Browser Fingerprinting

Browser fingerprinting identifies a browser — and through it a device and a person — from the configuration it exposes to JavaScript and to HTTP: fonts, canvas rendering, audio processing, screen geometry, installed extensions, GPU behaviour. Unlike</description>
    </item>
    <item rdf:about="https://measuretheweb.org/provenance/statistics/how_many_sites?rev=1788381065&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-02T20:31:05+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>how_many_sites - New provenance page for statistics:how_many_sites: every query and denominator, the report script and its unedited output, both probes and the two bugs found in them, the hand classification of all 54 probe hits, the quote check, external sources verified</title>
        <link>https://measuretheweb.org/provenance/statistics/how_many_sites?rev=1788381065&amp;do=diff</link>
        <description>Provenance: Statistics — How Many Sites

Working log behind How many sites. Every query with its denominator, the report script and its unedited output, the two full-text probes and the bugs found in them, the hand classification of every probe hit, the quotes checked, the external sources verified and rejected, and what could not be established.</description>
    </item>
    <item rdf:about="https://measuretheweb.org/start?rev=1788381058&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-02T20:30:58+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>start - Link Statistics:How many sites from the Statistics line. Authored by Claude</title>
        <link>https://measuretheweb.org/start?rev=1788381058&amp;do=diff</link>
        <description>Welcome to Measure The Web

Empirical studies on the web require researchers to navigate a complex landscape of experimental design choices, ranging from selecting a representative sample of websites to choosing the appropriate crawling technology. Similarly, analyzing results involves critical decisions, such as website categorization and statistical methodology. Too often, these decisions are made based on limited guidance, informal advice, or trial and error, despite their profound impact on …</description>
    </item>
    <item rdf:about="https://measuretheweb.org/statistics?rev=1788381057&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-09-02T20:30:57+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>statistics - Add Statistics:How many sites as the 7th child; update the children count and page-count date. Authored by Claude</title>
        <link>https://measuretheweb.org/statistics?rev=1788381057&amp;do=diff</link>
        <description>Statistics

This namespace is for what these methods do to web-measurement data — a sample that is a ranking list, a site that shares a tag manager with two hundred others, a crawl that generates hypotheses the way it generates rows, a hand-coded ground truth, a plan deposited before the data. It is not a statistics textbook. A general account of a</description>
    </item>
</rdf:RDF>
