User Tools

Site Tools


practices:ethics

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
practices:ethics [2026/08/21 14:50] – Boxes: <wrap> renders a span, use uppercase <WRAP>; one box per list, not per bullet karel.kubicek.claudepractices:ethics [2026/09/17 10:45] (current) – Clarify consentAction schema statistic; Authored by Claude karel.kubicek.claude
Line 252: Line 252:
 | ''humanAnnotation.agreementMetric'' | 3,318 papers that coded data by hand | 512 | 15.4% | | ''humanAnnotation.agreementMetric'' | 3,318 papers that coded data by hand | 512 | 15.4% |
 | ''crawlConfig.statefulness'' | 1,120 crawling papers | 219 | 19.6% | | ''crawlConfig.statefulness'' | 1,120 crawling papers | 219 | 19.6% |
-| ''crawlConfig.consentAction'' | 1,120 crawling papers | 349 | 31.2% |+| ''crawlConfig.consentAction'' (field populated; schema, not an audited paper claim) | 1,120 crawling papers | 349 | 31.2% | 
 + 
 +The consent row is a schema-completeness statistic: a populated non-sentinel field does not by itself show that the paper made the corresponding claim. The 2026-09-05 audit found **55 of 1,120 crawling papers (4.9%)** whose full text states a consent action; it also found **279 of 313 (89.1%)** ''no-interaction'' defaults unsupported. The two populations answer different questions.
  
 The reason it is unresolved rather than merely unreported is that the file does not answer the question a researcher has. `robots.txt` is a directive to *automated indexers*; a research crawl that loads a page once, in the way a browser would, and never republishes the content is not obviously the addressee, and a strict reading excludes exactly the pages a measurement is about (a `Disallow: /` on a site's tracking-heavy subpages removes the finding). There is no consensus in this literature, and inventing one here would be dishonest. What is not defensible is silence: **decide, do the same thing throughout the crawl, and write the sentence.** If you honour it, say what share of your sample it removed, because that is a bias in your denominator, not a footnote. If you do not, say why and what you did instead — most commonly a rate limit strict enough that the file's purpose (protecting the server) is served by other means. The reason it is unresolved rather than merely unreported is that the file does not answer the question a researcher has. `robots.txt` is a directive to *automated indexers*; a research crawl that loads a page once, in the way a browser would, and never republishes the content is not obviously the addressee, and a strict reading excludes exactly the pages a measurement is about (a `Disallow: /` on a site's tracking-heavy subpages removes the finding). There is no consensus in this literature, and inventing one here would be dishonest. What is not defensible is silence: **decide, do the same thing throughout the crawl, and write the sentence.** If you honour it, say what share of your sample it removed, because that is a bias in your denominator, not a footnote. If you do not, say why and what you did instead — most commonly a rate limit strict enough that the file's purpose (protecting the server) is served by other means.
Line 393: Line 395:
   * [[Programming:Crawler]] — where the rate limit and the User-Agent actually get set.   * [[Programming:Crawler]] — where the rate limit and the User-Agent actually get set.
   * [[Privacy:Consent]] — interacting with a banner is an interaction.   * [[Privacy:Consent]] — interacting with a banner is an interaction.
 +  * [[Privacy:Age assurance]] — **four questions this page does not currently answer**: crawling adult content, submitting a synthetic or a real identity document to a live vendor, publishing a working bypass, and measuring an age model without measuring children. That page states them as a gap here.
   * [[Design:Website selection]] — a robots.txt exclusion is a change to your sampling frame.   * [[Design:Website selection]] — a robots.txt exclusion is a change to your sampling frame.
   * [[:Artifacts]] — what you can release once the data has people in it.   * [[:Artifacts]] — what you can release once the data has people in it.
practices/ethics.1787323822.txt.gz · Last modified: by karel.kubicek.claude