User Tools

Site Tools


practices:ethics

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
practices:ethics [2026/09/15 17:16] – Add privacy:age_assurance to Related pages, naming the four ethics questions this page does not currently answer. Authored by Claude karel.kubicek.claudepractices:ethics [2026/09/17 10:45] (current) – Clarify consentAction schema statistic; Authored by Claude karel.kubicek.claude
Line 252: Line 252:
 | ''humanAnnotation.agreementMetric'' | 3,318 papers that coded data by hand | 512 | 15.4% | | ''humanAnnotation.agreementMetric'' | 3,318 papers that coded data by hand | 512 | 15.4% |
 | ''crawlConfig.statefulness'' | 1,120 crawling papers | 219 | 19.6% | | ''crawlConfig.statefulness'' | 1,120 crawling papers | 219 | 19.6% |
-| ''crawlConfig.consentAction'' | 1,120 crawling papers | 349 | 31.2% |+| ''crawlConfig.consentAction'' (field populated; schema, not an audited paper claim) | 1,120 crawling papers | 349 | 31.2% | 
 + 
 +The consent row is a schema-completeness statistic: a populated non-sentinel field does not by itself show that the paper made the corresponding claim. The 2026-09-05 audit found **55 of 1,120 crawling papers (4.9%)** whose full text states a consent action; it also found **279 of 313 (89.1%)** ''no-interaction'' defaults unsupported. The two populations answer different questions.
  
 The reason it is unresolved rather than merely unreported is that the file does not answer the question a researcher has. `robots.txt` is a directive to *automated indexers*; a research crawl that loads a page once, in the way a browser would, and never republishes the content is not obviously the addressee, and a strict reading excludes exactly the pages a measurement is about (a `Disallow: /` on a site's tracking-heavy subpages removes the finding). There is no consensus in this literature, and inventing one here would be dishonest. What is not defensible is silence: **decide, do the same thing throughout the crawl, and write the sentence.** If you honour it, say what share of your sample it removed, because that is a bias in your denominator, not a footnote. If you do not, say why and what you did instead — most commonly a rate limit strict enough that the file's purpose (protecting the server) is served by other means. The reason it is unresolved rather than merely unreported is that the file does not answer the question a researcher has. `robots.txt` is a directive to *automated indexers*; a research crawl that loads a page once, in the way a browser would, and never republishes the content is not obviously the addressee, and a strict reading excludes exactly the pages a measurement is about (a `Disallow: /` on a site's tracking-heavy subpages removes the finding). There is no consensus in this literature, and inventing one here would be dishonest. What is not defensible is silence: **decide, do the same thing throughout the crawl, and write the sentence.** If you honour it, say what share of your sample it removed, because that is a bias in your denominator, not a footnote. If you do not, say why and what you did instead — most commonly a rate limit strict enough that the file's purpose (protecting the server) is served by other means.
practices/ethics.1789492572.txt.gz · Last modified: by karel.kubicek.claude