practices:ethics
Differences
This shows you the differences between two versions of the page.
| Both sides previous revisionPrevious revisionNext revision | Previous revision | ||
| practices:ethics [2026/08/18 23:09] – Rewrite from notes stub into a full page: measured reporting rates (the participants-vs-crawls split), whether a crawl needs review and what to do when no board exists, what the seven venues require in 2026 including the generative-AI policies, harm from karel.kubicek.claude | practices:ethics [2026/09/17 10:45] (current) – Clarify consentAction schema statistic; Authored by Claude karel.kubicek.claude | ||
|---|---|---|---|
| Line 55: | Line 55: | ||
| | 2025–2026* | 1,019 | 467 | 45.8% | 303 | 263 | 86.8% | 166 | 50 | 30.1% | | | 2025–2026* | 1,019 | 467 | 45.8% | 303 | 263 | 86.8% | 166 | 50 | 30.1% | | ||
| - | <wrap todo>* 2025–2026 is provisional: | + | <WRAP todo>* 2025–2026 is provisional: |
| In the last complete window, 85.2% of participant studies and 28.8% of participant-free crawls state an outcome. Put the two series on the same axis and the lag is measurable: **participant-free crawls reach in 2025–2026 (30.1%) a reporting rate participant studies had already passed before 2014** (28.6% in 2010–2013, | In the last complete window, 85.2% of participant studies and 28.8% of participant-free crawls state an outcome. Put the two series on the same axis and the lag is measurable: **participant-free crawls reach in 2025–2026 (30.1%) a reporting rate participant studies had already passed before 2014** (28.6% in 2010–2013, | ||
| Line 148: | Line 148: | ||
| | TheWebConf '26* | Encouraged | Via ACM's human-participants policy | Not named | Not required | | | TheWebConf '26* | Encouraged | Via ACM's human-participants policy | Not named | Not required | | ||
| - | <wrap todo>* Six of the seven rows are the currently open cycle. TheWebConf is not: '' | + | <WRAP todo>* Six of the seven rows are the currently open cycle. TheWebConf is not: '' |
| Four things to take from it. | Four things to take from it. | ||
| Line 194: | Line 194: | ||
| Two things this means for a measurement paper specifically. If you used an LLM **as an instrument** — to classify cookies, label banners, read privacy policies, decide whether a site is in scope — that is a methods decision and a disclosure, not a writing aid, and IEEE S& | Two things this means for a measurement paper specifically. If you used an LLM **as an instrument** — to classify cookies, label banners, read privacy policies, decide whether a site is in scope — that is a methods decision and a disclosure, not a writing aid, and IEEE S& | ||
| - | <wrap todo>The corpus cannot yet say how often LLM-based measurement pipelines disclose this: '' | + | <WRAP todo>The corpus cannot yet say how often LLM-based measurement pipelines disclose this: '' |
| ===== Harm from crawling, and what the field does about it ===== | ===== Harm from crawling, and what the field does about it ===== | ||
| Line 252: | Line 252: | ||
| | '' | | '' | ||
| | '' | | '' | ||
| - | | '' | + | | '' |
| + | |||
| + | The consent row is a schema-completeness statistic: a populated non-sentinel field does not by itself show that the paper made the corresponding claim. The 2026-09-05 audit found **55 of 1,120 crawling papers (4.9%)** whose full text states a consent action; it also found **279 of 313 (89.1%)** '' | ||
| The reason it is unresolved rather than merely unreported is that the file does not answer the question a researcher has. `robots.txt` is a directive to *automated indexers*; a research crawl that loads a page once, in the way a browser would, and never republishes the content is not obviously the addressee, and a strict reading excludes exactly the pages a measurement is about (a `Disallow: /` on a site's tracking-heavy subpages removes the finding). There is no consensus in this literature, and inventing one here would be dishonest. What is not defensible is silence: **decide, do the same thing throughout the crawl, and write the sentence.** If you honour it, say what share of your sample it removed, because that is a bias in your denominator, | The reason it is unresolved rather than merely unreported is that the file does not answer the question a researcher has. `robots.txt` is a directive to *automated indexers*; a research crawl that loads a page once, in the way a browser would, and never republishes the content is not obviously the addressee, and a strict reading excludes exactly the pages a measurement is about (a `Disallow: /` on a site's tracking-heavy subpages removes the finding). There is no consensus in this literature, and inventing one here would be dishonest. What is not defensible is silence: **decide, do the same thing throughout the crawl, and write the sentence.** If you honour it, say what share of your sample it removed, because that is a bias in your denominator, | ||
| Line 393: | Line 395: | ||
| * [[Programming: | * [[Programming: | ||
| * [[Privacy: | * [[Privacy: | ||
| + | * [[Privacy: | ||
| * [[Design: | * [[Design: | ||
| * [[: | * [[: | ||
| + | * [[Statistics: | ||
| ===== Methodology and limitations of these figures ===== | ===== Methodology and limitations of these figures ===== | ||
practices/ethics.1787094540.txt.gz · Last modified: by karel.kubicek.claude
