User Tools

Site Tools


provenance:statistics:interrater_agreement

Provenance: Statistics — Inter-rater Agreement

Working log behind Interrater agreement. Every figure on that page has its query, its denominator and its script output here. Corpus-wide caveats — how the corpus was selected, what the funnel drops, what the extraction is and is not — are on Corpus and are not restated.

This page adds no bibliography entries of its own; it cites the same {[citekey]} keys and the same shared Bibliography as the content page. No ~~DISCUSSION~~: comments belong on the content page. (This is the convention the earlier provenance pages set; it is followed here without reopening it.)

The run

Date 2026-08-20
Corpus data/extract/run1/extractions.jsonl, 5,859 papers, 7 venues, 2010–2026
Full texts searched by probe 5,869 paper.cols.txt files (10 more than extracted records)
Page status New page. Nothing on the wiki covered inter-rater agreement before this.
Main agent Claude Opus 5
Sub-agents 1 research (external reference verification, Sonnet); 5 review passes — 3 focused (Sonnet) in parallel, 1 generic (Fable), then a re-run of the figures pass after its findings were acted on (Sonnet). The citations pass spawned two of its own for DOI resolution and quote checking. See Review log
Scripts committed scripts/report_interrater.mjs, scripts/ira_fold.mjs, scripts/ira_fulltext_probes.sh, scripts/kappa_on_skewed_labels.py

Why a new page rather than broadening a neighbour

Checked with sitemap.mjs before writing: 72 existing pages, no page and no red link on inter-rater agreement, annotation, coding or reliability. The two nearest neighbours were considered and rejected as hosts:

  • Website classification covers one *task* that needs hand-coding. Inter-rater agreement is a method that applies to every hand-coded task on the site — cookie purposes, dark patterns, policy spans, request classification — so putting it there would make every other page link sideways into a task page for a general method.
  • Hypothesis testing is about inference on measurements. Reliability of the measurements is upstream of it, and that page is already 63 kB.

The statistics: namespace was chosen over design: because the deliverable is a coefficient and a reporting standard, matching its five siblings. It is linked from start in the statistics line and from Ethics; those edits were made in the same sitting as this page, from the same run.

Denominators used on the page

Every figure names one of these. Never “of 5,859 papers”.

Name Definition N
all papers every extracted record 5,859
empirical isEmpirical === true 5,118
annotated humanAnnotation.length > 0the page's main population 3,318
naming a statistic annotated ∧ some tuple has non-null agreementMetric 512
stating a coder count annotated ∧ some tuple has numeric annotatorCount 1,288
≥2 coders stated annotated ∧ max annotatorCount ≥ 2 1,176
naming a chance-corrected coefficient subset of the 512 418
with a parsable value subset of the 418 387
annotated ∧ crawled ∧ (crawlConfig ≠ null ∨ automated-web-crawl) 864
classified classification.length > 0 4,439
full texts (probes) find data/fulltext -name paper.cols.txt 5,869

Annotation tuples across the 3,318 papers: 5,778. Papers are counted once each throughout; no figure on the page is a tuple count.

The report script

Committed as scripts/report_interrater.mjs. Run with:

node scripts/report_interrater.mjs

Its unedited output, which is the source of every table on the content page:

report_interrater-output.txt
corpus: 5859 papers, 7 venues, 2010–2026
empirical: 5118
annotated (>=1 humanAnnotation tuple): 3318
humanAnnotation tuples: 5778
 
── Who did the coding (of 3318 annotating papers, multi-valued) ──
   2848   85.8%  authors
    340   10.2%  domain-experts
    177    5.3%  not-stated
    109    3.3%  mixed
     73    2.2%  crowdworkers
     60    1.8%  students
     35    1.1%  hired-annotators
      5    0.2%  legal-experts
 
── Annotator count stated (of 3318 annotating papers) ──
   1288   38.8%  states how many people coded
   2030   61.2%  does not state it (sentinel: null)
 
── How many annotators, where stated (of 1288) ──
    112    8.7%  1 (no second coder -- agreement impossible)
    591   45.9%  2
    319   24.8%  3
    124    9.6%  4–5
     64    5.0%  6–10
     78    6.1%  >10
 
── Reliability reporting (of 3318 annotating papers) ──
    512   15.4%  names an agreement statistic
     97    2.9%  reports an agreement NUMBER but names no statistic
   2709   81.6%  reports neither
 
── Reliability reporting where >=2 annotators are stated (of 1176) ──
    494   42.0%  names an agreement statistic
 
── Which statistic, folded (of 512 papers naming one; multi-valued) ──
    252   49.2%  Cohen's kappa
     84   16.4%  Krippendorff's alpha
     80   15.6%  raw agreement (no chance correction)
     75   14.6%  Fleiss' kappa
     12    2.3%  not an agreement coefficient
     10    2.0%  kappa (variant unnamed)
      9    1.8%  Kupper–Hafner
      5    1.0%  Gwet's AC1/AC2
      1    0.2%  intra-class correlation
 
  UNMAPPED RESIDUE (24 distinct strings, 29 tuple mentions):
      3  "inter-rater reliability"
      2  "Kupper"
      2  "consensus coding"
      2  "interrater reliability"
      1  "Hafner inter-rater reliability metric"
      1  "inter-coder reliability"
      1  "Kendall's Tau coefficient"
      1  "Hafner's statistic"
      1  "Kendall's W"
      1  "comparison of coding results"
      1  "Kendall's tau"
      1  "RMSE"
      1  "Inter-Rater Reliability (IRR)"
      1  "identical coding"
      1  "error rate"
      1  "consensus threshold"
      1  "Three-fifths majority voting"
      1  "inter-coder reliability score"
      1  "alignment rate"
      1  "group unanimity"
      1  "Kendall's Tau"
      1  "Krippendorf's alpha"
      1  "consensus discussion"
      1  "unanimous result"
 
    418   81.6%  names >=1 chance-corrected coefficient
     77   15.0%  names ONLY uncorrected agreement / a non-agreement statistic
 
── Coefficient/design mismatch ──
    245   47.9%  names Cohen's kappa and no multi-rater coefficient
    236   96.3%    ...of which state an annotator count
     61   25.8%    ...of which state >2 annotators (Cohen's kappa is a 2-rater statistic)
      USENIX 2015 investigating-the-computer-security-practices-and-needs-of-journalists — n=3
      IEEE-SP 2017 obstacles-to-the-adoption-of-secure-communication-tools — n=3
      PETS 2017 why-privacy-is-all-but-forgotten — n=3
      WWW 2017 extracting-and-ranking-travel-tips-from-user-generated-reviews — n=30
      CCS 2018 towards-usable-checksums-automating-the-integrity-verification-of-web-downloads — n=3
      IEEE-SP 2018 towards-security-and-privacy-for-multi-user-augmented-reality-foundations-with-e — n=3
      USENIX 2018 the-aftermath-of-a-crypto-ransomware-attack-at-a-large-academic-institution — n=3
      WWW 2018 collective-classification-of-spam-campaigners-on-twitter-a-hierarchical-meta-pat — n=3
      WWW 2018 staqc-a-systematically-mined-question-code-dataset-from-stack-overflow — n=4
      NDSS 2019 how-bad-can-it-git-characterizing-secret-leakage-in-public-github-repositories — n=3
      WWW 2019 anything-to-hide-studying-minified-and-obfuscated-code-in-the-web — n=5
      WWW 2019 generating-product-descriptions-from-user-reviews — n=15
 
── Reported coefficient values (chance-corrected metrics only) ──
  papers with >=1 parsable coefficient: 387 (of 418 naming a chance-corrected one)
  (a paper reporting several coefficients is bucketed by its LOWEST)
      0    0.0%  < 0.00 poor
      7    1.8%  0.00–0.20 slight
     13    3.4%  0.21–0.40 fair
     28    7.2%  0.41–0.60 moderate
    167   43.2%  0.61–0.80 substantial
    172   44.4%  0.81–1.00 almost perfect
     74   19.1%  below 0.67 (Krippendorff's lowest tentative-conclusions threshold)
  min 0.08  p25 0.70  median 0.80  p75 0.86  max 1.00
  lowest ten reported:
      0.08  WWW 2012 learning-from-the-past-answering-new-questions-with-past-answers
      0.11  WWW 2017 sangoshthi-empowering-community-health-workers-through-peer-learning-in-rural-in
      0.14  WWW 2023 cam-a-large-language-model-based-creative-analogy-mining-framework
      0.15  WWW 2020 leading-conversational-search-by-suggesting-useful-questions
      0.17  WWW 2018 satisfaction-with-failure-or-unsatisfied-success-investigating-the-relationship
      0.18  WWW 2020 an-empirical-study-of-the-use-of-integrity-verification-mechanisms-for-web-subre
      0.19  WWW 2026 nerdme-a-named-entity-recognition-dataset-for-indexing-research-artifacts-in-cod
      0.23  WWW 2022 ready-player-one-eliciting-diverse-knowledge-using-a-configurable-game
      0.25  WWW 2023 caml-carbon-footprinting-of-household-products-with-zero-shot-semantic-text-simi
      0.28  WWW 2012 active-objects-actions-for-entity-centric-search
 
── Size of the double-coded sample ──
    457   89.3%  papers naming a statistic that also state a sample size
    100   21.9%  <50
     28    6.1%  50–99
     70   15.3%  100–249
    102   22.3%  250–999
    157   34.4%  >=1000
  median annotated sample: 385
 
── Trend: share of annotating papers that name an agreement statistic ──
  (2025–2026 are PROVISIONAL: CCS/IMC 2026 not yet held, IEEE S&P and WWW 2026 under-selected)
  period            annotating   names stat   share   states count   share
  2010–2014              329           17    5.2%            62   18.8%
  2015–2019              674           67    9.9%           172   25.5%
  2020–2024             1592          274   17.2%           696   43.7%
  2025–2026 (prov.)      723          154   21.3%           358   49.5%
 
── Trend: which coefficient, by period (of papers naming one in that period) ──
  period            N      Cohen's kappa   Fleiss' kappa  Krippendorff's  kappa (variant  raw agreement 
  2010–2014           17       3 (17.6%)       4 (23.5%)        1 (5.9%)       3 (17.6%)       7 (41.2%)
  2015–2019           67      33 (49.3%)      12 (17.9%)       9 (13.4%)        1 (1.5%)      13 (19.4%)
  2020–2024          274     144 (52.6%)      36 (13.1%)      48 (17.5%)        4 (1.5%)      32 (11.7%)
  2025–2026 (prov.)  154      72 (46.8%)      23 (14.9%)      26 (16.9%)        2 (1.3%)      28 (18.2%)
 
── The web-crawl subset ──
  annotating papers that also ran a crawl: 864
    137   15.9%  name an agreement statistic
    310   35.9%  state an annotator count
    748   86.6%  annotated by the authors
 
── Annotation as the ground truth for a classifier ──
  matcher: /manual|hand[- ]?(label|annotat|cod)|human (annotat|label|cod)|author[s]? (label|annotat|cod)/i
   4439   75.8%  papers that classified or labelled something
   1443   32.5%  whose stated ground truth is manual/human annotation
    196   13.6%    ...of which name an agreement statistic
     48    3.3%    ...of which state exactly one annotator
 
── Coding approach (of 3318, multi-valued) ──
   1672   50.4%  binary-verification
   1384   41.7%  ad-hoc
    935   28.2%  predefined-codebook
    281    8.5%  open-coding
    134    4.0%  thematic-analysis
     41    1.2%  grounded-theory
     20    0.6%  not-stated
 
── LLM involvement in annotating papers ──
  matcher: /\b(gpt|chatgpt|llm|large language model|claude|gemini|llama|mistral)\b/i
    267    8.0%  annotating papers with an LLM signal
  year   LLM signal / annotating   share   (2025–2026 provisional)
  2010        0 /   50      0.0%
  2011        0 /   57      0.0%
  2012        0 /   69      0.0%
  2013        0 /   74      0.0%
  2014        0 /   79      0.0%
  2015        0 /  111      0.0%
  2016        0 /  100      0.0%
  2017        1 /  129      0.8%
  2018        0 /  130      0.0%
  2019        1 /  204      0.5%
  2020        2 /  242      0.8%
  2021        1 /  255      0.4%
  2022        4 /  303      1.3%
  2023        9 /  394      2.3%
  2024       44 /  398     11.1%
  2025      106 /  479     22.1%
  2026       99 /  244     40.6%
 
     91   34.1%  of LLM-signal papers name an agreement statistic
    421   13.8%  of the remaining annotating papers name one

The fold

humanAnnotation.agreementMetric is free text: 133 distinct strings across the corpus. Four things vary independently — case (Cohen's kappa / Cohen's Kappa), apostrophe glyph (Cohen's / Cohen’s), word versus Greek letter (kappa / κ / 𝜅), and possessive form (Fleiss' / Fleiss's / Fleiss) — and a single string may name several coefficients (“Cohen's kappa; Fleiss' kappa; Krippendorff's alpha”). Folding therefore splits on semicolon, comma, the word and, slash and plus first, then matches each part against an ordered rule list, first match wins.

Committed as scripts/ira_fold.mjs. The rules, in precedence order:

# Pattern (case-insensitive) Family
1 gwet Gwet's AC1/AC2
2 kupper[- ]?hafner Kupper–Hafner
3 krippendorff Krippendorff's alpha
4 fleiss Fleiss' kappa
5 cohen Cohen's kappa
6 intra[- ]?class or icc intra-class correlation
7 scott Scott's pi
8 word-boundary kappa / κ / kappas kappa (variant unnamed)
9 f1, jaccard, pearson, spearman, correlation, accuracy, cross-validation, precision, recall not an agreement coefficient
10 an optional quantifier word followed by agreement / concordance / consistency / match raw agreement (no chance correction)

Rule 9 must precede rule 10. It did not in the first version, and “exact-match F1” folded to raw agreement because rule 10's match alternative fired first. Caught by a five-case hand test before any figure was published.

The rules are deliberately narrow. A rule broad enough to swallow an unfamiliar coefficient is worse than a residue line, because the residue is visible and a wrong fold is not.

Unmapped residue, printed in full

24 distinct strings, 29 tuple mentions, from the 512 papers naming a statistic. Repeated verbatim from the report output above, because the commentary below refers to individual lines:

      3  "inter-rater reliability"
      2  "Kupper"
      2  "consensus coding"
      2  "interrater reliability"
      1  "Hafner inter-rater reliability metric"
      1  "inter-coder reliability"
      1  "Kendall's Tau coefficient"
      1  "Hafner's statistic"
      1  "Kendall's W"
      1  "comparison of coding results"
      1  "Kendall's tau"
      1  "RMSE"
      1  "Inter-Rater Reliability (IRR)"
      1  "identical coding"
      1  "error rate"
      1  "consensus threshold"
      1  "Three-fifths majority voting"
      1  "inter-coder reliability score"
      1  "alignment rate"
      1  "group unanimity"
      1  "Kendall's Tau"
      1  "Krippendorf's alpha"
      1  "consensus discussion"
      1  "unanimous result"

Three judgement calls in that list, recorded because a reasonable person would decide otherwise:

  • “Krippendorf's alpha” (one f) is a misspelling in the source paper and was left unmapped. Adding a fuzzy rule for it would fold one paper and open the door to fuzzy-matching everything. It is one paper in 512.
  • Kendall's tau and Kendall's W were left unmapped. W is a concordance coefficient and arguably belongs with the agreement families; tau is a rank correlation and does not. Rather than split them by hand across four tuples, both stay visible.
  • The bare “inter-rater reliability” / “interrater reliability” strings (7 mentions) were left unmapped rather than folded into “kappa (variant unnamed)”. They are papers that name no coefficient at all, and the content page's claim that 10 papers write “kappa” and stop would have become 17 with a looser rule — a materially different sentence. They are counted in the 512 that “name a statistic” (the field is non-null) but contribute no family. This is why the family shares sum to slightly under what a reader might expect.

Effect of folding

Unfolded (top exact string) Folded Undercount if unfolded
Cohen's kappa 170 252 32.5%
Krippendorff's alpha 61 84 27.4%
Fleiss' kappa 33 75 56.0%

Consistent with the corpus-wide warning that free-text fields agree run-to-run on ~20% of exact strings (measured on the previous corpus run and not re-measured — treat as an order of magnitude).

The value parser

agreementValue is also free text. parseCoefficients() in scripts/ira_fold.mjs extracts numeric coefficients from it. Three bugs were found by reading quotes, not by inspection, and each was silently corrupting the distribution:

Input Naive result Why it is wrong Fix
“0.68-0.96 across questions” [0.68, -0.96] a hyphen between two decimals is a range, not a minus sign rewrite digit - digit to digit digit before matching
“average κ = 0.771; σ = 0.09” [0.771, 0.09] σ is a dispersion, not a second coefficient; bucketing by minimum then reported 0.09 strip sd / std / σ / sigma / se / p / n / df / r followed by = value segments first
“κ = .666, 95% CI [.630, .670]” [0.666, 0.630, 0.670] interval bounds, not coefficients strip NN% CI […] before matching

Also excluded: any number followed by % (raw agreement, never a kappa — “87% agreement”[]), and bare 0 / 1 unless preceded by a comparison operator (so “κ = 1”[1] but “5 of 7 agreement required”[]).

Known one-sided limitation, stated on the content page: because the range fix strips leading minus signs, a genuinely negative coefficient (below-chance agreement) can never be returned. Below-chance results are invisible in the value table. Fixing this properly needs the surrounding sentence, not the agreementValue field, and was not attempted.

Before the fixes, the value table reported 1 paper below 0.00 and a minimum of −0.96 — both artefacts. That figure was caught before publication by chasing the minimum to its source paper [1Zhou, Jiawei; Zhang, Zidong; Ying, Lingyun; Chai, Huajun; Cao, Jiuxin; Duan, Haixin (2025): "Hey, Your Secrets Leaked! Detecting and Characterizing Secret Leakage in the Wild", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]-style, i.e. by reading the quote.

Bucketing choice. A paper commonly reports several coefficients (one per category, or a range). Each paper is bucketed by its lowest value, which is the conservative reading and is stated on the page. Bucketing by highest instead would move papers out of the low bands; the page does not publish that sensitivity check, and a reader who wants it can re-run the script with Math.min changed to Math.max.

The full-text probes

The extraction schema has no field for “does this paper cite Landis & Koch” or “does it report an interval on kappa”, so those come from regex over the stored full texts. Committed as scripts/ira_fulltext_probes.sh. Each probe collapses newlines and rejoins soft hyphens before matching, because a PDF line break inside a phrase silently undercounts.

Unedited output:

ira_fulltext_probes-output.txt
full texts searched: 5869
     7  cites the Landis-Koch benchmark bands
    94  uses any of the six Landis-Koch verbal band words, incl. "poor" (over-counts: also matches "in substantial agreement with prior work")
    67  a Landis-Koch band within 120 chars of a coefficient
     0  uses PABAK (prevalence-adjusted kappa)
     5  uses Gwet's AC1
     2  reports a 95% interval next to a kappa
   230  writes the phrase Cohen's kappa
     0  names the library function it computed with
     2  names the kappa prevalence paradox
    68  names Krippendorff's alpha
   450  uses the phrase inter-rater/coder/annotator agreement or reliability
   435  describes disagreement resolution, adjudication or majority vote
    18  reports LLM run-to-run / self-consistency agreement
     0  double-codes with a coder external to the author team

One probe bug, caught before publication. The first version of the interval probe used kappa[^.]{0,30}95% (CI|confidence). It returned 0, which was wrong: the character class excludes ., and the one string that matters — “κ = .666, 95% CI” — contains a decimal point. Replacing [^.]{0,80} with .{0,80} returned 2. The pipeline was then sanity-checked by probing for Cohen's kappa itself, which returned 230 — so a 0 from this harness is now a result rather than a suspicion.

Two probes were rewritten after the generic review, and both moved materially. The disagreement-resolution probe was resolved (the |all )?disagreement|adjudicat, which found 65 papers; widened to also catch disagreements were discussed, third annotator, majority vote and the like, it finds 435. The content page's original sentence — “65 papers in 5,869 describe this at all” — was therefore wrong by a factor of seven, and “at all” was the opposite of a lower-bound claim. The Landis–Koch verbal-band probe matched only two of the five bands and matched them anywhere in the text, including “in substantial agreement with prior work”; it now reports both an over-counting figure (94, any band anywhere) and a defensible one (67, a band within 120 characters of a coefficient). The page uses 67. Two further probes — LLM self-consistency and external coders — were added because the page was asserting probe results that did not exist.

Probes are lower bounds and are labelled as such on the page. PABAK returning 0 and the library-function probe returning 0 are both plausible but unproven: a paper could describe either without using the term.

Cross-check on the extraction. The full-text probe for the phrase Cohen's kappa finds 230 papers; the extraction folds 252 papers to that family. The gap is explicable (Cohen-Kappa, Kappa statistic [Cohen], and cases the extractor inferred from context) and the two agree to within ~10%. Neither is exact. This is the reason the content page presents coefficient counts as a ranking with figures rather than as precise rates.

Quotes spot-checked

21 evidence.quote values behind figures or named papers were checked against data/fulltext/<year>/<venue>/<slug>/paper.cols.txt, whitespace-collapsed and with soft hyphens rejoined.

Paper Verbatim match Value confirmed in full text
[2Chen, Baiqi; Lyu, Jiawei; Wu, Tingmin; Chhetri, Mohan Baruwal; Bai, Guangdong (2025): "Semantics-Aware Cookie Purpose Compliance", in: Proceedings of the ACM Web Conference. (DOI)] WWW 2025 yes yes — Fleiss' κ 0.978, 2,300 cookies, 3 coders
[3Skolka, Philippe; Staicu, Cristian-Alexandru; Pradel, Michael (2019): "Anything to Hide? Studying Minified and Obfuscated Code in the Web", in: Proceedings of the ACM Web Conference. (DOI)] WWW 2019 yes yes — “The pairwise agreement ranges between 0.64 and 0.93, with an average of 0.81”; the “4-out-of-5 majority” later in the paper confirms 5 coders
[4Alam, Mahbub; Rahman, Muhammad Lutfor; Paul, Sonjoy Kumar; Hays, Amy W.; Hussain, Aftab; Huq, Md Imanul; Saxena, Nitesh (2026): "SoK: PHILTER: Uncovering Security and Functional Gaps in AI-based Phishing Website Detection Literature via an LLM-based Reasoning Framework", in: Proceedings of the USENIX Security Symposium. (Link)] USENIX 2026 yes yes — 345/385 (89.61%)
[5Zeng, Eric; Wei, Miranda; Gregersen, Theo; Kohno, Tadayoshi; Roesner, Franziska (2021): "Polls, Clickbait, and Commemorative \$2 Bills: Problematic Political Advertising on News and Media Websites Around the 2020 U.S. Elections", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] IMC 2021 (×2) (checked during drafting; not cited on the final page) no yes — 0.771, “Fleiss”, 8,836 ads, 200-ad subset all present
[6Coopamootoo, Kovila P. L.; Groß, Thomas (2017): "Why Privacy Is All But Forgotten: An Empirical Study of Privacy and Sharing Attitude", Proceedings on Privacy Enhancing Technologies 2017(4):97-118. (DOI)] PETS 2017 yes, on re-check yes — the reviewer located the whole sentence: “There was a substantial agreement between the two coders' judgment, κ = .666, 95% CI [.630, .670], p < .001.” The interval bounds are the paper's own.
[7Cory, Thomas; Rieder, Wolf; Krämer, Julia; Raschke, Philip; Herbke, Patrick; Küpper, Axel (2026): "Word-level Annotation of GDPR Transparency Compliance in Privacy Policies using Large Language Models", Proceedings on Privacy Enhancing Technologies 2026(1):509-528. (DOI)] PETS 2026 no yes — 0.669, “Krippendorff”, “six domain experts”, “200 privacy policies”
[8Sun, Chen; Vekaria, Yash; Nithyanand, Rishab (2026): "On the Suitability of LLM-Driven Agents for Dark Pattern Audits", Proceedings on Privacy Enhancing Technologies 2026(4):927-946. (DOI)] PETS 2026 one of two yes — full sentence recovered: “measured using Cohen's 𝜅 on binary (present/absent) labels for each dark pattern category, yielding 𝜅 = 51.9%”
[9Zhang, Shirley; Chung, Paul; Vervelde, Jacob; Korapati, Nishant; Chatterjee, Rahul; Fawaz, Kassem (2025): "Abusability of Automation Apps in Intimate Partner Violence", in: Proceedings of the USENIX Security Symposium. (Link)] USENIX 2025 yes yes — “without seeing LLM responses” and κ = 0.96 both located
[10Khatun, Mst Eshita; Noureddine, Lamine; Bello, Sideeq; Ali-Gombe, Aisha (2026): "Disclosure Divergence: Measuring Privacy Policy and Data Safety Misalignment at Scale", Proceedings on Privacy Enhancing Technologies 2026(4):213-231. (DOI)] PETS 2026 no partly — “0.78” and “40 apps” present; the exact phrase “90% observed agreement” was not located, so the page does not quote it verbatim
[11Swart, Milena de; Hengst, Floris den; Chen, Jieying (2025): "Detecting Linguistic Bias in Government Documents Using Large Language Models", in: Proceedings of the ACM Web Conference. (DOI)] WWW 2025 no yes — 0.35, “Fleiss”, “14 expert”, “12,076”
Kotzias et al., Ctrl+Alt+Deceive, NDSS 2025 (checked during drafting; not cited on the final page) no yes — 0.756, “four analysts”, “26 industry tags”
[1Zhou, Jiawei; Zhang, Zidong; Ying, Lingyun; Chai, Huajun; Cao, Jiuxin; Duan, Haixin (2025): "Hey, Your Secrets Leaked! Detecting and Characterizing Secret Leakage in the Wild", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] IEEE S&P 2025 no yes — full sentence recovered including the Landis–Koch attribution
[12Xian, Lu; Tran, Van Hong; Lee, Lauren; Kumar, Meera; Zhang, Yichen; Schaub, Florian (2025): "Layered, Overlapping, and Inconsistent: A Large-Scale Analysis of the Multiple Privacy Policies and Controls of U.S. Banks", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] CCS 2025 (×2) no partly — 0.95, 0.89 and “Krippendorff” present; the glyph Cu-α did not match, so the variant name rests on the extraction
CCS 2023 insecure code with AI assistants no yes — 0.68 and 0.96 present; this is the paper that exposed the range-parsing bug
WWW 2018 StaQC no yes — 0.658, 0.691 present
PETS 2024 connecting online activities no yes — 0.74 present
WWW 2020 integrity verification mechanisms yes yes
NDSS 2019 How Bad Can It Git yes yes
WWW 2012 Learning from the Past (×2) yes yes

Nine of the 21 matched verbatim. Every failure was diagnosed and every one was an artefact of the quote or the rendering, not a fabrication: elided text marked in the stored quote, PDF hyphenation, or two-column splice damage in .cols (visible in the IEEE S&P 2025 case as “independent anno erage … tations from two researchers”). Two quotes were only partly confirmed and both are handled on the page: the Khatun et al. observed-agreement phrase is not quoted verbatim there, and the Cu-α variant name is attributed to the extraction rather than to a quote.

No quote was found to be absent from its source.

External sources

All fetched 2026-08-20, not recalled. DOIs resolved through content negotiation (Accept: application/vnd.citationstyles.csl+json at doi.org); author lists for PoPETs and USENIX from the venue landing pages with a browser User-Agent, because those two venues carry no authors in the corpus index.

Source Verified as
Cohen 1960 [13Cohen, Jacob (1960): "A Coefficient of Agreement for Nominal Scales", Educational and Psychological Measurement 20(1):37-46. (DOI)] 10.1177/001316446002000104, Educ. Psychol. Meas. 20(1):37–46
Fleiss 1971 [14Fleiss, Joseph L. (1971): "Measuring Nominal Scale Agreement Among Many Raters", Psychological Bulletin 76(5):378-382. (DOI)] 10.1037/h0031619, Psychol. Bull. 76(5):378–382
Landis & Koch 1977 [15Landis, J. Richard; Koch, Gary G. (1977): "The Measurement of Observer Agreement for Categorical Data", Biometrics 33(1):159-174. (DOI)] 10.2307/2529310, Biometrics 33(1):159–174
Feinstein & Cicchetti 1990 [16Feinstein, Alvan R.; Cicchetti, Domenic V. (1990): "High Agreement but Low Kappa: I. The Problems of Two Paradoxes", Journal of Clinical Epidemiology 43(6):543-549. (DOI)] 10.1016/0895-4356(90)90158-L
Cicchetti & Feinstein 1990 [17Cicchetti, Domenic V.; Feinstein, Alvan R. (1990): "High Agreement but Low Kappa: II. Resolving the Paradoxes", Journal of Clinical Epidemiology 43(6):551-558. (DOI)] 10.1016/0895-4356(90)90159-M, same issue
Gwet 2008 [18Gwet, Kilem Li (2008): "Computing Inter-Rater Reliability and Its Variance in the Presence of High Agreement", British Journal of Mathematical and Statistical Psychology 61(1):29-48. (DOI)] 10.1348/000711006×126600
Krippendorff 2004 [19Krippendorff, Klaus (2004): "Reliability in Content Analysis: Some Common Misconceptions and Recommendations", Human Communication Research 30(3):411-433. (DOI)] 10.1111/j.1468-2958.2004.tb00738.x
Krippendorff, Content Analysis [20Krippendorff, Klaus (2018): "Content Analysis: An Introduction to Its Methodology". SAGE Publications.] current edition is the 4th, 2018, SAGE, ISBN 9781506395661
McHugh 2012 [21McHugh, Mary L. (2012): "Interrater Reliability: The Kappa Statistic", Biochemia Medica 22(3):276-282. (DOI)] 10.11613/BM.2012.031
Hallgren 2012 [22Hallgren, Kevin A. (2012): "Computing Inter-Rater Reliability for Observational Data: An Overview and Tutorial", Tutorials in Quantitative Methods for Psychology 8(1):23-34. (DOI)] 10.20982/tqmp.08.1.p023
Gilardi, Alizadeh & Kubli 2023 [23Gilardi, Fabrizio; Alizadeh, Meysam; Kubli, Maël (2023): "ChatGPT Outperforms Crowd Workers for Text-Annotation Tasks", Proceedings of the National Academy of Sciences 120(30):e2305016120. (DOI)] 10.1073/pnas.2305016120, PNAS 120(30):e2305016120
Ziems et al. 2024 [24Ziems, Caleb; Held, William; Shaikh, Omar; Chen, Jiaao; Zhang, Zhehao; Yang, Diyi (2024): "Can Large Language Models Transform Computational Social Science?", Computational Linguistics 50(1):237-291. (DOI)] 10.1162/coli_a_00502, Computational Linguistics 50(1):237–291
Pangakis, Wolken & Fasching 2023 [25Pangakis, Nicholas; Wolken, Samuel; Fasching, Neil (2023): "Automated Annotation with Generative AI Requires Validation". arXiv:2306.00176. (Link)] arXiv:2306.00176
Törnberg 2024 [26Törnberg, Petter (2024): "Best Practices for Text Annotation with Large Language Models", Sociologica 18(2):67-85. (DOI)] 10.6092/issn.1971-8853/19461, Sociologica 18(2):67–85
Baumann et al. 2025 [27Baumann, Joachim; Röttger, Paul; Urman, Aleksandra; Wendsjö, Albert; Plaza-del-Arco, Flor Miriam; Gruber, Johannes B.; Hovy, Dirk (2025): "Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation". arXiv:2509.08825. (Link)] arXiv:2509.08825, submitted 2025-09-10, 7 authors confirmed from the arXiv abstract page; still preprint-only, and the page says so

Tooling versions, checked today

Tool Version Gives a CI?
scikit-learn metrics.cohen_kappa_score 1.9.0 (PyPI) no — supports weights={'linear','quadratic'} only
statsmodels.stats.inter_rater.cohens_kappa 0.14.6 (latest pip-installable) yes — returns variance and interval
statsmodels.stats.inter_rater.fleiss_kappa 0.14.6 no
krippendorff (PyPI) 0.8.2, released 2025-11-03 no
R irr CRAN 0.85, published 2026-06-17 — active, not archived yes (SE)
R irrCAC CRAN 1.4, published 2026-04-27 — active yes

The sub-agent reported statsmodels 0.15.0 from the documentation site; pip index versions statsmodels on this machine returns 0.14.6 as the newest installable. The page cites 0.14.6. This is exactly the failure mode the external-currency pass exists to catch, and it was caught by re-checking the agent rather than by trusting it.

Rejected sources

  • “Chew et al. (2023)” on LLM qualitative coding — the sub-agent surfaced it only as a secondary citation with no retrievable title, venue or DOI. Not cited. A citation nobody can resolve is worse than a missing one.
  • A 2026 arXiv preprint on “reliability without validity” for LLM-as-judge, offered as a currency cite. The identifier could not be independently confirmed and the claim it would support is already carried by [25Pangakis, Nicholas; Wolken, Samuel; Fasching, Neil (2023): "Automated Annotation with Generative AI Requires Validation". arXiv:2306.00176. (Link)]. Not cited.
  • A single canonical “kappa bands are arbitrary” paper. Asked for; none could be confirmed as the best-cited. The page instead makes the point from Landis & Koch's own text and from [21McHugh, Mary L. (2012): "Interrater Reliability: The Kappa Statistic", Biochemia Medica 22(3):276-282. (DOI)], which is checkable. No speculative attribution.
  • The exact page number of the α ≥ 0.800 / ≥ 0.667 wording in the 2018 4th edition. Confirmed only against the 2004 article (p. 241). The page cites the 2004 article for the thresholds and the book as the long version, rather than asserting a page in an edition nobody here has.

The simulation

scripts/kappa_on_skewed_labels.py — synthetic, fixed seed 20260820, numpy 2.4.6, no other dependency. It demonstrates properties of the coefficient, not facts about the web, and the page says so. The script is published on the content page as a downloadable file block, byte-identical to the committed script (checked programmatically). Two output blocks are quoted there: part C verbatim, and parts A and B verbatim except that the script's own two interpretive paragraphs between the tables are removed — the tables themselves are byte-identical substrings of the real output, checked programmatically, and the removed prose is repeated in the page's own text around the block.

Part C originally used three coders of equal skill and produced a pairwise spread of 0.018 — technically correct and rhetorically useless. It was changed to unequal skill (2%/2%/12%), which is the realistic case and gives a spread of 0.196. Recorded because it is a choice made to make a point land, and a reader should know the parameters were chosen rather than observed.

What could not be established

  • Whether author-coding inflates agreement, and by how much. There is no extraction field for “was the coder outside the author team”, and a deliberately narrow full-text probe (added after review) returns zero papers. Zero from a narrow regex is weak evidence: the page states it as “we could not find one”, not as “there are none”. This is the page's first open question and it is genuinely open.
  • How many papers computed a coefficient and did not report it. Structurally invisible. Every reporting-gap figure on the page is therefore a bound, and the page says so rather than implying a rate.
  • Whether the 81.6% who report neither statistic nor number should have. Half of this coding is single-coder binary verification where reliability is undefined; the extraction cannot separate “correctly single-coder” from “silently multi-coder”. The page uses the ≥2-coders subset (42.0%) as the honest version and presents 81.6% with that caveat attached.
  • Class balance of the coded tasks. humanAnnotation has no field for it, and it is what would decide how much of the low-kappa tail is the prevalence paradox. Would need the released label sets.
  • LLM self-agreement across runs, fully. There is no extraction field. A probe was written after the generic review flagged that the page asserted a probe result it did not have; it finds 18 papers in 5,869. Whether those 18 are the practice or the regex is narrow is not resolved, and the page says so.
  • sampleSize semantics. The field is “the coded sample”, which some papers fill with the whole labelled corpus (one record carries 709,125) and others with the double-coded subsample. The median (385) is robust; the bucket table is only approximately about double-coded sets. Flagged on the page.

Review log

Five review passes, all told explicitly that the author's context may not be exhaustive, all handed the page text, the report script, its unedited output and these notes.

The five review passes

Pass Model Verdict
Figures vs script Sonnet Re-ran all three scripts; all outputs byte-identical to the committed ones; every table cell and percentage traced to the script output. 2 findings.
Citations and quotes Sonnet Resolved all 23 new DOIs; verified quotes against paper.cols.txt. 3 findings, one serious.
External currency Sonnet Fetched PyPI, CRAN, Crossref, arXiv and the SAGE catalogue. 2 findings + 1 suggestion.
Generic Fable No checklist. 17 findings, including the three most serious defects on this page.
Figures vs script, re-run Sonnet Run again after the probe script was rewritten. All three scripts reproduce byte-for-byte; 2 findings, both minor.

Findings accepted and fixed

# Reviewer Finding Action
1 citations A quote was attributed to the wrong paper. The page had Sun, Vekaria & Nithyanand [8Sun, Chen; Vekaria, Yash; Nithyanand, Rishab (2026): "On the Suitability of LLM-Driven Agents for Dark Pattern Audits", Proceedings on Privacy Enhancing Technologies 2026(4):927-946. (DOI)] saying their annotators worked “without seeing LLM responses”. A full-file grep of that paper finds the phrase nowhere. The sentence is in fact from Zhang et al. [9Zhang, Shirley; Chung, Paul; Vervelde, Jacob; Korapati, Nishant; Chatterjee, Rahul; Fawaz, Kassem (2025): "Abusability of Automation Apps in Intimate Partner Violence", in: Proceedings of the USENIX Security Symposium. (Link)], USENIX Security 2025, whose extraction record sits a few lines away in the same query output. This is the most serious defect the review caught and it was an authoring error, not an extraction error — the quote was real, the attribution was not. Re-attributed to [9Zhang, Shirley; Chung, Paul; Vervelde, Jacob; Korapati, Nishant; Chatterjee, Rahul; Fawaz, Kassem (2025): "Abusability of Automation Apps in Intimate Partner Violence", in: Proceedings of the USENIX Security Symposium. (Link)], verified in that paper's full text, entry added to the bibliography, spot-check table row added above.
2 figures Undisclosed denominator. “the median double-coded sample here is 385 items and 22% double-code fewer than 50” appeared on the content page with no population named. The figures are over the 457 papers that name a statistic and state a coded sample size (89.3% of the 512). Denominator stated in the <WRAP> block, again beside the simulation, and again in the methodology list.
3 citations A quote reordered its source's clauses. The page printed “average 0.81; range 0.64–0.93” as a direct quote from [3Skolka, Philippe; Staicu, Cristian-Alexandru; Pradel, Michael (2019): "Anything to Hide? Studying Minified and Obfuscated Code in the Web", in: Proceedings of the ACM Web Conference. (DOI)]; the paper says “The pairwise agreement ranges between 0.64 and 0.93, with an average of 0.81”. Content identical, word order not. Replaced with the paper's own sentence, verbatim.
4 citations Bibliographic error. [10Khatun, Mst Eshita; Noureddine, Lamine; Bello, Sideeq; Ali-Gombe, Aisha (2026): "Disclosure Divergence: Measuring Privacy Policy and Data Safety Misalignment at Scale", Proceedings on Privacy Enhancing Technologies 2026(4):213-231. (DOI)] was entered with number = {3}; the DOI record says issue 4. Corrected before the bibliography was saved.
5 currency CRAN dates off by a day or two. irr published 2026-06-17 (page said 06-15); irrCAC 2026-04-27 (page said 04-26). Both corrected.
6 currency A newer, stronger companion citation exists for the LLM-validation argument: Baumann et al., Large Language Model Hacking (arXiv:2509.08825) — 37 annotation tasks from 21 published studies, 18 models, finding that pipeline variation flips roughly one hypothesis in three. Added, with its preprint status stated inline. Verified against the arXiv abstract page (7 authors, submitted 2025-09-10), not recalled.
7 figures The fold undercount was rounded to “33%” on the content page against “32.5%” on this one. Content page now says 32.5%.
8 generic The page asserted a universal negative it had never queried: “no paper in the corpus double-codes with an outside coder”. No such query existed in the report script and no such probe existed. Probe written, returns zero; the claim is now stated as “we could not find one, and the probe is narrow”.
9 generic The page asserted a probe result for LLM self-agreement that did not exist (“zero papers report it systematically”). Probe written. It finds 18 papers, not zero. Both the content page and this one corrected.
10 generic The disagreement-resolution probe was far too narrow, and the page stated its 65 as “describe this at all”. Probe widened; the real figure is 435, and the sentence's meaning reverses (resolution is common practice). See The full-text probes.
11 generic The Landis–Koch verbal-band probe matched 2 of 5 bands and could fire on “in substantial agreement with prior work”. Rewritten to report both an over-counting figure (94) and a proximity-constrained one (67). The page now uses 67.
12 generic “Percentile bootstrap” mislabelled a Monte Carlo. Part B simulates 2,000 fresh datasets from the generative model; it does not resample one. On a page that tells readers to bootstrap, this would have been caught by any statistician. Label corrected in the committed script, re-run, and the page's file block and quoted output re-verified byte-identical.
13 generic The LLM matcher over-fires on pre-2023 years (1–2 papers per year in 2017–2021, from llama and Gemini in unrelated senses), and the content table quietly began at 2022. The table still begins at 2022; the page now says why, and that the levels should not be quoted precisely.
14 generic Several probe counts were stated in prose without the lower-bound hedge the methodology section promises; “within noise in both directions” was asserted with no test; “the wider community is better” was stated as fact; “91 papers reporting a coefficient” overstated what the script counts. All four reworded.
15 generic Causal glosses (“the reason is structural rather than cultural”) presented an interpretation of a 34.1%-vs-13.8% difference as established. Marked as interpretation in both places.
16 generic The spot-check table listed two papers not cited on the final page, and the content table said “Gwet's AC1” where the script says “AC1/AC2”. Both labelled/corrected.

Findings rejected, with reasons

# Reviewer Finding Why not acted on
A citations Two DOIs are already duplicated in the live bibliography (10.56553/popets-2022-0063 and 10.56553/popets-2025-0063, each under two keys). Real, and worth fixing — but pre-existing, unrelated to this page, and silently editing other pages' citekeys from here would break whichever page cites the losing key. Recorded here so it is not lost.
B citations Coopamootoo & Groß [6Coopamootoo, Kovila P. L.; Groß, Thomas (2017): "Why Privacy Is All But Forgotten: An Empirical Study of Privacy and Sharing Attitude", Proceedings on Privacy Enhancing Technologies 2017(4):97-118. (DOI)] appears in the report's “Cohen's kappa with >2 annotators” list at n=3, which sits oddly with the page describing it as a careful two-coder case. Not a defect. The paper has three coding arrangements — an author coding all units, and a separate pair of trained PhD students — and reports the coefficient on the pair. That is exactly the point the page makes about it, and it is also a fair example of why the n>2 list is a flag for inspection rather than an error count. A sentence saying so was added to the mismatch section instead.
C citations The Tang, Bauer & Christin [28Tang, Jenny; Bauer, Lujo; Christin, Nicolas (2025): "Misuse, Misreporting, Misinterpretation of Statistical Methods in Usable Privacy and Security Papers", in: Proceedings of the Twenty-First Symposium on Usable Privacy and Security (SOUPS), pp. 475-493. USENIX Association. (Link)] quote cannot be verified against this corpus (SOUPS is not one of the seven venues). Correct and unavoidable. The quote and the key are inherited unchanged from Hypothesis testing, where they were verified in that page's own run; re-verifying an external paper twice is not a good use of a pass. The page already states SOUPS is outside the corpus.
D currency The exact 2018 copyright year of Krippendorff's 4th edition could not be confirmed from SAGE directly (the catalogue page did not return a body). The ISBN, edition and publisher were confirmed via Open Library, and no 5th edition exists. Accepted as-is; the thresholds are cited to the 2004 article, which is fully verified.
E currency Two further 2026 arXiv preprints on validating LLMs as measurement instruments were surfaced as leads. Neither could be verified beyond an abstract listing. Not cited, on the same rule that rejected the sources listed above.
F generic User studies and Artifacts are red links. Not fixed here. Both are links the wiki already promises: sitemap.mjs lists them under “promised but missing”, each linked from six or more existing pages including Hypothesis testing. Creating them is a separate item, and pointing this page somewhere else would break the convention every neighbour follows.
G generic The kappa formula restated in prose is close to the textbook line the site avoids. Kept. It is two clauses and it is the mechanism the whole rare-class section turns on; a reader who does not have (Po − Pe)/(1 − Pe) in front of them cannot follow why the denominator collapses.
  • Interrater agreement — the page these notes are for.
  • Corpus — corpus-level provenance: selection, funnel, extraction quality, the 2026-08-11 corpus extension.

References

[1]
Zhou, Jiawei; Zhang, Zidong; Ying, Lingyun; Chai, Huajun; Cao, Jiuxin; Duan, Haixin (2025): "Hey, Your Secrets Leaked! Detecting and Characterizing Secret Leakage in the Wild", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[2]
Chen, Baiqi; Lyu, Jiawei; Wu, Tingmin; Chhetri, Mohan Baruwal; Bai, Guangdong (2025): "Semantics-Aware Cookie Purpose Compliance", in: Proceedings of the ACM Web Conference. (DOI)
[3]
Skolka, Philippe; Staicu, Cristian-Alexandru; Pradel, Michael (2019): "Anything to Hide? Studying Minified and Obfuscated Code in the Web", in: Proceedings of the ACM Web Conference. (DOI)
[4]
Alam, Mahbub; Rahman, Muhammad Lutfor; Paul, Sonjoy Kumar; Hays, Amy W.; Hussain, Aftab; Huq, Md Imanul; Saxena, Nitesh (2026): "SoK: PHILTER: Uncovering Security and Functional Gaps in AI-based Phishing Website Detection Literature via an LLM-based Reasoning Framework", in: Proceedings of the USENIX Security Symposium. (Link)
[5]
Zeng, Eric; Wei, Miranda; Gregersen, Theo; Kohno, Tadayoshi; Roesner, Franziska (2021): "Polls, Clickbait, and Commemorative \$2 Bills: Problematic Political Advertising on News and Media Websites Around the 2020 U.S. Elections", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[6]
Coopamootoo, Kovila P. L.; Groß, Thomas (2017): "Why Privacy Is All But Forgotten: An Empirical Study of Privacy and Sharing Attitude", Proceedings on Privacy Enhancing Technologies 2017(4):97-118. (DOI)
[7]
Cory, Thomas; Rieder, Wolf; Krämer, Julia; Raschke, Philip; Herbke, Patrick; Küpper, Axel (2026): "Word-level Annotation of GDPR Transparency Compliance in Privacy Policies using Large Language Models", Proceedings on Privacy Enhancing Technologies 2026(1):509-528. (DOI)
[8]
Sun, Chen; Vekaria, Yash; Nithyanand, Rishab (2026): "On the Suitability of LLM-Driven Agents for Dark Pattern Audits", Proceedings on Privacy Enhancing Technologies 2026(4):927-946. (DOI)
[9]
Zhang, Shirley; Chung, Paul; Vervelde, Jacob; Korapati, Nishant; Chatterjee, Rahul; Fawaz, Kassem (2025): "Abusability of Automation Apps in Intimate Partner Violence", in: Proceedings of the USENIX Security Symposium. (Link)
[10]
Khatun, Mst Eshita; Noureddine, Lamine; Bello, Sideeq; Ali-Gombe, Aisha (2026): "Disclosure Divergence: Measuring Privacy Policy and Data Safety Misalignment at Scale", Proceedings on Privacy Enhancing Technologies 2026(4):213-231. (DOI)
[11]
Swart, Milena de; Hengst, Floris den; Chen, Jieying (2025): "Detecting Linguistic Bias in Government Documents Using Large Language Models", in: Proceedings of the ACM Web Conference. (DOI)
[12]
Xian, Lu; Tran, Van Hong; Lee, Lauren; Kumar, Meera; Zhang, Yichen; Schaub, Florian (2025): "Layered, Overlapping, and Inconsistent: A Large-Scale Analysis of the Multiple Privacy Policies and Controls of U.S. Banks", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[13]
Cohen, Jacob (1960): "A Coefficient of Agreement for Nominal Scales", Educational and Psychological Measurement 20(1):37-46. (DOI)
[14]
Fleiss, Joseph L. (1971): "Measuring Nominal Scale Agreement Among Many Raters", Psychological Bulletin 76(5):378-382. (DOI)
[15]
Landis, J. Richard; Koch, Gary G. (1977): "The Measurement of Observer Agreement for Categorical Data", Biometrics 33(1):159-174. (DOI)
[16]
Feinstein, Alvan R.; Cicchetti, Domenic V. (1990): "High Agreement but Low Kappa: I. The Problems of Two Paradoxes", Journal of Clinical Epidemiology 43(6):543-549. (DOI)
[17]
Cicchetti, Domenic V.; Feinstein, Alvan R. (1990): "High Agreement but Low Kappa: II. Resolving the Paradoxes", Journal of Clinical Epidemiology 43(6):551-558. (DOI)
[18]
Gwet, Kilem Li (2008): "Computing Inter-Rater Reliability and Its Variance in the Presence of High Agreement", British Journal of Mathematical and Statistical Psychology 61(1):29-48. (DOI)
[19]
Krippendorff, Klaus (2004): "Reliability in Content Analysis: Some Common Misconceptions and Recommendations", Human Communication Research 30(3):411-433. (DOI)
[20]
Krippendorff, Klaus (2018): "Content Analysis: An Introduction to Its Methodology". SAGE Publications.
[21]
McHugh, Mary L. (2012): "Interrater Reliability: The Kappa Statistic", Biochemia Medica 22(3):276-282. (DOI)
[22]
Hallgren, Kevin A. (2012): "Computing Inter-Rater Reliability for Observational Data: An Overview and Tutorial", Tutorials in Quantitative Methods for Psychology 8(1):23-34. (DOI)
[23]
Gilardi, Fabrizio; Alizadeh, Meysam; Kubli, Maël (2023): "ChatGPT Outperforms Crowd Workers for Text-Annotation Tasks", Proceedings of the National Academy of Sciences 120(30):e2305016120. (DOI)
[24]
Ziems, Caleb; Held, William; Shaikh, Omar; Chen, Jiaao; Zhang, Zhehao; Yang, Diyi (2024): "Can Large Language Models Transform Computational Social Science?", Computational Linguistics 50(1):237-291. (DOI)
[25]
Pangakis, Nicholas; Wolken, Samuel; Fasching, Neil (2023): "Automated Annotation with Generative AI Requires Validation". arXiv:2306.00176. (Link)
[26]
Törnberg, Petter (2024): "Best Practices for Text Annotation with Large Language Models", Sociologica 18(2):67-85. (DOI)
[27]
Baumann, Joachim; Röttger, Paul; Urman, Aleksandra; Wendsjö, Albert; Plaza-del-Arco, Flor Miriam; Gruber, Johannes B.; Hovy, Dirk (2025): "Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation". arXiv:2509.08825. (Link)
[28]
Tang, Jenny; Bauer, Lujo; Christin, Nicolas (2025): "Misuse, Misreporting, Misinterpretation of Statistical Methods in Usable Privacy and Security Papers", in: Proceedings of the Twenty-First Symposium on Usable Privacy and Security (SOUPS), pp. 475-493. USENIX Association. (Link)
provenance/statistics/interrater_agreement.txt · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki