| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| statistics:hypothesis_testing [2026/08/13 12:54] – [Open Questions] formatting admin | statistics:hypothesis_testing [2026/09/04 08:29] (current) – Review pass: say 'at least a third' rather than 'mostly', and note that the corpus-wide figure is scored by a different, stricter rule. Authored by Claude karel.kubicek.claude |
|---|
| 38 of the 1,025 papers (3.7%) report a normality check — Shapiro–Wilk in most of them. Using one to //choose// between a //t//-test and a rank test is a documented mistake: the two-stage procedure distorts the type-I error rate of whatever runs second, and the pre-test's own power depends on //n// in the wrong direction, so it waves through non-normality in small samples and rejects trivial non-normality in large ones. Rochon, Gondan & Kieser {[rochon2012_totest]} and Rasch, Kubinger & Moder {[rasch2011_pretesting]} both conclude that pre-testing does not pay off. | 38 of the 1,025 papers (3.7%) report a normality check — Shapiro–Wilk in most of them. Using one to //choose// between a //t//-test and a rank test is a documented mistake: the two-stage procedure distorts the type-I error rate of whatever runs second, and the pre-test's own power depends on //n// in the wrong direction, so it waves through non-normality in small samples and rejects trivial non-normality in large ones. Rochon, Gondan & Kieser {[rochon2012_totest]} and Rasch, Kubinger & Moder {[rasch2011_pretesting]} both conclude that pre-testing does not pay off. |
| |
| <wrap todo> | <WRAP todo> |
| **Decide from the design and the estimand, before the data arrives, and say so.** A defensible sentence looks like: "Because our outcome is a count with a long right tail and we care about the typical site rather than the mean, we pre-specified rank-based tests." An indefensible one is "Shapiro–Wilk was significant, so we used Mann-Whitney." | **Decide from the design and the estimand, before the data arrives, and say so.** A defensible sentence looks like: "Because our outcome is a count with a long right tail and we care about the typical site rather than the mean, we pre-specified rank-based tests." An indefensible one is "Shapiro–Wilk was significant, so we used Mann-Whitney." |
| |
| The best example of doing this well in the corpus splits the choice **per outcome variable** and states the reason for each. Mai et al. {[mai2025_more]} use //"one-way repeated measures ANOVA as the omnibus test since the ad load data follow a normal distribution"// for one outcome and, //"for the hypothesis on predatory ad rates, since the rates do not approximately follow a normal distribution, we use Friedman test as the omnibus test"// for another, in the same paper.((Both sentences are in the paper's PDF and in **neither** of the corpus's stored text renderings of it — see [[provenance:statistics:hypothesis_testing]] for that discrepancy, which is a fact about the corpus and not about the paper.)) | The best example of doing this well in the corpus splits the choice **per outcome variable** and states the reason for each. Mai et al. {[mai2025_more]} use //"one-way repeated measures ANOVA as the omnibus test since the ad load data follow a normal distribution"// for one outcome and, //"for the hypothesis on predatory ad rates, since the rates do not approximately follow a normal distribution, we use Friedman test as the omnibus test"// for another, in the same paper. |
| </wrap> | </WRAP> |
| |
| ==== What Mann-Whitney U actually tests ==== | ==== What Mann-Whitney U actually tests ==== |
| * **A paper counts once**, never once per mention. 1,806 tuples across 1,025 papers. | * **A paper counts once**, never once per mention. 1,806 tuples across 1,025 papers. |
| * **''statistics.kind'' is a mid-band field**: two independent extraction runs over identical text agreed on it for 68% of papers, so a repeat run would move these percentages by a few points. The caveat applies to every share here and not to the folded rankings, which are rankings. That 68% figure was measured on the **previous** corpus run and has not been re-measured. | * **''statistics.kind'' is a mid-band field**: two independent extraction runs over identical text agreed on it for 68% of papers, so a repeat run would move these percentages by a few points. The caveat applies to every share here and not to the folded rankings, which are rankings. That 68% figure was measured on the **previous** corpus run and has not been re-measured. |
| * **Quotes were checked in bulk and by hand.** All 1,806 hypothesis-test quotes against the text the extractor read: 931 exact after whitespace normalisation (51.6%), 540 partial at ≥60% of five-word windows (29.9%), 335 below that (18.5%). **Below-threshold is not "unsupported"** — **five** of the 335 were chased by hand and all five were present, damaged by a dropped citation marker, a reflowed table caption or a two-column splice. The other 330 were not individually checked. Every quote published on this page was located by hand; one (Mai et al.) is in the paper's PDF but in **none** of the corpus's stored text renderings of it. Details on [[provenance:statistics:hypothesis_testing]]. | * **Quotes were checked in bulk and by hand.** All 1,806 hypothesis-test quotes against the text the extractor read: 931 exact after whitespace normalisation (51.6%), 541 partial at ≥60% of five-word windows (30.0%), 334 below that (18.5%). **At least a third of below-threshold is not the extraction's fault.** Re-checked on 2026-09-04 against an independent rendering of the same stored PDF, **118 of the 334 are present there** — the stored text is damaged, not the quote. That leaves 216 below threshold in both renderings, of which five were chased by hand and all five were present, damaged by a dropped citation marker, a reflowed table caption or a two-column splice; the other 211 were not individually checked. Every quote published on this page was located by hand. Details on [[provenance:statistics:hypothesis_testing]], and the corpus-wide version of the same defect — scored by a different, stricter rule, so the percentages are not comparable — on [[literature:corpus]]. |
| * **Sentinels are never answers, and ''statistics'' has none**, which is why the reporting-gap figures are bounds rather than rates. | * **Sentinels are never answers, and ''statistics'' has none**, which is why the reporting-gap figures are bounds rather than rates. |
| * **138 of the 5,859 records are posters** and 251 are ≤4 pages. A poster has its methodology compressed out, so it is systematically likelier to be scored "no test named" — the 5.7% is inflated by an unmeasured amount. | * **138 of the 5,859 records are posters** and 251 are ≤4 pages. A poster has its methodology compressed out, so it is systematically likelier to be scored "no test named" — the 5.7% is inflated by an unmeasured amount. |
| ===== Open Questions ===== | ===== Open Questions ===== |
| |
| | <WRAP todo> |
| * **Nobody has measured the real intra-class correlation of web-measurement outcomes.** The simulation on this page shows the false-positive rate depends almost entirely on the ICC, and no paper in this corpus reports one. Estimating the ICC of "sets a tracking cookie before consent" by tag manager, by CMS and by hosting provider is a small, self-contained, immediately useful study, and it would tell every crawl paper how badly its //p//-values are wrong. | * **Nobody has measured the real intra-class correlation of web-measurement outcomes.** The simulation on this page shows the false-positive rate depends almost entirely on the ICC, and no paper in this corpus reports one. Estimating the ICC of "sets a tracking cookie before consent" by tag manager, by CMS and by hosting provider is a small, self-contained, immediately useful study, and it would tell every crawl paper how badly its //p//-values are wrong. |
| * **How many published crawl findings survive a cluster-aware re-analysis?** Answerable on any paper that released per-site data (see [[:Artifacts]]), and nobody has done it. The three corpus papers that cluster are all 2023+, so essentially the whole literature is un-re-analysed. | * **How many published crawl findings survive a cluster-aware re-analysis?** Answerable on any paper that released per-site data (see [[:Artifacts]]), and nobody has done it. The three corpus papers that cluster are all 2023+, so essentially the whole literature is un-re-analysed. |
| * **No paper states the Mann-Whitney estimand.** 213 papers use the test; the probe for "stochastic superiority" or "stochastic dominance" in that sense returns zero. Whether authors know and do not write it, or write "median" because they believe it, is not answerable from text. | * **No paper states the Mann-Whitney estimand.** 213 papers use the test; the probe for "stochastic superiority" or "stochastic dominance" in that sense returns zero. Whether authors know and do not write it, or write "median" because they believe it, is not answerable from text. |
| * **None of the seven venues asks for a precise test name.** Tang et al. {[tang2025_misuse]} propose a minimum-reporting list for SOUPS; nobody has proposed it to IMC, PoPETs or a security venue, and their reviewer forms are not public, so the current state can only be read off the papers. | * **None of the seven venues asks for a precise test name.** Tang et al. {[tang2025_misuse]} propose a minimum-reporting list for SOUPS; nobody has proposed it to IMC, PoPETs or a security venue, and their reviewer forms are not public, so the current state can only be read off the papers. |
| | </WRAP> |
| ===== Related Pages ===== | ===== Related Pages ===== |
| |