| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| statistics:hypothesis_testing [2026/08/21 08:35] – [Open Questions] karelkubicek | statistics:hypothesis_testing [2026/09/04 08:29] (current) – Review pass: say 'at least a third' rather than 'mostly', and note that the corpus-wide figure is scored by a different, stricter rule. Authored by Claude karel.kubicek.claude |
|---|
| 38 of the 1,025 papers (3.7%) report a normality check — Shapiro–Wilk in most of them. Using one to //choose// between a //t//-test and a rank test is a documented mistake: the two-stage procedure distorts the type-I error rate of whatever runs second, and the pre-test's own power depends on //n// in the wrong direction, so it waves through non-normality in small samples and rejects trivial non-normality in large ones. Rochon, Gondan & Kieser {[rochon2012_totest]} and Rasch, Kubinger & Moder {[rasch2011_pretesting]} both conclude that pre-testing does not pay off. | 38 of the 1,025 papers (3.7%) report a normality check — Shapiro–Wilk in most of them. Using one to //choose// between a //t//-test and a rank test is a documented mistake: the two-stage procedure distorts the type-I error rate of whatever runs second, and the pre-test's own power depends on //n// in the wrong direction, so it waves through non-normality in small samples and rejects trivial non-normality in large ones. Rochon, Gondan & Kieser {[rochon2012_totest]} and Rasch, Kubinger & Moder {[rasch2011_pretesting]} both conclude that pre-testing does not pay off. |
| |
| <wrap todo> | <WRAP todo> |
| **Decide from the design and the estimand, before the data arrives, and say so.** A defensible sentence looks like: "Because our outcome is a count with a long right tail and we care about the typical site rather than the mean, we pre-specified rank-based tests." An indefensible one is "Shapiro–Wilk was significant, so we used Mann-Whitney." | **Decide from the design and the estimand, before the data arrives, and say so.** A defensible sentence looks like: "Because our outcome is a count with a long right tail and we care about the typical site rather than the mean, we pre-specified rank-based tests." An indefensible one is "Shapiro–Wilk was significant, so we used Mann-Whitney." |
| |
| The best example of doing this well in the corpus splits the choice **per outcome variable** and states the reason for each. Mai et al. {[mai2025_more]} use //"one-way repeated measures ANOVA as the omnibus test since the ad load data follow a normal distribution"// for one outcome and, //"for the hypothesis on predatory ad rates, since the rates do not approximately follow a normal distribution, we use Friedman test as the omnibus test"// for another, in the same paper.((Both sentences are in the paper's PDF and in **neither** of the corpus's stored text renderings of it — see [[provenance:statistics:hypothesis_testing]] for that discrepancy, which is a fact about the corpus and not about the paper.)) | The best example of doing this well in the corpus splits the choice **per outcome variable** and states the reason for each. Mai et al. {[mai2025_more]} use //"one-way repeated measures ANOVA as the omnibus test since the ad load data follow a normal distribution"// for one outcome and, //"for the hypothesis on predatory ad rates, since the rates do not approximately follow a normal distribution, we use Friedman test as the omnibus test"// for another, in the same paper. |
| </wrap> | </WRAP> |
| |
| ==== What Mann-Whitney U actually tests ==== | ==== What Mann-Whitney U actually tests ==== |
| * **A paper counts once**, never once per mention. 1,806 tuples across 1,025 papers. | * **A paper counts once**, never once per mention. 1,806 tuples across 1,025 papers. |
| * **''statistics.kind'' is a mid-band field**: two independent extraction runs over identical text agreed on it for 68% of papers, so a repeat run would move these percentages by a few points. The caveat applies to every share here and not to the folded rankings, which are rankings. That 68% figure was measured on the **previous** corpus run and has not been re-measured. | * **''statistics.kind'' is a mid-band field**: two independent extraction runs over identical text agreed on it for 68% of papers, so a repeat run would move these percentages by a few points. The caveat applies to every share here and not to the folded rankings, which are rankings. That 68% figure was measured on the **previous** corpus run and has not been re-measured. |
| * **Quotes were checked in bulk and by hand.** All 1,806 hypothesis-test quotes against the text the extractor read: 931 exact after whitespace normalisation (51.6%), 540 partial at ≥60% of five-word windows (29.9%), 335 below that (18.5%). **Below-threshold is not "unsupported"** — **five** of the 335 were chased by hand and all five were present, damaged by a dropped citation marker, a reflowed table caption or a two-column splice. The other 330 were not individually checked. Every quote published on this page was located by hand; one (Mai et al.) is in the paper's PDF but in **none** of the corpus's stored text renderings of it. Details on [[provenance:statistics:hypothesis_testing]]. | * **Quotes were checked in bulk and by hand.** All 1,806 hypothesis-test quotes against the text the extractor read: 931 exact after whitespace normalisation (51.6%), 541 partial at ≥60% of five-word windows (30.0%), 334 below that (18.5%). **At least a third of below-threshold is not the extraction's fault.** Re-checked on 2026-09-04 against an independent rendering of the same stored PDF, **118 of the 334 are present there** — the stored text is damaged, not the quote. That leaves 216 below threshold in both renderings, of which five were chased by hand and all five were present, damaged by a dropped citation marker, a reflowed table caption or a two-column splice; the other 211 were not individually checked. Every quote published on this page was located by hand. Details on [[provenance:statistics:hypothesis_testing]], and the corpus-wide version of the same defect — scored by a different, stricter rule, so the percentages are not comparable — on [[literature:corpus]]. |
| * **Sentinels are never answers, and ''statistics'' has none**, which is why the reporting-gap figures are bounds rather than rates. | * **Sentinels are never answers, and ''statistics'' has none**, which is why the reporting-gap figures are bounds rather than rates. |
| * **138 of the 5,859 records are posters** and 251 are ≤4 pages. A poster has its methodology compressed out, so it is systematically likelier to be scored "no test named" — the 5.7% is inflated by an unmeasured amount. | * **138 of the 5,859 records are posters** and 251 are ≤4 pages. A poster has its methodology compressed out, so it is systematically likelier to be scored "no test named" — the 5.7% is inflated by an unmeasured amount. |