User Tools

Site Tools


statistics:hypothesis_testing

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
statistics:hypothesis_testing [2026/09/04 08:12] – Remove the retracted footnote about Mai et al.'s quotes being absent from every stored rendering (they are in paper.cols.txt), and refresh the quote-check figures: 931/541/118 rescued/216 below in both. Authored by Claude karel.kubicek.claudestatistics:hypothesis_testing [2026/09/04 08:29] (current) – Review pass: say 'at least a third' rather than 'mostly', and note that the corpus-wide figure is scored by a different, stricter rule. Authored by Claude karel.kubicek.claude
Line 605: Line 605:
   * **A paper counts once**, never once per mention. 1,806 tuples across 1,025 papers.   * **A paper counts once**, never once per mention. 1,806 tuples across 1,025 papers.
   * **''statistics.kind'' is a mid-band field**: two independent extraction runs over identical text agreed on it for 68% of papers, so a repeat run would move these percentages by a few points. The caveat applies to every share here and not to the folded rankings, which are rankings. That 68% figure was measured on the **previous** corpus run and has not been re-measured.   * **''statistics.kind'' is a mid-band field**: two independent extraction runs over identical text agreed on it for 68% of papers, so a repeat run would move these percentages by a few points. The caveat applies to every share here and not to the folded rankings, which are rankings. That 68% figure was measured on the **previous** corpus run and has not been re-measured.
-  * **Quotes were checked in bulk and by hand.** All 1,806 hypothesis-test quotes against the text the extractor read: 931 exact after whitespace normalisation (51.6%), 541 partial at ≥60% of five-word windows (30.0%), 334 below that (18.5%). **Below-threshold mostly is not the extraction's fault.** Re-checked on 2026-09-04 against an independent rendering of the same stored PDF, **118 of the 334 are present there** — the stored text is damaged, not the quote. That leaves 216 below threshold in both renderings, of which five were chased by hand and all five were present, damaged by a dropped citation marker, a reflowed table caption or a two-column splice; the other 211 were not individually checked. Every quote published on this page was located by hand. Details on [[provenance:statistics:hypothesis_testing]], and the corpus-wide version of the same measurement on [[literature:corpus]].+  * **Quotes were checked in bulk and by hand.** All 1,806 hypothesis-test quotes against the text the extractor read: 931 exact after whitespace normalisation (51.6%), 541 partial at ≥60% of five-word windows (30.0%), 334 below that (18.5%). **At least a third of below-threshold is not the extraction's fault.** Re-checked on 2026-09-04 against an independent rendering of the same stored PDF, **118 of the 334 are present there** — the stored text is damaged, not the quote. That leaves 216 below threshold in both renderings, of which five were chased by hand and all five were present, damaged by a dropped citation marker, a reflowed table caption or a two-column splice; the other 211 were not individually checked. Every quote published on this page was located by hand. Details on [[provenance:statistics:hypothesis_testing]], and the corpus-wide version of the same defect — scored by a different, stricter rule, so the percentages are not comparable — on [[literature:corpus]].
   * **Sentinels are never answers, and ''statistics'' has none**, which is why the reporting-gap figures are bounds rather than rates.   * **Sentinels are never answers, and ''statistics'' has none**, which is why the reporting-gap figures are bounds rather than rates.
   * **138 of the 5,859 records are posters** and 251 are ≤4 pages. A poster has its methodology compressed out, so it is systematically likelier to be scored "no test named" — the 5.7% is inflated by an unmeasured amount.   * **138 of the 5,859 records are posters** and 251 are ≤4 pages. A poster has its methodology compressed out, so it is systematically likelier to be scored "no test named" — the 5.7% is inflated by an unmeasured amount.
statistics/hypothesis_testing.1788509544.txt.gz · Last modified: by karel.kubicek.claude