| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| provenance:statistics:pvalue_corrections [2026/08/13 07:09] – Add Q23-Q26 (full-text bound on stating k, Fisher tests on the period trend, venue-mix control, preregistration field-vs-hand-count), fill in the generic review log with all 15 findings and their dispositions, state how many of the 62 below-threshold quot karel.kubicek.claude | provenance:statistics:pvalue_corrections [2026/09/21 13:34] (current) – Quote-check refresh 2026-09-21: re-ran quote_check.mjs --statistics multiple-comparison-correction with the pypdf fallback; 62 below threshold -> 40 rescued + 22 below in both, plus a cross-map showing 10 (not 22) unread. Authored by Claude karel.kubicek.claude |
|---|
| | Report script | ''scripts/report_pvalue_corrections.mjs'' | | | Report script | ''scripts/report_pvalue_corrections.mjs'' | |
| | Free-text fold it depends on | ''scripts/mcc_fold.mjs'' — 11 procedure families plus two non-correction buckets, with a self-test | | | Free-text fold it depends on | ''scripts/mcc_fold.mjs'' — 11 procedure families plus two non-correction buckets, with a self-test | |
| | Quote verification | ''scripts/quote_check.mjs --statistics multiple-comparison-correction'' (the ''--statistics'' filter was added for this page) | | | Quote verification | ''scripts/quote_check.mjs %%--statistics%% multiple-comparison-correction'' (the ''%%--statistics%%'' filter was added for this page) | |
| | Stale-number guard | ''scripts/check_page_numbers.mjs'', whole-page, with ''--code'' | | | Stale-number guard | ''scripts/check_page_numbers.mjs'', whole-page, with ''%%--code%%'' | |
| | Non-corpus number ledger | ''out/pvalue-external-facts.txt'' | | | Non-corpus number ledger | ''out/pvalue-external-facts.txt'' | |
| | Runnable code published on the page | ''out/adjust_pvalues.py'' | | | Runnable code published on the page | ''out/adjust_pvalues.py'' | |
| | Written | 2026-08-13, against the corpus as extended on 2026-08-11 (commit ''8a6b843'') | | | Written | 2026-08-13, against the corpus as extended on 2026-08-11 (commit ''8a6b843'') | |
| |
| **Creating, not extending, and not overlapping.** ''statistics:pvalue_corrections'' was a red link promised from [[start]]. The only neighbouring page that existed when this was written is [[statistics:study_preregistration]], which touches multiplicity in one subsection ("For scale: the rest of the inference-hygiene stack is thin too") and cites this page forward. That subsection's figures (15.3% of ''inferential'', 24.5% of hypothesis-test papers) agree with this page's; this page keeps the ''hypothesisTest'' denominator throughout and adds the corrected-vs-tuple distinction the preregistration page did not need. **Nothing was overwritten and no neighbouring figure was contradicted.** | **Creating, not extending, and not overlapping.** ''statistics:pvalue_corrections'' was a red link promised from [[:start]]. The only neighbouring page that existed when this was written is [[statistics:study_preregistration]], which touches multiplicity in one subsection ("For scale: the rest of the inference-hygiene stack is thin too") and cites this page forward. That subsection's figures (15.3% of ''inferential'', 24.5% of hypothesis-test papers) agree with this page's; this page keeps the ''hypothesisTest'' denominator throughout and adds the corrected-vs-tuple distinction the preregistration page did not need. **Nothing was overwritten and no neighbouring figure was contradicted.** |
| |
| **Reachability** needed no work: [[start]] already lists ''[[Statistics:Pvalue corrections|P-value corrections]]'' in its outline, and [[statistics:study_preregistration]] links it twice. | **Reachability** needed no work: [[:start]] already lists ''[[Statistics:Pvalue corrections|P-value corrections]]'' in its outline, and [[statistics:study_preregistration]] links it twice. |
| |
| ===== 2. The number the task started from, and what happened to it ===== | ===== 2. The number the task started from, and what happened to it ===== |
| node scripts/report_pvalue_corrections.mjs # every figure on the page | node scripts/report_pvalue_corrections.mjs # every figure on the page |
| node scripts/report_pvalue_corrections.mjs --wiki # DokuWiki tables | node scripts/report_pvalue_corrections.mjs --wiki # DokuWiki tables |
| node scripts/quote_check.mjs --statistics multiple-comparison-correction --show 62 | node scripts/quote_check.mjs --statistics multiple-comparison-correction --show 400 |
| python3 out/adjust_pvalues.py --demo # the published code's output | python3 out/adjust_pvalues.py --demo # the published code's output |
| cat out/pvalue-report.txt out/adjust_pvalues_demo.txt \ | cat out/pvalue-report.txt out/adjust_pvalues_demo.txt \ |
| | Q21 | Is the partial-pooling alternative present? | 5,869 | 65 papers fit a mixed-effects/multilevel/hierarchical model. **Not** checked paper by paper for whether any frames it as a multiplicity strategy — see §8 | | | Q21 | Is the partial-pooling alternative present? | 5,869 | 65 papers fit a mixed-effects/multilevel/hierarchical model. **Not** checked paper by paper for whether any frames it as a multiplicity strategy — see §8 | |
| | Q22 | Permutation tests as a multiplicity device? | 5,869 | 26 papers, 8 read; every one uses it as the test itself, not as a max-//T// adjustment | | | Q22 | Permutation tests as a multiplicity device? | 5,869 | 26 papers, 8 read; every one uses it as the test itself, not as a max-//T// adjustment | |
| | Q23 | Does the **full text** state a family size near the correction? | the 265 that corrected | a count of comparisons/tests/hypotheses within ±1,500 characters of a procedure mention: **41 (15.5%)**. This is the **upper** bound; Q10's 4.9% of tuples is the lower one | | | Q23 | Does the **full text** state a family size near the correction? | the 265 that corrected | unguarded regex: 41 hits, **mostly χ² notation** (see §5.4). With three guards: 20. **Hand-reading all 20: 15 genuine = 5.7%.** Q10's tuple figure is 4.9%; the two agree to within a point | |
| | Q24 | Is the period trend real? | the 1,025 | Fisher's exact, two-sided, hand-implemented in the report script and matched against ''scipy.stats.fisher_exact'': 2010–2014 vs 2020–2024 //p// = **2.1 × 10⁻⁵**; 2020–2024 vs 2025–2026 //p// = **0.0328**; 2015–2019 vs 2025–2026 //p// = **0.5689** | | | Q24 | Is the period trend real? | the 1,025 | Fisher's exact, two-sided, hand-implemented in the report script and matched against ''scipy.stats.fisher_exact'': 2010–2014 vs 2020–2024 //p// = **2.1 × 10⁻⁵**; 2020–2024 vs 2025–2026 //p// = **0.0328**; 2015–2019 vs 2025–2026 //p// = **0.5689** | |
| | Q25 | Does venue composition explain the 2025–2026 fall? | the 1,025 | **No.** Expected rate from each period's venue mix alone: 18.8% / 22.1% / 25.4% / **25.6%**, against observed 9.3% / 23.3% / 28.7% / **21.0%**. The mix barely moved between the last two buckets | | | Q25 | Does venue composition explain the 2025–2026 fall? | the 1,025 | **No.** Expected rate from each period's venue mix alone: 18.8% / 22.1% / 25.4% / **25.6%**, against observed 9.3% / 23.3% / 28.7% / **21.0%**. The mix barely moved between the last two buckets | |
| ^ Bucket ^ Papers ^ What it actually is ^ | ^ Bucket ^ Papers ^ What it actually is ^ |
| | NEGATIVE | 3 | The paper states it did **not** correct. All three justify it as exploratory: {[pu2016_model]}, {[goetzen2022_ctrl]}, {[naji2025_responsibility]} | | | NEGATIVE | 3 | The paper states it did **not** correct. All three justify it as exploratory: {[pu2016_model]}, {[goetzen2022_ctrl]}, {[naji2025_responsibility]} | |
| | NOT-A-MCC | 2 | Greenhouse–Geisser, a **sphericity** correction to the //F//-test's degrees of freedom. One is {[boettger2025_regional]} | | | NOT-A-MCC | 2 | Greenhouse–Geisser, a **sphericity** correction to the //F//-test's degrees of freedom. One is {[bottger2025_regional]} | |
| |
| A query that counts ''kind == "multiple-comparison-correction"'' and stops there reports 269 where the answer is 265, and reports three papers that explicitly declined as having complied. | A query that counts ''kind == "multiple-comparison-correction"'' and stops there reports 269 where the answer is 265, and reports three papers that explicitly declined as having complied. |
| * ''correction factor of two'' (USENIX 2024), ''p-value threshold adjustment for six repeated tests'' (IEEE S&P 2024), ''conservative α = 0.002'' (PoPETs 2024) — a numerator and no procedure | * ''correction factor of two'' (USENIX 2024), ''p-value threshold adjustment for six repeated tests'' (IEEE S&P 2024), ''conservative α = 0.002'' (PoPETs 2024) — a numerator and no procedure |
| * ''multiple test procedures'' (NDSS 2019) — feature selection, arguably a false positive of the schema, kept in the bucket rather than removed by hand | * ''multiple test procedures'' (NDSS 2019) — feature selection, arguably a false positive of the schema, kept in the bucket rather than removed by hand |
| | |
| | ==== 5.4 The family-size guards and their hand list ==== |
| | |
| | Q23's full-text pass needs its own fold, for the same reason ''mcc_fold.mjs'' does: the naive pattern is dominated by a homograph. ''/\d+\s*tests?/'' matches the degrees-of-freedom digit in **χ² notation**, which is everywhere in this literature. |
| | |
| | ^ Guard ^ Pattern applied to the 40 characters before the digit ^ Matches rejected ^ |
| | | chi-squared notation | ''/[χ𝜒Xx]\s*$/'' | **77** | |
| | | exponent | ''/[eE]\s*[-−+]?\s*$/'' | 3 | |
| | | venue page furniture | ''/(Symposium%%|%%Proceedings%%|%%Association%%|%%Conference%%|%%USENIX%%|%%pp\.)\s*$/i'' | 1 | |
| | |
| | Twenty matches survive the guards and all twenty were read. Four are still artefacts and are a hand list in the report script, keyed on slug so a future corpus prints an unclassified paper rather than silently bucketing it: |
| | |
| | ^ Paper ^ Matched ^ Why it is not a family size ^ |
| | | ''how-does-your-password-measure-up…'' | ''5 tests'' | χ² degrees of freedom as a subscript: //"χ 2 5 tests"// | |
| | | ''measuring-password-guessability-for-an-entire-university'' | ''1 test'' | a test name: //"the G1 test"// | |
| | | ''sunlight-fine-grained-targeting-detection…'' | ''836 hypothesis'' | counts //discoveries//, not the family: //"836 hypothesis below 5%"// | |
| | | ''e-vote-your-conscience…'' | ''2 Pairwise test'' | figure axis labels spliced together: //"3 3 2 Pairwise test"// | |
| | | ''a-multi-region-investigation-of-the-perceptions…'' | ''12 comparisons'' | running header and table cells interleaved. The sentence is //"Dunn's tests (multiple pairwise comparisons)"// with no count; the 12 is page furniture. **Found by reading the source, after the re-review, not by a reviewer** | |
| | |
| | One further hit is genuine as a //paper// and wrong as a //number//, and is annotated rather than rejected: ''deliberate-exposure-to-opposing-views…'' states its family as //"we make 3 comparisons with each population"// and the regex matched a later ''12 comparisons'', the total across populations. The paper does state //k//, so it counts. |
| | |
| | **15 genuine of 265 papers = 5.7%.** The report prints all twenty with their surrounding sentences, so the figure can be discounted line by line rather than taken on trust. |
| |
| ===== 6. Quotes checked against source ===== | ===== 6. Quotes checked against source ===== |
| |
| Bulk pass over every correction quote, with the ''--statistics'' filter added to ''quote_check.mjs'' for this page: | Bulk pass over every correction quote, with the ''%%--statistics%%'' filter added to ''quote_check.mjs'' for this page. **Re-run 2026-09-21** with the PDF fallback — see //Quote-check refresh, 2026-09-21// at the foot of this page: |
| |
| <code> | <code> |
| $ node scripts/quote_check.mjs --statistics multiple-comparison-correction | $ node scripts/quote_check.mjs --statistics multiple-comparison-correction --show 400 |
| 292 quotes checked: 167 exact, 63 partial (>=60% of 5-word windows), 62 below threshold, 0 with no full text on disk. | 292 quotes checked: 167 exact, 63 partial (>=60% of 5-word windows), 40 rescued from the PDF, 22 below threshold in both renderings, 0 with no full text on disk. |
| </code> | </code> |
| |
| 62 below threshold is 21% and higher than this tool usually reports. **The first 40 of the 62 were read** (the tool's ''--show 40'' default output, kept in ''out/pvalue-quotecheck.txt''); the remaining 22 were **not** read. In all 40 read, the cause is not fabrication: correction quotes are short (median well under 20 words), so one dropped citation marker or one extractor ellipsis destroys a large fraction of the five-word windows. ''"Fisher's binomial proportion test ... with a Bonferroni correction."'' scores 0% and is real. **The 22 unread ones are an open item** — nothing on the content page depends on them, since every quote the page uses was checked by hand (below), but they are not evidence for anything either. | **The old figure was 62 below threshold, i.e. 21%, and it was higher than this tool usually reported because the tool was only reading one rendering.** 40 of the 62 are located in an independent ''pypdf'' rendering of the same ''paper.pdf'' and are a defect in the stored text, not in the extraction. **22** are below threshold in both, which is **7.5%** of 292. ''exact'' (167) and ''partial'' (63) did not move. |
| | |
| | **The 40/22 split is not the 40-read/22-unread split, and the coincidence is a trap.** The August run printed only the first 40 of the 62 (''%%--show%% 40'' was its default; that output is kept in ''out/pvalue-quotecheck.txt''), and those 40 were read by hand. Cross-mapping the two runs on paper/label/section: |
| | |
| | * **28** of the 40 read by hand are now //rescued// — the hand read and the PDF fallback agree. |
| | * **12** of the 40 read by hand are still below threshold in both renderings. All 12 were found present when read; the hand read remains the only evidence for them. |
| | * **10** of the 22 now below threshold in both were never printed by the August run and have **not** been read. These are the open item, not 22. |
| | * The remaining 12 //rescued// come from the 22 the August run never printed, so the fallback has closed 12 of those 22 without anyone reading them. |
| | |
| | In all 40 read, the cause was not fabrication: correction quotes are short (median well under 20 words), so one dropped citation marker or one extractor ellipsis destroys a large fraction of the five-word windows. ''"Fisher's binomial proportion test ... with a Bonferroni correction."'' scores 0% and is real. **The 10 unread ones are the open item** — nothing on the content page depends on them, since every quote the page uses was checked by hand (below), but they are not evidence for anything either. |
| |
| **Every quote used on the content page was then located by hand** in ''data/fulltext/<year>/<venue>/<slug>/paper.cols.txt''. All 18 checked out: | **Every quote used on the content page was then located by hand** in ''data/fulltext/<year>/<venue>/<slug>/paper.cols.txt''. All 18 checked out: |
| | {[maass2021_effective]} | "45 significance tests" | found verbatim | | | {[maass2021_effective]} | "45 significance tests" | found verbatim | |
| | {[valapu2025_binary]} | "family-wise error rate" | found verbatim | | | {[valapu2025_binary]} | "family-wise error rate" | found verbatim | |
| | {[nenadic2026_overcoming]} | "Benjamini-Hochberg" (4 occurrences) | found; the "parallel hypothesis tests increases the risk of false positives" sentence and the two-families split both verbatim | | | {[nenadic2026_swiss]} | "Benjamini-Hochberg" (4 occurrences) | found; the "parallel hypothesis tests increases the risk of false positives" sentence and the two-families split both verbatim | |
| | {[liu2024_opted]} | "16 personas" | found verbatim, including "original value multiplied by 16" | | | {[liu2024_opted]} | "16 personas" | found verbatim, including "original value multiplied by 16" | |
| | {[bouhoula2024_automated]} | "Holm-Bonferroni" (2 occurrences) | found verbatim | | | {[bouhoula2024_automated]} | "Holm-Bonferroni" (2 occurrences) | found verbatim | |
| | {[naji2025_responsibility]} | "multiple testing correction" | found verbatim; the extraction dropped the "[78]" citation marker, which is why it scored low in the bulk pass | | | {[naji2025_responsibility]} | "multiple testing correction" | found verbatim; the extraction dropped the "[78]" citation marker, which is why it scored low in the bulk pass | |
| | {[pu2016_model]} | "multiple testing problem" | **the extractor's quote contains an ellipsis and does not match**. Located by searching "exploratory" instead: the full sentence adds "…where multiplicity adjustments are neither mandatory, nor important [7]", and the page quotes the expanded version | | | {[pu2016_model]} | "multiple testing problem" | **the extractor's quote contains an ellipsis and does not match**. Located by searching "exploratory" instead: the full sentence adds "…where multiplicity adjustments are neither mandatory, nor important [7]", and the page quotes the expanded version | |
| | {[boettger2025_regional]} | "Greenhouse-Geisser" | found verbatim | | | {[bottger2025_regional]} | "Greenhouse-Geisser" | found verbatim | |
| | {[bobek2026_community]} | "simultaneous confidence band" | found verbatim | | | {[bobek2026_community]} | "simultaneous confidence band" | found verbatim | |
| | {[despres2024_best]} | "conservative" | found verbatim, including "(based on a maximum of 25 hypotheses tested per outcome)" | | | {[despres2024_best]} | "conservative" | found verbatim, including "(based on a maximum of 25 hypotheses tested per outcome)" | |
| | {[ho2025_efficacy]} | "multiple comparison" | found verbatim, including "N = 2,693 per group" | | | {[ho2025_efficacy]} | "multiple comparison" | found verbatim, including "N = 2,693 per group" | |
| |
| <wrap todo> | <WRAP todo> |
| One thing found while checking and **not** put on the page. {[kablo2023_privacy]} writes //"we adjusted the standard threshold of 0.05 to a new threshold of 0.00047 (0.005 / 106)"//. 0.05/106 = 0.00047, so the printed threshold is right and the parenthetical divisor has a typo, or is a PDF-extraction artefact of "0.05". It is a one-character slip in an otherwise exemplary paper and calling it out on the page would be a gotcha, not a lesson. Recorded here instead. | One thing found while checking and **not** put on the page. {[kablo2023_privacy]} writes //"we adjusted the standard threshold of 0.05 to a new threshold of 0.00047 (0.005 / 106)"//. 0.05/106 = 0.00047, so the printed threshold is right and the parenthetical divisor has a typo, or is a PDF-extraction artefact of "0.05". It is a one-character slip in an otherwise exemplary paper and calling it out on the page would be a gotcha, not a lesson. Recorded here instead. |
| </wrap> | </WRAP> |
| |
| ===== 7. External sources, and how each was verified ===== | ===== 7. External sources, and how each was verified ===== |
| ==== 7.4 The published code was run, and its Holm implementation cross-checked ==== | ==== 7.4 The published code was run, and its Holm implementation cross-checked ==== |
| |
| ''numpy'' 2.4.6, ''scipy'' 1.17.1, ''statsmodels'' 0.14.6, installed into the sandbox with ''pip install --break-system-packages''. The page publishes the unedited ''--demo'' output. | ''numpy'' 2.4.6, ''scipy'' 1.17.1, ''statsmodels'' 0.14.6, installed into the sandbox with ''pip install %%--break-system-packages%%''. The page publishes the unedited ''%%--demo%%'' output. |
| |
| <code> | <code> |
| **36 of the 38 recruited human participants.** The two that did not are ''IEEE-SP/2020/are-anonymity-seekers-just-like-everybody-else'' and ''IEEE-SP/2024/a-picture-is-worth-500-labels''. **Exactly one of the 38 ran a crawl**: ''PETS/2024/what-does-it-mean-to-be-creepy'', whose ''studyTypes'' includes ''automated-web-crawl'' and which also recruited participants. So the population that sizes a study with a correction in the calculation is, with one partial exception, a user-study population. | **36 of the 38 recruited human participants.** The two that did not are ''IEEE-SP/2020/are-anonymity-seekers-just-like-everybody-else'' and ''IEEE-SP/2024/a-picture-is-worth-500-labels''. **Exactly one of the 38 ran a crawl**: ''PETS/2024/what-does-it-mean-to-be-creepy'', whose ''studyTypes'' includes ''automated-web-crawl'' and which also recruited participants. So the population that sizes a study with a correction in the calculation is, with one partial exception, a user-study population. |
| |
| <wrap todo> | <WRAP todo> |
| An earlier version of this section said "37 of the 38" and named ''IMC/2022/what-factors-affect-targeting-and-bids-in-online-advertising'' as the exception and as a crawl. **All three parts of that were wrong** — the count is 36, that paper has ''participants'' tuples, and its ''studyTypes'' are ''['user-study', 'manual-audit']'' with ''crawlConfig === null'', so it is not a crawl by the definition in §3. It was caught by the figures-vs-script reviewer and is recorded here rather than silently corrected, because it is the clearest example on this page of a claim written from memory instead of from a query. | An earlier version of this section said "37 of the 38" and named ''IMC/2022/what-factors-affect-targeting-and-bids-in-online-advertising'' as the exception and as a crawl. **All three parts of that were wrong** — the count is 36, that paper has ''participants'' tuples, and its ''studyTypes'' are ''['user-study', 'manual-audit']'' with ''crawlConfig === null'', so it is not a crawl by the definition in §3. It was caught by the figures-vs-script reviewer and is recorded here rather than silently corrected, because it is the clearest example on this page of a claim written from memory instead of from a query. |
| </wrap> | </WRAP> |
| |
| ===== 10. Judgement calls ===== | ===== 10. Judgement calls ===== |
| Four reviewers, all told explicitly that the author's context may not be exhaustive, and all handed the page text, the report script, its unedited output, and these notes. The three focused passes ran in parallel first; the generic pass ran after their findings were applied. | Four reviewers, all told explicitly that the author's context may not be exhaustive, and all handed the page text, the report script, its unedited output, and these notes. The three focused passes ran in parallel first; the generic pass ran after their findings were applied. |
| |
| ==== 11.1 ''sonnet'' — figures against the script ==== | ==== 11.1 sonnet — figures against the script ==== |
| |
| Re-ran all three scripts (byte-identical to the committed output), re-derived the 1,025 / 269 / 141 anchor counts from ''extractions.jsonl'' in an independent Python re-implementation of the fold, and cross-checked ''holm()''/''bonferroni()'' against ''statsmodels'' on 20 random families plus edge cases (all-ones, single value, ties): max abs diff 0.0. | Re-ran all three scripts (byte-identical to the committed output), re-derived the 1,025 / 269 / 141 anchor counts from ''extractions.jsonl'' in an independent Python re-implementation of the fold, and cross-checked ''holm()''/''bonferroni()'' against ''statsmodels'' on 20 random families plus edge cases (all-ones, single value, ties): max abs diff 0.0. |
| One correction to the reviewer, recorded for the record: it reported that ''IMC/2022/what-factors-affect-targeting-and-bids'' //"has 286 participants (recruited via Prolific)"//. The ''participants'' array on that record has **2 tuples**, not 286 people; the reviewer conflated a tuple count with a headcount. Its substantive point — that the paper has participants and is not a crawl — is correct and was the basis for the fix. | One correction to the reviewer, recorded for the record: it reported that ''IMC/2022/what-factors-affect-targeting-and-bids'' //"has 286 participants (recruited via Prolific)"//. The ''participants'' array on that record has **2 tuples**, not 286 people; the reviewer conflated a tuple count with a headcount. Its substantive point — that the paper has participants and is not a crawl — is correct and was the basis for the fix. |
| |
| ==== 11.2 ''sonnet'' — citations and quotes ==== | ==== 11.2 sonnet — citations and quotes ==== |
| |
| Verified all 27 citekeys resolve with no collisions, all 17 corpus entries against ''data/corpus2/.meta'' and the papers' own front matter, all 8 external DOIs through ''doi.org'', and all 18 quotations against ''paper.cols.txt''. | Verified all 27 citekeys resolve with no collisions, all 17 corpus entries against ''data/corpus2/.meta'' and the papers' own front matter, all 8 external DOIs through ''doi.org'', and all 18 quotations against ''paper.cols.txt''. |
| | Kablo & Cabarcos's //"(0.005 / 106)"// is confirmed as an error in the paper itself, not an extraction artefact | **no change needed** | Stays off the content page for the reason in §6 | | | Kablo & Cabarcos's //"(0.005 / 106)"// is confirmed as an error in the paper itself, not an extraction artefact | **no change needed** | Stays off the content page for the reason in §6 | |
| |
| ==== 11.3 ''sonnet'' — external currency, as of 2026-08-13 ==== | ==== 11.3 sonnet — external currency, as of 2026-08-13 ==== |
| |
| ^ Finding ^ Accepted? ^ What was done ^ | ^ Finding ^ Accepted? ^ What was done ^ |
| | **CONSORT 2025** {[hopewell2025_consort]} added a multiplicity-reporting requirement, published 2025 and therefore newer than anything else on the page | **yes** | Re-verified independently before writing: DOI ''10.1136/bmj-2024-081124'' resolves; the sentence //"Any methods used to mitigate or account for multiplicity should be described…"// was read verbatim from the explanation-and-elaboration document at ''pmc.ncbi.nlm.nih.gov/articles/PMC11995452''. Added to the currency box and to //What to Report//, framed as a reporting standard in an adjacent field, not as a change of procedure | | | **CONSORT 2025** {[hopewell2025_consort]} added a multiplicity-reporting requirement, published 2025 and therefore newer than anything else on the page | **yes, then corrected in the re-review** | The DOI and the quote were verified before writing (''10.1136/bmj-2024-081124'' resolves; the sentence was read verbatim from the E&E document at ''pmc.ncbi.nlm.nih.gov/articles/PMC11995452''). **The novelty claim was not verified and was wrong** — see §11.5 | |
| | ''multipletests'' accepts ''bonferroni, sidak, holm-sidak, holm, simes-hochberg, hommel, fdr_bh, fdr_by, fdr_tsbh, fdr_tsbky'' as of 0.14.6 | **yes** | The exact method strings are now on the page, which is more useful than "does all four plus Šidák" | | | ''multipletests'' accepts ''bonferroni, sidak, holm-sidak, holm, simes-hochberg, hommel, fdr_bh, fdr_by, fdr_tsbh, fdr_tsbky'' as of 0.14.6 | **yes** | The exact method strings are now on the page, which is more useful than "does all four plus Šidák" | |
| | e-values / e-BH and selective inference are active research but have no adopted standing and zero presence in this corpus, so //"nothing has been overturned"// holds | **yes, sharpened** | The currency box now says so explicitly instead of leaving it implicit. Rejecting e-BH as a recommendation is recorded in §7.5 | | | e-values / e-BH and selective inference are active research but have no adopted standing and zero presence in this corpus, so //"nothing has been overturned"// holds | **yes, sharpened** | The currency box now says so explicitly instead of leaving it implicit. Rejecting e-BH as a recommendation is recorded in §7.5 | |
| | All 17 corpus DOIs and all 3 USENIX URLs return HTTP 200; SciPy 1.18.0 and statsmodels 0.14.6 are current; ''false_discovery_control'' is unchanged between 1.17.1 and 1.18.0 | **no change needed** | — | | | All 17 corpus DOIs and all 3 USENIX URLs return HTTP 200; SciPy 1.18.0 and statsmodels 0.14.6 are current; ''false_discovery_control'' is unchanged between 1.17.1 and 1.18.0 | **no change needed** | — | |
| |
| ==== 11.4 ''fable'' — generic ==== | ==== 11.4 fable — generic ==== |
| |
| No checklist; run after the three focused passes were applied. It returned fifteen findings and **fourteen were accepted**, which is the highest accept rate of the four reviewers and the reason this slot is worth keeping. | No checklist; run after the three focused passes were applied. It returned fifteen findings and **fourteen were accepted**, which is the highest accept rate of the four reviewers and the reason this slot is worth keeping. |
| | //"The only mechanism that addresses it"// is contradicted twice on the same page (the "cheap version" paragraph, and the multiverse open question) | **yes** | Now "the standard mechanism", with the alternatives named | | | //"The only mechanism that addresses it"// is contradicted twice on the same page (the "cheap version" paragraph, and the multiverse open question) | **yes** | Now "the standard mechanism", with the alternatives named | |
| | Provenance §11.4 was published as an empty placeholder while §12 counted a ''fable'' pass among the reviewers | **yes** | This section. The placeholder should not have been saved; recording a review before it exists is the same defect as publishing a figure before checking it | | | Provenance §11.4 was published as an empty placeholder while §12 counted a ''fable'' pass among the reviewers | **yes** | This section. The placeholder should not have been saved; recording a review before it exists is the same defect as publishing a figure before checking it | |
| | Provenance §6 says //"Reading them, the cause is not fabrication"// about the 62 below-threshold quotes without saying how many were read, on the one bucket where fabrication would hide | **yes** | §6 now states 40 read, 22 not, and says the 22 are evidence for nothing | | | Provenance §6 says //"Reading them, the cause is not fabrication"// about the 62 below-threshold quotes without saying how many were read, on the one bucket where fabrication would hide | **yes** | §6 stated 40 read, 22 not, and said the 22 are evidence for nothing. **Superseded 2026-09-21:** the 62 is now 40 //rescued// + 22 below threshold in both, and §6 gives the cross-map — 10 quotes, not 22, are unread | |
| | Provenance arithmetic drift: the content page said "nine DOIs" where §7.1 listed eight, and §12 said "26 entries" next to "27 keys" | **yes** | CONSORT's DOI added to §7.1; §12 corrected to 27 entries and 28 keys | | | Provenance arithmetic drift: the content page said "nine DOIs" where §7.1 listed eight, and §12 said "26 entries" next to "27 keys" | **yes** | CONSORT's DOI added to §7.1; §12 corrected to 27 entries and 28 keys | |
| | The headline box's three bullets omit the //both// cell (34.7%), which is the one that sharpens the story | **yes** | Fourth bullet added, and it is now the point the box makes: the two participant rows are the two high ones | | | The headline box's three bullets omit the //both// cell (34.7%), which is the one that sharpens the story | **yes** | Fourth bullet added, and it is now the point the box makes: the two participant rows are the two high ones | |
| |
| **Rejected: none outright.** The one partial is the ''gelman2014_statistical'' page range from §11.2, left at ''460'' because no primary source for an end page could be found. | **Rejected: none outright.** The one partial is the ''gelman2014_statistical'' page range from §11.2, left at ''460'' because no primary source for an end page could be found. |
| | |
| | ==== 11.5 Re-review, both sonnet passes, after the fixes ==== |
| | |
| | The two focused reviewers whose findings were acted on were re-run against the edited pages, as the workflow requires. Both found real defects in the **new** material, which is the argument for re-running them. |
| | |
| | ^ Finding ^ Accepted? ^ What was done ^ |
| | | **Q23's regex is dominated by false positives.** Of the 41 unguarded hits, 24 are χ²-test notation — the digit in ''χ2 tests'' sits immediately before the word — plus exponents (''1.82e−06 Pairwise test'') and USENIX page furniture (''…Symposium 3589 test''). 15.5% is inflated roughly 2.4× | **yes** | The worst defect in this run and it was introduced //by a fix// for a previous finding. Three guards added, each inspecting the 40 characters before the digit, with **every rejected match counted by guard** (77 χ², 3 exponent, 1 page furniture). That leaves 20, and all 20 were then **read**: 15 genuine, 5 artefacts named with their reason in a hand list keyed on slug. The figure is now **15 of 265 = 5.7%**, and the page states it alongside the tuple figure's 4.9% instead of presenting a range. Every accepted match is printed with its sentence in the report output | |
| | | The complement, //"84.5% of corrections come without a nearby k"//, inherits the same error | **yes** | Now 94.3%, and the Open Questions bullet no longer implies CONSORT asks for //k// | |
| | | The page's //"upper bound"// framing was wrong in both directions: the window misses a count stated in a distant table caption //and// can admit an unrelated number | **yes** | The script and the methodology bullet now say explicitly that **neither pass is a bound** | |
| | | **CONSORT 2025 did not "add multiplicity".** CONSORT 2010's elaboration already discouraged multiple primary outcomes //"because of the problems of interpretation associated with multiplicity of analyses"//, item 7b already required stating adjustment for interim analyses, and item 20 already named //"multiplicity of analyses"// in Limitations. What 2025 adds is the explicit obligation to report that //no// method was used | **yes** | Rewritten to claim only the negative obligation, with the CONSORT 2010 history stated. **This is the clearest overclaim of the run**: the DOI and the quote were checked, and the word "added" was not | |
| | | The quoted sentence is in the **explanation and elaboration** (''…081124''), not in the statement/checklist (''…081123''), which contains no occurrence of "multiplicity" in its checklist text | **yes** | The citation was already to the E&E, which is correct; a footnote now says so and says the statement does not carry the sentence | |
| | | //"Nenadic et al. correct separately to the two families because they make two claims"// puts an inference in the paper's mouth: the two families are two contrast definitions (CH vs. EU, CH & EU vs. EU) of **one** research question, and the paper's stated reason is that it estimates separate models | **yes** | Both mentions rewritten to the paper's own reason and its own description of the split. The sentence is now more useful as well as more accurate | |
| | | **''nenadic2026_overcoming'' duplicates ''nenadic2026_swiss''** — same paper, same DOI, two keys, the second already in the bibliography from [[design:sampling]] | **yes** | The duplicate was removed from ''literature:bibliography'' and both pages now cite ''nenadic2026_swiss''. ''bibgen.mjs'' derives a key from title words and does not check the live bibliography for the same DOI; **checking the DOI, not just the key, is the lesson** | |
| | | §11.3 presented the CONSORT novelty claim as verified when only the DOI and the quote were checked | **yes** | §11.3 amended to say which parts were verified and to point here | |
| | | The Bonferroni surviving-uses claims were checked against the literature: Holm-compatible simultaneous confidence regions were an open problem until Strassburger & Bretz (2008) and Guilbaud (2008), and are noted there as often non-informative, while Bonferroni gives intervals at 1 − α/k by a union bound; and Bonferroni's threshold depends only on //k// where Holm's depends on each //p//-value's rank | **no change needed** | Both claims stand as written. This was the fix most likely to have overreached and it did not | |
| | | The Fisher implementation, the venue-mix control, the preregistration cross-check, the family-arithmetic, the headline box, the demo cross-references and every previously-checked figure reproduce exactly, including on 30 random 2×2 tables against ''scipy.stats.fisher_exact'' (max abs diff ≈ 5 × 10⁻¹⁴) | **no change needed** | — | |
| |
| ===== 12. The run itself ===== | ===== 12. The run itself ===== |
| | Reviewers | three ''sonnet'' passes (figures-vs-script, citations-and-quotes, external currency) and one ''fable'' generic pass | | | Reviewers | three ''sonnet'' passes (figures-vs-script, citations-and-quotes, external currency) and one ''fable'' generic pass | |
| | New scripts | ''scripts/mcc_fold.mjs'', ''scripts/report_pvalue_corrections.mjs'' (which carries its own two-sided Fisher exact test, checked against ''scipy.stats.fisher_exact'' to four significant figures) | | | New scripts | ''scripts/mcc_fold.mjs'', ''scripts/report_pvalue_corrections.mjs'' (which carries its own two-sided Fisher exact test, checked against ''scipy.stats.fisher_exact'' to four significant figures) | |
| | Modified scripts | ''scripts/quote_check.mjs'' — added a ''--statistics <kind>'' filter, which did not exist | | | Modified scripts | ''scripts/quote_check.mjs'' — added a ''%%--statistics%% <kind>'' filter, which did not exist | |
| | New artifacts | ''out/pvalue-report.txt'', ''out/pvalue-quotecheck.txt'', ''out/pvalue-external-facts.txt'', ''out/adjust_pvalues.py'', ''out/adjust_pvalues_demo.txt'', ''out/bib_additions_pvalue.bib'' | | | New artifacts | ''out/pvalue-report.txt'', ''out/pvalue-quotecheck.txt'', ''out/pvalue-external-facts.txt'', ''out/adjust_pvalues.py'', ''out/adjust_pvalues_demo.txt'', ''out/bib_additions_pvalue.bib'' | |
| | Bibliography | **27 entries appended**: 17 corpus papers via ''bibgen.mjs'', 9 external statistics references, and ''hopewell2025_consort'' added after the currency review. No duplicate keys; all 28 keys used on the page resolve and the rendered page reports zero bibtex warnings | | | Bibliography | **27 entries appended**: 17 corpus papers via ''bibgen.mjs'', 9 external statistics references, and ''hopewell2025_consort'' added after the currency review. No duplicate keys; all 28 keys used on the page resolve and the rendered page reports zero bibtex warnings | |
| * [[literature:corpus]] — the corpus-level provenance: how the 5,859 papers were selected, extracted and validated. | * [[literature:corpus]] — the corpus-level provenance: how the 5,859 papers were selected, extracted and validated. |
| * [[provenance:statistics:study_preregistration]] — the sibling ''statistics:'' provenance page, and the source of this page's format. | * [[provenance:statistics:study_preregistration]] — the sibling ''statistics:'' provenance page, and the source of this page's format. |
| | |
| | ===== Amendment, 2026-09-04: citekey consolidation ===== |
| | |
| | * ''boettger2025_regional'' was one of two keys for the same paper in [[:literature:bibliography]]. The wiki-wide consolidation of 2026-09-04 (drain item ''dedup-regional-filter-lists-bibkey'') kept ''bottger2025_regional'' and deleted the other entry. |
| | * 1 citation marker on [[:statistics:pvalue_corrections]] and 2 citation markers on this page were repointed to the kept key. No prose on either page changed, and no figure moved. Statements above that name the deleted key describe the state when they were written. Full query log and the invariants checked before saving: [[:provenance:literature:bibliography]]. |
| |
| ====== References ====== | ====== References ====== |
| |
| <bibtex bibliography></bibtex> | <bibtex bibliography></bibtex> |
| | |
| | ===== Markup sweep, 2026-09-17 ===== |
| | |
| | Mechanical rendering repair only: a fresh live raw/XHTML export of 188 pages was checked with ''check_wrap.mjs'' and ''check_typography.mjs''. Affected plugin tags, CLI flags and heading markup were repaired; no figures or substantive prose were changed. The resulting source and rendered DOM were re-checked after saving. |
| | |
| | ===== Link-hygiene sweep, 2026-09-17 ===== |
| | |
| | The fresh rendered-DOM sweep found this page's two bare ''start'' links resolved to ''provenance:statistics:start''. They are now root-anchored; the post-save DOM was re-checked for red links. No figures or citations changed. |
| | |
| | ===== Quote-check refresh, 2026-09-21 ===== |
| | |
| | The 2026-09-04 ''cols''-vs-PDF audit on [[:provenance:literature:corpus]] showed that 73.1% of evidence quotes that cannot be located in ''paper.cols.txt'' **are** present in an independent ''pypdf'' rendering of the same ''paper.pdf''. ''scripts/quote_check.mjs'' was patched the same day to re-check everything below threshold against that second rendering and report a fourth verdict, **RESCUED**. §6's figure predates the patch and overstated this page's quote-failure rate by nearly a factor of three. Re-run, unedited first line: |
| | |
| | <code> |
| | $ node scripts/quote_check.mjs --statistics multiple-comparison-correction --show 400 |
| | 292 quotes checked: 167 exact, 63 partial (>=60% of 5-word windows), 40 rescued from the PDF, 22 below threshold in both renderings, 0 with no full text on disk. |
| | </code> |
| | |
| | ^ Figure ^ Was ^ Is ^ Why ^ |
| | | quotes checked | 292 | 292 | population unchanged — the corpus has not moved | |
| | | exact | 167 | 167 | unchanged | |
| | | partial (≥60% of 5-word windows) | 63 | 63 | unchanged | |
| | | rescued from the PDF | — | **40** | new verdict; these were inside the old 62 | |
| | | below threshold | **62** | **22** (in both renderings) | 62 = 40 + 22 exactly; nothing else moved | |
| | | below-threshold rate | 21% | **7.5%** | 22 of 292 | |
| | | unread below-threshold quotes | 22 | **10** | cross-mapped against ''out/pvalue-quotecheck.txt'' — see §6 | |
| | |
| | **What this does and does not say.** It does not say 40 extractions were wrong and are now right — the quotes were always in the papers, and 28 of these 40 had already been read by hand and found present. It says the //stored text// could not locate them and a second rendering of the same PDF can, so counting them as quote failures measured ''decolumn.mjs'', not the extraction. |
| | |
| | **The one number that needed more than the script.** ''62 = 40 + 22'' and ''62 = 40 read + 22 unread'' are two different partitions of the same 62, and reading the new output alone would let a future editor merge them. §6 now carries the cross-map: 28 of the 40 hand-read are //rescued//, 12 are still below in both, and 10 of the 22 now-below have never been read. |
| | |
| | **Scope of this edit.** §6, the §3 command line (''%%--show 62%%'' → ''%%--show 400%%'', since 62 no longer sizes the output), the review-log row that quoted the old disposition, and the content page's quote bullet. ''report_pvalue_corrections.mjs'' and ''mcc_fold.mjs'' were **not** re-run in this pass; no fold, rate, Fisher //p//-value or citation was touched, and every other figure on [[:statistics:pvalue_corrections]] and this page stands as published. |
| | |
| | ^ Item ^ Value ^ |
| | | Date | 2026-09-21, unsupervised | |
| | | Command | ''%%node scripts/quote_check.mjs --statistics multiple-comparison-correction --show 400%%'' | |
| | | Artifacts | ''out/qc0921/stat_mcc.txt'' (full run, 40 RESCUED rows and 22 below-threshold rows listed); the August output ''out/pvalue-quotecheck.txt'' was kept and used for the cross-map | |
| | | Script changes | none — ''quote_check.mjs'' was already patched on 2026-09-04 | |
| | | Reviewers | one ''sonnet'' figures-vs-script pass over this page and [[:statistics:pvalue_corrections]] | |
| | | Pages saved | this page, [[:statistics:pvalue_corrections]] | |
| |