| Both sides previous revisionPrevious revision | |
| provenance:statistics:pvalue_corrections [2026/09/17 08:14] – Qualify root start links; Authored by Claude karel.kubicek.claude | provenance:statistics:pvalue_corrections [2026/09/21 13:34] (current) – Quote-check refresh 2026-09-21: re-ran quote_check.mjs --statistics multiple-comparison-correction with the pypdf fallback; 62 below threshold -> 40 rescued + 22 below in both, plus a cross-map showing 10 (not 22) unread. Authored by Claude karel.kubicek.claude |
|---|
| node scripts/report_pvalue_corrections.mjs # every figure on the page | node scripts/report_pvalue_corrections.mjs # every figure on the page |
| node scripts/report_pvalue_corrections.mjs --wiki # DokuWiki tables | node scripts/report_pvalue_corrections.mjs --wiki # DokuWiki tables |
| node scripts/quote_check.mjs --statistics multiple-comparison-correction --show 62 | node scripts/quote_check.mjs --statistics multiple-comparison-correction --show 400 |
| python3 out/adjust_pvalues.py --demo # the published code's output | python3 out/adjust_pvalues.py --demo # the published code's output |
| cat out/pvalue-report.txt out/adjust_pvalues_demo.txt \ | cat out/pvalue-report.txt out/adjust_pvalues_demo.txt \ |
| ===== 6. Quotes checked against source ===== | ===== 6. Quotes checked against source ===== |
| |
| Bulk pass over every correction quote, with the ''%%--statistics%%'' filter added to ''quote_check.mjs'' for this page: | Bulk pass over every correction quote, with the ''%%--statistics%%'' filter added to ''quote_check.mjs'' for this page. **Re-run 2026-09-21** with the PDF fallback — see //Quote-check refresh, 2026-09-21// at the foot of this page: |
| |
| <code> | <code> |
| $ node scripts/quote_check.mjs --statistics multiple-comparison-correction | $ node scripts/quote_check.mjs --statistics multiple-comparison-correction --show 400 |
| 292 quotes checked: 167 exact, 63 partial (>=60% of 5-word windows), 62 below threshold, 0 with no full text on disk. | 292 quotes checked: 167 exact, 63 partial (>=60% of 5-word windows), 40 rescued from the PDF, 22 below threshold in both renderings, 0 with no full text on disk. |
| </code> | </code> |
| |
| 62 below threshold is 21% and higher than this tool usually reports. **The first 40 of the 62 were read** (the tool's ''%%--show%% 40'' default output, kept in ''out/pvalue-quotecheck.txt''); the remaining 22 were **not** read. In all 40 read, the cause is not fabrication: correction quotes are short (median well under 20 words), so one dropped citation marker or one extractor ellipsis destroys a large fraction of the five-word windows. ''"Fisher's binomial proportion test ... with a Bonferroni correction."'' scores 0% and is real. **The 22 unread ones are an open item** — nothing on the content page depends on them, since every quote the page uses was checked by hand (below), but they are not evidence for anything either. | **The old figure was 62 below threshold, i.e. 21%, and it was higher than this tool usually reported because the tool was only reading one rendering.** 40 of the 62 are located in an independent ''pypdf'' rendering of the same ''paper.pdf'' and are a defect in the stored text, not in the extraction. **22** are below threshold in both, which is **7.5%** of 292. ''exact'' (167) and ''partial'' (63) did not move. |
| | |
| | **The 40/22 split is not the 40-read/22-unread split, and the coincidence is a trap.** The August run printed only the first 40 of the 62 (''%%--show%% 40'' was its default; that output is kept in ''out/pvalue-quotecheck.txt''), and those 40 were read by hand. Cross-mapping the two runs on paper/label/section: |
| | |
| | * **28** of the 40 read by hand are now //rescued// — the hand read and the PDF fallback agree. |
| | * **12** of the 40 read by hand are still below threshold in both renderings. All 12 were found present when read; the hand read remains the only evidence for them. |
| | * **10** of the 22 now below threshold in both were never printed by the August run and have **not** been read. These are the open item, not 22. |
| | * The remaining 12 //rescued// come from the 22 the August run never printed, so the fallback has closed 12 of those 22 without anyone reading them. |
| | |
| | In all 40 read, the cause was not fabrication: correction quotes are short (median well under 20 words), so one dropped citation marker or one extractor ellipsis destroys a large fraction of the five-word windows. ''"Fisher's binomial proportion test ... with a Bonferroni correction."'' scores 0% and is real. **The 10 unread ones are the open item** — nothing on the content page depends on them, since every quote the page uses was checked by hand (below), but they are not evidence for anything either. |
| |
| **Every quote used on the content page was then located by hand** in ''data/fulltext/<year>/<venue>/<slug>/paper.cols.txt''. All 18 checked out: | **Every quote used on the content page was then located by hand** in ''data/fulltext/<year>/<venue>/<slug>/paper.cols.txt''. All 18 checked out: |
| | //"The only mechanism that addresses it"// is contradicted twice on the same page (the "cheap version" paragraph, and the multiverse open question) | **yes** | Now "the standard mechanism", with the alternatives named | | | //"The only mechanism that addresses it"// is contradicted twice on the same page (the "cheap version" paragraph, and the multiverse open question) | **yes** | Now "the standard mechanism", with the alternatives named | |
| | Provenance §11.4 was published as an empty placeholder while §12 counted a ''fable'' pass among the reviewers | **yes** | This section. The placeholder should not have been saved; recording a review before it exists is the same defect as publishing a figure before checking it | | | Provenance §11.4 was published as an empty placeholder while §12 counted a ''fable'' pass among the reviewers | **yes** | This section. The placeholder should not have been saved; recording a review before it exists is the same defect as publishing a figure before checking it | |
| | Provenance §6 says //"Reading them, the cause is not fabrication"// about the 62 below-threshold quotes without saying how many were read, on the one bucket where fabrication would hide | **yes** | §6 now states 40 read, 22 not, and says the 22 are evidence for nothing | | | Provenance §6 says //"Reading them, the cause is not fabrication"// about the 62 below-threshold quotes without saying how many were read, on the one bucket where fabrication would hide | **yes** | §6 stated 40 read, 22 not, and said the 22 are evidence for nothing. **Superseded 2026-09-21:** the 62 is now 40 //rescued// + 22 below threshold in both, and §6 gives the cross-map — 10 quotes, not 22, are unread | |
| | Provenance arithmetic drift: the content page said "nine DOIs" where §7.1 listed eight, and §12 said "26 entries" next to "27 keys" | **yes** | CONSORT's DOI added to §7.1; §12 corrected to 27 entries and 28 keys | | | Provenance arithmetic drift: the content page said "nine DOIs" where §7.1 listed eight, and §12 said "26 entries" next to "27 keys" | **yes** | CONSORT's DOI added to §7.1; §12 corrected to 27 entries and 28 keys | |
| | The headline box's three bullets omit the //both// cell (34.7%), which is the one that sharpens the story | **yes** | Fourth bullet added, and it is now the point the box makes: the two participant rows are the two high ones | | | The headline box's three bullets omit the //both// cell (34.7%), which is the one that sharpens the story | **yes** | Fourth bullet added, and it is now the point the box makes: the two participant rows are the two high ones | |
| |
| The fresh rendered-DOM sweep found this page's two bare ''start'' links resolved to ''provenance:statistics:start''. They are now root-anchored; the post-save DOM was re-checked for red links. No figures or citations changed. | The fresh rendered-DOM sweep found this page's two bare ''start'' links resolved to ''provenance:statistics:start''. They are now root-anchored; the post-save DOM was re-checked for red links. No figures or citations changed. |
| | |
| | ===== Quote-check refresh, 2026-09-21 ===== |
| | |
| | The 2026-09-04 ''cols''-vs-PDF audit on [[:provenance:literature:corpus]] showed that 73.1% of evidence quotes that cannot be located in ''paper.cols.txt'' **are** present in an independent ''pypdf'' rendering of the same ''paper.pdf''. ''scripts/quote_check.mjs'' was patched the same day to re-check everything below threshold against that second rendering and report a fourth verdict, **RESCUED**. §6's figure predates the patch and overstated this page's quote-failure rate by nearly a factor of three. Re-run, unedited first line: |
| | |
| | <code> |
| | $ node scripts/quote_check.mjs --statistics multiple-comparison-correction --show 400 |
| | 292 quotes checked: 167 exact, 63 partial (>=60% of 5-word windows), 40 rescued from the PDF, 22 below threshold in both renderings, 0 with no full text on disk. |
| | </code> |
| | |
| | ^ Figure ^ Was ^ Is ^ Why ^ |
| | | quotes checked | 292 | 292 | population unchanged — the corpus has not moved | |
| | | exact | 167 | 167 | unchanged | |
| | | partial (≥60% of 5-word windows) | 63 | 63 | unchanged | |
| | | rescued from the PDF | — | **40** | new verdict; these were inside the old 62 | |
| | | below threshold | **62** | **22** (in both renderings) | 62 = 40 + 22 exactly; nothing else moved | |
| | | below-threshold rate | 21% | **7.5%** | 22 of 292 | |
| | | unread below-threshold quotes | 22 | **10** | cross-mapped against ''out/pvalue-quotecheck.txt'' — see §6 | |
| | |
| | **What this does and does not say.** It does not say 40 extractions were wrong and are now right — the quotes were always in the papers, and 28 of these 40 had already been read by hand and found present. It says the //stored text// could not locate them and a second rendering of the same PDF can, so counting them as quote failures measured ''decolumn.mjs'', not the extraction. |
| | |
| | **The one number that needed more than the script.** ''62 = 40 + 22'' and ''62 = 40 read + 22 unread'' are two different partitions of the same 62, and reading the new output alone would let a future editor merge them. §6 now carries the cross-map: 28 of the 40 hand-read are //rescued//, 12 are still below in both, and 10 of the 22 now-below have never been read. |
| | |
| | **Scope of this edit.** §6, the §3 command line (''%%--show 62%%'' → ''%%--show 400%%'', since 62 no longer sizes the output), the review-log row that quoted the old disposition, and the content page's quote bullet. ''report_pvalue_corrections.mjs'' and ''mcc_fold.mjs'' were **not** re-run in this pass; no fold, rate, Fisher //p//-value or citation was touched, and every other figure on [[:statistics:pvalue_corrections]] and this page stands as published. |
| | |
| | ^ Item ^ Value ^ |
| | | Date | 2026-09-21, unsupervised | |
| | | Command | ''%%node scripts/quote_check.mjs --statistics multiple-comparison-correction --show 400%%'' | |
| | | Artifacts | ''out/qc0921/stat_mcc.txt'' (full run, 40 RESCUED rows and 22 below-threshold rows listed); the August output ''out/pvalue-quotecheck.txt'' was kept and used for the cross-map | |
| | | Script changes | none — ''quote_check.mjs'' was already patched on 2026-09-04 | |
| | | Reviewers | one ''sonnet'' figures-vs-script pass over this page and [[:statistics:pvalue_corrections]] | |
| | | Pages saved | this page, [[:statistics:pvalue_corrections]] | |
| |