User Tools

Site Tools


provenance:statistics:annotation

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

provenance:statistics:annotation [2026/09/11 02:59] – New provenance page for statistics:annotation: every query with its denominator, the ground-truth fold with its full residue, four hand-audit maps, the quote checks, external sources including the rejected ones, and the full four-pass review log with disp karel.kubicek.claudeprovenance:statistics:annotation [2026/09/11 03:17] (current) – Add the fifth review pass, which found that two independently-accepted fixes composed into a false claim, and record the DokuWiki heading-markup defect the run caught itself. Regenerated from the current script output. Authored by Claude karel.kubicek.claude
Line 814: Line 814:
   ['5,859 / 4,439 / 3,318 / 1,120 / 1,762', 'data/extract/OVERVIEW.md corpus populations'],   ['5,859 / 4,439 / 3,318 / 1,120 / 1,762', 'data/extract/OVERVIEW.md corpus populations'],
   ['15.4% agreement-metric base rate', 'computed here in G, and matches statistics:interrater_agreement'],   ['15.4% agreement-metric base rate', 'computed here in G, and matches statistics:interrater_agreement'],
-  ['252 papers folding to Cohen\'s kappa, 18 self-agreement papers', 'statistics:interrater_agreement and its report script — NOT re-derived here'],+  ['252 papers folding to Cohen\'s kappa, 18 self-agreement papers, 34.1% vs 13.8%', 'statistics:interrater_agreement and its report script — NOT re-derived here; its LLM-signal matcher is full-text where this page\'s is the classification.method enum'],
   ['19.4% of 175 name a resolvable model', 'design:website_classification / scripts/report_llm_currency.mjs'],   ['19.4% of 175 name a resolvable model', 'design:website_classification / scripts/report_llm_currency.mjs'],
   ['37.3% validation rate for topic classification', 'design:website_classification, population 330 not 4,439'],   ['37.3% validation rate for topic classification', 'design:website_classification, population 330 not 4,439'],
Line 1251: Line 1251:
  
 Figure                                                                                      Where it is from Figure                                                                                      Where it is from
-------------------------------------------------------------------------------------------  ---------------------------------------------------------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------  --------------------------------------------------------------------------------------------------------------------------------------------------------------------
 0.99 — "published F1 scores of up to 0.99"                                                  pendlebury2019_tesseract, abstract, verbatim in paper.cols.txt 0.99 — "published F1 scores of up to 0.99"                                                  pendlebury2019_tesseract, abstract, verbatim in paper.cols.txt
 20,000 requests heuristic-labelled and expert-validated as ground truth                     xiong2026_tgnn, appendix, verbatim 20,000 requests heuristic-labelled and expert-validated as ground truth                     xiong2026_tgnn, appendix, verbatim
Line 1260: Line 1260:
 5,859 / 4,439 / 3,318 / 1,120 / 1,762                                                       data/extract/OVERVIEW.md corpus populations 5,859 / 4,439 / 3,318 / 1,120 / 1,762                                                       data/extract/OVERVIEW.md corpus populations
 15.4% agreement-metric base rate                                                            computed here in G, and matches statistics:interrater_agreement 15.4% agreement-metric base rate                                                            computed here in G, and matches statistics:interrater_agreement
-252 papers folding to Cohen's kappa, 18 self-agreement papers                               statistics:interrater_agreement and its report script — NOT re-derived here+252 papers folding to Cohen's kappa, 18 self-agreement papers, 34.1% vs 13.8%               statistics:interrater_agreement and its report script — NOT re-derived here; its LLM-signal matcher is full-text where this page's is the classification.method enum
 19.4% of 175 name a resolvable model                                                        design:website_classification / scripts/report_llm_currency.mjs 19.4% of 175 name a resolvable model                                                        design:website_classification / scripts/report_llm_currency.mjs
 37.3% validation rate for topic classification                                              design:website_classification, population 330 not 4,439 37.3% validation rate for topic classification                                              design:website_classification, population 330 not 4,439
Line 4732: Line 4732:
  
 The same pass also named what it thought was good, which is worth recording because it is the part a checklist reviewer cannot produce: the temperature and prompt audits (probe count, audited count, //and why they differ//), the independence box, using the extraction's own error as the worked warning, the Wilson self-test checking against a defining equation rather than a constant, and the per-target table pointing each row at the page a reader with that target is on. The same pass also named what it thought was good, which is worth recording because it is the part a checklist reviewer cannot produce: the temperature and prompt audits (probe count, audited count, //and why they differ//), the independence box, using the extraction's own error as the worked warning, the Wilson self-test checking against a defining equation rather than a constant, and the per-target table pointing each row at the page a reader with that target is on.
 +
 +==== Pass 5 — figures again, after the fixes (model: sonnet) ====
 +
 +Re-run because the first pass's findings were acted on and the base-rate correction touched every table. The brief said explicitly to assume the page had acquired **new** defects while old ones were fixed, and it had.
 +
 +^ # ^ Finding ^ Disposition ^
 +| 5.1 | **HIGH, and it is a compound of two accepted fixes.** Pass 4 finding 4.2 changed //"three of the four are below the base rate"// to //"all four"//, which was **true of the 70.1% figure it was checked against**. Pass 4 finding 4.4 then replaced that base rate with the like-for-like **60.1%**, and mobile apps at 66.3% moved above it — leaving a sentence that was correct when it was written and false by the time it was published. | **Accepted.** Reverted to three of the four, with mobile apps named as the exception and the 60.1% stated in the sentence so the next edit cannot repeat this. |
 +| 5.2 | **LOW.** "all around half the 60.1% base rate" is right for cookies (53%) and IP addresses (56%) and wrong for domains (68%). | **Accepted**; the two are now stated separately. |
 +| 5.3 | **LOW.** 83.4% in the headline and 83.6% in the per-method table read as a mismatch unless the reader digs into this page. | **Accepted.** Both are right — 175 versus 177, for two different questions — and a footnote now says so at the point of first contact rather than only here. |
 +| 5.4 | **LOW.** "about 15%" is a generous rounding of the sibling page's 13.8%. | **Accepted**; both pairs of figures are now given exactly. |
 +| — | Verified clean by re-running every script from scratch and diffing against the committed outputs: the whole "Which kind, though" table cell by cell and all three readings under it; the 39.1%-of-151 comparison and 175 − 151 = 24; the joint reproducibility distribution (confirmed a joint over the three named items, 106/42/25/2 summing to 175, median 0); the residue paragraph, **including that all four quoted residue strings really are in the residue with the counts claimed** — the prior HIGH defect did not recur; and every superlative on the page. | — |
 +
 +**One defect this run found without a reviewer**, recorded because it is cheap and site-wide: the heading ''%%==== Which kind, though — and this part //has// moved ====%%'' rendered with its italic markers **visible**, and its anchor id silently collapsed to ''parthasmoved''. DokuWiki headings ignore inline markup. Every link check passed, because the ids on both sides matched each other. The page now has no markup in any heading; the wiki at large has many.
  
 ==== What the review layer cost and bought ==== ==== What the review layer cost and bought ====
  
-**Four passes, twenty-two findings, twenty-one accepted and one rejected.** The three focused passes found eight distinct defects (two of them the same citekey typo, found independently); the unchecklisted fourth pass found fourteen, including the four most serious things wrong with the page: a mis-scoped base rate that inflated every gap the page argued for, a retracted claim still live in a section nobody re-read, a false claim that a labelled corpus did not exist, and five confident numbers with nothing behind them.+**Five passes, twenty-six findings, twenty-five accepted and one rejected.** The three focused passes found eight distinct defects (two of them the same citekey typo, found independently); the unchecklisted fourth pass found fourteen, including the four most serious things wrong with the page: a mis-scoped base rate that inflated every gap the page argued for, a retracted claim still live in a section nobody re-read, a false claim that a labelled corpus did not exist, and five confident numbers with nothing behind them.
  
-**Two lessons for the next run, both cheap.**+**Three lessons for the next run, all cheap.**
  
 First, ''check_page_numbers.mjs'' returned OK through every one of these. It traces digits, and almost every defect here was about what a number //refers to// — which population, which denominator, which noun. The one class of it that //is// mechanically catchable is the one that bit hardest: the report script printed three residue examples from the author's memory instead of from the residue Map it had just built. A one-line assertion that a printed example is a member of the set it illustrates would have caught it, and ''gt_fold.mjs'' and ''annotation_audit.mjs'' now carry that assertion. **If a script prints an example, it should read it out of the thing it is an example of.** First, ''check_page_numbers.mjs'' returned OK through every one of these. It traces digits, and almost every defect here was about what a number //refers to// — which population, which denominator, which noun. The one class of it that //is// mechanically catchable is the one that bit hardest: the report script printed three residue examples from the author's memory instead of from the residue Map it had just built. A one-line assertion that a printed example is a member of the set it illustrates would have caught it, and ''gt_fold.mjs'' and ''annotation_audit.mjs'' now carry that assertion. **If a script prints an example, it should read it out of the thing it is an example of.**
  
-Second, the three focused briefs each named the specific places the author thought they were most likely wrong, and they came back with exactly those places plus little else. The unchecklisted pass, given only the reader and the "no textbook" rule, came back with the page's actual argument. **Both are needed and they are not substitutes**; a run that skipped the fourth pass would have shipped a page whose headline figure was wrong in the reader's favour.+Second, **a fix is not safe until it is re-checked against the other fixes from the same round.** Two findings were accepted independently and correctly, and their composition was false: one changed a sentence to match a base rate, the other changed the base rate. Nothing caught it but a re-run of the whole figures pass, which is why step 9 of the workflow says to re-run any reviewer whose findings you acted on. It earned its slot here at the first attempt. 
 + 
 +Third, the three focused briefs each named the specific places the author thought they were most likely wrong, and they came back with exactly those places plus little else. The unchecklisted pass, given only the reader and the "no textbook" rule, came back with the page's actual argument. **Both are needed and they are not substitutes**; a run that skipped the fourth pass would have shipped a page whose headline figure was wrong in the reader's favour.
  
 ===== Related ===== ===== Related =====
provenance/statistics/annotation.txt · Last modified: by karel.kubicek.claude