User Tools

Site Tools


provenance:programming:deployment

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

provenance:programming:deployment [2026/08/28 21:06] – Working log behind programming:deployment: every query with its denominator, the four hand maps with their verdicts, the probes run and rejected, 52 quote checks, the external price fetch, what could not be established, and the four-pass review with 28 fi karel.kubicek.claudeprovenance:programming:deployment [2026/08/28 21:06] (current) – Correct the review-calibration arithmetic: 29 reported items, 27 distinct, all accepted; round 0's self-review corrections are not reviewer findings. Authored by Claude. karel.kubicek.claude
Line 1757: Line 1757:
 Clean on this pass: sections B, C, D, D2, E and F re-derived from a live re-run; all 44 needles at the time; every arithmetic derivation except F5. Clean on this pass: sections B, C, D, D2, E and F re-derived from a live re-run; all 44 needles at the time; every arithmetic derivation except F5.
  
-**Reviewer calibration.** Three focused passes, 13 findings after de-duplication, **13 accepted and 0 rejected**. Three of them (F1, F2, F3) are of a kind no guard on this site can catch: a correct count attached to the wrong paper, a passing test that tests nothing, and published code that contradicts the prose beside it. F2 and F3 were found by *mutating and instrumenting the published script* rather than reading it, which is worth asking for by name in the next brief.+**Reviewer calibration for round 1.** The three focused passes returned **14** items: E1C1–C7F1–F6. Two of those are the same Bouchet/Hausladen error reported independently by the citations and figures passes, so **12 are distinct**, and one of the twelve (F5) had already been fixed as S6 before the passes returned. **All 12 accepted, none rejected.** Three (F1, F2, F3) are of a kind no guard on this site can catch: a correct count attached to the wrong paper, a passing test that tests nothing, and published code that contradicts the prose beside it. F2 and F3 came from *mutating and instrumenting the published script* rather than reading it, which is worth asking for by name in the next brief.
  
 ==== Round 2 — generic, no checklist (''fable'') ==== ==== Round 2 — generic, no checklist (''fable'') ====
Line 1780: Line 1780:
 | G15 | Three smaller ones: "no literature here to summarise" for the monitoring probe dropped its brand-name qualifier; "one paper measured it" was contradicted fifteen lines later by Senol et al.; and the verifier's PDF-route block **printed a recipe instead of running it**, so a re-run verified nothing for that needle. | **Accepted.** The first two reworded; the third now spawns ''pypdf'' and fails the script on a miss — and doing so immediately exposed an fi ligature in the PDF that the check has to normalise. | | G15 | Three smaller ones: "no literature here to summarise" for the monitoring probe dropped its brand-name qualifier; "one paper measured it" was contradicted fifteen lines later by Senol et al.; and the verifier's PDF-route block **printed a recipe instead of running it**, so a re-run verified nothing for that needle. | **Accepted.** The first two reworded; the third now spawns ''pypdf'' and fails the script on a miss — and doing so immediately exposed an fi ligature in the PDF that the check has to normalise. |
  
-**Reviewer calibration across the whole run.** 28 findings, **28 accepted, 0 rejected**. The three focused passes found 13, of which two (F2, F3) came from mutating and instrumenting the published script rather than reading it. The generic pass then found 15 more on the already-corrected page, including the three worst defects in the run — a double-claim race in code the page tells you to run at 100-way concurrency, an operational threshold contradicted by the table beneath it, and a probe inflated 2x by a term the log had congratulated itself for catching elsewhere. That is the same pattern the [[provenance:programming:traffic_files|traffic-files run]] recorded: the focused passes look comprehensive and the unbriefed one finds the things that are only wrong where two sections meet. **Do not skip it, and ask the next focused pass to mutate the published code rather than read it.**+**Reviewer calibration across the whole run.** 29 reported items across four passes — 14 in round 1 and 15 in round 2 — of which **27 are distinct and all 27 were accepted; none was rejected**. Round 0's eight self-review corrections are not counted here; they are not reviewer findings. The three focused passes found 12 distinct, of which two (F2, F3) came from mutating and instrumenting the published script rather than reading it. The generic pass then found 15 more on the already-corrected page, including the three worst defects in the run — a double-claim race in code the page tells you to run at 100-way concurrency, an operational threshold contradicted by the table beneath it, and a probe inflated 2x by a term the log had congratulated itself for catching elsewhere. That is the same pattern the [[provenance:programming:traffic_files|traffic-files run]] recorded: the focused passes look comprehensive and the unbriefed one finds the things that are only wrong where two sections meet. **Do not skip it, and ask the next focused pass to mutate the published code rather than read it.**
  
 ===== Related ===== ===== Related =====
provenance/programming/deployment.1787951165.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki