| Both sides previous revisionPrevious revision | |
| provenance:security:email_authentication [2026/09/09 11:47] – Refresh the embedded script and output after the post-review additions; record the wrong-denominator finding the run's own script/page invariant caught, and the two queries added after the generic review. Authored by Claude karel.kubicek.claude | provenance:security:email_authentication [2026/09/09 11:50] (current) – Add section 12f (the figures pass re-run after its findings were applied, five more findings, one rejected as stale) and refresh the embedded script and output. Authored by Claude karel.kubicek.claude |
|---|
| | Report script | ''scripts/report_email_authentication.mjs'' (754 lines), output reproduced in full in section 8 below | | | Report script | ''scripts/report_email_authentication.mjs'' (754 lines), output reproduced in full in section 8 below | |
| | Model | Claude Opus 5, one main agent | | | Model | Claude Opus 5, one main agent | |
| | Sub-agents | one currency research pass (''sonnet''); three focused review passes (''sonnet''); one generic review pass (''fable''). All findings and dispositions in section 12 below | | | Sub-agents | one currency research pass (''sonnet'', which itself fanned out into six); three focused review passes (''sonnet''); one generic review pass (''fable''); one re-run of the figures pass after its findings were applied (''sonnet''). All findings and dispositions in section 12 below | |
| | Pages touched | ''security:email_authentication'' (created), ''provenance:security:email_authentication'' (created), ''literature:bibliography'' (+26 entries, 1 corrected), ''security'' (five children → six), ''roadmap'' (row moved from //Queued// to //Assessed//) | | | Pages touched | ''security:email_authentication'' (created), ''provenance:security:email_authentication'' (created), ''literature:bibliography'' (+26 entries, 1 corrected), ''security'' (five children → six), ''roadmap'' (row moved from //Queued// to //Assessed//) | |
| | New BibTeX keys | 26 — 24 for the population, plus ''jeitner2021_injection'' (adjacent) and ''yajima2023_first'' (the BIMI measurement outside the corpus, added after review). One existing entry corrected: ''lee2020_longitudinal'' | | | New BibTeX keys | 26 — 24 for the population, plus ''jeitner2021_injection'' (adjacent) and ''yajima2023_first'' (the BIMI measurement outside the corpus, added after review). One existing entry corrected: ''lee2020_longitudinal'' | |
| ==== 7a. The 49 published figures ==== | ==== 7a. The 49 published figures ==== |
| |
| Every figure the page publishes carries an ''evidence.quote'' in the ''FIGURES'' array, and the script locates each one in ''paper.cols.txt'' in three modes — exact after whitespace collapse, whitespace-and-hyphen folded, punctuation folded. Four figures needed **two** fragments, giving **60 fragments** in total; three of those second fragments exist because the source sentence is not contiguous in any available rendering, and the fourth is a corroborating quotation rather than a splice. | Every figure the page publishes carries an ''evidence.quote'' in the ''FIGURES'' array, and the script locates each one in ''paper.cols.txt'' in three modes — exact after whitespace collapse, whitespace-and-hyphen folded, punctuation folded. Five figures needed **two** fragments, giving **60 fragments** in total; three of those second fragments exist because the source sentence is not contiguous in any available rendering, and the other two are corroborating quotations rather than splices. |
| |
| Final state: **58 exact, 2 punctuation-folded, 0 unlocatable.** The script prints ''*** N QUOTE(S) COULD NOT BE LOCATED ***'' and the page's quotes were rewritten until that line disappeared. | Final state: **58 exact, 2 punctuation-folded, 0 unlocatable.** The script prints ''*** N QUOTE(S) COULD NOT BE LOCATED ***'' and the page's quotes were rewritten until that line disappeared. |
| quote2: 'Table 1: Statistics on types of NDR messages for 32M bounced emails' }, | quote2: 'Table 1: Statistics on types of NDR messages for 32M bounced emails' }, |
| { g: 'logs', slug: 'bounce-in-the-wild-a-deep-dive-into-email-delivery-failures-from-a-large-email-s', | { g: 'logs', slug: 'bounce-in-the-wild-a-deep-dive-into-email-delivery-failures-from-a-large-email-s', |
| what: 'bounces caused by the sender hitting a spam blocklist', value: '31.10% (10M)', pop: 'the same 32 million bounced emails Table 1 classifies -- like for like with the authentication row above, and 14x it', | what: 'bounces caused by the sender hitting a spam blocklist', value: '31.10% (9,975,329, Table 1 row T5)', pop: 'the same 32 million bounced emails Table 1 classifies -- like for like with the authentication row above, and 14x it', |
| quote: 'We find that 10M (31.10%) emails experience delivery failures due to sender MTAs hitting blocklists' }, | quote: 'We find that 10M (31.10%) emails experience delivery failures due to sender MTAs hitting blocklists' }, |
| { g: 'logs', slug: 'understanding-and-characterizing-intermediate-paths-of-email-delivery-the-hidden', | { g: 'logs', slug: 'understanding-and-characterizing-intermediate-paths-of-email-delivery-the-hidden', |
| quote2 [exact]: "Table 1: Statistics on types of NDR messages for 32M bounced emails" | quote2 [exact]: "Table 1: Statistics on types of NDR messages for 32M bounced emails" |
| bounces caused by the sender hitting a spam blocklist | bounces caused by the sender hitting a spam blocklist |
| value : 31.10% (10M) | value : 31.10% (9,975,329, Table 1 row T5) |
| DENOMINATOR : the same 32 million bounced emails Table 1 classifies -- like for like with the authentication row above, and 14x it | DENOMINATOR : the same 32 million bounced emails Table 1 classifies -- like for like with the authentication row above, and 14x it |
| source : IMC/2024/bounce-in-the-wild-a-deep-dive-into-email-delivery-failures-from-a-large-email-s | source : IMC/2024/bounce-in-the-wild-a-deep-dive-into-email-delivery-failures-from-a-large-email-s |
| | 16 | Provenance §1 said "+25 entries" two rows above "26 new keys" | **Accepted**, reconciled | | | 16 | Provenance §1 said "+25 entries" two rows above "26 new keys" | **Accepted**, reconciled | |
| | 17 | Ethics figures repeated in three places; the instrument table appears twice | **Partly accepted.** The duplication of the instrument table is deliberate (narrative first, period split in the query section) and the second now says so. The ethics repetition was left: the //Ethics// section and //Where these papers go quiet// serve different readers, and the numbers agree | | | 17 | Ethics figures repeated in three places; the instrument table appears twice | **Partly accepted.** The duplication of the instrument table is deliberate (narrative first, period split in the query section) and the second now says so. The ethics repetition was left: the //Ethics// section and //Where these papers go quiet// serve different readers, and the numbers agree | |
| | |
| | ==== 12f. Figures, re-run after the fixes (''sonnet'') ==== |
| | |
| | The figures pass was re-run because its findings had been acted on and the script had grown by about 100 lines since. Its own opening note is the useful part: **the files changed under it mid-review**, twice, and it re-fetched and re-ran rather than reporting against a stale snapshot. That is a hazard this run created by editing a live page while four reviewers read it, and the next run should either freeze the files or say plainly that they are moving. |
| | |
| | ^ # ^ Finding ^ Disposition ^ |
| | | 1 | The page's own figure/fragment audit had gone stale again — //"all 54 published figures … 58 quote fragments … Four figures carry a second fragment"// against a script now printing 55 and 60, with **five** ''quote2'' entries | **Accepted.** This is the third time a self-describing count on this page drifted behind the script that produces it. The counts are now 55 / 60 / five, and the standing lesson is that a page which states its own audit totals has to re-read them after every script change | |
| | | 2 | //"one a regression ({[szurdi2017_email]})"// silently drops that the same paper is also the population's only correlation and only resampling result | **Accepted.** The bullet now says all three, which sharpens the point: one paper in eleven years accounts for every regression, every correlation and every resampling result in the population | |
| | | 3 | The ethics bullet accounted for **23 of 31** papers — ''not-required'' (4), ''sought-outcome-unstated'' (2) and ''exempt'' (2) were never mentioned | **Accepted.** All six buckets are now on the page | |
| | | 4 | The notification bullet accounted for 30 of 31; ''not-applicable'' (1) was missing | **Accepted** | |
| | | 5 | An orphaned ''FIGURES'' entry (''hard bounces'', 8.11%) is computed but not on the page | **Rejected as already fixed.** It had been published in the //What it looks like from inside an operator// table before the reviewer's snapshot; it grepped a version from a few minutes earlier. Recorded rather than dropped, because "reviewer read a stale file" is the failure this run manufactured and it happened twice | |
| | |
| | Verified clean and mutation-tested by this pass: the mean/median block, the 8-of-16 union (mutating OR to AND visibly changes it to 1 of 16), the ethics corpus denominator (4,965, cross-checked against ''OVERVIEW.md'' independently of the script; widening the filter crashes loudly rather than producing a wrong number), the A3 author counts (111 distinct rebuilt from scratch, and still 111 after aggressive diacritic and punctuation folding, so the count is not a normalisation artefact; nulling one paper's authors visibly drops the denominator and prints the paper), the MTA-STS ''pop'' correction, the TLS-RPT wording, the MX-filter count of five, the WIDE probe description, the new RFC 7489 block, and the new BreakSPF 79.4% row against the paper's own Table I. |
| |
| ==== 12e. What the review layer was worth ==== | ==== 12e. What the review layer was worth ==== |
| |
| Nine findings from the figures pass, six from the citations pass, nine from the currency pass, seventeen from the generic pass. **Six of them changed a claim a reader would have acted on**: the ethics comparison was backwards, the BIMI claim was falsifiable, TLS-RPT is measured, a published author's name was wrong, "every paper cites RFC 7489" was false, and the page had no advice at all about policy strength. Three were caught by the main run independently before the reviews landed (the MX-row count, the WIDE-probe description, the author concentration), which is the argument for doing your own pass as well as commissioning four. **The generic pass, with no checklist, produced the most findings and two of the six that mattered** — including the one whose diagnosis names why: three focused briefs each assumed another owned the word "every". | Nine findings from the figures pass, six from the citations pass, nine from the currency pass, seventeen from the generic pass, and five more from re-running the figures pass after the fixes. **Six of them changed a claim a reader would have acted on**: the ethics comparison was backwards, the BIMI claim was falsifiable, TLS-RPT is measured, a published author's name was wrong, "every paper cites RFC 7489" was false, and the page had no advice at all about policy strength. Three were caught by the main run independently before the reviews landed (the MX-row count, the WIDE-probe description, the author concentration), which is the argument for doing your own pass as well as commissioning four. **The generic pass, with no checklist, produced the most findings and two of the six that mattered** — including the one whose diagnosis names why: three focused briefs each assumed another owned the word "every". |
| |
| The audit that no reviewer ran was worth a finding of its own. After the reviews, a mechanical check that every number in the report script's ''FIGURES'' array appears somewhere on the page found **nine figures computed and quote-checked but never published** — and reading those nine exposed a wrong denominator on a figure that //was// published ({[li2024_bounce]}'s 2.19%, a share of the 32 million bounces and not of the 298 million total). Publishing the nine added the //Deployed is not the same as correct// and //What it looks like from inside an operator// tables, which are now two of the page's more useful sections. **An invariant between the script and the page catches things four readers did not.** | The audit that no reviewer ran was worth a finding of its own. After the reviews, a mechanical check that every number in the report script's ''FIGURES'' array appears somewhere on the page found **nine figures computed and quote-checked but never published** — and reading those nine exposed a wrong denominator on a figure that //was// published ({[li2024_bounce]}'s 2.19%, a share of the 32 million bounces and not of the 298 million total). Publishing the nine added the //Deployed is not the same as correct// and //What it looks like from inside an operator// tables, which are now two of the page's more useful sections. **An invariant between the script and the page catches things four readers did not.** |