User Tools

Site Tools


provenance:security:email_authentication

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Next revision
Previous revision
provenance:security:email_authentication [2026/09/09 11:40] – New page: the working log behind security:email_authentication - every query with its denominator, the hand maps, both folds with residue in full, the quote/splice analysis, the accepted and rejected external sources, and all four review passes with dispo karel.kubicek.claudeprovenance:security:email_authentication [2026/09/09 11:50] (current) – Add section 12f (the figures pass re-run after its findings were applied, five more findings, one rejected as stale) and refresh the embedded script and output. Authored by Claude karel.kubicek.claude
Line 13: Line 13:
 | Report script | ''scripts/report_email_authentication.mjs'' (754 lines), output reproduced in full in section 8 below | | Report script | ''scripts/report_email_authentication.mjs'' (754 lines), output reproduced in full in section 8 below |
 | Model | Claude Opus 5, one main agent | | Model | Claude Opus 5, one main agent |
-| Sub-agents | one currency research pass (''sonnet''); three focused review passes (''sonnet''); one generic review pass (''fable''). All findings and dispositions in section 12 below |+| Sub-agents | one currency research pass (''sonnet'', which itself fanned out into six); three focused review passes (''sonnet''); one generic review pass (''fable''); one re-run of the figures pass after its findings were applied (''sonnet''). All findings and dispositions in section 12 below |
 | Pages touched | ''security:email_authentication'' (created), ''provenance:security:email_authentication'' (created), ''literature:bibliography'' (+26 entries, 1 corrected), ''security'' (five children → six), ''roadmap'' (row moved from //Queued// to //Assessed//) | | Pages touched | ''security:email_authentication'' (created), ''provenance:security:email_authentication'' (created), ''literature:bibliography'' (+26 entries, 1 corrected), ''security'' (five children → six), ''roadmap'' (row moved from //Queued// to //Assessed//) |
 | New BibTeX keys | 26 — 24 for the population, plus ''jeitner2021_injection'' (adjacent) and ''yajima2023_first'' (the BIMI measurement outside the corpus, added after review). One existing entry corrected: ''lee2020_longitudinal'' | | New BibTeX keys | 26 — 24 for the population, plus ''jeitner2021_injection'' (adjacent) and ''yajima2023_first'' (the BIMI measurement outside the corpus, added after review). One existing entry corrected: ''lee2020_longitudinal'' |
Line 102: Line 102:
 | Q15 | Is the population complete? | all 5,859 papers, two probes | see section 3 above | H | | Q15 | Is the population complete? | all 5,859 papers, two probes | see section 3 above | H |
 | Q16 | What tools do these papers use? | the 26 of the 31 naming at least one tool as ''used'' | Postfix 10, OpenSSL/pyOpenSSL 5, a PGP implementation 5, an SPF validator library 4, ZMap 4, Dovecot 3, a spam/blocklist service 3 … residue in section 6 | I | | Q16 | What tools do these papers use? | the 26 of the 31 naming at least one tool as ''used'' | Postfix 10, OpenSSL/pyOpenSSL 5, a PGP implementation 5, an SPF validator library 4, ZMap 4, Dovecot 3, a spam/blocklist service 3 … residue in section 6 | I |
 +| Q17 | Do these papers cite RFC 7489, the DMARC RFC that RFC 9989 obsoleted in May 2026? | the 31, full text | **18 cite it; 13 do not, and 7 of those never mention DMARC at all.** Added after the generic review falsified the page's claim that //every// paper cites it | G |
 +| Q18 | Does the filter change the answer? | the Tranco top million in {[wang2024_breakspf]}, one scan | SPF is **60.9%** of the top million and **79.4%** of the 738,310 that have an MX record or answer with an SMTP banner on port 25. Same paper, same day, **18.5 points** | E |
  
 The full per-paper lists behind Q3, Q4 and Q16 are printed in the report output; they are not reproduced here because they are the same 31 slugs eleven times over. The full per-paper lists behind Q3, Q4 and Q16 are printed in the report output; they are not reproduced here because they are the same 31 slugs eleven times over.
Line 107: Line 109:
 ==== Figures published on the page, with denominators ==== ==== Figures published on the page, with denominators ====
  
-53 figures, each hand-keyed in the ''FIGURES'' array of the report script with its population and its evidence quote, and each printed by report section E. They are not repeated here — section E of the output embedded in section 8 below is the authoritative list, and it is generated, so it cannot drift from the script. Read it rather than this page if you are checking a number.+55 figures, each hand-keyed in the ''FIGURES'' array of the report script with its population and its evidence quote, and each printed by report section E. They are not repeated here — section E of the output embedded in section 8 below is the authoritative list, and it is generated, so it cannot drift from the script. Read it rather than this page if you are checking a number
 + 
 +**One figure was published against the wrong denominator and caught by the run's own audit, not by a reviewer.** {[li2024_bounce]} reports that 701,347 (2.19%) messages hard-bounced on sender authentication failure. The first version of this page labelled that "the same 298 million emails; share of all, not of bounces" — which is exactly backwards. The paper's Table 1 is captioned //"Statistics on types of NDR messages for 32M bounced emails"// and 701,347 / 32M = 2.19%; against the 298 million total the same count is 0.24%. It was found by a mechanical check that every ''FIGURES'' entry's numbers appear somewhere on the page, which surfaced nine figures that were computed and quote-checked but never published, and reading those nine is what exposed the label. The page now carries both ''li2024_bounce'' bounce-reason figures against the 32 million, side by side, and says so.
  
 Three figures on the page are **derived** rather than quoted, and are marked here because a derived percentage is where an inferred denominator hides: Three figures on the page are **derived** rather than quoted, and are marked here because a derived percentage is where an inferred denominator hides:
Line 167: Line 171:
 ==== 7a. The 49 published figures ==== ==== 7a. The 49 published figures ====
  
-Every figure the page publishes carries an ''evidence.quote'' in the ''FIGURES'' array, and the script locates each one in ''paper.cols.txt'' in three modes — exact after whitespace collapse, whitespace-and-hyphen folded, punctuation folded. Four figures needed **two** fragments, giving **57 fragments** in total; three of those second fragments exist because the source sentence is not contiguous in any available rendering, and the fourth is a corroborating quotation rather than a splice.+Every figure the page publishes carries an ''evidence.quote'' in the ''FIGURES'' array, and the script locates each one in ''paper.cols.txt'' in three modes — exact after whitespace collapse, whitespace-and-hyphen folded, punctuation folded. Five figures needed **two** fragments, giving **60 fragments** in total; three of those second fragments exist because the source sentence is not contiguous in any available rendering, and the other two are corroborating quotations rather than splices.
  
-Final state: **55 exact, 2 punctuation-folded, 0 unlocatable.** The script prints ''*** N QUOTE(S) COULD NOT BE LOCATED ***'' and the page's quotes were rewritten until that line disappeared.+Final state: **58 exact, 2 punctuation-folded, 0 unlocatable.** The script prints ''*** N QUOTE(S) COULD NOT BE LOCATED ***'' and the page's quotes were rewritten until that line disappeared.
  
 **That is a selection effect and it should be stated as one.** Reaching zero meant changing the quote, never the figure. Ten quotes were rewritten in the first pass and one in the last; in every case the //number and the claim// stayed and only the string moved, usually by shortening to the contiguous part of the same sentence. The full before/after list: **That is a selection effect and it should be stated as one.** Reaching zero meant changing the quote, never the figure. Ten quotes were rewritten in the first pass and one in the last; in every case the //number and the claim// stayed and only the string moved, usually by shortening to the contiguous part of the same sentence. The full before/after list:
Line 470: Line 474:
     what: 'SPF', value: '60.9% published, 55.9% valid', pop: 'Tranco top million domains, 2023–24',     what: 'SPF', value: '60.9% published, 55.9% valid', pop: 'Tranco top million domains, 2023–24',
     quote: '60.9% of the top million domains have deployed SPF records, and 55.9% have deployed valid SPF records' },     quote: '60.9% of the top million domains have deployed SPF records, and 55.9% have deployed valid SPF records' },
 +  { g: 'adoption', slug: 'breakspf-how-shared-infrastructures-magnify-spf-vulnerabilities-across-the-internet',
 +    what: 'SPF, the SAME scan filtered to mail-serving domains', value: '79.4% published, 72.7% valid',
 +    pop: 'the 738,310 of the Tranco top million that have an MX record OR answer with an SMTP banner on port 25 -- the same paper, the same day, an 18.5-point difference from its unfiltered figure',
 +    quote: 'The adoption and valid rates among email domains are 79.4% and 72.7%, respectively' },
   { g: 'adoption', slug: 'unraveling-the-complexities-of-mta-sts-deployment-and-management-in-securing-ema',   { g: 'adoption', slug: 'unraveling-the-complexities-of-mta-sts-deployment-and-management-in-securing-ema',
     what: 'MTA-STS', value: '68,030 domains (0.07%–0.13%)', pop: 'domains WITH AN MX RECORD in the .com/.net/.org/.se zone files (86.8M; Table 1 column header reads "Domains with MX Records"), 29 Sep 2024',     what: 'MTA-STS', value: '68,030 domains (0.07%–0.13%)', pop: 'domains WITH AN MX RECORD in the .com/.net/.org/.se zone files (86.8M; Table 1 column header reads "Domains with MX Records"), 29 Sep 2024',
Line 564: Line 572:
     quote: '259M (87.07%) are non-bounced, 14M (4.82%) are soft-bounced, and 24M (8.11%) are hard-bounced' },     quote: '259M (87.07%) are non-bounced, 14M (4.82%) are soft-bounced, and 24M (8.11%) are hard-bounced' },
   { g: 'logs', slug: 'bounce-in-the-wild-a-deep-dive-into-email-delivery-failures-from-a-large-email-s',   { g: 'logs', slug: 'bounce-in-the-wild-a-deep-dive-into-email-delivery-failures-from-a-large-email-s',
-    what: 'bounces caused by sender authentication failure', value: '2.19% (701K)', pop: 'the same 298 million emails; share of allnot of bounces', +    what: 'bounces caused by sender authentication failure', value: '2.19% (701,347)', 
-    quote: 'We find that 701K (2.19%) emails are hard-bounced due to sender authentication failure' },+    pop: 'the 32 MILLION BOUNCED emails Table 1 classifies -- NOT the 298M total. 701,347/32M = 2.19%; against the 298M it would be 0.24%. This page published the wrong denominator for it once', 
 +    quote: 'We find that 701K (2.19%) emails are hard-bounced due to sender authentication failure', 
 +    quote2: 'Table 1: Statistics on types of NDR messages for 32M bounced emails' }, 
 +  { g: 'logs', slug: 'bounce-in-the-wild-a-deep-dive-into-email-delivery-failures-from-a-large-email-s', 
 +    what: 'bounces caused by the sender hitting a spam blocklist', value: '31.10% (9,975,329, Table 1 row T5)', pop: 'the same 32 million bounced emails Table 1 classifies -- like for like with the authentication row above, and 14x it', 
 +    quote: 'We find that 10M (31.10%) emails experience delivery failures due to sender MTAs hitting blocklists' },
   { g: 'logs', slug: 'understanding-and-characterizing-intermediate-paths-of-email-delivery-the-hidden',   { g: 'logs', slug: 'understanding-and-characterizing-intermediate-paths-of-email-delivery-the-hidden',
     what: 'paths relying entirely on third-party relaying', value: '82.7% (86.9M)', pop: '105 million emails with a reconstructable intermediate path',     what: 'paths relying entirely on third-party relaying', value: '82.7% (86.9M)', pop: '105 million emails with a reconstructable intermediate path',
Line 1602: Line 1615:
       source      : NDSS/2024/breakspf-how-shared-infrastructures-magnify-spf-vulnerabilities-across-the-internet       source      : NDSS/2024/breakspf-how-shared-infrastructures-magnify-spf-vulnerabilities-across-the-internet
       quote [exact]: "60.9% of the top million domains have deployed SPF records, and 55.9% have deployed valid SPF records"       quote [exact]: "60.9% of the top million domains have deployed SPF records, and 55.9% have deployed valid SPF records"
 +  SPF, the SAME scan filtered to mail-serving domains
 +      value       : 79.4% published, 72.7% valid
 +      DENOMINATOR : the 738,310 of the Tranco top million that have an MX record OR answer with an SMTP banner on port 25 -- the same paper, the same day, an 18.5-point difference from its unfiltered figure
 +      source      : NDSS/2024/breakspf-how-shared-infrastructures-magnify-spf-vulnerabilities-across-the-internet
 +      quote [exact]: "The adoption and valid rates among email domains are 79.4% and 72.7%, respectively"
   MTA-STS   MTA-STS
       value       : 68,030 domains (0.07%–0.13%)       value       : 68,030 domains (0.07%–0.13%)
Line 1749: Line 1767:
       quote [exact]: "259M (87.07%) are non-bounced, 14M (4.82%) are soft-bounced, and 24M (8.11%) are hard-bounced"       quote [exact]: "259M (87.07%) are non-bounced, 14M (4.82%) are soft-bounced, and 24M (8.11%) are hard-bounced"
   bounces caused by sender authentication failure   bounces caused by sender authentication failure
-      value       : 2.19% (701K+      value       : 2.19% (701,347
-      DENOMINATOR : the same 298 million emails; share of allnot of bounces+      DENOMINATOR : the 32 MILLION BOUNCED emails Table 1 classifies -- NOT the 298M total. 701,347/32M = 2.19%; against the 298M it would be 0.24%. This page published the wrong denominator for it once
       source      : IMC/2024/bounce-in-the-wild-a-deep-dive-into-email-delivery-failures-from-a-large-email-s       source      : IMC/2024/bounce-in-the-wild-a-deep-dive-into-email-delivery-failures-from-a-large-email-s
       quote [exact]: "We find that 701K (2.19%) emails are hard-bounced due to sender authentication failure"       quote [exact]: "We find that 701K (2.19%) emails are hard-bounced due to sender authentication failure"
 +      quote2 [exact]: "Table 1: Statistics on types of NDR messages for 32M bounced emails"
 +  bounces caused by the sender hitting a spam blocklist
 +      value       : 31.10% (9,975,329, Table 1 row T5)
 +      DENOMINATOR : the same 32 million bounced emails Table 1 classifies -- like for like with the authentication row above, and 14x it
 +      source      : IMC/2024/bounce-in-the-wild-a-deep-dive-into-email-delivery-failures-from-a-large-email-s
 +      quote [exact]: "We find that 10M (31.10%) emails experience delivery failures due to sender MTAs hitting blocklists"
   paths relying entirely on third-party relaying   paths relying entirely on third-party relaying
       value       : 82.7% (86.9M)       value       : 82.7% (86.9M)
Line 1830: Line 1854:
       quote [exact]: "we are able to send emails with valid SPF entries from 26 095 domains"       quote [exact]: "we are able to send emails with valid SPF entries from 26 095 domains"
  
-  53 figures, 57 quote fragments. Location modes: {"exact":55,"punctuation-folded":2}+  55 figures, 60 quote fragments. Location modes: {"exact":58,"punctuation-folded":2}
   All figure quotes located in paper.cols.txt.   All figure quotes located in paper.cols.txt.
  
Line 2202: Line 2226:
 K. DONE K. DONE
 ============================================================================== ==============================================================================
-figures published: 53+figures published: 55
 population       : 31 papers population       : 31 papers
 generated        : run this file, do not retype its numbers generated        : run this file, do not retype its numbers
Line 2353: Line 2377:
 | 16 | Provenance §1 said "+25 entries" two rows above "26 new keys" | **Accepted**, reconciled | | 16 | Provenance §1 said "+25 entries" two rows above "26 new keys" | **Accepted**, reconciled |
 | 17 | Ethics figures repeated in three places; the instrument table appears twice | **Partly accepted.** The duplication of the instrument table is deliberate (narrative first, period split in the query section) and the second now says so. The ethics repetition was left: the //Ethics// section and //Where these papers go quiet// serve different readers, and the numbers agree | | 17 | Ethics figures repeated in three places; the instrument table appears twice | **Partly accepted.** The duplication of the instrument table is deliberate (narrative first, period split in the query section) and the second now says so. The ethics repetition was left: the //Ethics// section and //Where these papers go quiet// serve different readers, and the numbers agree |
 +
 +==== 12f. Figures, re-run after the fixes (''sonnet'') ====
 +
 +The figures pass was re-run because its findings had been acted on and the script had grown by about 100 lines since. Its own opening note is the useful part: **the files changed under it mid-review**, twice, and it re-fetched and re-ran rather than reporting against a stale snapshot. That is a hazard this run created by editing a live page while four reviewers read it, and the next run should either freeze the files or say plainly that they are moving.
 +
 +^ # ^ Finding ^ Disposition ^
 +| 1 | The page's own figure/fragment audit had gone stale again — //"all 54 published figures … 58 quote fragments … Four figures carry a second fragment"// against a script now printing 55 and 60, with **five** ''quote2'' entries | **Accepted.** This is the third time a self-describing count on this page drifted behind the script that produces it. The counts are now 55 / 60 / five, and the standing lesson is that a page which states its own audit totals has to re-read them after every script change |
 +| 2 | //"one a regression ({[szurdi2017_email]})"// silently drops that the same paper is also the population's only correlation and only resampling result | **Accepted.** The bullet now says all three, which sharpens the point: one paper in eleven years accounts for every regression, every correlation and every resampling result in the population |
 +| 3 | The ethics bullet accounted for **23 of 31** papers — ''not-required'' (4), ''sought-outcome-unstated'' (2) and ''exempt'' (2) were never mentioned | **Accepted.** All six buckets are now on the page |
 +| 4 | The notification bullet accounted for 30 of 31; ''not-applicable'' (1) was missing | **Accepted** |
 +| 5 | An orphaned ''FIGURES'' entry (''hard bounces'', 8.11%) is computed but not on the page | **Rejected as already fixed.** It had been published in the //What it looks like from inside an operator// table before the reviewer's snapshot; it grepped a version from a few minutes earlier. Recorded rather than dropped, because "reviewer read a stale file" is the failure this run manufactured and it happened twice |
 +
 +Verified clean and mutation-tested by this pass: the mean/median block, the 8-of-16 union (mutating OR to AND visibly changes it to 1 of 16), the ethics corpus denominator (4,965, cross-checked against ''OVERVIEW.md'' independently of the script; widening the filter crashes loudly rather than producing a wrong number), the A3 author counts (111 distinct rebuilt from scratch, and still 111 after aggressive diacritic and punctuation folding, so the count is not a normalisation artefact; nulling one paper's authors visibly drops the denominator and prints the paper), the MTA-STS ''pop'' correction, the TLS-RPT wording, the MX-filter count of five, the WIDE probe description, the new RFC 7489 block, and the new BreakSPF 79.4% row against the paper's own Table I.
  
 ==== 12e. What the review layer was worth ==== ==== 12e. What the review layer was worth ====
  
-Nine findings from the figures pass, six from the citations pass, nine from the currency pass, seventeen from the generic pass. **Six of them changed a claim a reader would have acted on**: the ethics comparison was backwards, the BIMI claim was falsifiable, TLS-RPT is measured, a published author's name was wrong, "every paper cites RFC 7489" was false, and the page had no advice at all about policy strength. Three were caught by the main run independently before the reviews landed (the MX-row count, the WIDE-probe description, the author concentration), which is the argument for doing your own pass as well as commissioning four. **The generic pass, with no checklist, produced the most findings and two of the six that mattered** — including the one whose diagnosis names why: three focused briefs each assumed another owned the word "every".+Nine findings from the figures pass, six from the citations pass, nine from the currency pass, seventeen from the generic pass, and five more from re-running the figures pass after the fixes. **Six of them changed a claim a reader would have acted on**: the ethics comparison was backwards, the BIMI claim was falsifiable, TLS-RPT is measured, a published author's name was wrong, "every paper cites RFC 7489" was false, and the page had no advice at all about policy strength. Three were caught by the main run independently before the reviews landed (the MX-row count, the WIDE-probe description, the author concentration), which is the argument for doing your own pass as well as commissioning four. **The generic pass, with no checklist, produced the most findings and two of the six that mattered** — including the one whose diagnosis names why: three focused briefs each assumed another owned the word "every". 
 + 
 +The audit that no reviewer ran was worth a finding of its own. After the reviews, a mechanical check that every number in the report script's ''FIGURES'' array appears somewhere on the page found **nine figures computed and quote-checked but never published** — and reading those nine exposed a wrong denominator on a figure that //was// published ({[li2024_bounce]}'s 2.19%, a share of the 32 million bounces and not of the 298 million total). Publishing the nine added the //Deployed is not the same as correct// and //What it looks like from inside an operator// tables, which are now two of the page's more useful sections. **An invariant between the script and the page catches things four readers did not.**
  
 One rejection is worth as much as the accepts: the citations reviewer reported the SIDN Labs figures as unverifiable because the dashboard is JS-rendered, and a different agent reached the JSON behind it and confirmed them. "I could not fetch it" is not "it is wrong", and treating the two as the same would have cost the page its only current external deployment statistic. One rejection is worth as much as the accepts: the citations reviewer reported the SIDN Labs figures as unverifiable because the dashboard is JS-rendered, and a different agent reached the JSON behind it and confirmed them. "I could not fetch it" is not "it is wrong", and treating the two as the same would have cost the page its only current external deployment statistic.
provenance/security/email_authentication.1788954042.txt.gz · Last modified: by karel.kubicek.claude