User Tools

Site Tools


practices:notifying_websites

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
practices:notifying_websites [2026/08/13 19:58] – Review fixes: correct Sasaki 2022 (telephone campaign, right denominators), Qin 2024 (859 sent, no significant treatment effect), Cetin control rate, Utz effect size; add phone/imprint channel box; flag Alexa as defunct; split RIPE/ARIN validation claims; karel.kubicek.claudepractices:notifying_websites [2026/09/04 13:10] (current) – Generic review: lead box no longer says all network RCTs failed; Extended Hell(o) row gets its funnel and comparison group; outcome kinds named; security.txt series disagreement stated; three figures corrected; methodology caveat rebalanced. Authored by C karel.kubicek.claude
Line 1: Line 1:
 ====== Notifying Websites ====== ====== Notifying Websites ======
  
-You crawled 100k sites and found that 8,000 of them leak something. Before you write the paper you have to tell those 8,000 operators, and every venue in this field now expects you to say so in the paper.((See //Disclosure timelines and what venues expect// below: IMC, USENIX Security, IEEE S&P and CCS all require it in some form as of their 2026 calls.)) This page is about the mechanics of doing that at scale: **how you get a contact for a party you have never met, what response rate to plan for, how long to wait before publishing, and what to report so a reviewer accepts the result.**+You crawled 100k sites and found that 8,000 of them leak something. Before you write the paper you have to decide what to do about those 8,000 operators, and the major venues now expect the paper to say what you decided.((See //Disclosure timelines and what venues expect// below: IMC, USENIX Security, IEEE S&P and CCS all require it in some form as of their 2026 calls; PoPETs sets out ethical principles but no disclosure requirement, and NDSS and TheWebConf are not addressed there.)) This page is about the mechanics of doing that at scale: **how you get a contact for a party you have never met, what response rate to plan for, how long to wait before publishing, and what to report so a reviewer accepts the result.**
  
 It is not about one-off coordinated disclosure to a named vendor. Emailing Google's security team is a solved problem with a queue and an SLA; emailing 8,000 strangers is a measurement in its own right, with a delivery funnel, a control group and a p-value. Treat it as one. It is not about one-off coordinated disclosure to a named vendor. Emailing Google's security team is a solved problem with a queue and an SLA; emailing 8,000 strangers is a measurement in its own right, with a delivery funnel, a control group and a p-value. Treat it as one.
  
 <WRAP important> <WRAP important>
-The single most consequential design decision on this page is the **control group**. Sites get fixed for reasons that have nothing to do with you — automatic updates, unrelated maintenance, going offline. Every notification study without a control arm has attributed some of that background remediation to its own notifications — and **both of the properly randomised, control-arm experiments in this corpus found no significant effect from any treatment**, including treatments that observational studies had reported as effective {[lone2022_sav,qin2024_rov]}. Randomise, hold back an arm, and report both. Note what this does and does not mean: those two results are about network operators making an expensive configuration change, and {[maass2021_effective]}, which also had a control groupmeasured 56.6% against 9.2%. The lesson is not that notification never works — it is that a study without a control arm cannot tell you which case it is in.+The single most consequential design decision on this page is the **control group**. Sites get fixed for reasons that have nothing to do with you — automatic updates, unrelated maintenance, going offline. Every notification study without a control arm has attributed some of that background remediation to its own notifications — and **the two most recent randomised, control-arm experiments on network operators in this corpus — on deploying SAV and ROV, expensive and low-visibility configuration changes — found no significant effect from any treatment**, including treatments that observational studies had reported as effective {[lone2022_sav,qin2024_rov]}. The earlier randomised experiments on network operators, on Heartbleed patching and firewall misconfiguration {[durumeric2014_heartbleed,li2016_youve]}, did find one. Randomise, hold back an arm, and report both. Note what that does and does not mean: {[maass2021_effective]} is also a randomised controlled experiment — a full factorial design with a control group — and it measured 56.6% against 9.2% on German website owners. The lesson is not that notification never works. It is that a study without a control arm cannot tell you which of those two cases it is in.
 </WRAP> </WRAP>
  
Line 28: Line 28:
 | ran a network scan or probe | 858 | 37.6% | 23.4% | 4.2% | 10.7% | 24.0% | | ran a network scan or probe | 858 | 37.6% | 23.4% | 4.2% | 10.7% | 24.0% |
 | assessed compliance with a law | 376 | 42.0% | 14.6% | 9.8% | 21.5% | 12.0% | | assessed compliance with a law | 376 | 42.0% | 14.6% | 9.8% | 21.5% | 12.0% |
 +| **crawled ∧ assessed a law** | **123** | 35.0% | 14.6% | **14.6%** | 15.4% | 20.3% |
  
-Two things to take from this table. **Network measurement is ahead of web measurement**: scanning papers notify at 61.0% (''yes''+''partial'') against 43.5% for crawling papers, and are less than half as likely to leave the question unanswered. The scanning community built the norm first, largely because Internet-wide scanning provoked complaints that forced the question. **Legal-compliance papers state a position most often** (only 12.0% ''not-stated'') and are also the most likely to say outright that they did not notify — a GDPR-violation finding across thousands of sites is case where authors have decided, and said, that individual notification is not the right instrument.+Two things to take from this table. **Network measurement is ahead of web measurement**: scanning papers notify at 61.0% (''yes''+''partial'') against 43.5% for crawling papers, and are less than half as likely to say outright that they did not (4.2% vs 9.4%), as well as meaningfully less likely to leave it unanswered (24.0% vs 31.1%). The scanning community built the norm first, largely because Internet-wide scanning provoked complaints that forced the question. **Legal-compliance papers state a position most often** (only 12.0% ''not-stated'')and the bottom row is the population closest to whoever is reading this: 123 papers that both crawled the web and assessed a law, i.e. found a compliance violation across many sites. **14.6% of them say outright that they did not notify** — more than three times the corpus-wide 4.3%. That is usually decision rather than negligence. A GDPR-violation finding across thousands of sites is the case where authors most often decide, and say, that individual notification is the wrong instrument, and route the finding to a regulator or to publication instead. If that is your situation, //deciding not to notify is a defensible position that you have to argue in the paper//, not an omission you can leave to the reader (two worked examples are under //Deciding not to run a campaign// below).
  
 The direction of travel is unambiguous: The direction of travel is unambiguous:
  
 ^ Indicator ^ Denominator ^ 2010–2013 ^ 2014–2017 ^ 2018–2021 ^ 2022–2024 ^ 2025–2026* ^ ^ Indicator ^ Denominator ^ 2010–2013 ^ 2014–2017 ^ 2018–2021 ^ 2022–2024 ^ 2025–2026* ^
-| Notified (''yes''\|''partial'') | empirical ∧ ethics | 21.3% | 37.7% | 45.9% | 52.8% | 60.3% |+| Notified (''yes'' or ''partial'') | empirical ∧ ethics | 21.3% | 37.7% | 45.9% | 52.8% | 60.3% |
 | Said nothing (''not-stated'') | empirical ∧ ethics | 61.5% | 40.2% | 26.5% | 16.9% | 13.1% | | Said nothing (''not-stated'') | empirical ∧ ethics | 61.5% | 40.2% | 26.5% | 16.9% | 13.1% |
 | Notified, crawlers only | crawled ∧ ethics | 14.6% | 34.7% | 41.4% | 50.6% | 54.3% | | Notified, crawlers only | crawled ∧ ethics | 14.6% | 34.7% | 41.4% | 50.6% | 54.3% |
Line 41: Line 42:
 | Contacted a regulator or CERT | empirical ∧ ethics | 0.3% | 3.0% | 4.7% | 3.4% | 2.8% | | Contacted a regulator or CERT | empirical ∧ ethics | 0.3% | 3.0% | 4.7% | 3.4% | 2.8% |
  
-<wrap todo>* 2025–2026 is provisional: CCS 2026 and IMC 2026 have not been held, and IEEE S&P 2026 and TheWebConf 2026 abstracts are not in OpenAlex, so those venue-years are under-represented by construction. See [[Literature:Corpus]].</wrap>+<WRAP todo>* 2025–2026 is provisional: CCS 2026 and IMC 2026 have not been held, and IEEE S&P 2026 and TheWebConf 2026 abstracts are not in OpenAlex, so those venue-years are under-represented by construction. See [[Literature:Corpus]].</WRAP>
  
-The one flat line is the interesting one. **Regulator and CERT contact has not grown and remains rare — 3.3% of the 4,472 say yes, and only 7.2% of the 376 papers that assessed a law.** Given that a CERT is often the only party that can reach a whole constituency, and that EU law now obliges every member state to run one for exactly this purpose (below), this is the largest gap between what is available and what is used.+The one line that is not rising is the interesting one. **Regulator and CERT contact has not grown and remains rare — 3.3% of the 4,472 say yes, and only 7.2% of the 376 papers that assessed a law.** Given that a CERT is often the only party that can reach a whole constituency, and that EU law now obliges every member state to run one for exactly this purpose (below), this is the largest gap between what is available and what is used
 + 
 +And when the intermediary //is// used it is used **as well as**, not instead of: of the 147 papers that contacted a regulator or CERT, **135 (91.8%) also notified the affected party directly**. Nobody in this corpus treats "we told the CERT" as discharging the obligation. Plan the intermediary as an extra arm, not as the cheap way out of finding contacts.
  
 ==== Saying you notified is not saying how ==== ==== Saying you notified is not saying how ====
  
-2,870 of the 4,472 (64.2%) give some free-text detail about their disclosure. Folding that text for any mention of a channel — the fold rules and their residue are on the provenance page — **2,303 of those 2,870 (80.2%) name no channel at all**, and 2,155 (75.1%) name no outcome. Where a channel is named it is usually a single large platform (11.7% mention Google, Apple, Meta, Microsoft, Amazon or Mozilla by name); a generic email to the operator or developer is 3.9%, a CERT 1.4%, a hosting provider 1.3%, a bug-bounty programme 1.3%, WHOIS 0.6%, a data protection authority 0.6%.+2,870 of the 4,472 (64.2%) give some free-text detail about their disclosure, and **all 2,160 that said they notified are among them**Scoping to those 2,160 — a channel is only a fair question of a paper that says it notified somebody((Of the 710 papers that gave a detail without saying they notified, only 273 (38.5%) are participant-consent or debriefing notes; the rest are data-handling and harm-limitation notes, and a handful even name a channel. So the exclusion is about the question being ill-posed for that group, not about its text being all one kind.)) — and folding the text for any mention of a channel (the rules and their residue are on the provenance page): 
 + 
 +**1,633 of the 2,160 (75.6%) do not say through what channel.** Where a channel does appear it is usually a single large platform named outright (14.7% mention Google, Apple, Meta, Microsoft, Amazon or Mozilla); a generic email to the operator or developer is 4.6%, a CERT 1.8%, a bug-bounty programme 1.6%, a hosting provider 1.6%, WHOIS 0.7%, a data protection authority 0.6%. And 2,155 of the 2,870 (75.1%) name no outcome either.
  
-And **only 7 of the 2,870 report a countable response ratio in that text** — figures like "Contacted 318 of 559 sensitive organizations; received 11 responses by submission" or "Reported 110 malicious extensions to Google; 62.7% were removed". Every rate in the next section had to be read out of the papers' own results sections, not out of their ethics sections. That is the reporting gap this page exists to close: the modal paper in this corpus says //we disclosed responsibly// and stops.+Worse for anyone trying to plan a campaign: **only 7 of the 2,870 report a countable response ratio in that text** — figures like "Contacted 318 of 559 sensitive organizations; received 11 responses by submission" or "Reported 110 malicious extensions to Google; 62.7% were removed". Every rate in the next section had to be read out of the papers' own results sections, not out of their ethics sections. That is the reporting gap this page exists to close: the modal paper in this corpus says //we disclosed responsibly// and stops.
  
 ===== Getting a contact: what still works in 2026 ===== ===== Getting a contact: what still works in 2026 =====
Line 71: Line 76:
  
 ^ Channel ^ How you get it ^ What it reaches ^ Status and evidence ^ ^ Channel ^ How you get it ^ What it reaches ^ Status and evidence ^
-| **''security.txt''** (RFC 9116) | ''GET https://<domain>/.well-known/security.txt'' | the security team, if there is one | RFC 9116, Informational, April 2022 — current, not obsoleted.((Verified against [[https://www.rfc-editor.org/rfc/rfc9116.txt|rfc-editor.org/rfc/rfc9116.txt]] on 2026-08-13: "Category: Informational … April 2022". Three errata exist, none touching the field list.)) Adoption is the problem: 11–16% of the Alexa top 100, 8–10% of the top 1K, "3–4% for the top 10K sites, and only a percent for the top 100K" {[poteat2021_securitytxt]}. **That measurement is from 2021 and its ranking frame no longer exists** — Alexa was discontinued in 2022 (see [[Design:Website selection]]), the paper predates RFC 9116, and **nothing in this corpus re-measures security.txt adoption since**. Assume it has risenyou have no citeable current figure +| **''security.txt''** (RFC 9116) | ''GET https://<domain>/.well-known/security.txt'' | the security team, if there is one | RFC 9116, Informational, April 2022 — current, not obsoleted.((Verified against [[https://www.rfc-editor.org/rfc/rfc9116.txt|rfc-editor.org/rfc/rfc9116.txt]] on 2026-08-13: "Category: Informational … April 2022". Three errata exist, none touching the field list.)) Adoption is the problem, and it is steep in rank: 11–16% of the Alexa top 100, 8–10% of the top 1K, "3–4% for the top 10K sites, and only a percent for the top 100K" {[poteat2021_securitytxt]}. That is 2021 measurement on a ranking that was discontinued in 2022 (see [[Design:Website selection]]) and predates RFC 9116**no paper in the seven corpus venues re-measures adoption since.** Outside them, two do: 0.49% of the Tranco top million and 1.6% of the top 100K in October–November 2021 {[findlay2022_securitytxt]}, and 42 successful weekly scans of the Tranco top million over the 55 weeks from December 2021 to January 2023 showing 32.0→34.0% (top 100), 16.1→18.8% (top 1K), 7.9→9.8% (top 10K), 2.4→3.2% (top 100K) and 0.7→1.0% (top 1M) {[hilbig2023_securitytxt]}.((Both are outside the corpus's seven venues: MADweb is an NDSS workshop, DTRAP an ACM journal. A vendor scan of gTLD zone files reports 573,123 of 241 million domains (0.24%) with a ''security.txt'' in early 2026, up from 0.05% in 2021 — [[https://blog.iotdef.com/the-state-of-security-txt-adoption-an-analysis-of-240-million-domains-in-2026/|blog.iotdef.com]], fetched 2026-09-04; methodology stated, data and code not published, so treat it as an order of magnitude, not a citation.)) The two peer-reviewed series disagree by about 2× at the top ranks (Alexa against Tranco, and a Kaplan–Meier deployment estimate against weekly snapshots); the corpus's own incidental counts sit with the Tranco series: 23 domains in {[drakonakis2020_cookie]} (2020, "the most ineffective" of its four channels; no denominator given), 25 of 256 in {[roth2022_security]} (2022), 34 of the 100 most popular services in {[bhattacharya2026_asi]} (2025). Six corpus papers that notified **used** it as one channel among several — {[drakonakis2020_cookie,ghasemisharif2022_saat,roth2022_security,czybik2023_lazy,innocenti2025_only,bhattacharya2026_asi]} — and **none separates its delivery or response from the other channels**, so whether a listed contact actually answers is still unmeasured; {[hilbig2023_securitytxt]} names exactly that as future work 
-| **RIR abuse contact** | RIPE ''abuse-c:'' / ARIN Abuse POC, resolved from the IP; RIPEstat's abuse-contact-finder covers all five RIRs over HTTPS | the **hosting provider or CDN**, not the site owner | The one source that scales — free and with no meaningful rate limit. ARIN verifies its Abuse POC annually;((ARIN NRPM §3.6: "Each of the following Points of Contact are to be verified annually … Admin, Tech, NOC, Abuse", [[https://www.arin.net/participate/policy/nrpm/|arin.net]], fetched 2026-08-13. RIPE states it works to keep abuse contacts valid but no validation cadence could be found on a RIPE primary source on 2026-08-13 — do not assume annual.)) RIPE states it keeps abuse contacts valid but publishes no cadence could verify. Still the workhorse: "WHOIS was the most frequently mentioned channel (seven HPOs)" among 24 interviewed providers in 2026, who "confirmed that their contact is available via WHOIS, for example through ARIN or RIPE databases" {[stivala2026_behind]}. Note this is the //network// registry, which GDPR redaction did not touch — not the //domain// registrant record, which it did |+| **RIR abuse contact** | RIPE ''abuse-c:'' / ARIN Abuse POC, resolved from the IP; RIPEstat's abuse-contact-finder covers all five RIRs over HTTPS | the **hosting provider or CDN**, not the site owner | The one source that scales — free and with no meaningful rate limit. ARIN verifies its Abuse POC annually;((ARIN NRPM §3.6: "Each of the following Points of Contact are to be verified annually … Admin, Tech, NOC, Abuse", [[https://www.arin.net/participate/policy/nrpm/|arin.net]], fetched 2026-08-13. RIPE states it works to keep abuse contacts valid but no validation cadence could be found on a RIPE primary source on 2026-08-13 — do not assume annual.)) RIPE states it keeps abuse contacts valid but publishes no cadence that could be verified from a RIPE source. Still the workhorse: "WHOIS was the most frequently mentioned channel (seven HPOs)" among 24 interviewed providers in 2026, who "confirmed that their contact is available via WHOIS, for example through ARIN or RIPE databases" {[stivala2026_behind]}. Note this is the //network// registry, which GDPR redaction did not touch — not the //domain// registrant record, which it did. It is not clean either: 384 of 552 RIR-WHOIS mails bounced in {[gilad2017_rpki]} ("whois entries are often outdated"), and 302 of 1,385 (21.8%) were undeliverable in {[li2026_rpkiinvalid]}. For DNS resolvers there is one more machine-readable source, the ''RNAME'' of the zone's SOA record; {[deccio2020_closeddoors]} used it to reach 43 administrators and heard back from five, three of whom the authors already knew |
 | **PeeringDB technical contact** | PeeringDB API, per AS | the network operator | Preferred over WHOIS by {[lone2022_sav]}: "We preferred peeringDB because it has been used in previous studies and they found the database up-to-date" — i.e. it is relaying prior work's assessment, not measuring it | | **PeeringDB technical contact** | PeeringDB API, per AS | the network operator | Preferred over WHOIS by {[lone2022_sav]}: "We preferred peeringDB because it has been used in previous studies and they found the database up-to-date" — i.e. it is relaying prior work's assessment, not measuring it |
 | **The site's own imprint / privacy-policy contact** | read it off the page | the legally responsible party — often the actual decision-maker | The highest-yield channel measured. Addresses parsed from privacy-policy and contact pages achieved **87.8% delivery against 33.8% for RFC 2142 aliases** {[utz2023_comparing]}; 9 of 11 organisations reached via a privacy-policy address resolved the issue {[elyadmani2025_keys]}. {[maass2021_effective]} collected German ''Impressum'' addresses by hand, three researchers per site | | **The site's own imprint / privacy-policy contact** | read it off the page | the legally responsible party — often the actual decision-maker | The highest-yield channel measured. Addresses parsed from privacy-policy and contact pages achieved **87.8% delivery against 33.8% for RFC 2142 aliases** {[utz2023_comparing]}; 9 of 11 organisations reached via a privacy-policy address resolved the issue {[elyadmani2025_keys]}. {[maass2021_effective]} collected German ''Impressum'' addresses by hand, three researchers per site |
Line 80: Line 85:
 | **Established notification operators** | Shadowserver Foundation, CERT-BUND and equivalents | operators who already trust the sender | Shadowserver still sends free daily reports to 201 national CSIRTs across 175 countries.((Fetched from [[https://www.shadowserver.org/what-we-do/network-reporting/|shadowserver.org]] on 2026-08-13.)) {[munteanu2025_catch22]} ran its whole campaign this way and thanks Shadowserver and CERT-BUND for it — the current worked example | | **Established notification operators** | Shadowserver Foundation, CERT-BUND and equivalents | operators who already trust the sender | Shadowserver still sends free daily reports to 201 national CSIRTs across 175 countries.((Fetched from [[https://www.shadowserver.org/what-we-do/network-reporting/|shadowserver.org]] on 2026-08-13.)) {[munteanu2025_catch22]} ran its whole campaign this way and thanks Shadowserver and CERT-BUND for it — the current worked example |
 | **Platform consoles** | the platform's own channel to its registered operators | only operators who pre-registered | The strongest single effect in the literature, and the least available to you: a Search Console message lifted remediation to 82.4% / 76.8% against 54.6% / 43.4% otherwise {[li2016_remedying]}. Only 22–32% of affected sites had a registered webmaster, and you are not Google | | **Platform consoles** | the platform's own channel to its registered operators | only operators who pre-registered | The strongest single effect in the literature, and the least available to you: a Search Console message lifted remediation to 82.4% / 76.8% against 54.6% / 43.4% otherwise {[li2016_remedying]}. Only 22–32% of affected sites had a registered webmaster, and you are not Google |
 +| **A bulk report to the platform** | the store's or platform's abuse channel, with the whole list attached | the platform, which removes or contacts | Not a notification of operators at all, but the instrument that produced the highest removal figures in this corpus: 813 of 1,095 apps reported to Google were removed {[roundy2020_creepware]}; 1,316 of 1,804 GPTs reported to OpenAI were gone within two weeks {[shen2025_gptracker]}. {[nan2023_spying]} fell back to it after a 1.5% developer response; {[edu2022_alexa]} reported 675 skills to Amazon alongside the developers. It reaches only what the platform hosts, and the platform decides what counts as a violation |
 | **HackerOne Disclosure Assistance** | ''hackerone.com/disclosure-assistance'' | organisations with no disclosure policy at all | Newer than the literature: HackerOne "will work with friendly hackers on a **best-effort basis**" to verify a bug, find someone at the affected organisation and relay it.((Fetched from [[https://docs.hackerone.com/en/articles/8466632-disclosure-assistance|docs.hackerone.com]] on 2026-08-13; page dated 11 June 2024.)) It updates {[stock2016_hey]}, which discarded reward programmes because they "usually only accept and forward reports for their customers" — that premise no longer holds for HackerOne. But it is a per-report human service gated on having exhausted other options, so it is **not** a bulk channel | | **HackerOne Disclosure Assistance** | ''hackerone.com/disclosure-assistance'' | organisations with no disclosure policy at all | Newer than the literature: HackerOne "will work with friendly hackers on a **best-effort basis**" to verify a bug, find someone at the affected organisation and relay it.((Fetched from [[https://docs.hackerone.com/en/articles/8466632-disclosure-assistance|docs.hackerone.com]] on 2026-08-13; page dated 11 June 2024.)) It updates {[stock2016_hey]}, which discarded reward programmes because they "usually only accept and forward reports for their customers" — that premise no longer holds for HackerOne. But it is a per-report human service gated on having exhausted other options, so it is **not** a bulk channel |
  
Line 92: Line 98:
 ==== A tested contact-discovery cascade ==== ==== A tested contact-discovery cascade ====
  
-Three of those channels are machine-queryable. This runs them in cost order and prints per-source coverage, because a pooled "we found contacts for N% of domains" is the figure a reviewer sends back. Failures are printed rather than swallowed: a silent zero and a missing contact look identical, and the difference understates your own coverage.+Two of those channels are machine-queryable per domain, and RDAP is worth querying even though it is mostly redacted, because what it returns tells you //which kind// of redaction you are up against. (PeeringDB is machine-queryable too, but per AS — use it when your unit is a network operator rather than a website.This runs the three in cost order and prints per-source coverage, because a pooled "we found contacts for N% of domains" is the figure a reviewer sends back. Failures are printed rather than swallowed: a silent zero and a missing contact look identical, and the difference understates your own coverage.
  
 <file python find_contacts.py> <file python find_contacts.py>
Line 113: Line 119:
   * scraping mailto: links off the page -- highest yield, worst measured   * scraping mailto: links off the page -- highest yield, worst measured
     delivery, and it collects addresses nobody published for this purpose.     delivery, and it collects addresses nobody published for this purpose.
 +
 +Why RIPEstat rather than the Abusix Abuse Contact DB for the RIR abuse contact:
 +Abusix is a DNS TXT lookup, free and unmetered, and is what the 2016 USENIX
 +campaign used, so it is the better choice at real scale. RIPEstat is used here
 +because it needs nothing but the standard library over HTTPS, which keeps this
 +script runnable -- and therefore testable -- anywhere. Swap in Abusix once you have
 +a DNS client and thousands of domains.
  
 Failures are printed, never swallowed: a silent zero looks identical to a Failures are printed, never swallowed: a silent zero looks identical to a
-missing contact and would understate your own coverage.+missing contact and would understate your own coverage. The one exception is the 
 +IANA bootstrap fetch, which raises: it is not a per-domain zero, it makes RDAP 
 +coverage unmeasurable for the whole batch.
 """ """
  
Line 257: Line 272:
 </file> </file>
  
-Real output, run 2026-08-13 on five hand-picked domains — **not a sample, and not representative**: these are large, well-resourced organisations, so ''security.txt'' coverage here is an order of magnitude above the 34% {[poteat2021_securitytxt]} measured at the top 10K.+Real output, run 2026-08-13 on five hand-picked domains — **not a sample, and not representative**: these are large, well-resourced organisations, so ''security.txt'' coverage here is well above any measured tier — 3234for the top 100 in {[hilbig2023_securitytxt]}, 3–4% at the top 10K in {[poteat2021_securitytxt]}.
  
 <code> <code>
Line 264: Line 279:
 google.com               security.txt  mailto:security@google.com               expires=2030-04-01T00:00:00z google.com               security.txt  mailto:security@google.com               expires=2030-04-01T00:00:00z
 google.com               RDAP          —                                        response carried no email — redacted, or role-only google.com               RDAP          —                                        response carried no email — redacted, or role-only
-google.com               RIR abuse     network-abuse@google.com                 ip=74.125.29.113 rir=arin+google.com               RIR abuse     network-abuse@google.com                 ip=172.217.208.139 rir=arin
 ethz.ch                  security.txt  mailto:security@ethz.ch                  expires=2028-01-31T07:00:00.000Z ethz.ch                  security.txt  mailto:security@ethz.ch                  expires=2028-01-31T07:00:00.000Z
 ethz.ch                  RDAP          —                                        .ch has no RDAP service in the IANA bootstrap ethz.ch                  RDAP          —                                        .ch has no RDAP service in the IANA bootstrap
Line 274: Line 289:
 bbc.co.uk                security.txt  mailto:security@bbc.co.uk                expires=2038-01-19T03:14:07Z bbc.co.uk                security.txt  mailto:security@bbc.co.uk                expires=2038-01-19T03:14:07Z
 bbc.co.uk                RDAP          redacted@nominet.uk                      role=registrant bbc.co.uk                RDAP          redacted@nominet.uk                      role=registrant
-bbc.co.uk                RIR abuse     abuse@fastly.com                         ip=151.101.192.81 rir=arin+bbc.co.uk                RIR abuse     abuse@fastly.com                         ip=151.101.128.81 rir=arin
 wikipedia.org            security.txt  mailto:security@wikimedia.org            expires=2029-03-31T09:00:00.000Z wikipedia.org            security.txt  mailto:security@wikimedia.org            expires=2029-03-31T09:00:00.000Z
 wikipedia.org            RDAP          —                                        response carried no email — redacted, or role-only wikipedia.org            RDAP          —                                        response carried no email — redacted, or role-only
Line 297: Line 312:
 {[sasaki2022_ics]} collects the comparison itself, and it is the shortest statement of the range: "its remediation rate was approximately 18% … Our remediation rate is higher than most previous notification experiments: it was approximately 40% for cross-site scripting and a WordPress vulnerability, 33%–42% for different WordPress vulnerability, and less than 20% for DNS zone poisoning. The only campaigns that reported similar remediation rates were on publicly accessible Git repositories (78%–81%) and Heartbleed (approximately 40%–90%)." {[sasaki2022_ics]} collects the comparison itself, and it is the shortest statement of the range: "its remediation rate was approximately 18% … Our remediation rate is higher than most previous notification experiments: it was approximately 40% for cross-site scripting and a WordPress vulnerability, 33%–42% for different WordPress vulnerability, and less than 20% for DNS zone poisoning. The only campaigns that reported similar remediation rates were on publicly accessible Git repositories (78%–81%) and Heartbleed (approximately 40%–90%)."
  
-These are the campaigns in the corpus that reported a quotable outcome, each with **its own denominator** — they are not comparable to each other, because "response", "remediation" and "fix" are defined differently in each and the populations are wildly different. Plan against the range, not the mean. The hand pass classified 22 campaigns in total; 16 are below, and the six not shown reported a volume sent but no response or remediation figure. They are listed on [[provenance:practices:notifying_websites]], along with the one row below — {[stock2018_didnt]} — whose full text is missing from the corpus and was read from the publisher's PDF instead.+These are the campaigns in the corpus that reported a quotable outcome, each with **its own denominator** — they are not comparable to each other, because "response", "remediation" and "fix" are defined differently in each and the populations are wildly different. Plan against the range, not the mean. The Outcome column mixes four different quantities and says which each row reports: **remediation** measured by re-scan (Maass, Czybik), **response** counts (Nguyen, Li 2024), **delivery** alone (Roth 2022, the 2026 validation survey), and **removal by a platform** (Edu). Pick the rows that measure what you will measure. The hand pass classified 32 campaigns in total; 26 are below, and the six not shown reported a volume sent but no response or remediation figure. Two rows — {[munteanu2025_catch22]} and {[liao2016_seeking]} — are here for their channel rather than a rate, because routing a campaign through an established notification operator or a national CERT is the pattern, not the number. The six are listed on [[provenance:practices:notifying_websites]], along with the one row below — {[stock2018_didnt]} — whose full text is missing from the corpus and was read from the publisher's PDF instead. Rows marked * are from the provisional 2025–2026 slice.
  
 ^ Study ^ Notified ^ Channel ^ Outcome, in the paper's own terms ^ Control ^ ^ Study ^ Notified ^ Channel ^ Outcome, in the paper's own terms ^ Control ^
Line 305: Line 320:
 | {[li2016_youve]} USENIX Sec 2016 | 2,563 ICS + 3,536 IPv6 + 5,960 amplifier contacts | WHOIS abuse contact, national CERTs, US-CERT | "at most 18% of the population remediating" under the best regimen; direct verbose 9.8% vs national CERT 3.1% vs US-CERT 1.4% (IPv6, two days) | US-CERT arm statistically indistinguishable from control | | {[li2016_youve]} USENIX Sec 2016 | 2,563 ICS + 3,536 IPv6 + 5,960 amplifier contacts | WHOIS abuse contact, national CERTs, US-CERT | "at most 18% of the population remediating" under the best regimen; direct verbose 9.8% vs national CERT 3.1% vs US-CERT 1.4% (IPv6, two days) | US-CERT arm statistically indistinguishable from control |
 | {[li2016_remedying]} TheWebConf 2016 | 760,935 hijacking incidents | browser interstitial, search warning, Search Console message, WHOIS admin email | 59.5% of incidents resolved over 11 months; **82.4% / 76.8% where a Search Console alert reached a pre-registered webmaster** vs 54.6% / 43.4% otherwise | — (channel comparison) | | {[li2016_remedying]} TheWebConf 2016 | 760,935 hijacking incidents | browser interstitial, search warning, Search Console message, WHOIS admin email | 59.5% of incidents resolved over 11 months; **82.4% / 76.8% where a Search Console alert reached a pre-registered webmaster** vs 54.6% / 43.4% otherwise | — (channel comparison) |
-| {[stock2018_didnt]} NDSS 2018 | >24,000 domains, seven arms of ~4,000 | six email variants (plain, HTML, tracking, mailbot, S/MIME, friendly tone) | notified 24% (Git) and 17% (WordPress) fixed; **74.4% / 33.3% among those who opened the report** | 13% (Git), 14% (WordPress) |+| {[liao2016_seeking]} IEEE S&P 2016 | over 120 FQDNs to US-CERT and 136 to CCERT (China) | two national CERTs; CCERT then "notified all related organizations" — "it is difficult for us to directly contact the victims" | "27 responded and fixed their problems"; the number CCERT notified is not stated, so no rate | — | 
 +| {[gilad2017_rpki]} NDSS 2017 | 552 victims and offenders of RPKI misconfiguration, over six months | RIR WHOIS contact, sent by the ROAlert system | **only 168 of 552 emails did not bounce**; "Over 42%" of bad-ROA alerts fixed a month later; 19% for loose-ROA alerts; 52 administrators engaged, 40 reported fixing | "about 15%" fixed among operators ROAlert could not reach — a comparison group the paper reports, not a randomised control | 
 +| {[stock2018_didnt]} NDSS 2018 | >24,000 domains, seven arms of ~4,000 | six email variants (plain, HTML, tracking, mailbot, S/MIME, friendly tone) | notified 24% (Git) and 17% (WordPress) fixed; **74.4% / 33.3% among those who //viewed// the report** | 13% (Git), 14% (WordPress) |
 | {[cetin2019_cleaning]} NDSS 2019 | ISP customers with Mirai infections | ISP walled garden (quarantine + notification) vs email-only vs nothing | walled garden "remediates 92% of the infections within 14 days"; **"Email-only notifications have no observable impact compared to a control group"** | **"The control group achieved the lowest cleanup rate (74%)"** with no notification at all — the cautionary figure on this whole page | | {[cetin2019_cleaning]} NDSS 2019 | ISP customers with Mirai infections | ISP walled garden (quarantine + notification) vs email-only vs nothing | walled garden "remediates 92% of the infections within 14 days"; **"Email-only notifications have no observable impact compared to a control group"** | **"The control group achieved the lowest cleanup rate (74%)"** with no notification at all — the cautionary figure on this whole page |
 +| {[roth2020complex]} NDSS 2020 | 2,699 sites with a broken or double-framing-prone ''X-Frame-Options'' header | RFC 2142 aliases (info, security, webmaster) plus WHOIS contact, from a researcher's own address | "most emails bounced"; **117 non-automated responses**; sites deploying ''frame-ancestors'' rose from 511 to 554 (+43) in eight days, 62 responders said they would deploy it | — |
 | {[maass2021_effective]} USENIX Sec 2021 | 4,594 German site owners, 18 arms + control | postal letter and email, addresses read by hand from each ''Impressum'' | **"56.6 % of all notified operators remediating within two months"**; 76.3% for a legal-research-group letter citing fines, 33.9% for a computer-science email citing privacy | **9.2%** | | {[maass2021_effective]} USENIX Sec 2021 | 4,594 German site owners, 18 arms + control | postal letter and email, addresses read by hand from each ''Impressum'' | **"56.6 % of all notified operators remediating within two months"**; 76.3% for a legal-research-group letter citing fines, 33.9% for a computer-science email citing privacy | **9.2%** |
 | {[nguyen2021_sharefirst]} USENIX Sec 2021 | 11,914 app developers | contact address from the Play Store listing | **448 responses (3.8%)** | — | | {[nguyen2021_sharefirst]} USENIX Sec 2021 | 11,914 app developers | contact address from the Play Store listing | **448 responses (3.8%)** | — |
 | {[squarcina2021_subdomain]} USENIX Sec 2021 | sites with dangling subdomain records | direct contact vs the authors' national CERT | direct 31% / 22% fixed vs national CERT 10% / 14% | — | | {[squarcina2021_subdomain]} USENIX Sec 2021 | sites with dangling subdomain records | direct contact vs the authors' national CERT | direct 31% / 22% fixed vs national CERT 10% / 14% | — |
 +| {[edu2022_alexa]} TheWebConf 2022 | 675 Alexa skills reported to Amazon and to "all affected developers" — "whenever we have their contact details"; 246 with broken traceability | developer contact details where available, plus the platform | of the 246, **111 (45.12%) "no longer pose a threat"** about a year later (45 removed, 24 lost the permission, 41 fixed — the parts sum to 110); 107 (43.5%) still broken. No delivery or reply counts | — |
 +| {[roth2022_security]} USENIX Sec 2022 | 256 domains serving inconsistent security headers | ''security@'' and ''webmaster@''; only 25 of the 256 hosted a ''security.txt'' | **delivery failed for 197 of 256**; 21 non-automatic answers, 7 confirmed and fixed or explained. A second round with 105 hand-curated addresses had 4 failures | — |
 | {[lone2022_sav]} IEEE S&P 2022 | 2,320 network operators, RCT | 8 treatments: direct email (PeeringDB→WHOIS→abuse), national CERT/NIC.br, NOG mailing lists, × nudges | **"none of the notification treatments significantly improved SAV deployment compared to the control group"** | remediation observed in control too | | {[lone2022_sav]} IEEE S&P 2022 | 2,320 network operators, RCT | 8 treatments: direct email (PeeringDB→WHOIS→abuse), national CERT/NIC.br, NOG mailing lists, × nudges | **"none of the notification treatments significantly improved SAV deployment compared to the control group"** | remediation observed in control too |
 | {[sasaki2022_ics]} IEEE S&P 2022 | 160 operators of 317 exposed ICS devices | **manual telephone calls**, email only for scheduling — "the first study to directly contact the organization operating the device" | reached the person in charge for 212 devices; **"50% of the persons in charge … stated that they mitigated or will mitigate"**, and follow-up scans confirmed devices "reduced by 58% when we were able to contact the persons in charge" | 13% decrease among un-notified devices, χ² p<0.0001 | | {[sasaki2022_ics]} IEEE S&P 2022 | 160 operators of 317 exposed ICS devices | **manual telephone calls**, email only for scheduling — "the first study to directly contact the organization operating the device" | reached the person in charge for 212 devices; **"50% of the persons in charge … stated that they mitigated or will mitigate"**, and follow-up scans confirmed devices "reduced by 58% when we were able to contact the persons in charge" | 13% decrease among un-notified devices, χ² p<0.0001 |
 | {[bennett2022_spfail]} IMC 2022 | 6,488 mail-server notifications | ''postmaster@'' per the SMTP spec | 31.6% undelivered; of 4,434 delivered, 512 (12%) opened, 177 (4%) patched, **9 (<1%) between private and public disclosure** | — | | {[bennett2022_spfail]} IMC 2022 | 6,488 mail-server notifications | ''postmaster@'' per the SMTP spec | 31.6% undelivered; of 4,434 delivered, 512 (12%) opened, 177 (4%) patched, **9 (<1%) between private and public disclosure** | — |
 +| {[czybik2023_lazy]} IMC 2023 | **111,951** domain operators with erroneous SPF records, May 2023 | ''postmaster@'' and ''security@'', plus "the contact named in security.txt, if available"; own mail server throttled to one mail a second | **6,931 of 211,018 errors fixed two weeks later — "a success rate of 3.3 %"**; 300 thank-you mails, 3 spam complaints; bounces "large" but uncounted | — |
 +| {[nan2023_spying]} USENIX Sec 2023 | 1,381 developers of IoT companion apps whose privacy policies omit exposed data | developer email from the app store, tracked by a mail-merge tool | after a month: **381 opened, 850 unopened, 150 bounced**; 21 of the 381 openers changed their policy (1.5% of sent); the list was then escalated to Google Play, 360 Store and APKPure | — |
 +| {[blechschmidt2023_hello]} USENIX Sec 2023 | 4,484 domains with obviously misconfigured mail servers | email to the operator | **"at least 2,700 were not delivered and bounced"** (≥60%, the worst delivery figure on this page); 26 non-automated replies; "only 1,076 were still misconfigured in November, which is a decrease of 39.7%" among the domains whose mail was delivered | the 2,700 bounced domains fell 20.7% over the same period — a comparison group the paper reports while saying it "cannot conclude a causal relation", since deliverability itself may drive fixing |
 | {[utz2023_comparing]} PoPETs 2023 | 159,035 domains, 4 privacy issues + 1 security issue | parsed privacy-policy/contact addresses vs RFC 2142 aliases | **87.8% vs 33.8% delivery**; remediation effects significant but small — the paper puts them at "0–1" to "1–2 percentage points" over control (Fisher's exact, Holm–Bonferroni), with a few later-date and generic-alias cells around 3 pp | yes, per issue | | {[utz2023_comparing]} PoPETs 2023 | 159,035 domains, 4 privacy issues + 1 security issue | parsed privacy-policy/contact addresses vs RFC 2142 aliases | **87.8% vs 33.8% delivery**; remediation effects significant but small — the paper puts them at "0–1" to "1–2 percentage points" over control (Fisher's exact, Holm–Bonferroni), with a few later-date and generic-alias cells around 3 pp | yes, per issue |
 | {[qin2024_rov]} NDSS 2024 | 1,012 non-deploying ASes randomised into 6 treatment arms + control; **859 emails sent** | operator email from PeeringDB, falling back to WHOIS; nudge variants (baseline, social norms, authority, reminder, elicitation) and native language | 824 of 859 delivered, **4.07% bounce rate** against the ">50%" it cites from prior work — and **"none of the notification treatments has a significant effect"** (survival analysis; relative risk 0.46–1.35, every confidence interval spanning 1) | 11 of 138 remediated | | {[qin2024_rov]} NDSS 2024 | 1,012 non-deploying ASes randomised into 6 treatment arms + control; **859 emails sent** | operator email from PeeringDB, falling back to WHOIS; nudge variants (baseline, social norms, authority, reminder, elicitation) and native language | 824 of 859 delivered, **4.07% bounce rate** against the ">50%" it cites from prior work — and **"none of the notification treatments has a significant effect"** (survival analysis; relative risk 0.46–1.35, every confidence interval spanning 1) | 11 of 138 remediated |
-| {[elyadmani2025_keys]} IEEE S&P 2025 | 160 organisations leaking cloud-bucket secrets | leaked-file contents, OSINT, disclosure programmes, privacy-policy addresses — routed through a CSIRT partner | **95/160 (59.4%) acted**; only 20 organisations replied at all | — | +| {[li2024_wellinformed]} CCS 2024 | 4,399 emails to developers of apps lacking GDPR runtime privacy notices | developer email from the Google Play listing; IRB "minimal risk" | 4,169 delivered; **821 unique replies, 60 of them human, 13 substantive**; no re-scan, so no remediation figure | — | 
-| {[munteanu2025_catch22]} USENIX Sec 2025 | operators of compromised hosts | Shadowserver Foundation and CERT-BUND ran the campaign | the current worked example of delegating delivery | — |+| {[elyadmani2025_keys]} IEEE S&P 2025| 160 organisations leaking cloud-bucket secrets | leaked-file contents, OSINT, disclosure programmes, privacy-policy addresses — routed through a CSIRT partner | **95/160 (59.4%) acted**; only 20 organisations replied at all | — | 
 +| {[munteanu2025_catch22]} USENIX Sec 2025| operators of compromised hosts | Shadowserver Foundation and CERT-BUND ran the campaign | the current worked example of delegating delivery | — | 
 +| {[li2026_rpkiinvalid]} NDSS 2026* | owners of 1,731 RPKI-invalid prefixes, 1,385 email addresses | abuse and technical contacts from RIR WHOIS records; the mail asked whether the prefix was misconfigured or hijacked, i.e. a validation survey rather than a fix request | **302 of 1,385 undeliverable (21.8%)**; 174 organisations answered, 16.1% of the delivered | — |
  
 ==== The loss is in delivery, not in willingness ==== ==== The loss is in delivery, not in willingness ====
  
-The delivery funnel is the story, and it is why headline remediation rates look so bad. {[stock2016_hey]}: **5.8% of reports received**, but ~40% remediation among those actually read. {[bennett2022_spfail]}: 31.6% undelivered → 12% of the delivered opened → 4% of the openers patched. {[stock2018_didnt]}: 74.4% fixed among Git operators who opened the report, against a 13% control. Operators who read your report largely act on it. Most never read it.+The delivery funnel is the story, and it is why headline remediation rates look so bad. {[stock2016_hey]}: **5.8% of reports received**, but ~40% remediation among those actually read. {[bennett2022_spfail]}: 31.6% undelivered → 12% of the delivered opened → 4% of the openers patched. {[stock2018_didnt]}: 74.4% fixed among Git operators who //viewed// the report, against a 13% control. {[nan2023_spying]} instrumented the same funnel for app developers: 1,381 sent → 150 bounced → 381 opened → 21 acted. {[gilad2017_rpki]} lost 384 of 552 RIR-WHOIS mails to bounces before anyone could read them; {[roth2022_security]} lost 197 of 256 alias mails, and cut the failure to 4 of 105 once the addresses were curated by hand. Operators who read your report largely act on it. Most never read it. The largest fully automated email campaign in the corpus since {[utz2023_comparing]}, {[czybik2023_lazy]}, is the sobering baseline for a fully automated email campaign in 2023: 111,951 mails, 3.3% of the errors fixed two weeks later — and the fix rate tracked how easy the fix was, 5.7% for syntax errors against 1.6% for DNS-lookup limits.
  
 So the highest-leverage thing you can do is not writing a better email. It is: So the highest-leverage thing you can do is not writing a better email. It is:
Line 333: Line 358:
 ^ Factor ^ Finding ^ Source ^ ^ Factor ^ Finding ^ Source ^
 | Sender identity | Legal research group beats computer-science group: 59.7% vs 54% remediation (p<0.05) | {[maass2021_effective]} | | Sender identity | Legal research group beats computer-science group: 59.7% vs 54% remediation (p<0.05) | {[maass2021_effective]} |
-| Framing | Legal-compliance-plus-fine beats plain GDPR beats privacy-harm; survival 50.1% / 56.6% / 69.6%, all differences significant | {[maass2021_effective]} |+| Framing | Legal-compliance-plus-fine beats plain GDPR beats privacy-harm; survival 50.1% / 56.6% / 69.6%, all differences significant. //Survival// throughout this table is the share **still non-compliant**, so lower is better | {[maass2021_effective]} |
 | Framing | **Contradicted.** A fine warning had **no** significant effect once the notification arrived; only //receiving// it mattered | {[utz2023_comparing]} | | Framing | **Contradicted.** A fine warning had **no** significant effect once the notification arrived; only //receiving// it mattered | {[utz2023_comparing]} |
 | Medium | A **postal letter** beats email: survival 55.6% vs 66.3% (p<0.0001), "increasing the remediation rate by between 3.9 and 17.9 percentage points (mean: 11.1)" — for "around 5000 € on domestic postage" across 2,660 letters | {[maass2021_effective]} | | Medium | A **postal letter** beats email: survival 55.6% vs 66.3% (p<0.0001), "increasing the remediation rate by between 3.9 and 17.9 percentage points (mean: 11.1)" — for "around 5000 € on domestic postage" across 2,660 letters | {[maass2021_effective]} |
Line 349: Line 374:
 | Persistence | 7 months later, 3.5% of remediated sites had regressed — "long-term effectiveness of approximately 95%" | {[maass2021_effective]} | | Persistence | 7 months later, 3.5% of remediated sites had regressed — "long-term effectiveness of approximately 95%" | {[maass2021_effective]} |
 | Public disclosure | Private notification "made little difference"; the **public CVE 60 days later** correlated with a much larger drop in vulnerable servers | {[bennett2022_spfail]} | | Public disclosure | Private notification "made little difference"; the **public CVE 60 days later** correlated with a much larger drop in vulnerable servers | {[bennett2022_spfail]} |
 +| Address source | Generic aliases: delivery failed for 197 of 256 domains. Hand-curated addresses for the 184 still-affected sites: 4 failures of 105 | {[roth2022_security]} |
 +| Fix difficulty | Errors an operator can fix by editing one line were fixed at 5.7%; errors needing a change at an external provider at 1.6% — same template; the authors' explanation ("We assume that these are often non-trivial to fix"), not a tested one | {[czybik2023_lazy]} |
 +| Escalation | When 1.5% of notified developers acted, the authors handed the list to the app stores; Google Play "responded to our request quickly", two other stores did not | {[nan2023_spying]} |
  
 Read the two contradictions as the state of the art rather than as noise. **Message-content tuning is a dead end**: three studies varied wording, format, signing, tone and translation, and the effects are small, inconsistent, or vanish under multiple-comparison correction. **Sender identity, medium and reachability are the real variables** — and even those may not survive a control group, which is what {[lone2022_sav]} demonstrated on the one population where somebody randomised properly. Read the two contradictions as the state of the art rather than as noise. **Message-content tuning is a dead end**: three studies varied wording, format, signing, tone and translation, and the effects are small, inconsistent, or vanish under multiple-comparison correction. **Sender identity, medium and reachability are the real variables** — and even those may not survive a control group, which is what {[lone2022_sav]} demonstrated on the one population where somebody randomised properly.
Line 387: Line 415:
   * **False positives reach real people.** Twelve IPv6 contacts in {[li2016_youve]} rebutted the vulnerability claim. Your detector's precision is now somebody else's inbox; measure it before, not after.   * **False positives reach real people.** Twelve IPv6 contacts in {[li2016_youve]} rebutted the vulnerability claim. Your detector's precision is now somebody else's inbox; measure it before, not after.
   * **Ethics review is not automatic and not uniform.** {[stock2016_hey]} recorded that its institutions "neither mandate nor provide an IRB approval before conducting such experiments". {[maass2021_effective]} got approval from the ethics committees of two of three institutions and a dean's approval from the third. See [[Practices:Ethics]], and {[hantke2024_redlines]} for what server operators themselves consider acceptable.   * **Ethics review is not automatic and not uniform.** {[stock2016_hey]} recorded that its institutions "neither mandate nor provide an IRB approval before conducting such experiments". {[maass2021_effective]} got approval from the ethics committees of two of three institutions and a dean's approval from the third. See [[Practices:Ethics]], and {[hantke2024_redlines]} for what server operators themselves consider acceptable.
 +
 +==== The receiving end ====
 +
 +Four papers look at what happens after a report arrives — {[stivala2026_behind]} above, on how hosting providers triage, and three more here — and none of them is encouraging about either the unsolicited individual email or the generic disclosure form. {[bijmans2026_tickets]} analysed 1.3 million abuse reports seized by Dutch law enforcement from one hosting provider with a bulletproof reputation: **2.6% of reports were ever linked to a notification to the customer**. The rate was not about the abuse; it was about the reporter. Netcraft's phishing reports led to a customer notification 72% of the time and Spamhaus listings 60–84%, while "individual abuse reporting is often easily ignored" — one DMCA service filed 126,085 reports that produced four notifications, and one honeypot operator's 30,467 automated reports produced none. The authors are explicit that one abusive provider does not generalise; the direction of the effect is still the point: **a report is acted on in proportion to what the reporter can do to the recipient's business**, which a research group cannot do at all. On the vendor side, {[hastings2016_weakkeys]} followed the 2012 disclosure of weak-key generation to 61 device vendors — of the 37 with weak RSA keys, 5 issued a public advisory, 11 answered privately, 3 sent an auto-reply, 18 nothing — and found that "vendor notification, positive vendor responses, and even vendor-produced public security advisories appear to have little correlation with end-user security": vulnerable populations kept growing for years at vendors that had published an advisory. And when the recipient does have a formal channel, it may be the wrong shape: {[bhattacharya2026_asi]} briefed 58 of the 100 most popular services about account-security flaws through their disclosure programmes, got 38 acknowledgements and 25 substantive replies, most classified "Not applicable" or "Out of scope", and notes that "HackerOne rate limits disclosures, which slowed our disclosure process immensely"; {[roth2022_security]} found that "many of those that answered instructed us to contact HackerOne" although the message never used the word vulnerability.
 +
 +==== Deciding not to run a campaign ====
 +
 +Two recent papers reasoned their way out of a campaign in print, and both are worth citing when you do the same. {[rao2024_unfiltered]} found domains whose cloud email filters could be bypassed, cited {[bennett2022_spfail]} ("over 80% of the domains contacted were unresponsive"), and "elected to work directly and closely with filtering service providers to update their documentation, notify their customers (with whom they do have an existing business relationship) and resolve the issues identified" — seven filtering vendors instead of the 1,262 misconfigured domains. {[ryan2023_passivessh]} recovered 189 unique SSH host keys from a scan of millions of devices, disclosed to four manufacturers and CERT/CC, and wrote: "We considered notifying operators of affected devices whose keys we had recovered, but we determined this would be infeasible." Neither is a shortcut; both state the alternative they took and why.
  
 ==== Covert notification ==== ==== Covert notification ====
Line 407: Line 443:
   - **Your detector's precision**, because false positives went to real inboxes.   - **Your detector's precision**, because false positives went to real inboxes.
   - **The disclosure timeline you followed** and which convention it came from.   - **The disclosure timeline you followed** and which convention it came from.
-  - **The message itself**, in an appendix or [[Artifacts|artefact]]. Every study above that reports a framing effect published its text; you cannot replicate a framing result without it.+  - **The message itself**, in an appendix or [[:artifacts|artefact]]. Every study above that reports a framing effect published its text; you cannot replicate a framing result without it.
   - **The ethics decision**: review body and outcome, whether the study was covert, how you debriefed, the opt-out mechanism, and how many opted out.   - **The ethics decision**: review body and outcome, whether the study was covert, how you debriefed, the opt-out mechanism, and how many opted out.
   - **Everything hostile that happened.** Legal threats, complaints to your institution, suspended customers. This is the part later researchers most need and the part most often missing.   - **Everything hostile that happened.** Legal threats, complaints to your institution, suspended customers. This is the part later researchers most need and the part most often missing.
Line 413: Line 449:
 ===== Papers to read first ===== ===== Papers to read first =====
  
-If you read four: {[stock2016_hey]} for the channel survey and the delivery funnel, {[li2016_youve]} for the arm-by-arm comparison and the decay curve, {[maass2021_effective]} for the only large factorial design on framing and medium, and {[lone2022_sav]} for the control group that undermines the rest. Then {[utz2023_comparing]} for privacy versus security and the delivery figures that supersede the 2016 channel advice, {[stock2018_didnt]} for why message tuning is a dead end, {[li2016_remedying]} for what a platform can do that you cannot, and {[stivala2026_behind]} for the receiving end.+If you read four: {[stock2016_hey]} for the channel survey and the delivery funnel, {[li2016_youve]} for the arm-by-arm comparison and the decay curve, {[maass2021_effective]} for the only large factorial design on framing and medium, and {[lone2022_sav]} for the control group that undermines the rest. Then {[utz2023_comparing]} for privacy versus security and the delivery figures that supersede the 2016 channel advice, {[stock2018_didnt]} for why message tuning is a dead end, {[li2016_remedying]} for what a platform can do that you cannot, {[czybik2023_lazy]} for what a 100,000-mail automated campaign actually yields in 2023, and {[stivala2026_behind]} and {[bijmans2026_tickets]} for the receiving end.
  
 ===== Related pages ===== ===== Related pages =====
Line 421: Line 457:
   * [[Practices:Public relations]] — the other post-publication channel.   * [[Practices:Public relations]] — the other post-publication channel.
   * [[Statistics:Hypothesis testing]] and [[Statistics:Pvalue corrections]] — for the arm comparisons.   * [[Statistics:Hypothesis testing]] and [[Statistics:Pvalue corrections]] — for the arm comparisons.
-  * [[Artifacts]] — publishing the notification text and the contact-discovery code.+  * [[:artifacts|Artifacts]] — publishing the notification text and the contact-discovery code.
   * [[Design:Website selection]] — your notified population is your sample, with the same biases.   * [[Design:Website selection]] — your notified population is your sample, with the same biases.
  
Line 429: Line 465:
  
   * **Seven venues only.** EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS are absent, and CHI/SOUPS are where a good deal of the operator-facing usable-security work appears. NDSS 2016 and NDSS 2018 full text was not retrieved at all, which is why {[stock2018_didnt]} — the canonical follow-up study — was read from the publisher's PDF rather than from the corpus.   * **Seven venues only.** EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS are absent, and CHI/SOUPS are where a good deal of the operator-facing usable-security work appears. NDSS 2016 and NDSS 2018 full text was not retrieved at all, which is why {[stock2018_didnt]} — the canonical follow-up study — was read from the publisher's PDF rather than from the corpus.
-  * ''ethics.notifiedAffectedParties'' **agreed with an independent re-extraction on 67% of papers**, so treat the percentages as accurate to a few points, not to the decimal. ''ethics.disclosureDetail'' is free text capped at 20 words and is reported below only as folded families with the residue printed+  * ''ethics.notifiedAffectedParties'' **agreed with an independent re-extraction on 67% of papers**, so treat every percentage above as accurate to a few points, not to the decimal.((That 67% was measured on 100 papers of the previous, 4,322-paper extraction run and has not been re-measured on the current corpus. Treat it as the right order of magnitude.)) ''ethics.disclosureDetail'' is free text capped at 20 words, so it is reported only as folded families and only as a ranking; the fold rules and the unmapped residue are on the provenance page
-  * **The 22 campaigns are hand-classified, not a census.** A regex over 5,869 full texts produced 179 candidates; 22 were classified as campaigns, 16 rejected with a stated reason, and **144 were never read**. The table above is a curated reading list, and the unreviewed residue is published in full.+  * **The 32 campaigns are hand-classified, not a census.** Two sentence-level rules run over 5,869 full-text files (ten more than the 5,859 extraction records — a few papers have text but no record). The narrow rule produced 179 candidates: 28 campaigns and 151 rejections with a stated reason. A wider recall probeadded on 2026-09-04 after the narrow rule was found to miss a 111,951-mail campaign, produced 78 more: 3 campaigns and 75 rejections. One campaign, {[roth2022_security]}, was caught by neither rule and found through a full-text search for ''security.txt''; three rejections carry over from a wider first-pass scan. Nothing is unreviewed, **but most rejections are screens, not reads**: 126 on every sentence carrying a notification verb, 75 on the probe-matching sentences alone, 28 on a full read of the relevant passages or paper; the provenance page says whichA third rule would find more; the table above is a curated reading list, not a census.
  
 The complete query log, the report script with its unedited output, the folding rules with their unmapped residue, the spot-checked quotes, and every external source that was rejected are on **[[provenance:practices:notifying_websites]]**. Corpus-wide caveats are on [[Literature:Corpus]]. The complete query log, the report script with its unedited output, the folding rules with their unmapped residue, the spot-checked quotes, and every external source that was rejected are on **[[provenance:practices:notifying_websites]]**. Corpus-wide caveats are on [[Literature:Corpus]].
practices/notifying_websites.1786651138.txt.gz · Last modified: by karel.kubicek.claude