| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| practices:notifying_websites [2026/08/21 14:50] – Boxes: <wrap> renders a span, use uppercase <WRAP>; one box per list, not per bullet karel.kubicek.claude | practices:notifying_websites [2026/09/04 13:10] (current) – Generic review: lead box no longer says all network RCTs failed; Extended Hell(o) row gets its funnel and comparison group; outcome kinds named; security.txt series disagreement stated; three figures corrected; methodology caveat rebalanced. Authored by C karel.kubicek.claude |
|---|
| |
| <WRAP important> | <WRAP important> |
| The single most consequential design decision on this page is the **control group**. Sites get fixed for reasons that have nothing to do with you — automatic updates, unrelated maintenance, going offline. Every notification study without a control arm has attributed some of that background remediation to its own notifications — and **both randomised, control-arm experiments on network operators in this corpus found no significant effect from any treatment**, including treatments that observational studies had reported as effective {[lone2022_sav,qin2024_rov]}. Randomise, hold back an arm, and report both. Note what that does and does not mean: {[maass2021_effective]} is also a randomised controlled experiment — a full factorial design with a control group — and it measured 56.6% against 9.2% on German website owners. The lesson is not that notification never works. It is that a study without a control arm cannot tell you which of those two cases it is in. | The single most consequential design decision on this page is the **control group**. Sites get fixed for reasons that have nothing to do with you — automatic updates, unrelated maintenance, going offline. Every notification study without a control arm has attributed some of that background remediation to its own notifications — and **the two most recent randomised, control-arm experiments on network operators in this corpus — on deploying SAV and ROV, expensive and low-visibility configuration changes — found no significant effect from any treatment**, including treatments that observational studies had reported as effective {[lone2022_sav,qin2024_rov]}. The earlier randomised experiments on network operators, on Heartbleed patching and firewall misconfiguration {[durumeric2014_heartbleed,li2016_youve]}, did find one. Randomise, hold back an arm, and report both. Note what that does and does not mean: {[maass2021_effective]} is also a randomised controlled experiment — a full factorial design with a control group — and it measured 56.6% against 9.2% on German website owners. The lesson is not that notification never works. It is that a study without a control arm cannot tell you which of those two cases it is in. |
| </WRAP> | </WRAP> |
| |
| | **crawled ∧ assessed a law** | **123** | 35.0% | 14.6% | **14.6%** | 15.4% | 20.3% | | | **crawled ∧ assessed a law** | **123** | 35.0% | 14.6% | **14.6%** | 15.4% | 20.3% | |
| |
| Two things to take from this table. **Network measurement is ahead of web measurement**: scanning papers notify at 61.0% (''yes''+''partial'') against 43.5% for crawling papers, and are less than half as likely to say outright that they did not (4.2% vs 9.4%), as well as meaningfully less likely to leave it unanswered (24.0% vs 31.1%). The scanning community built the norm first, largely because Internet-wide scanning provoked complaints that forced the question. **Legal-compliance papers state a position most often** (only 12.0% ''not-stated''), and the bottom row is the population closest to whoever is reading this: 123 papers that both crawled the web and assessed a law, i.e. found a compliance violation across many sites. **14.6% of them say outright that they did not notify** — more than three times the corpus-wide 4.3%. That is usually a decision rather than negligence. A GDPR-violation finding across thousands of sites is the case where authors most often decide, and say, that individual notification is the wrong instrument, and route the finding to a regulator or to publication instead. If that is your situation, //deciding not to notify is a defensible position that you have to argue in the paper//, not an omission you can leave to the reader. | Two things to take from this table. **Network measurement is ahead of web measurement**: scanning papers notify at 61.0% (''yes''+''partial'') against 43.5% for crawling papers, and are less than half as likely to say outright that they did not (4.2% vs 9.4%), as well as meaningfully less likely to leave it unanswered (24.0% vs 31.1%). The scanning community built the norm first, largely because Internet-wide scanning provoked complaints that forced the question. **Legal-compliance papers state a position most often** (only 12.0% ''not-stated''), and the bottom row is the population closest to whoever is reading this: 123 papers that both crawled the web and assessed a law, i.e. found a compliance violation across many sites. **14.6% of them say outright that they did not notify** — more than three times the corpus-wide 4.3%. That is usually a decision rather than negligence. A GDPR-violation finding across thousands of sites is the case where authors most often decide, and say, that individual notification is the wrong instrument, and route the finding to a regulator or to publication instead. If that is your situation, //deciding not to notify is a defensible position that you have to argue in the paper//, not an omission you can leave to the reader (two worked examples are under //Deciding not to run a campaign// below). |
| |
| The direction of travel is unambiguous: | The direction of travel is unambiguous: |
| |
| ^ Channel ^ How you get it ^ What it reaches ^ Status and evidence ^ | ^ Channel ^ How you get it ^ What it reaches ^ Status and evidence ^ |
| | **''security.txt''** (RFC 9116) | ''GET https://<domain>/.well-known/security.txt'' | the security team, if there is one | RFC 9116, Informational, April 2022 — current, not obsoleted.((Verified against [[https://www.rfc-editor.org/rfc/rfc9116.txt|rfc-editor.org/rfc/rfc9116.txt]] on 2026-08-13: "Category: Informational … April 2022". Three errata exist, none touching the field list.)) Adoption is the problem: 11–16% of the Alexa top 100, 8–10% of the top 1K, "3–4% for the top 10K sites, and only a percent for the top 100K" {[poteat2021_securitytxt]}. **That measurement is from 2021 and its ranking frame no longer exists** — Alexa was discontinued in 2022 (see [[Design:Website selection]]), the paper predates RFC 9116, and **nothing in this corpus re-measures security.txt adoption since**. Assume it has risen; you have no citeable current figure | | | **''security.txt''** (RFC 9116) | ''GET https://<domain>/.well-known/security.txt'' | the security team, if there is one | RFC 9116, Informational, April 2022 — current, not obsoleted.((Verified against [[https://www.rfc-editor.org/rfc/rfc9116.txt|rfc-editor.org/rfc/rfc9116.txt]] on 2026-08-13: "Category: Informational … April 2022". Three errata exist, none touching the field list.)) Adoption is the problem, and it is steep in rank: 11–16% of the Alexa top 100, 8–10% of the top 1K, "3–4% for the top 10K sites, and only a percent for the top 100K" {[poteat2021_securitytxt]}. That is a 2021 measurement on a ranking that was discontinued in 2022 (see [[Design:Website selection]]) and predates RFC 9116; **no paper in the seven corpus venues re-measures adoption since.** Outside them, two do: 0.49% of the Tranco top million and 1.6% of the top 100K in October–November 2021 {[findlay2022_securitytxt]}, and 42 successful weekly scans of the Tranco top million over the 55 weeks from December 2021 to January 2023 showing 32.0→34.0% (top 100), 16.1→18.8% (top 1K), 7.9→9.8% (top 10K), 2.4→3.2% (top 100K) and 0.7→1.0% (top 1M) {[hilbig2023_securitytxt]}.((Both are outside the corpus's seven venues: MADweb is an NDSS workshop, DTRAP an ACM journal. A vendor scan of gTLD zone files reports 573,123 of 241 million domains (0.24%) with a ''security.txt'' in early 2026, up from 0.05% in 2021 — [[https://blog.iotdef.com/the-state-of-security-txt-adoption-an-analysis-of-240-million-domains-in-2026/|blog.iotdef.com]], fetched 2026-09-04; methodology stated, data and code not published, so treat it as an order of magnitude, not a citation.)) The two peer-reviewed series disagree by about 2× at the top ranks (Alexa against Tranco, and a Kaplan–Meier deployment estimate against weekly snapshots); the corpus's own incidental counts sit with the Tranco series: 23 domains in {[drakonakis2020_cookie]} (2020, "the most ineffective" of its four channels; no denominator given), 25 of 256 in {[roth2022_security]} (2022), 34 of the 100 most popular services in {[bhattacharya2026_asi]} (2025). Six corpus papers that notified **used** it as one channel among several — {[drakonakis2020_cookie,ghasemisharif2022_saat,roth2022_security,czybik2023_lazy,innocenti2025_only,bhattacharya2026_asi]} — and **none separates its delivery or response from the other channels**, so whether a listed contact actually answers is still unmeasured; {[hilbig2023_securitytxt]} names exactly that as future work | |
| | **RIR abuse contact** | RIPE ''abuse-c:'' / ARIN Abuse POC, resolved from the IP; RIPEstat's abuse-contact-finder covers all five RIRs over HTTPS | the **hosting provider or CDN**, not the site owner | The one source that scales — free and with no meaningful rate limit. ARIN verifies its Abuse POC annually;((ARIN NRPM §3.6: "Each of the following Points of Contact are to be verified annually … Admin, Tech, NOC, Abuse", [[https://www.arin.net/participate/policy/nrpm/|arin.net]], fetched 2026-08-13. RIPE states it works to keep abuse contacts valid but no validation cadence could be found on a RIPE primary source on 2026-08-13 — do not assume annual.)) RIPE states it keeps abuse contacts valid but publishes no cadence that could be verified from a RIPE source. Still the workhorse: "WHOIS was the most frequently mentioned channel (seven HPOs)" among 24 interviewed providers in 2026, who "confirmed that their contact is available via WHOIS, for example through ARIN or RIPE databases" {[stivala2026_behind]}. Note this is the //network// registry, which GDPR redaction did not touch — not the //domain// registrant record, which it did | | | **RIR abuse contact** | RIPE ''abuse-c:'' / ARIN Abuse POC, resolved from the IP; RIPEstat's abuse-contact-finder covers all five RIRs over HTTPS | the **hosting provider or CDN**, not the site owner | The one source that scales — free and with no meaningful rate limit. ARIN verifies its Abuse POC annually;((ARIN NRPM §3.6: "Each of the following Points of Contact are to be verified annually … Admin, Tech, NOC, Abuse", [[https://www.arin.net/participate/policy/nrpm/|arin.net]], fetched 2026-08-13. RIPE states it works to keep abuse contacts valid but no validation cadence could be found on a RIPE primary source on 2026-08-13 — do not assume annual.)) RIPE states it keeps abuse contacts valid but publishes no cadence that could be verified from a RIPE source. Still the workhorse: "WHOIS was the most frequently mentioned channel (seven HPOs)" among 24 interviewed providers in 2026, who "confirmed that their contact is available via WHOIS, for example through ARIN or RIPE databases" {[stivala2026_behind]}. Note this is the //network// registry, which GDPR redaction did not touch — not the //domain// registrant record, which it did. It is not clean either: 384 of 552 RIR-WHOIS mails bounced in {[gilad2017_rpki]} ("whois entries are often outdated"), and 302 of 1,385 (21.8%) were undeliverable in {[li2026_rpkiinvalid]}. For DNS resolvers there is one more machine-readable source, the ''RNAME'' of the zone's SOA record; {[deccio2020_closeddoors]} used it to reach 43 administrators and heard back from five, three of whom the authors already knew | |
| | **PeeringDB technical contact** | PeeringDB API, per AS | the network operator | Preferred over WHOIS by {[lone2022_sav]}: "We preferred peeringDB because it has been used in previous studies and they found the database up-to-date" — i.e. it is relaying prior work's assessment, not measuring it | | | **PeeringDB technical contact** | PeeringDB API, per AS | the network operator | Preferred over WHOIS by {[lone2022_sav]}: "We preferred peeringDB because it has been used in previous studies and they found the database up-to-date" — i.e. it is relaying prior work's assessment, not measuring it | |
| | **The site's own imprint / privacy-policy contact** | read it off the page | the legally responsible party — often the actual decision-maker | The highest-yield channel measured. Addresses parsed from privacy-policy and contact pages achieved **87.8% delivery against 33.8% for RFC 2142 aliases** {[utz2023_comparing]}; 9 of 11 organisations reached via a privacy-policy address resolved the issue {[elyadmani2025_keys]}. {[maass2021_effective]} collected German ''Impressum'' addresses by hand, three researchers per site | | | **The site's own imprint / privacy-policy contact** | read it off the page | the legally responsible party — often the actual decision-maker | The highest-yield channel measured. Addresses parsed from privacy-policy and contact pages achieved **87.8% delivery against 33.8% for RFC 2142 aliases** {[utz2023_comparing]}; 9 of 11 organisations reached via a privacy-policy address resolved the issue {[elyadmani2025_keys]}. {[maass2021_effective]} collected German ''Impressum'' addresses by hand, three researchers per site | |
| | **Established notification operators** | Shadowserver Foundation, CERT-BUND and equivalents | operators who already trust the sender | Shadowserver still sends free daily reports to 201 national CSIRTs across 175 countries.((Fetched from [[https://www.shadowserver.org/what-we-do/network-reporting/|shadowserver.org]] on 2026-08-13.)) {[munteanu2025_catch22]} ran its whole campaign this way and thanks Shadowserver and CERT-BUND for it — the current worked example | | | **Established notification operators** | Shadowserver Foundation, CERT-BUND and equivalents | operators who already trust the sender | Shadowserver still sends free daily reports to 201 national CSIRTs across 175 countries.((Fetched from [[https://www.shadowserver.org/what-we-do/network-reporting/|shadowserver.org]] on 2026-08-13.)) {[munteanu2025_catch22]} ran its whole campaign this way and thanks Shadowserver and CERT-BUND for it — the current worked example | |
| | **Platform consoles** | the platform's own channel to its registered operators | only operators who pre-registered | The strongest single effect in the literature, and the least available to you: a Search Console message lifted remediation to 82.4% / 76.8% against 54.6% / 43.4% otherwise {[li2016_remedying]}. Only 22–32% of affected sites had a registered webmaster, and you are not Google | | | **Platform consoles** | the platform's own channel to its registered operators | only operators who pre-registered | The strongest single effect in the literature, and the least available to you: a Search Console message lifted remediation to 82.4% / 76.8% against 54.6% / 43.4% otherwise {[li2016_remedying]}. Only 22–32% of affected sites had a registered webmaster, and you are not Google | |
| | | **A bulk report to the platform** | the store's or platform's abuse channel, with the whole list attached | the platform, which removes or contacts | Not a notification of operators at all, but the instrument that produced the highest removal figures in this corpus: 813 of 1,095 apps reported to Google were removed {[roundy2020_creepware]}; 1,316 of 1,804 GPTs reported to OpenAI were gone within two weeks {[shen2025_gptracker]}. {[nan2023_spying]} fell back to it after a 1.5% developer response; {[edu2022_alexa]} reported 675 skills to Amazon alongside the developers. It reaches only what the platform hosts, and the platform decides what counts as a violation | |
| | **HackerOne Disclosure Assistance** | ''hackerone.com/disclosure-assistance'' | organisations with no disclosure policy at all | Newer than the literature: HackerOne "will work with friendly hackers on a **best-effort basis**" to verify a bug, find someone at the affected organisation and relay it.((Fetched from [[https://docs.hackerone.com/en/articles/8466632-disclosure-assistance|docs.hackerone.com]] on 2026-08-13; page dated 11 June 2024.)) It updates {[stock2016_hey]}, which discarded reward programmes because they "usually only accept and forward reports for their customers" — that premise no longer holds for HackerOne. But it is a per-report human service gated on having exhausted other options, so it is **not** a bulk channel | | | **HackerOne Disclosure Assistance** | ''hackerone.com/disclosure-assistance'' | organisations with no disclosure policy at all | Newer than the literature: HackerOne "will work with friendly hackers on a **best-effort basis**" to verify a bug, find someone at the affected organisation and relay it.((Fetched from [[https://docs.hackerone.com/en/articles/8466632-disclosure-assistance|docs.hackerone.com]] on 2026-08-13; page dated 11 June 2024.)) It updates {[stock2016_hey]}, which discarded reward programmes because they "usually only accept and forward reports for their customers" — that premise no longer holds for HackerOne. But it is a per-report human service gated on having exhausted other options, so it is **not** a bulk channel | |
| |
| </file> | </file> |
| |
| Real output, run 2026-08-13 on five hand-picked domains — **not a sample, and not representative**: these are large, well-resourced organisations, so ''security.txt'' coverage here is an order of magnitude above the 3–4% {[poteat2021_securitytxt]} measured at the top 10K. | Real output, run 2026-08-13 on five hand-picked domains — **not a sample, and not representative**: these are large, well-resourced organisations, so ''security.txt'' coverage here is well above any measured tier — 32–34% for the top 100 in {[hilbig2023_securitytxt]}, 3–4% at the top 10K in {[poteat2021_securitytxt]}. |
| |
| <code> | <code> |
| {[sasaki2022_ics]} collects the comparison itself, and it is the shortest statement of the range: "its remediation rate was approximately 18% … Our remediation rate is higher than most previous notification experiments: it was approximately 40% for cross-site scripting and a WordPress vulnerability, 33%–42% for different WordPress vulnerability, and less than 20% for DNS zone poisoning. The only campaigns that reported similar remediation rates were on publicly accessible Git repositories (78%–81%) and Heartbleed (approximately 40%–90%)." | {[sasaki2022_ics]} collects the comparison itself, and it is the shortest statement of the range: "its remediation rate was approximately 18% … Our remediation rate is higher than most previous notification experiments: it was approximately 40% for cross-site scripting and a WordPress vulnerability, 33%–42% for different WordPress vulnerability, and less than 20% for DNS zone poisoning. The only campaigns that reported similar remediation rates were on publicly accessible Git repositories (78%–81%) and Heartbleed (approximately 40%–90%)." |
| |
| These are the campaigns in the corpus that reported a quotable outcome, each with **its own denominator** — they are not comparable to each other, because "response", "remediation" and "fix" are defined differently in each and the populations are wildly different. Plan against the range, not the mean. The hand pass classified 22 campaigns in total; 16 are below, and the six not shown reported a volume sent but no response or remediation figure. One row below — {[munteanu2025_catch22]} — is here for its channel rather than a rate, because delegating a campaign to an established notification operator is the pattern, not the number. They are listed on [[provenance:practices:notifying_websites]], along with the one row below — {[stock2018_didnt]} — whose full text is missing from the corpus and was read from the publisher's PDF instead. | These are the campaigns in the corpus that reported a quotable outcome, each with **its own denominator** — they are not comparable to each other, because "response", "remediation" and "fix" are defined differently in each and the populations are wildly different. Plan against the range, not the mean. The Outcome column mixes four different quantities and says which each row reports: **remediation** measured by re-scan (Maass, Czybik), **response** counts (Nguyen, Li 2024), **delivery** alone (Roth 2022, the 2026 validation survey), and **removal by a platform** (Edu). Pick the rows that measure what you will measure. The hand pass classified 32 campaigns in total; 26 are below, and the six not shown reported a volume sent but no response or remediation figure. Two rows — {[munteanu2025_catch22]} and {[liao2016_seeking]} — are here for their channel rather than a rate, because routing a campaign through an established notification operator or a national CERT is the pattern, not the number. The six are listed on [[provenance:practices:notifying_websites]], along with the one row below — {[stock2018_didnt]} — whose full text is missing from the corpus and was read from the publisher's PDF instead. Rows marked * are from the provisional 2025–2026 slice. |
| |
| ^ Study ^ Notified ^ Channel ^ Outcome, in the paper's own terms ^ Control ^ | ^ Study ^ Notified ^ Channel ^ Outcome, in the paper's own terms ^ Control ^ |
| | {[li2016_youve]} USENIX Sec 2016 | 2,563 ICS + 3,536 IPv6 + 5,960 amplifier contacts | WHOIS abuse contact, national CERTs, US-CERT | "at most 18% of the population remediating" under the best regimen; direct verbose 9.8% vs national CERT 3.1% vs US-CERT 1.4% (IPv6, two days) | US-CERT arm statistically indistinguishable from control | | | {[li2016_youve]} USENIX Sec 2016 | 2,563 ICS + 3,536 IPv6 + 5,960 amplifier contacts | WHOIS abuse contact, national CERTs, US-CERT | "at most 18% of the population remediating" under the best regimen; direct verbose 9.8% vs national CERT 3.1% vs US-CERT 1.4% (IPv6, two days) | US-CERT arm statistically indistinguishable from control | |
| | {[li2016_remedying]} TheWebConf 2016 | 760,935 hijacking incidents | browser interstitial, search warning, Search Console message, WHOIS admin email | 59.5% of incidents resolved over 11 months; **82.4% / 76.8% where a Search Console alert reached a pre-registered webmaster** vs 54.6% / 43.4% otherwise | — (channel comparison) | | | {[li2016_remedying]} TheWebConf 2016 | 760,935 hijacking incidents | browser interstitial, search warning, Search Console message, WHOIS admin email | 59.5% of incidents resolved over 11 months; **82.4% / 76.8% where a Search Console alert reached a pre-registered webmaster** vs 54.6% / 43.4% otherwise | — (channel comparison) | |
| | | {[liao2016_seeking]} IEEE S&P 2016 | over 120 FQDNs to US-CERT and 136 to CCERT (China) | two national CERTs; CCERT then "notified all related organizations" — "it is difficult for us to directly contact the victims" | "27 responded and fixed their problems"; the number CCERT notified is not stated, so no rate | — | |
| | | {[gilad2017_rpki]} NDSS 2017 | 552 victims and offenders of RPKI misconfiguration, over six months | RIR WHOIS contact, sent by the ROAlert system | **only 168 of 552 emails did not bounce**; "Over 42%" of bad-ROA alerts fixed a month later; 19% for loose-ROA alerts; 52 administrators engaged, 40 reported fixing | "about 15%" fixed among operators ROAlert could not reach — a comparison group the paper reports, not a randomised control | |
| | {[stock2018_didnt]} NDSS 2018 | >24,000 domains, seven arms of ~4,000 | six email variants (plain, HTML, tracking, mailbot, S/MIME, friendly tone) | notified 24% (Git) and 17% (WordPress) fixed; **74.4% / 33.3% among those who //viewed// the report** | 13% (Git), 14% (WordPress) | | | {[stock2018_didnt]} NDSS 2018 | >24,000 domains, seven arms of ~4,000 | six email variants (plain, HTML, tracking, mailbot, S/MIME, friendly tone) | notified 24% (Git) and 17% (WordPress) fixed; **74.4% / 33.3% among those who //viewed// the report** | 13% (Git), 14% (WordPress) | |
| | {[cetin2019_cleaning]} NDSS 2019 | ISP customers with Mirai infections | ISP walled garden (quarantine + notification) vs email-only vs nothing | walled garden "remediates 92% of the infections within 14 days"; **"Email-only notifications have no observable impact compared to a control group"** | **"The control group achieved the lowest cleanup rate (74%)"** with no notification at all — the cautionary figure on this whole page | | | {[cetin2019_cleaning]} NDSS 2019 | ISP customers with Mirai infections | ISP walled garden (quarantine + notification) vs email-only vs nothing | walled garden "remediates 92% of the infections within 14 days"; **"Email-only notifications have no observable impact compared to a control group"** | **"The control group achieved the lowest cleanup rate (74%)"** with no notification at all — the cautionary figure on this whole page | |
| | | {[roth2020complex]} NDSS 2020 | 2,699 sites with a broken or double-framing-prone ''X-Frame-Options'' header | RFC 2142 aliases (info, security, webmaster) plus WHOIS contact, from a researcher's own address | "most emails bounced"; **117 non-automated responses**; sites deploying ''frame-ancestors'' rose from 511 to 554 (+43) in eight days, 62 responders said they would deploy it | — | |
| | {[maass2021_effective]} USENIX Sec 2021 | 4,594 German site owners, 18 arms + control | postal letter and email, addresses read by hand from each ''Impressum'' | **"56.6 % of all notified operators remediating within two months"**; 76.3% for a legal-research-group letter citing fines, 33.9% for a computer-science email citing privacy | **9.2%** | | | {[maass2021_effective]} USENIX Sec 2021 | 4,594 German site owners, 18 arms + control | postal letter and email, addresses read by hand from each ''Impressum'' | **"56.6 % of all notified operators remediating within two months"**; 76.3% for a legal-research-group letter citing fines, 33.9% for a computer-science email citing privacy | **9.2%** | |
| | {[nguyen2021_sharefirst]} USENIX Sec 2021 | 11,914 app developers | contact address from the Play Store listing | **448 responses (3.8%)** | — | | | {[nguyen2021_sharefirst]} USENIX Sec 2021 | 11,914 app developers | contact address from the Play Store listing | **448 responses (3.8%)** | — | |
| | {[squarcina2021_subdomain]} USENIX Sec 2021 | sites with dangling subdomain records | direct contact vs the authors' national CERT | direct 31% / 22% fixed vs national CERT 10% / 14% | — | | | {[squarcina2021_subdomain]} USENIX Sec 2021 | sites with dangling subdomain records | direct contact vs the authors' national CERT | direct 31% / 22% fixed vs national CERT 10% / 14% | — | |
| | | {[edu2022_alexa]} TheWebConf 2022 | 675 Alexa skills reported to Amazon and to "all affected developers" — "whenever we have their contact details"; 246 with broken traceability | developer contact details where available, plus the platform | of the 246, **111 (45.12%) "no longer pose a threat"** about a year later (45 removed, 24 lost the permission, 41 fixed — the parts sum to 110); 107 (43.5%) still broken. No delivery or reply counts | — | |
| | | {[roth2022_security]} USENIX Sec 2022 | 256 domains serving inconsistent security headers | ''security@'' and ''webmaster@''; only 25 of the 256 hosted a ''security.txt'' | **delivery failed for 197 of 256**; 21 non-automatic answers, 7 confirmed and fixed or explained. A second round with 105 hand-curated addresses had 4 failures | — | |
| | {[lone2022_sav]} IEEE S&P 2022 | 2,320 network operators, RCT | 8 treatments: direct email (PeeringDB→WHOIS→abuse), national CERT/NIC.br, NOG mailing lists, × nudges | **"none of the notification treatments significantly improved SAV deployment compared to the control group"** | remediation observed in control too | | | {[lone2022_sav]} IEEE S&P 2022 | 2,320 network operators, RCT | 8 treatments: direct email (PeeringDB→WHOIS→abuse), national CERT/NIC.br, NOG mailing lists, × nudges | **"none of the notification treatments significantly improved SAV deployment compared to the control group"** | remediation observed in control too | |
| | {[sasaki2022_ics]} IEEE S&P 2022 | 160 operators of 317 exposed ICS devices | **manual telephone calls**, email only for scheduling — "the first study to directly contact the organization operating the device" | reached the person in charge for 212 devices; **"50% of the persons in charge … stated that they mitigated or will mitigate"**, and follow-up scans confirmed devices "reduced by 58% when we were able to contact the persons in charge" | 13% decrease among un-notified devices, χ² p<0.0001 | | | {[sasaki2022_ics]} IEEE S&P 2022 | 160 operators of 317 exposed ICS devices | **manual telephone calls**, email only for scheduling — "the first study to directly contact the organization operating the device" | reached the person in charge for 212 devices; **"50% of the persons in charge … stated that they mitigated or will mitigate"**, and follow-up scans confirmed devices "reduced by 58% when we were able to contact the persons in charge" | 13% decrease among un-notified devices, χ² p<0.0001 | |
| | {[bennett2022_spfail]} IMC 2022 | 6,488 mail-server notifications | ''postmaster@'' per the SMTP spec | 31.6% undelivered; of 4,434 delivered, 512 (12%) opened, 177 (4%) patched, **9 (<1%) between private and public disclosure** | — | | | {[bennett2022_spfail]} IMC 2022 | 6,488 mail-server notifications | ''postmaster@'' per the SMTP spec | 31.6% undelivered; of 4,434 delivered, 512 (12%) opened, 177 (4%) patched, **9 (<1%) between private and public disclosure** | — | |
| | | {[czybik2023_lazy]} IMC 2023 | **111,951** domain operators with erroneous SPF records, May 2023 | ''postmaster@'' and ''security@'', plus "the contact named in security.txt, if available"; own mail server throttled to one mail a second | **6,931 of 211,018 errors fixed two weeks later — "a success rate of 3.3 %"**; 300 thank-you mails, 3 spam complaints; bounces "large" but uncounted | — | |
| | | {[nan2023_spying]} USENIX Sec 2023 | 1,381 developers of IoT companion apps whose privacy policies omit exposed data | developer email from the app store, tracked by a mail-merge tool | after a month: **381 opened, 850 unopened, 150 bounced**; 21 of the 381 openers changed their policy (1.5% of sent); the list was then escalated to Google Play, 360 Store and APKPure | — | |
| | | {[blechschmidt2023_hello]} USENIX Sec 2023 | 4,484 domains with obviously misconfigured mail servers | email to the operator | **"at least 2,700 were not delivered and bounced"** (≥60%, the worst delivery figure on this page); 26 non-automated replies; "only 1,076 were still misconfigured in November, which is a decrease of 39.7%" among the domains whose mail was delivered | the 2,700 bounced domains fell 20.7% over the same period — a comparison group the paper reports while saying it "cannot conclude a causal relation", since deliverability itself may drive fixing | |
| | {[utz2023_comparing]} PoPETs 2023 | 159,035 domains, 4 privacy issues + 1 security issue | parsed privacy-policy/contact addresses vs RFC 2142 aliases | **87.8% vs 33.8% delivery**; remediation effects significant but small — the paper puts them at "0–1" to "1–2 percentage points" over control (Fisher's exact, Holm–Bonferroni), with a few later-date and generic-alias cells around 3 pp | yes, per issue | | | {[utz2023_comparing]} PoPETs 2023 | 159,035 domains, 4 privacy issues + 1 security issue | parsed privacy-policy/contact addresses vs RFC 2142 aliases | **87.8% vs 33.8% delivery**; remediation effects significant but small — the paper puts them at "0–1" to "1–2 percentage points" over control (Fisher's exact, Holm–Bonferroni), with a few later-date and generic-alias cells around 3 pp | yes, per issue | |
| | {[qin2024_rov]} NDSS 2024 | 1,012 non-deploying ASes randomised into 6 treatment arms + control; **859 emails sent** | operator email from PeeringDB, falling back to WHOIS; nudge variants (baseline, social norms, authority, reminder, elicitation) and native language | 824 of 859 delivered, **4.07% bounce rate** against the ">50%" it cites from prior work — and **"none of the notification treatments has a significant effect"** (survival analysis; relative risk 0.46–1.35, every confidence interval spanning 1) | 11 of 138 remediated | | | {[qin2024_rov]} NDSS 2024 | 1,012 non-deploying ASes randomised into 6 treatment arms + control; **859 emails sent** | operator email from PeeringDB, falling back to WHOIS; nudge variants (baseline, social norms, authority, reminder, elicitation) and native language | 824 of 859 delivered, **4.07% bounce rate** against the ">50%" it cites from prior work — and **"none of the notification treatments has a significant effect"** (survival analysis; relative risk 0.46–1.35, every confidence interval spanning 1) | 11 of 138 remediated | |
| | {[elyadmani2025_keys]} IEEE S&P 2025 | 160 organisations leaking cloud-bucket secrets | leaked-file contents, OSINT, disclosure programmes, privacy-policy addresses — routed through a CSIRT partner | **95/160 (59.4%) acted**; only 20 organisations replied at all | — | | | {[li2024_wellinformed]} CCS 2024 | 4,399 emails to developers of apps lacking GDPR runtime privacy notices | developer email from the Google Play listing; IRB "minimal risk" | 4,169 delivered; **821 unique replies, 60 of them human, 13 substantive**; no re-scan, so no remediation figure | — | |
| | {[munteanu2025_catch22]} USENIX Sec 2025 | operators of compromised hosts | Shadowserver Foundation and CERT-BUND ran the campaign | the current worked example of delegating delivery | — | | | {[elyadmani2025_keys]} IEEE S&P 2025* | 160 organisations leaking cloud-bucket secrets | leaked-file contents, OSINT, disclosure programmes, privacy-policy addresses — routed through a CSIRT partner | **95/160 (59.4%) acted**; only 20 organisations replied at all | — | |
| | | {[munteanu2025_catch22]} USENIX Sec 2025* | operators of compromised hosts | Shadowserver Foundation and CERT-BUND ran the campaign | the current worked example of delegating delivery | — | |
| | | {[li2026_rpkiinvalid]} NDSS 2026* | owners of 1,731 RPKI-invalid prefixes, 1,385 email addresses | abuse and technical contacts from RIR WHOIS records; the mail asked whether the prefix was misconfigured or hijacked, i.e. a validation survey rather than a fix request | **302 of 1,385 undeliverable (21.8%)**; 174 organisations answered, 16.1% of the delivered | — | |
| |
| ==== The loss is in delivery, not in willingness ==== | ==== The loss is in delivery, not in willingness ==== |
| |
| The delivery funnel is the story, and it is why headline remediation rates look so bad. {[stock2016_hey]}: **5.8% of reports received**, but ~40% remediation among those actually read. {[bennett2022_spfail]}: 31.6% undelivered → 12% of the delivered opened → 4% of the openers patched. {[stock2018_didnt]}: 74.4% fixed among Git operators who //viewed// the report, against a 13% control. Operators who read your report largely act on it. Most never read it. | The delivery funnel is the story, and it is why headline remediation rates look so bad. {[stock2016_hey]}: **5.8% of reports received**, but ~40% remediation among those actually read. {[bennett2022_spfail]}: 31.6% undelivered → 12% of the delivered opened → 4% of the openers patched. {[stock2018_didnt]}: 74.4% fixed among Git operators who //viewed// the report, against a 13% control. {[nan2023_spying]} instrumented the same funnel for app developers: 1,381 sent → 150 bounced → 381 opened → 21 acted. {[gilad2017_rpki]} lost 384 of 552 RIR-WHOIS mails to bounces before anyone could read them; {[roth2022_security]} lost 197 of 256 alias mails, and cut the failure to 4 of 105 once the addresses were curated by hand. Operators who read your report largely act on it. Most never read it. The largest fully automated email campaign in the corpus since {[utz2023_comparing]}, {[czybik2023_lazy]}, is the sobering baseline for a fully automated email campaign in 2023: 111,951 mails, 3.3% of the errors fixed two weeks later — and the fix rate tracked how easy the fix was, 5.7% for syntax errors against 1.6% for DNS-lookup limits. |
| |
| So the highest-leverage thing you can do is not writing a better email. It is: | So the highest-leverage thing you can do is not writing a better email. It is: |
| | Persistence | 7 months later, 3.5% of remediated sites had regressed — "long-term effectiveness of approximately 95%" | {[maass2021_effective]} | | | Persistence | 7 months later, 3.5% of remediated sites had regressed — "long-term effectiveness of approximately 95%" | {[maass2021_effective]} | |
| | Public disclosure | Private notification "made little difference"; the **public CVE 60 days later** correlated with a much larger drop in vulnerable servers | {[bennett2022_spfail]} | | | Public disclosure | Private notification "made little difference"; the **public CVE 60 days later** correlated with a much larger drop in vulnerable servers | {[bennett2022_spfail]} | |
| | | Address source | Generic aliases: delivery failed for 197 of 256 domains. Hand-curated addresses for the 184 still-affected sites: 4 failures of 105 | {[roth2022_security]} | |
| | | Fix difficulty | Errors an operator can fix by editing one line were fixed at 5.7%; errors needing a change at an external provider at 1.6% — same template; the authors' explanation ("We assume that these are often non-trivial to fix"), not a tested one | {[czybik2023_lazy]} | |
| | | Escalation | When 1.5% of notified developers acted, the authors handed the list to the app stores; Google Play "responded to our request quickly", two other stores did not | {[nan2023_spying]} | |
| |
| Read the two contradictions as the state of the art rather than as noise. **Message-content tuning is a dead end**: three studies varied wording, format, signing, tone and translation, and the effects are small, inconsistent, or vanish under multiple-comparison correction. **Sender identity, medium and reachability are the real variables** — and even those may not survive a control group, which is what {[lone2022_sav]} demonstrated on the one population where somebody randomised properly. | Read the two contradictions as the state of the art rather than as noise. **Message-content tuning is a dead end**: three studies varied wording, format, signing, tone and translation, and the effects are small, inconsistent, or vanish under multiple-comparison correction. **Sender identity, medium and reachability are the real variables** — and even those may not survive a control group, which is what {[lone2022_sav]} demonstrated on the one population where somebody randomised properly. |
| * **False positives reach real people.** Twelve IPv6 contacts in {[li2016_youve]} rebutted the vulnerability claim. Your detector's precision is now somebody else's inbox; measure it before, not after. | * **False positives reach real people.** Twelve IPv6 contacts in {[li2016_youve]} rebutted the vulnerability claim. Your detector's precision is now somebody else's inbox; measure it before, not after. |
| * **Ethics review is not automatic and not uniform.** {[stock2016_hey]} recorded that its institutions "neither mandate nor provide an IRB approval before conducting such experiments". {[maass2021_effective]} got approval from the ethics committees of two of three institutions and a dean's approval from the third. See [[Practices:Ethics]], and {[hantke2024_redlines]} for what server operators themselves consider acceptable. | * **Ethics review is not automatic and not uniform.** {[stock2016_hey]} recorded that its institutions "neither mandate nor provide an IRB approval before conducting such experiments". {[maass2021_effective]} got approval from the ethics committees of two of three institutions and a dean's approval from the third. See [[Practices:Ethics]], and {[hantke2024_redlines]} for what server operators themselves consider acceptable. |
| | |
| | ==== The receiving end ==== |
| | |
| | Four papers look at what happens after a report arrives — {[stivala2026_behind]} above, on how hosting providers triage, and three more here — and none of them is encouraging about either the unsolicited individual email or the generic disclosure form. {[bijmans2026_tickets]} analysed 1.3 million abuse reports seized by Dutch law enforcement from one hosting provider with a bulletproof reputation: **2.6% of reports were ever linked to a notification to the customer**. The rate was not about the abuse; it was about the reporter. Netcraft's phishing reports led to a customer notification 72% of the time and Spamhaus listings 60–84%, while "individual abuse reporting is often easily ignored" — one DMCA service filed 126,085 reports that produced four notifications, and one honeypot operator's 30,467 automated reports produced none. The authors are explicit that one abusive provider does not generalise; the direction of the effect is still the point: **a report is acted on in proportion to what the reporter can do to the recipient's business**, which a research group cannot do at all. On the vendor side, {[hastings2016_weakkeys]} followed the 2012 disclosure of weak-key generation to 61 device vendors — of the 37 with weak RSA keys, 5 issued a public advisory, 11 answered privately, 3 sent an auto-reply, 18 nothing — and found that "vendor notification, positive vendor responses, and even vendor-produced public security advisories appear to have little correlation with end-user security": vulnerable populations kept growing for years at vendors that had published an advisory. And when the recipient does have a formal channel, it may be the wrong shape: {[bhattacharya2026_asi]} briefed 58 of the 100 most popular services about account-security flaws through their disclosure programmes, got 38 acknowledgements and 25 substantive replies, most classified "Not applicable" or "Out of scope", and notes that "HackerOne rate limits disclosures, which slowed our disclosure process immensely"; {[roth2022_security]} found that "many of those that answered instructed us to contact HackerOne" although the message never used the word vulnerability. |
| | |
| | ==== Deciding not to run a campaign ==== |
| | |
| | Two recent papers reasoned their way out of a campaign in print, and both are worth citing when you do the same. {[rao2024_unfiltered]} found domains whose cloud email filters could be bypassed, cited {[bennett2022_spfail]} ("over 80% of the domains contacted were unresponsive"), and "elected to work directly and closely with filtering service providers to update their documentation, notify their customers (with whom they do have an existing business relationship) and resolve the issues identified" — seven filtering vendors instead of the 1,262 misconfigured domains. {[ryan2023_passivessh]} recovered 189 unique SSH host keys from a scan of millions of devices, disclosed to four manufacturers and CERT/CC, and wrote: "We considered notifying operators of affected devices whose keys we had recovered, but we determined this would be infeasible." Neither is a shortcut; both state the alternative they took and why. |
| |
| ==== Covert notification ==== | ==== Covert notification ==== |
| - **Your detector's precision**, because false positives went to real inboxes. | - **Your detector's precision**, because false positives went to real inboxes. |
| - **The disclosure timeline you followed** and which convention it came from. | - **The disclosure timeline you followed** and which convention it came from. |
| - **The message itself**, in an appendix or [[Artifacts|artefact]]. Every study above that reports a framing effect published its text; you cannot replicate a framing result without it. | - **The message itself**, in an appendix or [[:artifacts|artefact]]. Every study above that reports a framing effect published its text; you cannot replicate a framing result without it. |
| - **The ethics decision**: review body and outcome, whether the study was covert, how you debriefed, the opt-out mechanism, and how many opted out. | - **The ethics decision**: review body and outcome, whether the study was covert, how you debriefed, the opt-out mechanism, and how many opted out. |
| - **Everything hostile that happened.** Legal threats, complaints to your institution, suspended customers. This is the part later researchers most need and the part most often missing. | - **Everything hostile that happened.** Legal threats, complaints to your institution, suspended customers. This is the part later researchers most need and the part most often missing. |
| ===== Papers to read first ===== | ===== Papers to read first ===== |
| |
| If you read four: {[stock2016_hey]} for the channel survey and the delivery funnel, {[li2016_youve]} for the arm-by-arm comparison and the decay curve, {[maass2021_effective]} for the only large factorial design on framing and medium, and {[lone2022_sav]} for the control group that undermines the rest. Then {[utz2023_comparing]} for privacy versus security and the delivery figures that supersede the 2016 channel advice, {[stock2018_didnt]} for why message tuning is a dead end, {[li2016_remedying]} for what a platform can do that you cannot, and {[stivala2026_behind]} for the receiving end. | If you read four: {[stock2016_hey]} for the channel survey and the delivery funnel, {[li2016_youve]} for the arm-by-arm comparison and the decay curve, {[maass2021_effective]} for the only large factorial design on framing and medium, and {[lone2022_sav]} for the control group that undermines the rest. Then {[utz2023_comparing]} for privacy versus security and the delivery figures that supersede the 2016 channel advice, {[stock2018_didnt]} for why message tuning is a dead end, {[li2016_remedying]} for what a platform can do that you cannot, {[czybik2023_lazy]} for what a 100,000-mail automated campaign actually yields in 2023, and {[stivala2026_behind]} and {[bijmans2026_tickets]} for the receiving end. |
| |
| ===== Related pages ===== | ===== Related pages ===== |
| * [[Practices:Public relations]] — the other post-publication channel. | * [[Practices:Public relations]] — the other post-publication channel. |
| * [[Statistics:Hypothesis testing]] and [[Statistics:Pvalue corrections]] — for the arm comparisons. | * [[Statistics:Hypothesis testing]] and [[Statistics:Pvalue corrections]] — for the arm comparisons. |
| * [[Artifacts]] — publishing the notification text and the contact-discovery code. | * [[:artifacts|Artifacts]] — publishing the notification text and the contact-discovery code. |
| * [[Design:Website selection]] — your notified population is your sample, with the same biases. | * [[Design:Website selection]] — your notified population is your sample, with the same biases. |
| |
| * **Seven venues only.** EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS are absent, and CHI/SOUPS are where a good deal of the operator-facing usable-security work appears. NDSS 2016 and NDSS 2018 full text was not retrieved at all, which is why {[stock2018_didnt]} — the canonical follow-up study — was read from the publisher's PDF rather than from the corpus. | * **Seven venues only.** EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS are absent, and CHI/SOUPS are where a good deal of the operator-facing usable-security work appears. NDSS 2016 and NDSS 2018 full text was not retrieved at all, which is why {[stock2018_didnt]} — the canonical follow-up study — was read from the publisher's PDF rather than from the corpus. |
| * ''ethics.notifiedAffectedParties'' **agreed with an independent re-extraction on 67% of papers**, so treat every percentage above as accurate to a few points, not to the decimal.((That 67% was measured on 100 papers of the previous, 4,322-paper extraction run and has not been re-measured on the current corpus. Treat it as the right order of magnitude.)) ''ethics.disclosureDetail'' is free text capped at 20 words, so it is reported only as folded families and only as a ranking; the fold rules and the unmapped residue are on the provenance page. | * ''ethics.notifiedAffectedParties'' **agreed with an independent re-extraction on 67% of papers**, so treat every percentage above as accurate to a few points, not to the decimal.((That 67% was measured on 100 papers of the previous, 4,322-paper extraction run and has not been re-measured on the current corpus. Treat it as the right order of magnitude.)) ''ethics.disclosureDetail'' is free text capped at 20 words, so it is reported only as folded families and only as a ranking; the fold rules and the unmapped residue are on the provenance page. |
| * **The 22 campaigns are hand-classified, not a census.** A regex over 5,869 full-text files (ten more than the 5,859 extraction records — a few papers have text but no record) produced 179 candidates; 22 were classified as campaigns, 13 rejected with a stated reason, and **144 were never read**. Three further rejections came from a wider first-pass scan and fall outside the committed regex, which is why the reasoned rejections total 16. The table above is a curated reading list, and the unreviewed residue is published in full. | * **The 32 campaigns are hand-classified, not a census.** Two sentence-level rules run over 5,869 full-text files (ten more than the 5,859 extraction records — a few papers have text but no record). The narrow rule produced 179 candidates: 28 campaigns and 151 rejections with a stated reason. A wider recall probe, added on 2026-09-04 after the narrow rule was found to miss a 111,951-mail campaign, produced 78 more: 3 campaigns and 75 rejections. One campaign, {[roth2022_security]}, was caught by neither rule and found through a full-text search for ''security.txt''; three rejections carry over from a wider first-pass scan. Nothing is unreviewed, **but most rejections are screens, not reads**: 126 on every sentence carrying a notification verb, 75 on the probe-matching sentences alone, 28 on a full read of the relevant passages or paper; the provenance page says which. A third rule would find more; the table above is a curated reading list, not a census. |
| |
| The complete query log, the report script with its unedited output, the folding rules with their unmapped residue, the spot-checked quotes, and every external source that was rejected are on **[[provenance:practices:notifying_websites]]**. Corpus-wide caveats are on [[Literature:Corpus]]. | The complete query log, the report script with its unedited output, the folding rules with their unmapped residue, the spot-checked quotes, and every external source that was rejected are on **[[provenance:practices:notifying_websites]]**. Corpus-wide caveats are on [[Literature:Corpus]]. |