User Tools

Site Tools


practices:notifying_websites

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
practices:notifying_websites [2026/09/04 12:58] – Citations review: weak-keys Table 2 split corrected to 5/11/3/18; Hilbig et al. 42 successful scans over 55 weeks. Authored by Claude karel.kubicek.claudepractices:notifying_websites [2026/09/04 13:10] (current) – Generic review: lead box no longer says all network RCTs failed; Extended Hell(o) row gets its funnel and comparison group; outcome kinds named; security.txt series disagreement stated; three figures corrected; methodology caveat rebalanced. Authored by C karel.kubicek.claude
Line 6: Line 6:
  
 <WRAP important> <WRAP important>
-The single most consequential design decision on this page is the **control group**. Sites get fixed for reasons that have nothing to do with you — automatic updates, unrelated maintenance, going offline. Every notification study without a control arm has attributed some of that background remediation to its own notifications — and **both randomised, control-arm experiments on network operators in this corpus found no significant effect from any treatment**, including treatments that observational studies had reported as effective {[lone2022_sav,qin2024_rov]}. Randomise, hold back an arm, and report both. Note what that does and does not mean: {[maass2021_effective]} is also a randomised controlled experiment — a full factorial design with a control group — and it measured 56.6% against 9.2% on German website owners. The lesson is not that notification never works. It is that a study without a control arm cannot tell you which of those two cases it is in.+The single most consequential design decision on this page is the **control group**. Sites get fixed for reasons that have nothing to do with you — automatic updates, unrelated maintenance, going offline. Every notification study without a control arm has attributed some of that background remediation to its own notifications — and **the two most recent randomised, control-arm experiments on network operators in this corpus — on deploying SAV and ROV, expensive and low-visibility configuration changes — found no significant effect from any treatment**, including treatments that observational studies had reported as effective {[lone2022_sav,qin2024_rov]}. The earlier randomised experiments on network operators, on Heartbleed patching and firewall misconfiguration {[durumeric2014_heartbleed,li2016_youve]}, did find one. Randomise, hold back an arm, and report both. Note what that does and does not mean: {[maass2021_effective]} is also a randomised controlled experiment — a full factorial design with a control group — and it measured 56.6% against 9.2% on German website owners. The lesson is not that notification never works. It is that a study without a control arm cannot tell you which of those two cases it is in.
 </WRAP> </WRAP>
  
Line 30: Line 30:
 | **crawled ∧ assessed a law** | **123** | 35.0% | 14.6% | **14.6%** | 15.4% | 20.3% | | **crawled ∧ assessed a law** | **123** | 35.0% | 14.6% | **14.6%** | 15.4% | 20.3% |
  
-Two things to take from this table. **Network measurement is ahead of web measurement**: scanning papers notify at 61.0% (''yes''+''partial'') against 43.5% for crawling papers, and are less than half as likely to say outright that they did not (4.2% vs 9.4%), as well as meaningfully less likely to leave it unanswered (24.0% vs 31.1%). The scanning community built the norm first, largely because Internet-wide scanning provoked complaints that forced the question. **Legal-compliance papers state a position most often** (only 12.0% ''not-stated''), and the bottom row is the population closest to whoever is reading this: 123 papers that both crawled the web and assessed a law, i.e. found a compliance violation across many sites. **14.6% of them say outright that they did not notify** — more than three times the corpus-wide 4.3%. That is usually a decision rather than negligence. A GDPR-violation finding across thousands of sites is the case where authors most often decide, and say, that individual notification is the wrong instrument, and route the finding to a regulator or to publication instead. If that is your situation, //deciding not to notify is a defensible position that you have to argue in the paper//, not an omission you can leave to the reader.+Two things to take from this table. **Network measurement is ahead of web measurement**: scanning papers notify at 61.0% (''yes''+''partial'') against 43.5% for crawling papers, and are less than half as likely to say outright that they did not (4.2% vs 9.4%), as well as meaningfully less likely to leave it unanswered (24.0% vs 31.1%). The scanning community built the norm first, largely because Internet-wide scanning provoked complaints that forced the question. **Legal-compliance papers state a position most often** (only 12.0% ''not-stated''), and the bottom row is the population closest to whoever is reading this: 123 papers that both crawled the web and assessed a law, i.e. found a compliance violation across many sites. **14.6% of them say outright that they did not notify** — more than three times the corpus-wide 4.3%. That is usually a decision rather than negligence. A GDPR-violation finding across thousands of sites is the case where authors most often decide, and say, that individual notification is the wrong instrument, and route the finding to a regulator or to publication instead. If that is your situation, //deciding not to notify is a defensible position that you have to argue in the paper//, not an omission you can leave to the reader (two worked examples are under //Deciding not to run a campaign// below).
  
 The direction of travel is unambiguous: The direction of travel is unambiguous:
Line 76: Line 76:
  
 ^ Channel ^ How you get it ^ What it reaches ^ Status and evidence ^ ^ Channel ^ How you get it ^ What it reaches ^ Status and evidence ^
-| **''security.txt''** (RFC 9116) | ''GET https://<domain>/.well-known/security.txt'' | the security team, if there is one | RFC 9116, Informational, April 2022 — current, not obsoleted.((Verified against [[https://www.rfc-editor.org/rfc/rfc9116.txt|rfc-editor.org/rfc/rfc9116.txt]] on 2026-08-13: "Category: Informational … April 2022". Three errata exist, none touching the field list.)) Adoption is the problem, and it is steep in rank: 11–16% of the Alexa top 100, 8–10% of the top 1K, "3–4% for the top 10K sites, and only a percent for the top 100K" {[poteat2021_securitytxt]}. That is a 2021 measurement on a ranking that was discontinued in 2022 (see [[Design:Website selection]]) and predates RFC 9116; **no paper in the seven corpus venues re-measures adoption since.** Outside them, two do: 0.49% of the Tranco top million and 1.6% of the top 100K in October–November 2021 {[findlay2022_securitytxt]}, and 42 successful weekly scans of the Tranco top million over the 55 weeks from December 2021 to January 2023 showing 32.0→34.0% (top 100), 16.1→18.8% (top 1K), 7.9→9.8% (top 10K), 2.4→3.2% (top 100K) and 0.7→1.0% (top 1M) {[hilbig2023_securitytxt]}.((Both are outside the corpus's seven venues: MADweb is an NDSS workshop, DTRAP an ACM journal. A vendor scan of gTLD zone files reports 573,123 of 241 million domains (0.24%) with a ''security.txt'' in early 2026, up from 0.05% in 2021 — [[https://blog.iotdef.com/the-state-of-security-txt-adoption-an-analysis-of-240-million-domains-in-2026/|blog.iotdef.com]], fetched 2026-09-04; methodology stated, data and code not published, so treat it as an order of magnitude, not a citation.)) The corpus's own incidental counts agree: 23 of the vulnerable domains in {[drakonakis2020_cookie]} (2020, "the most ineffective" of its four channels), 25 of 256 in {[roth2022_security]} (2022), 34 of the 100 most popular services in {[bhattacharya2026_asi]} (2025). Six campaigns in the corpus **used** it as one channel among several — {[drakonakis2020_cookie,ghasemisharif2022_saat,roth2022_security,czybik2023_lazy,innocenti2025_only,bhattacharya2026_asi]} — and **none separates its delivery or response from the other channels**, so whether a listed contact actually answers is still unmeasured; {[hilbig2023_securitytxt]} names exactly that as future work |+| **''security.txt''** (RFC 9116) | ''GET https://<domain>/.well-known/security.txt'' | the security team, if there is one | RFC 9116, Informational, April 2022 — current, not obsoleted.((Verified against [[https://www.rfc-editor.org/rfc/rfc9116.txt|rfc-editor.org/rfc/rfc9116.txt]] on 2026-08-13: "Category: Informational … April 2022". Three errata exist, none touching the field list.)) Adoption is the problem, and it is steep in rank: 11–16% of the Alexa top 100, 8–10% of the top 1K, "3–4% for the top 10K sites, and only a percent for the top 100K" {[poteat2021_securitytxt]}. That is a 2021 measurement on a ranking that was discontinued in 2022 (see [[Design:Website selection]]) and predates RFC 9116; **no paper in the seven corpus venues re-measures adoption since.** Outside them, two do: 0.49% of the Tranco top million and 1.6% of the top 100K in October–November 2021 {[findlay2022_securitytxt]}, and 42 successful weekly scans of the Tranco top million over the 55 weeks from December 2021 to January 2023 showing 32.0→34.0% (top 100), 16.1→18.8% (top 1K), 7.9→9.8% (top 10K), 2.4→3.2% (top 100K) and 0.7→1.0% (top 1M) {[hilbig2023_securitytxt]}.((Both are outside the corpus's seven venues: MADweb is an NDSS workshop, DTRAP an ACM journal. A vendor scan of gTLD zone files reports 573,123 of 241 million domains (0.24%) with a ''security.txt'' in early 2026, up from 0.05% in 2021 — [[https://blog.iotdef.com/the-state-of-security-txt-adoption-an-analysis-of-240-million-domains-in-2026/|blog.iotdef.com]], fetched 2026-09-04; methodology stated, data and code not published, so treat it as an order of magnitude, not a citation.)) The two peer-reviewed series disagree by about 2× at the top ranks (Alexa against Tranco, and a Kaplan–Meier deployment estimate against weekly snapshots); the corpus's own incidental counts sit with the Tranco series: 23 domains in {[drakonakis2020_cookie]} (2020, "the most ineffective" of its four channels; no denominator given), 25 of 256 in {[roth2022_security]} (2022), 34 of the 100 most popular services in {[bhattacharya2026_asi]} (2025). Six corpus papers that notified **used** it as one channel among several — {[drakonakis2020_cookie,ghasemisharif2022_saat,roth2022_security,czybik2023_lazy,innocenti2025_only,bhattacharya2026_asi]} — and **none separates its delivery or response from the other channels**, so whether a listed contact actually answers is still unmeasured; {[hilbig2023_securitytxt]} names exactly that as future work |
 | **RIR abuse contact** | RIPE ''abuse-c:'' / ARIN Abuse POC, resolved from the IP; RIPEstat's abuse-contact-finder covers all five RIRs over HTTPS | the **hosting provider or CDN**, not the site owner | The one source that scales — free and with no meaningful rate limit. ARIN verifies its Abuse POC annually;((ARIN NRPM §3.6: "Each of the following Points of Contact are to be verified annually … Admin, Tech, NOC, Abuse", [[https://www.arin.net/participate/policy/nrpm/|arin.net]], fetched 2026-08-13. RIPE states it works to keep abuse contacts valid but no validation cadence could be found on a RIPE primary source on 2026-08-13 — do not assume annual.)) RIPE states it keeps abuse contacts valid but publishes no cadence that could be verified from a RIPE source. Still the workhorse: "WHOIS was the most frequently mentioned channel (seven HPOs)" among 24 interviewed providers in 2026, who "confirmed that their contact is available via WHOIS, for example through ARIN or RIPE databases" {[stivala2026_behind]}. Note this is the //network// registry, which GDPR redaction did not touch — not the //domain// registrant record, which it did. It is not clean either: 384 of 552 RIR-WHOIS mails bounced in {[gilad2017_rpki]} ("whois entries are often outdated"), and 302 of 1,385 (21.8%) were undeliverable in {[li2026_rpkiinvalid]}. For DNS resolvers there is one more machine-readable source, the ''RNAME'' of the zone's SOA record; {[deccio2020_closeddoors]} used it to reach 43 administrators and heard back from five, three of whom the authors already knew | | **RIR abuse contact** | RIPE ''abuse-c:'' / ARIN Abuse POC, resolved from the IP; RIPEstat's abuse-contact-finder covers all five RIRs over HTTPS | the **hosting provider or CDN**, not the site owner | The one source that scales — free and with no meaningful rate limit. ARIN verifies its Abuse POC annually;((ARIN NRPM §3.6: "Each of the following Points of Contact are to be verified annually … Admin, Tech, NOC, Abuse", [[https://www.arin.net/participate/policy/nrpm/|arin.net]], fetched 2026-08-13. RIPE states it works to keep abuse contacts valid but no validation cadence could be found on a RIPE primary source on 2026-08-13 — do not assume annual.)) RIPE states it keeps abuse contacts valid but publishes no cadence that could be verified from a RIPE source. Still the workhorse: "WHOIS was the most frequently mentioned channel (seven HPOs)" among 24 interviewed providers in 2026, who "confirmed that their contact is available via WHOIS, for example through ARIN or RIPE databases" {[stivala2026_behind]}. Note this is the //network// registry, which GDPR redaction did not touch — not the //domain// registrant record, which it did. It is not clean either: 384 of 552 RIR-WHOIS mails bounced in {[gilad2017_rpki]} ("whois entries are often outdated"), and 302 of 1,385 (21.8%) were undeliverable in {[li2026_rpkiinvalid]}. For DNS resolvers there is one more machine-readable source, the ''RNAME'' of the zone's SOA record; {[deccio2020_closeddoors]} used it to reach 43 administrators and heard back from five, three of whom the authors already knew |
 | **PeeringDB technical contact** | PeeringDB API, per AS | the network operator | Preferred over WHOIS by {[lone2022_sav]}: "We preferred peeringDB because it has been used in previous studies and they found the database up-to-date" — i.e. it is relaying prior work's assessment, not measuring it | | **PeeringDB technical contact** | PeeringDB API, per AS | the network operator | Preferred over WHOIS by {[lone2022_sav]}: "We preferred peeringDB because it has been used in previous studies and they found the database up-to-date" — i.e. it is relaying prior work's assessment, not measuring it |
Line 272: Line 272:
 </file> </file>
  
-Real output, run 2026-08-13 on five hand-picked domains — **not a sample, and not representative**: these are large, well-resourced organisations, so ''security.txt'' coverage here is an order of magnitude above the 34% {[poteat2021_securitytxt]} measured at the top 10K.+Real output, run 2026-08-13 on five hand-picked domains — **not a sample, and not representative**: these are large, well-resourced organisations, so ''security.txt'' coverage here is well above any measured tier — 3234for the top 100 in {[hilbig2023_securitytxt]}, 3–4% at the top 10K in {[poteat2021_securitytxt]}.
  
 <code> <code>
Line 312: Line 312:
 {[sasaki2022_ics]} collects the comparison itself, and it is the shortest statement of the range: "its remediation rate was approximately 18% … Our remediation rate is higher than most previous notification experiments: it was approximately 40% for cross-site scripting and a WordPress vulnerability, 33%–42% for different WordPress vulnerability, and less than 20% for DNS zone poisoning. The only campaigns that reported similar remediation rates were on publicly accessible Git repositories (78%–81%) and Heartbleed (approximately 40%–90%)." {[sasaki2022_ics]} collects the comparison itself, and it is the shortest statement of the range: "its remediation rate was approximately 18% … Our remediation rate is higher than most previous notification experiments: it was approximately 40% for cross-site scripting and a WordPress vulnerability, 33%–42% for different WordPress vulnerability, and less than 20% for DNS zone poisoning. The only campaigns that reported similar remediation rates were on publicly accessible Git repositories (78%–81%) and Heartbleed (approximately 40%–90%)."
  
-These are the campaigns in the corpus that reported a quotable outcome, each with **its own denominator** — they are not comparable to each other, because "response", "remediation" and "fix" are defined differently in each and the populations are wildly different. Plan against the range, not the mean. The hand pass classified 32 campaigns in total; 26 are below, and the six not shown reported a volume sent but no response or remediation figure. Two rows — {[munteanu2025_catch22]} and {[liao2016_seeking]} — are here for their channel rather than a rate, because routing a campaign through an established notification operator or a national CERT is the pattern, not the number. The six are listed on [[provenance:practices:notifying_websites]], along with the one row below — {[stock2018_didnt]} — whose full text is missing from the corpus and was read from the publisher's PDF instead. Rows marked * are from the provisional 2025–2026 slice.+These are the campaigns in the corpus that reported a quotable outcome, each with **its own denominator** — they are not comparable to each other, because "response", "remediation" and "fix" are defined differently in each and the populations are wildly different. Plan against the range, not the mean. The Outcome column mixes four different quantities and says which each row reports: **remediation** measured by re-scan (Maass, Czybik), **response** counts (Nguyen, Li 2024), **delivery** alone (Roth 2022, the 2026 validation survey), and **removal by a platform** (Edu). Pick the rows that measure what you will measure. The hand pass classified 32 campaigns in total; 26 are below, and the six not shown reported a volume sent but no response or remediation figure. Two rows — {[munteanu2025_catch22]} and {[liao2016_seeking]} — are here for their channel rather than a rate, because routing a campaign through an established notification operator or a national CERT is the pattern, not the number. The six are listed on [[provenance:practices:notifying_websites]], along with the one row below — {[stock2018_didnt]} — whose full text is missing from the corpus and was read from the publisher's PDF instead. Rows marked * are from the provisional 2025–2026 slice.
  
 ^ Study ^ Notified ^ Channel ^ Outcome, in the paper's own terms ^ Control ^ ^ Study ^ Notified ^ Channel ^ Outcome, in the paper's own terms ^ Control ^
Line 335: Line 335:
 | {[czybik2023_lazy]} IMC 2023 | **111,951** domain operators with erroneous SPF records, May 2023 | ''postmaster@'' and ''security@'', plus "the contact named in security.txt, if available"; own mail server throttled to one mail a second | **6,931 of 211,018 errors fixed two weeks later — "a success rate of 3.3 %"**; 300 thank-you mails, 3 spam complaints; bounces "large" but uncounted | — | | {[czybik2023_lazy]} IMC 2023 | **111,951** domain operators with erroneous SPF records, May 2023 | ''postmaster@'' and ''security@'', plus "the contact named in security.txt, if available"; own mail server throttled to one mail a second | **6,931 of 211,018 errors fixed two weeks later — "a success rate of 3.3 %"**; 300 thank-you mails, 3 spam complaints; bounces "large" but uncounted | — |
 | {[nan2023_spying]} USENIX Sec 2023 | 1,381 developers of IoT companion apps whose privacy policies omit exposed data | developer email from the app store, tracked by a mail-merge tool | after a month: **381 opened, 850 unopened, 150 bounced**; 21 of the 381 openers changed their policy (1.5% of sent); the list was then escalated to Google Play, 360 Store and APKPure | — | | {[nan2023_spying]} USENIX Sec 2023 | 1,381 developers of IoT companion apps whose privacy policies omit exposed data | developer email from the app store, tracked by a mail-merge tool | after a month: **381 opened, 850 unopened, 150 bounced**; 21 of the 381 openers changed their policy (1.5% of sent); the list was then escalated to Google Play, 360 Store and APKPure | — |
-| {[blechschmidt2023_hello]} USENIX Sec 2023 | operators of misconfigured mail servers (STARTTLS/confidentiality) | email notification to the operator | "only 1,076 were still misconfigured in November, which is a decrease of 39.7%" among the domains whose notification did not bounce | — |+| {[blechschmidt2023_hello]} USENIX Sec 2023 | 4,484 domains with obviously misconfigured mail servers | email to the operator | **"at least 2,700 were not delivered and bounced"** (≥60%, the worst delivery figure on this page); 26 non-automated replies; "only 1,076 were still misconfigured in November, which is a decrease of 39.7%" among the domains whose mail was delivered the 2,700 bounced domains fell 20.7% over the same period — a comparison group the paper reports while saying it "cannot conclude a causal relation", since deliverability itself may drive fixing |
 | {[utz2023_comparing]} PoPETs 2023 | 159,035 domains, 4 privacy issues + 1 security issue | parsed privacy-policy/contact addresses vs RFC 2142 aliases | **87.8% vs 33.8% delivery**; remediation effects significant but small — the paper puts them at "0–1" to "1–2 percentage points" over control (Fisher's exact, Holm–Bonferroni), with a few later-date and generic-alias cells around 3 pp | yes, per issue | | {[utz2023_comparing]} PoPETs 2023 | 159,035 domains, 4 privacy issues + 1 security issue | parsed privacy-policy/contact addresses vs RFC 2142 aliases | **87.8% vs 33.8% delivery**; remediation effects significant but small — the paper puts them at "0–1" to "1–2 percentage points" over control (Fisher's exact, Holm–Bonferroni), with a few later-date and generic-alias cells around 3 pp | yes, per issue |
 | {[qin2024_rov]} NDSS 2024 | 1,012 non-deploying ASes randomised into 6 treatment arms + control; **859 emails sent** | operator email from PeeringDB, falling back to WHOIS; nudge variants (baseline, social norms, authority, reminder, elicitation) and native language | 824 of 859 delivered, **4.07% bounce rate** against the ">50%" it cites from prior work — and **"none of the notification treatments has a significant effect"** (survival analysis; relative risk 0.46–1.35, every confidence interval spanning 1) | 11 of 138 remediated | | {[qin2024_rov]} NDSS 2024 | 1,012 non-deploying ASes randomised into 6 treatment arms + control; **859 emails sent** | operator email from PeeringDB, falling back to WHOIS; nudge variants (baseline, social norms, authority, reminder, elicitation) and native language | 824 of 859 delivered, **4.07% bounce rate** against the ">50%" it cites from prior work — and **"none of the notification treatments has a significant effect"** (survival analysis; relative risk 0.46–1.35, every confidence interval spanning 1) | 11 of 138 remediated |
Line 345: Line 345:
 ==== The loss is in delivery, not in willingness ==== ==== The loss is in delivery, not in willingness ====
  
-The delivery funnel is the story, and it is why headline remediation rates look so bad. {[stock2016_hey]}: **5.8% of reports received**, but ~40% remediation among those actually read. {[bennett2022_spfail]}: 31.6% undelivered → 12% of the delivered opened → 4% of the openers patched. {[stock2018_didnt]}: 74.4% fixed among Git operators who //viewed// the report, against a 13% control. {[nan2023_spying]} instrumented the same funnel for app developers: 1,381 sent → 150 bounced → 381 opened → 21 acted. {[gilad2017_rpki]} lost 384 of 552 RIR-WHOIS mails to bounces before anyone could read them; {[roth2022_security]} lost 197 of 256 alias mails, and cut the failure to 4 of 105 once the addresses were curated by hand. Operators who read your report largely act on it. Most never read it. The one campaign in the corpus at the 100,000 scale, {[czybik2023_lazy]}, is the sobering baseline for a fully automated email campaign in 2023: 111,951 mails, 3.3% of the errors fixed two weeks later — and the fix rate tracked how easy the fix was, 5.7% for syntax errors against 1.6% for DNS-lookup limits.+The delivery funnel is the story, and it is why headline remediation rates look so bad. {[stock2016_hey]}: **5.8% of reports received**, but ~40% remediation among those actually read. {[bennett2022_spfail]}: 31.6% undelivered → 12% of the delivered opened → 4% of the openers patched. {[stock2018_didnt]}: 74.4% fixed among Git operators who //viewed// the report, against a 13% control. {[nan2023_spying]} instrumented the same funnel for app developers: 1,381 sent → 150 bounced → 381 opened → 21 acted. {[gilad2017_rpki]} lost 384 of 552 RIR-WHOIS mails to bounces before anyone could read them; {[roth2022_security]} lost 197 of 256 alias mails, and cut the failure to 4 of 105 once the addresses were curated by hand. Operators who read your report largely act on it. Most never read it. The largest fully automated email campaign in the corpus since {[utz2023_comparing]}, {[czybik2023_lazy]}, is the sobering baseline for a fully automated email campaign in 2023: 111,951 mails, 3.3% of the errors fixed two weeks later — and the fix rate tracked how easy the fix was, 5.7% for syntax errors against 1.6% for DNS-lookup limits.
  
 So the highest-leverage thing you can do is not writing a better email. It is: So the highest-leverage thing you can do is not writing a better email. It is:
Line 375: Line 375:
 | Public disclosure | Private notification "made little difference"; the **public CVE 60 days later** correlated with a much larger drop in vulnerable servers | {[bennett2022_spfail]} | | Public disclosure | Private notification "made little difference"; the **public CVE 60 days later** correlated with a much larger drop in vulnerable servers | {[bennett2022_spfail]} |
 | Address source | Generic aliases: delivery failed for 197 of 256 domains. Hand-curated addresses for the 184 still-affected sites: 4 failures of 105 | {[roth2022_security]} | | Address source | Generic aliases: delivery failed for 197 of 256 domains. Hand-curated addresses for the 184 still-affected sites: 4 failures of 105 | {[roth2022_security]} |
-| Fix difficulty | Errors an operator can fix by editing one line were fixed at 5.7%; errors needing a change at an external provider at 1.6% — the message was identical | {[czybik2023_lazy]} |+| Fix difficulty | Errors an operator can fix by editing one line were fixed at 5.7%; errors needing a change at an external provider at 1.6% — same template; the authors' explanation ("We assume that these are often non-trivial to fix"), not a tested one | {[czybik2023_lazy]} |
 | Escalation | When 1.5% of notified developers acted, the authors handed the list to the app stores; Google Play "responded to our request quickly", two other stores did not | {[nan2023_spying]} | | Escalation | When 1.5% of notified developers acted, the authors handed the list to the app stores; Google Play "responded to our request quickly", two other stores did not | {[nan2023_spying]} |
  
Line 418: Line 418:
 ==== The receiving end ==== ==== The receiving end ====
  
-Three papers look at what happens after a report arrives, and none of them is encouraging about the unsolicited individual email. {[bijmans2026_tickets]} analysed 1.3 million abuse reports seized by Dutch law enforcement from one hosting provider with a bulletproof reputation: **2.6% of reports were ever linked to a notification to the customer**. The rate was not about the abuse; it was about the reporter. Netcraft's phishing reports led to a customer notification 72% of the time and Spamhaus listings 60–84%, while "individual abuse reporting is often easily ignored" — one DMCA service filed 126,085 reports that produced four notifications, and one honeypot operator's 30,467 automated reports produced none. The authors are explicit that one abusive provider does not generalise; the direction of the effect is still the point: **a report is acted on in proportion to what the reporter can do to the recipient's business**, which a research group cannot do at all. On the vendor side, {[hastings2016_weakkeys]} followed the 2012 disclosure of weak-key generation to 61 device vendors — of the 37 with weak RSA keys, 5 issued a public advisory, 11 answered privately, 3 sent an auto-reply, 18 nothing — and found that "vendor notification, positive vendor responses, and even vendor-produced public security advisories appear to have little correlation with end-user security": vulnerable populations kept growing for years at vendors that had published an advisory. And when the recipient does have a formal channel, it may be the wrong shape: {[bhattacharya2026_asi]} briefed 58 of the 100 most popular services about account-security flaws through their disclosure programmes, got 38 acknowledgements and 25 substantive replies, most classified "Not applicable" or "Out of scope", and notes that "HackerOne rate limits disclosures, which slowed our disclosure process immensely"; {[roth2022_security]} found that "many of those that answered instructed us to contact HackerOne" although the message never used the word vulnerability.+Four papers look at what happens after a report arrives — {[stivala2026_behind]} aboveon how hosting providers triage, and three more here — and none of them is encouraging about either the unsolicited individual email or the generic disclosure form. {[bijmans2026_tickets]} analysed 1.3 million abuse reports seized by Dutch law enforcement from one hosting provider with a bulletproof reputation: **2.6% of reports were ever linked to a notification to the customer**. The rate was not about the abuse; it was about the reporter. Netcraft's phishing reports led to a customer notification 72% of the time and Spamhaus listings 60–84%, while "individual abuse reporting is often easily ignored" — one DMCA service filed 126,085 reports that produced four notifications, and one honeypot operator's 30,467 automated reports produced none. The authors are explicit that one abusive provider does not generalise; the direction of the effect is still the point: **a report is acted on in proportion to what the reporter can do to the recipient's business**, which a research group cannot do at all. On the vendor side, {[hastings2016_weakkeys]} followed the 2012 disclosure of weak-key generation to 61 device vendors — of the 37 with weak RSA keys, 5 issued a public advisory, 11 answered privately, 3 sent an auto-reply, 18 nothing — and found that "vendor notification, positive vendor responses, and even vendor-produced public security advisories appear to have little correlation with end-user security": vulnerable populations kept growing for years at vendors that had published an advisory. And when the recipient does have a formal channel, it may be the wrong shape: {[bhattacharya2026_asi]} briefed 58 of the 100 most popular services about account-security flaws through their disclosure programmes, got 38 acknowledgements and 25 substantive replies, most classified "Not applicable" or "Out of scope", and notes that "HackerOne rate limits disclosures, which slowed our disclosure process immensely"; {[roth2022_security]} found that "many of those that answered instructed us to contact HackerOne" although the message never used the word vulnerability.
  
 ==== Deciding not to run a campaign ==== ==== Deciding not to run a campaign ====
  
-Two recent papers reasoned their way out of a campaign in print, and both are worth citing when you do the same. {[rao2024_unfiltered]} found domains whose cloud email filters could be bypassed, cited the ">80% unresponsive" of {[bennett2022_spfail]}, and "elected to work directly and closely with filtering service providers to update their documentation, notify their customers (with whom they do have an existing business relationship) and resolve the issues identified" — seven vendors instead of thousands of domains. {[ryan2023_passivessh]} recovered SSH host keys from millions of devices, disclosed to four manufacturers and CERT/CC, and wrote: "We considered notifying operators of affected devices whose keys we had recovered, but we determined this would be infeasible." Neither is a shortcut; both state the alternative they took and why.+Two recent papers reasoned their way out of a campaign in print, and both are worth citing when you do the same. {[rao2024_unfiltered]} found domains whose cloud email filters could be bypassed, cited {[bennett2022_spfail]} ("over 80% of the domains contacted were unresponsive"), and "elected to work directly and closely with filtering service providers to update their documentation, notify their customers (with whom they do have an existing business relationship) and resolve the issues identified" — seven filtering vendors instead of the 1,262 misconfigured domains. {[ryan2023_passivessh]} recovered 189 unique SSH host keys from a scan of millions of devices, disclosed to four manufacturers and CERT/CC, and wrote: "We considered notifying operators of affected devices whose keys we had recovered, but we determined this would be infeasible." Neither is a shortcut; both state the alternative they took and why.
  
 ==== Covert notification ==== ==== Covert notification ====
Line 466: Line 466:
   * **Seven venues only.** EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS are absent, and CHI/SOUPS are where a good deal of the operator-facing usable-security work appears. NDSS 2016 and NDSS 2018 full text was not retrieved at all, which is why {[stock2018_didnt]} — the canonical follow-up study — was read from the publisher's PDF rather than from the corpus.   * **Seven venues only.** EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS are absent, and CHI/SOUPS are where a good deal of the operator-facing usable-security work appears. NDSS 2016 and NDSS 2018 full text was not retrieved at all, which is why {[stock2018_didnt]} — the canonical follow-up study — was read from the publisher's PDF rather than from the corpus.
   * ''ethics.notifiedAffectedParties'' **agreed with an independent re-extraction on 67% of papers**, so treat every percentage above as accurate to a few points, not to the decimal.((That 67% was measured on 100 papers of the previous, 4,322-paper extraction run and has not been re-measured on the current corpus. Treat it as the right order of magnitude.)) ''ethics.disclosureDetail'' is free text capped at 20 words, so it is reported only as folded families and only as a ranking; the fold rules and the unmapped residue are on the provenance page.   * ''ethics.notifiedAffectedParties'' **agreed with an independent re-extraction on 67% of papers**, so treat every percentage above as accurate to a few points, not to the decimal.((That 67% was measured on 100 papers of the previous, 4,322-paper extraction run and has not been re-measured on the current corpus. Treat it as the right order of magnitude.)) ''ethics.disclosureDetail'' is free text capped at 20 words, so it is reported only as folded families and only as a ranking; the fold rules and the unmapped residue are on the provenance page.
-  * **The 32 campaigns are hand-classified, not a census.** Two sentence-level rules run over 5,869 full-text files (ten more than the 5,859 extraction records — a few papers have text but no record). The narrow rule produced 179 candidates: 28 campaigns and 151 rejections with a stated reason. A wider recall probe, added on 2026-09-04 after the narrow rule was found to miss a 111,951-mail campaign, produced 78 more: 3 campaigns and 75 rejections. One campaign, {[roth2022_security]}, was caught by neither rule and found through a full-text search for ''security.txt''; three rejections carry over from a wider first-pass scan. **Nothing is unreviewed**but 201 of the 229 rejections were screened on every sentence carrying a notification verb rather than on a full read, and the provenance page says which. A third rule would find more; the table above is a curated reading list, not a census.+  * **The 32 campaigns are hand-classified, not a census.** Two sentence-level rules run over 5,869 full-text files (ten more than the 5,859 extraction records — a few papers have text but no record). The narrow rule produced 179 candidates: 28 campaigns and 151 rejections with a stated reason. A wider recall probe, added on 2026-09-04 after the narrow rule was found to miss a 111,951-mail campaign, produced 78 more: 3 campaigns and 75 rejections. One campaign, {[roth2022_security]}, was caught by neither rule and found through a full-text search for ''security.txt''; three rejections carry over from a wider first-pass scan. Nothing is unreviewed**but most rejections are screens, not reads**: 126 on every sentence carrying a notification verb, 75 on the probe-matching sentences alone, 28 on a full read of the relevant passages or paper; the provenance page says which. A third rule would find more; the table above is a curated reading list, not a census.
  
 The complete query log, the report script with its unedited output, the folding rules with their unmapped residue, the spot-checked quotes, and every external source that was rejected are on **[[provenance:practices:notifying_websites]]**. Corpus-wide caveats are on [[Literature:Corpus]]. The complete query log, the report script with its unedited output, the folding rules with their unmapped residue, the spot-checked quotes, and every external source that was rejected are on **[[provenance:practices:notifying_websites]]**. Corpus-wide caveats are on [[Literature:Corpus]].
practices/notifying_websites.1788526690.txt.gz · Last modified: by karel.kubicek.claude