User Tools

Site Tools


provenance:privacy:email_tracking

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
provenance:privacy:email_tracking [2026/09/02 07:14] – Substantiate the quote-check claim: name the eight below-threshold quotes read by hand with their longest contiguous run, and correct the claim about the four lowest-coverage papers (two were not independently checked; verify_email_figures.mjs gained Reav karel.kubicek.claudeprovenance:privacy:email_tracking [2026/09/02 07:41] (current) – Add sections 14.6-14.8: the generic (Fable) pass, the re-verification pass, and the prose-mode attribution guard that came out of it. Fix the provenance page's own three-way disagreement about its figure count, the stale script annotations it republishes, karel.kubicek.claude
Line 56: Line 56:
 ===== 3. The population rule ===== ===== 3. The population rule =====
  
-**There is no query for this page.** No field in the extraction means "the measured object is a message". The population is a published candidate pool plus one hand-written verdict per paperBoth are in ''scripts/msg_fold.mjs'' and both are printed by the report.+**There is no query for this page.** No field in the extraction means "the measured object is a message". The population is a published candidate pool plus a verdict per paper — but **not** a hand-written verdict for all 354 of them, and an earlier draft of this page and of the content page both said soThe real accounting, which the report now prints: 
 + 
 +^ ^ Papers ^ 
 +in the hand map (''MAP'') | **183** | 
 +| ... of which also carry a written reason | 125 | 
 +| ... the rest are ''PHISH''''SOC'' / ''WEBPIXEL'' rows where the verdict is the reason | 58 | 
 +| **not** in the hand map, ''OFF'' by default | **171** | 
 +... given a reason by a named ''OFF_FAMILIES'' rule | 46 | 
 +| ... carrying only "not individually annotated" | 125 | 
 + 
 +All 173 non-''OFF'' papers are hand-mapped with a reason. The 125 defaulted papers were screened at **title** level and nothing further was recorded about them individually. Until 2026-09-02 the default reason string read //"read and judged off-topic: the message vocabulary is incidental to the object measured"// and was printed per paper in the published residue — a canned sentence asserting a per-paper reading that did not happen, and one that would have hidden a typo'd ''MAP'' key as a silent ''OFF''. It now says what is true.
  
 ==== 3.1 Pool signals ==== ==== 3.1 Pool signals ====
Line 140: Line 150:
 ''verify_email_figures.mjs'' checks every literal per-paper figure quoted on the content page against that paper's own ''paper.cols.txt'', **not** against ''evidence.quote'' — ''detection[].prevalence'' is a model summary, so a figure can be right in the paper and wrong in the extraction. Whitespace is collapsed on both sides. ''verify_email_figures.mjs'' checks every literal per-paper figure quoted on the content page against that paper's own ''paper.cols.txt'', **not** against ''evidence.quote'' — ''detection[].prevalence'' is a model summary, so a figure can be right in the paper and wrong in the extraction. Whitespace is collapsed on both sides.
  
-**60 figures across 22 papers, all present.** The pass earned its keep three times:+**78 figures across 24 papers, all present.** (It was 60 across 22 when this section was first written and grew twice during review; the count here is now taken from the script's own summary line in §12, not retyped.) The pass earned its keep three times:
  
   - **{[agarwal2025_fishing]}: the mathematical-italic //k//.** The paper typesets "27.7𝑘 smishing messages, 19.3𝑘 sender IDs, and 20𝑘 URLs" with U+1D458, not ASCII ''k''. An obvious ASCII needle reported three present figures as missing. The needles now carry the real codepoint and a comment saying why.   - **{[agarwal2025_fishing]}: the mathematical-italic //k//.** The paper typesets "27.7𝑘 smishing messages, 19.3𝑘 sender IDs, and 20𝑘 URLs" with U+1D458, not ASCII ''k''. An obvious ASCII needle reported three present figures as missing. The needles now carry the real codepoint and a comment saying why.
Line 172: Line 182:
 | {[li2025_hades]} | 53% | 10 words, //"mail servers of popular ESPs and websites have been included"// | present; the 39,201 (76.88%) is immediately before it | | {[li2025_hades]} | 53% | 10 words, //"mail servers of popular ESPs and websites have been included"// | present; the 39,201 (76.88%) is immediately before it |
  
-The pattern is uniform enough to be worth stating as a rule: **a coverage figure in the 40–60% band on a two-column paper is a column splice, not a fabrication.** A coverage figure below about 25% is worth opening. Four of the 62 are below that — {[tu2016_security]} (23%), {[utz2023_comparing]} (20%), {[reaves2016_sending]} (20%) and {[agarwal2025_fishing]} (18%). Checking that claim is how ''verify_email_figures.mjs'' gained two more papers: {[agarwal2025_fishing]} was already covered, {[reaves2016_sending]} was not and the page quotes three of its figures, so it was added (386,327 messages, 522 containing email addresses, 14 months, over 400 numbers — all present); {[jiang2013_greystar]} was added in the same pass. **The page quotes no figure from the remaining two**, so nothing on it rests on those two quotes.+The pattern is uniform enough to be worth stating as a rule: **a coverage figure in the 40–60% band on a two-column paper is a column splice, not a fabrication.** A coverage figure below about 25% is worth opening, and the rule above was inferred from the middle of the distribution, so the tail was checked separately**The two lowest were read by hand as well, and both are present:** 
 + 
 +^ Paper ^ Coverage ^ What the paper says ^ 
 +| {[tu2016_security]} | 23% | //"victim account's activity logs, as shown in Figure 8(d), confirms … that those three attack actions are successful"// — the quote's own ellipsis covers the figure reference, and a column break falls inside it | 
 +| {[utz2023_comparing]} | 20% | //"No HTTPS was rarest, / with 2.85 % of sites"// — verbatim, split across the column boundary | 
 + 
 +Four of the 62 are below 25% — {[tu2016_security]} (23%), {[utz2023_comparing]} (20%), {[reaves2016_sending]} (20%) and {[agarwal2025_fishing]} (18%). Checking that claim is how ''verify_email_figures.mjs'' gained two more papers: {[agarwal2025_fishing]} was already covered, {[reaves2016_sending]} was not and the page quotes three of its figures, so it was added (386,327 messages, 522 containing email addresses, 14 months, over 400 numbers — all present); {[jiang2013_greystar]} was added in the same pass. **The page quotes no figure from the remaining two**, so nothing on it rests on those two quotes.
  
 ===== 9. External sources, and how each was verified ===== ===== 9. External sources, and how each was verified =====
Line 232: Line 248:
   admitted by hand          : 3   admitted by hand          : 3
 hand-mapped verdicts        : 183 hand-mapped verdicts        : 183
 +  ... of which carry a written reason      : 125 (182 MAP entries; the rest are
 +                                             PHISH / SOC / WEBPIXEL rows where the verdict IS the reason)
 +NOT in the hand map, verdict OFF by default: 171
 +  ... given a reason by an OFF_FAMILIES rule: 46
 +  ... carrying only "not individually annotated": 125
 +  These were screened at TITLE level. Nothing further was recorded per paper.
 ON the page                 : 70 ON the page                 : 70
 rejected outside the pool   : 9 (documented, see section H) rejected outside the pool   : 9 (documented, see section H)
Line 373: Line 395:
 artifacts.availability is computed over the EMPIRICAL corpus, as OVERVIEW.md does, artifacts.availability is computed over the EMPIRICAL corpus, as OVERVIEW.md does,
 because a non-empirical paper releasing nothing is not a reporting gap. because a non-empirical paper releasing nothing is not a reporting gap.
 +
 +--- Study shape: platforms, study types, crawl configuration, participants
 +Platform              Papers  Share of 70
 +--------------------  ------  -----------
 +other-online-service  54      77.1%      
 +web                   33      47.1%      
 +mobile                18      25.7%      
 +offline                     5.7%       
 +iot                         1.4%       
 +Study type                  Papers  Share of 70
 +--------------------------  ------  -----------
 +manual-audit                42      60.0%      
 +system-or-defence-proposal  35      50.0%      
 +existing-dataset-analysis   33      47.1%      
 +automated-web-crawl         28      40.0%      
 +network-scan-or-probe       15      21.4%      
 +code-or-binary-analysis     12      17.1%      
 +mobile-app-analysis         11      15.7%      
 +interview-or-survey         11      15.7%      
 +user-study                  10      14.3%      
 +simulation-or-theory-only         1.4%       
 +ran an automated web crawl (crawlConfig OR studyTypes): 29/70 = 41.4%
 +have a recorded crawlConfig object              : 27
 +crawlConfig field  Stated (of 27)  Share
 +-----------------  --------------  -----
 +consentAction      14              51.9%
 +statefulness                     25.9%
 +headless                         18.5%
 +browsers           16              59.3%
 +interactionDepth   22              81.5%
 +authentication     22              81.5%
 +recruited human participants: 12/70 = 17.1%
 +temporal.mode (papers, multi-valued)  Papers
 +------------------------------------  ------
 +active-probing                        27    
 +live-crawl                            26    
 +passive-collection                    25    
 +existing-dataset                      19    
  
 --- Ethics and legal assessment, on-page papers against the corpus base rate --- Ethics and legal assessment, on-page papers against the corpus base rate
Line 1283: Line 1343:
 --- OFF — 181 papers — homograph or unrelated --- OFF — 181 papers — homograph or unrelated
   IMC/2010/estimating-and-sampling-graphs-with-multidimensional-random-walks   IMC/2010/estimating-and-sampling-graphs-with-multidimensional-random-walks
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2010/inference-and-analysis-of-formal-models-of-botnet-command-and-control-protocols   CCS/2010/inference-and-analysis-of-formal-models-of-botnet-command-and-control-protocols
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2010/dissecting-one-click-frauds   CCS/2010/dissecting-one-click-frauds
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2010/building-a-dynamic-reputation-system-for-dns   USENIX/2010/building-a-dynamic-reputation-system-for-dns
       network-layer work that uses spam as a label source       network-layer work that uses spam as a label source
   USENIX/2010/searching-the-searchers-with-searchaudit   USENIX/2010/searching-the-searchers-with-searchaudit
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IMC/2011/gq-practical-containment-for-measuring-modern-malware-systems   IMC/2011/gq-practical-containment-for-measuring-modern-malware-systems
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
Line 1297: Line 1357:
       network-layer work that uses spam as a label source       network-layer work that uses spam as a label source
   USENIX/2011/dirty-jobs-the-role-of-freelance-labor-in-web-service-abuse   USENIX/2011/dirty-jobs-the-role-of-freelance-labor-in-web-service-abuse
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2011/forensic-triage-for-mobile-phones-with-dec0de   USENIX/2011/forensic-triage-for-mobile-phones-with-dec0de
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2011/measuring-and-analyzing-search-redirection-attacks-in-the-illicit-online-prescri   USENIX/2011/measuring-and-analyzing-search-redirection-attacks-in-the-illicit-online-prescri
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2012/detecting-money-stealing-apps-in-alternative-android-markets   CCS/2012/detecting-money-stealing-apps-in-alternative-android-markets
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
Line 1307: Line 1367:
       legitimacy of webmail accounts; the unit is the account, not the message       legitimacy of webmail accounts; the unit is the account, not the message
   CCS/2012/mobile-data-charging-new-attacks-and-countermeasures   CCS/2012/mobile-data-charging-new-attacks-and-countermeasures
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2012/guess-who-s-texting-you-evaluating-the-security-of-smartphone-messaging-applicat   NDSS/2012/guess-who-s-texting-you-evaluating-the-security-of-smartphone-messaging-applicat
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2012/hey-you-get-off-of-my-market-detecting-malicious-apps-in-official-and-alternativ   NDSS/2012/hey-you-get-off-of-my-market-detecting-malicious-apps-in-official-and-alternativ
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2013/cross-origin-pixel-stealing-timing-attacks-using-css-filters   CCS/2013/cross-origin-pixel-stealing-timing-attacks-using-css-filters
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2013/whyper-towards-automating-risk-assessment-of-mobile-applications   USENIX/2013/whyper-towards-automating-risk-assessment-of-mobile-applications
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2014/dialing-back-abuse-on-phone-verified-accounts   CCS/2014/dialing-back-abuse-on-phone-verified-accounts
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2014/are-you-ready-to-lock   CCS/2014/are-you-ready-to-lock
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2013/detecting-passive-content-leaks-and-pollution-in-android-applications   NDSS/2013/detecting-passive-content-leaks-and-pollution-in-android-applications
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
   USENIX/2013/jekyll-on-ios-when-benign-apps-become-evil   USENIX/2013/jekyll-on-ios-when-benign-apps-become-evil
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2013/automatically-inferring-the-evolution-of-malicious-activity-on-the-internet   NDSS/2013/automatically-inferring-the-evolution-of-malicious-activity-on-the-internet
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2013/protecting-sensitive-web-content-from-client-side-vulnerabilities-with-cryptons   CCS/2013/protecting-sensitive-web-content-from-client-side-vulnerabilities-with-cryptons
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2013/vetting-undesirable-behaviors-in-android-apps-with-permission-use-analysis   CCS/2013/vetting-undesirable-behaviors-in-android-apps-with-permission-use-analysis
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
   CCS/2014/consequences-of-connectivity-characterizing-account-hijacking-on-twitter   CCS/2014/consequences-of-connectivity-characterizing-account-hijacking-on-twitter
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   WWW/2013/two-years-of-short-urls-internet-measurement-security-threats-and-countermeasure   WWW/2013/two-years-of-short-urls-internet-measurement-security-threats-and-countermeasure
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2014/autocog-measuring-the-description-to-permission-fidelity-in-android-applications   CCS/2014/autocog-measuring-the-description-to-permission-fidelity-in-android-applications
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
   CCS/2014/real-threats-to-your-data-bills-security-loopholes-and-defenses-in-mobile-data-c   CCS/2014/real-threats-to-your-data-bills-security-loopholes-and-defenses-in-mobile-data-c
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2014/understanding-the-dark-side-of-domain-parking   USENIX/2014/understanding-the-dark-side-of-domain-parking
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IMC/2014/handcrafted-fraud-and-extortion-manual-account-hijacking-in-the-wild   IMC/2014/handcrafted-fraud-and-extortion-manual-account-hijacking-in-the-wild
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2015/leakage-abuse-attacks-against-searchable-encryption   CCS/2015/leakage-abuse-attacks-against-searchable-encryption
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2015/copperdroid-automatic-reconstruction-of-android-malware-behaviors   NDSS/2015/copperdroid-automatic-reconstruction-of-android-malware-behaviors
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
Line 1349: Line 1409:
       network-layer work that uses spam as a label source       network-layer work that uses spam as a label source
   NDSS/2015/what-s-in-your-dongle-and-bank-account-mandatory-and-discretionary-protection-of   NDSS/2015/what-s-in-your-dongle-and-bank-account-mandatory-and-discretionary-protection-of
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2015/android-permissions-remystified-a-field-study-on-contextual-integrity   USENIX/2015/android-permissions-remystified-a-field-study-on-contextual-integrity
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
   USENIX/2015/cloudy-with-a-chance-of-breach-forecasting-cyber-security-incidents   USENIX/2015/cloudy-with-a-chance-of-breach-forecasting-cyber-security-incidents
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2014/on-the-mismanagement-and-maliciousness-of-networks   NDSS/2014/on-the-mismanagement-and-maliciousness-of-networks
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2014/man-vs-machine-practical-adversarial-detection-of-malicious-crowdsourcing-worker   USENIX/2014/man-vs-machine-practical-adversarial-detection-of-malicious-crowdsourcing-worker
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2014/dspin-detecting-automatically-spun-content-on-the-web   NDSS/2014/dspin-detecting-automatically-spun-content-on-the-web
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IMC/2015/affiliate-crookies-characterizing-affiliate-marketing-abuse   IMC/2015/affiliate-crookies-characterizing-affiliate-marketing-abuse
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IMC/2015/detecting-malicious-activity-with-dns-backscatter   IMC/2015/detecting-malicious-activity-with-dns-backscatter
       network-layer work that uses spam as a label source       network-layer work that uses spam as a label source
   IMC/2015/leveraging-internet-background-radiation-for-opportunistic-network-analysis   IMC/2015/leveraging-internet-background-radiation-for-opportunistic-network-analysis
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IMC/2015/the-dark-menace-characterizing-network-based-attacks-in-the-cloud   IMC/2015/the-dark-menace-characterizing-network-based-attacks-in-the-cloud
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2015/evilcohort-detecting-communities-of-malicious-accounts-on-online-services   USENIX/2015/evilcohort-detecting-communities-of-malicious-accounts-on-online-services
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2016/featuresmith-automatically-engineering-features-for-malware-detection-by-mining   CCS/2016/featuresmith-automatically-engineering-features-for-malware-detection-by-mining
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
Line 1375: Line 1435:
       interconnect (SIM-box) bypass fraud: the object is call routing revenue, not a message delivered to a user       interconnect (SIM-box) bypass fraud: the object is call routing revenue, not a message delivered to a user
   CCS/2016/predator-proactive-recognition-and-elimination-of-domain-abuse-at-time-of-regist   CCS/2016/predator-proactive-recognition-and-elimination-of-domain-abuse-at-time-of-regist
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IMC/2016/bdrmap-inference-of-borders-between-ip-networks   IMC/2016/bdrmap-inference-of-borders-between-ip-networks
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   PETS/2016/tales-from-the-dark-side-privacy-dark-strategies-and-privacy-dark-patterns   PETS/2016/tales-from-the-dark-side-privacy-dark-strategies-and-privacy-dark-patterns
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2016/the-ever-changing-labyrinth-a-large-scale-analysis-of-wildcard-dns-powered-black   USENIX/2016/the-ever-changing-labyrinth-a-large-scale-analysis-of-wildcard-dns-powered-black
       network-layer work that uses spam as a label source       network-layer work that uses spam as a label source
   IEEE-SP/2017/obstacles-to-the-adoption-of-secure-communication-tools   IEEE-SP/2017/obstacles-to-the-adoption-of-secure-communication-tools
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IMC/2017/understanding-the-role-of-registrars-in-dnssec-deployment   IMC/2017/understanding-the-role-of-registrars-in-dnssec-deployment
       network-layer work that uses spam as a label source       network-layer work that uses spam as a label source
Line 1389: Line 1449:
       pub/sub or SDN "unsubscribe" — protocol verb, not marketing mail       pub/sub or SDN "unsubscribe" — protocol verb, not marketing mail
   USENIX/2017/characterizing-the-nature-and-dynamics-of-tor-exit-blocking   USENIX/2017/characterizing-the-nature-and-dynamics-of-tor-exit-blocking
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2017/malton-towards-on-device-non-invasive-mobile-malware-analysis-for-art   USENIX/2017/malton-towards-on-device-non-invasive-mobile-malware-analysis-for-art
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
   CCS/2018/detecting-attacks-against-robotic-vehicles-a-control-invariant-approach   CCS/2018/detecting-attacks-against-robotic-vehicles-a-control-invariant-approach
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2018/towards-paving-the-way-for-large-scale-windows-malware-analysis-generic-binary-u   CCS/2018/towards-paving-the-way-for-large-scale-windows-malware-analysis-generic-binary-u
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
   IEEE-SP/2018/computer-security-and-privacy-for-refugees-in-the-united-states   IEEE-SP/2018/computer-security-and-privacy-for-refugees-in-the-united-states
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IMC/2018/is-the-web-ready-for-ocsp-must-staple   IMC/2018/is-the-web-ready-for-ocsp-must-staple
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   PETS/2018/toward-distribution-estimation-under-local-differential-privacy-with-small-sampl   PETS/2018/toward-distribution-estimation-under-local-differential-privacy-with-small-sampl
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2018/from-patching-delays-to-infection-symptoms-using-risk-profiles-for-an-early-disc   USENIX/2018/from-patching-delays-to-infection-symptoms-using-risk-profiles-for-an-early-disc
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2018/o-single-sign-off-where-art-thou-an-empirical-analysis-of-single-sign-on-account   USENIX/2018/o-single-sign-off-where-art-thou-an-empirical-analysis-of-single-sign-on-account
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2019/a-systematic-framework-to-generate-invariants-for-anomaly-detection-in-industrial-control-systems   NDSS/2019/a-systematic-framework-to-generate-invariants-for-anomaly-detection-in-industrial-control-systems
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2019/cleaning-up-the-internet-of-evil-things-real-world-evidence-on-isp-and-consumer-efforts-to-remove-mirai   NDSS/2019/cleaning-up-the-internet-of-evil-things-real-world-evidence-on-isp-and-consumer-efforts-to-remove-mirai
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2019/privacy-attacks-to-the-4g-and-5g-cellular-paging-protocols-using-side-channel-information   NDSS/2019/privacy-attacks-to-the-4g-and-5g-cellular-paging-protocols-using-side-channel-information
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2019/a-billion-open-interfaces-for-eve-and-mallory-mitm-dos-and-tracking-attacks-on-i   USENIX/2019/a-billion-open-interfaces-for-eve-and-mallory-mitm-dos-and-tracking-attacks-on-i
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2019/policylint-investigating-internal-privacy-policy-contradictions-on-google-play   USENIX/2019/policylint-investigating-internal-privacy-policy-contradictions-on-google-play
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2019/reading-the-tea-leaves-a-comparative-analysis-of-threat-intelligence   USENIX/2019/reading-the-tea-leaves-a-comparative-analysis-of-threat-intelligence
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   WWW/2019/evaluating-login-challenges-as-adefense-against-account-takeover   WWW/2019/evaluating-login-challenges-as-adefense-against-account-takeover
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   WWW/2019/exploring-user-behavior-in-email-re-finding-tasks   WWW/2019/exploring-user-behavior-in-email-re-finding-tasks
       information retrieval inside a mailbox — no privacy or security measurement       information retrieval inside a mailbox — no privacy or security measurement
Line 1425: Line 1485:
       what a notification should say, not what it discloses to third parties       what a notification should say, not what it discloses to third parties
   WWW/2019/understanding-the-evolution-of-mobile-app-ecosystems-a-longitudinal-measurement   WWW/2019/understanding-the-evolution-of-mobile-app-ecosystems-a-longitudinal-measurement
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2020/towards-attribution-in-mobile-markets-identifying-developer-account-polymorphism   CCS/2020/towards-attribution-in-mobile-markets-identifying-developer-account-polymorphism
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IEEE-SP/2020/the-many-kinds-of-creepware-used-for-interpersonal-attacks   IEEE-SP/2020/the-many-kinds-of-creepware-used-for-interpersonal-attacks
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IEEE-SP/2020/karonte-detecting-insecure-multi-binary-interactions-in-embedded-firmware   IEEE-SP/2020/karonte-detecting-insecure-multi-binary-interactions-in-embedded-firmware
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
   IMC/2020/revisiting-transactional-statistics-of-high-scalability-blockchains   IMC/2020/revisiting-transactional-statistics-of-high-scalability-blockchains
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2020/blag-improving-the-accuracy-of-blacklists   NDSS/2020/blag-improving-the-accuracy-of-blacklists
       network-layer work that uses spam as a label source       network-layer work that uses spam as a label source
Line 1439: Line 1499:
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
   NDSS/2020/surfingattack-interactive-hidden-attack-on-voice-assistants-using-ultrasonic-guided-waves   NDSS/2020/surfingattack-interactive-hidden-attack-on-voice-assistants-using-ultrasonic-guided-waves
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   PETS/2020/sok-anatomy-of-data-breaches   PETS/2020/sok-anatomy-of-data-breaches
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2020/security-analysis-of-unified-payments-interface-and-payment-apps-in-india   USENIX/2020/security-analysis-of-unified-payments-interface-and-payment-apps-in-india
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2020/teerex-discovery-and-exploitation-of-memory-corruption-vulnerabilities-in-sgx-en   USENIX/2020/teerex-discovery-and-exploitation-of-memory-corruption-vulnerabilities-in-sgx-en
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   WWW/2020/fairrec-two-sided-fairness-for-personalized-recommendations-in-two-sided-platfor   WWW/2020/fairrec-two-sided-fairness-for-personalized-recommendations-in-two-sided-platfor
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2021/dont-forget-the-stuffing-revisiting-the-security-impact-of-typo-tolerant-passwor   CCS/2021/dont-forget-the-stuffing-revisiting-the-security-impact-of-typo-tolerant-passwor
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2021/alchemist-fusing-application-and-audit-logs-for-precise-attack-provenance-without-instrumentation   NDSS/2021/alchemist-fusing-application-and-audit-logs-for-precise-attack-provenance-without-instrumentation
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2021/from-whois-to-whowas-a-large-scale-measurement-study-of-domain-registration-privacy-under-the-gdpr   NDSS/2021/from-whois-to-whowas-a-large-scale-measurement-study-of-domain-registration-privacy-under-the-gdpr
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   PETS/2021/the-motivated-can-encrypt-even-with-pgp   PETS/2021/the-motivated-can-encrypt-even-with-pgp
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   PETS/2021/we-three-brothers-have-always-known-everything-of-each-other-a-cross-cultural-st   PETS/2021/we-three-brothers-have-always-known-everything-of-each-other-a-cross-cultural-st
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   PETS/2021/defining-privacy-how-users-interpret-technical-terms-in-privacy-policies   PETS/2021/defining-privacy-how-users-interpret-technical-terms-in-privacy-policies
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2021/alpaca-application-layer-protocol-confusion-analyzing-and-mitigating-cracks-in-t   USENIX/2021/alpaca-application-layer-protocol-confusion-analyzing-and-mitigating-cracks-in-t
       network-layer work that uses spam as a label source       network-layer work that uses spam as a label source
   USENIX/2021/effect-of-mood-location-trust-and-presence-of-others-on-video-based-social-authe   USENIX/2021/effect-of-mood-location-trust-and-presence-of-others-on-video-based-social-authe
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2021/now-im-a-bit-angry-individuals-awareness-perception-and-responses-to-data-breach   USENIX/2021/now-im-a-bit-angry-individuals-awareness-perception-and-responses-to-data-breach
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2021/privatedrop-practical-privacy-preserving-authentication-for-apple-airdrop   USENIX/2021/privatedrop-practical-privacy-preserving-authentication-for-apple-airdrop
       contact discovery, but the contribution is a protocol; no measurement of an address space       contact discovery, but the contribution is a protocol; no measurement of an address space
   USENIX/2021/strategies-and-perceived-risks-of-sending-sensitive-documents   USENIX/2021/strategies-and-perceived-risks-of-sending-sensitive-documents
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2021/the-hijackers-guide-to-the-galaxy-off-path-taking-over-internet-resources   USENIX/2021/the-hijackers-guide-to-the-galaxy-off-path-taking-over-internet-resources
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2021/why-older-adults-dont-use-password-managers   USENIX/2021/why-older-adults-dont-use-password-managers
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   WWW/2021/an-investigation-of-identity-account-inconsistency-in-single-sign-on   WWW/2021/an-investigation-of-identity-account-inconsistency-in-single-sign-on
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   WWW/2021/privacy-policies-over-time-curation-and-analysis-of-a-million-document-dataset   WWW/2021/privacy-policies-over-time-curation-and-analysis-of-a-million-document-dataset
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2022/stolenencoder-stealing-pre-trained-encoders-in-self-supervised-learning   CCS/2022/stolenencoder-stealing-pre-trained-encoders-in-self-supervised-learning
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IMC/2022/stop-drop-and-roa-effectiveness-of-defenses-through-the-lens-of-drop   IMC/2022/stop-drop-and-roa-effectiveness-of-defenses-through-the-lens-of-drop
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2022/auto-draft-238   NDSS/2022/auto-draft-238
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   PETS/2022/athena-probabilistic-verification-of-machine-unlearning   PETS/2022/athena-probabilistic-verification-of-machine-unlearning
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   PETS/2022/user-perceptions-of-gmail-s-confidential-mode   PETS/2022/user-perceptions-of-gmail-s-confidential-mode
       usability of one provider feature       usability of one provider feature
Line 1493: Line 1553:
       network-layer work that uses spam as a label source       network-layer work that uses spam as a label source
   USENIX/2022/pre-hijacked-accounts-an-empirical-study-of-security-failures-in-user-account-cr   USENIX/2022/pre-hijacked-accounts-an-empirical-study-of-security-failures-in-user-account-cr
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2022/sgxfuzz-efficiently-synthesizing-nested-structures-for-sgx-enclave-fuzzing   USENIX/2022/sgxfuzz-efficiently-synthesizing-nested-structures-for-sgx-enclave-fuzzing
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
   CCS/2023/black-ostrich-web-application-scanning-with-string-solvers   CCS/2023/black-ostrich-web-application-scanning-with-string-solvers
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2023/comprehension-from-chaos-towards-informed-consent-for-private-computation   CCS/2023/comprehension-from-chaos-towards-informed-consent-for-private-computation
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2023/fetchbench-systematic-identification-and-characterization-of-proprietary-prefetc   CCS/2023/fetchbench-systematic-identification-and-characterization-of-proprietary-prefetc
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2023/i-still-know-what-you-watched-last-sunday-privacy-of-the-hbbtv-protocol-in-the-european-smart-tv-landscape   NDSS/2023/i-still-know-what-you-watched-last-sunday-privacy-of-the-hbbtv-protocol-in-the-european-smart-tv-landscape
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   PETS/2023/designing-a-location-trace-anonymization-contest   PETS/2023/designing-a-location-trace-anonymization-contest
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2023/weve-disabled-mfa-for-you-an-evaluation-of-the-security-and-usability-of-multi-f   CCS/2023/weve-disabled-mfa-for-you-an-evaluation-of-the-security-and-usability-of-multi-f
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   PETS/2023/no-privacy-among-spies-assessing-the-functionality-and-insecurity-of-consumer-an   PETS/2023/no-privacy-among-spies-assessing-the-functionality-and-insecurity-of-consumer-an
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2023/a-large-scale-measurement-of-website-login-policies   USENIX/2023/a-large-scale-measurement-of-website-login-policies
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2023/a-two-decade-retrospective-analysis-of-a-universitys-vulnerability-to-attacks-ex   USENIX/2023/a-two-decade-retrospective-analysis-of-a-universitys-vulnerability-to-attacks-ex
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2023/security-and-privacy-failures-in-popular-2fa-apps   USENIX/2023/security-and-privacy-failures-in-popular-2fa-apps
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2023/basecomp-a-comparative-analysis-for-integrity-protection-in-cellular-baseband-so   USENIX/2023/basecomp-a-comparative-analysis-for-integrity-protection-in-cellular-baseband-so
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
   USENIX/2023/work-from-home-and-covid-19-trajectories-of-endpoint-security-management-in-a-se   USENIX/2023/work-from-home-and-covid-19-trajectories-of-endpoint-security-management-in-a-se
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2023/anatomy-of-a-high-profile-data-breach-dissecting-the-aftermath-of-a-crypto-walle   USENIX/2023/anatomy-of-a-high-profile-data-breach-dissecting-the-aftermath-of-a-crypto-walle
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2023/instructions-unclear-undefined-behaviour-in-cellular-network-specifications   USENIX/2023/instructions-unclear-undefined-behaviour-in-cellular-network-specifications
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
   USENIX/2023/notice-the-imposter-a-study-on-user-tag-spoofing-attack-in-mobile-apps   USENIX/2023/notice-the-imposter-a-study-on-user-tag-spoofing-attack-in-mobile-apps
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2023/policycomp-counterpart-comparison-of-privacy-policies-uncovers-overbroad-persona   USENIX/2023/policycomp-counterpart-comparison-of-privacy-policies-uncovers-overbroad-persona
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2024/maginot-line-assessing-a-new-cross-app-threat-to-pii-as-factor-authentication-in-chinese-mobile-apps   NDSS/2024/maginot-line-assessing-a-new-cross-app-threat-to-pii-as-factor-authentication-in-chinese-mobile-apps
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2024/on-precisely-detecting-censorship-circumvention-in-real-world-networks   NDSS/2024/on-precisely-detecting-censorship-circumvention-in-real-world-networks
       network-layer work that uses spam as a label source       network-layer work that uses spam as a label source
   WWW/2023/unsupervised-anomaly-detection-on-microservice-traces-through-graph-vae   WWW/2023/unsupervised-anomaly-detection-on-microservice-traces-through-graph-vae
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2024/a-first-look-at-security-and-privacy-risks-in-the-rapidapi-ecosystem   CCS/2024/a-first-look-at-security-and-privacy-risks-in-the-rapidapi-ecosystem
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IEEE-SP/2024/combing-for-credentials-active-pattern-extraction-from-smart-reply   IEEE-SP/2024/combing-for-credentials-active-pattern-extraction-from-smart-reply
       ML-privacy work whose training data happens to be mail       ML-privacy work whose training data happens to be mail
   IEEE-SP/2024/deeptheft-stealing-dnn-model-architectures-through-power-side-channel   IEEE-SP/2024/deeptheft-stealing-dnn-model-architectures-through-power-side-channel
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2024/a-mixed-methods-study-on-user-experiences-and-challenges-of-recovery-codes-for-a   USENIX/2024/a-mixed-methods-study-on-user-experiences-and-challenges-of-recovery-codes-for-a
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2024/guardians-of-the-galaxy-content-moderation-in-the-interplanetary-file-system   USENIX/2024/guardians-of-the-galaxy-content-moderation-in-the-interplanetary-file-system
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2024/landscape-exploring-ldap-weaknesses-and-data-leaks-at-internet-scale   USENIX/2024/landscape-exploring-ldap-weaknesses-and-data-leaks-at-internet-scale
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2024/malla-demystifying-real-world-large-language-model-integrated-malicious-services   USENIX/2024/malla-demystifying-real-world-large-language-model-integrated-malicious-services
       LLM-for-hire services, one output of which is spam; the unit is the service       LLM-for-hire services, one output of which is spam; the unit is the service
   USENIX/2024/pixel-thief-exploiting-svg-filter-leakage-in-firefox-and-chrome   USENIX/2024/pixel-thief-exploiting-svg-filter-leakage-in-firefox-and-chrome
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2024/simurai-slicing-through-the-complexity-of-sim-card-security-research   USENIX/2024/simurai-slicing-through-the-complexity-of-sim-card-security-research
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
Line 1555: Line 1615:
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
   IEEE-SP/2016/the-cracked-cookie-jar-http-cookie-hijacking-and-the-exposure-of-private-informa   IEEE-SP/2016/the-cracked-cookie-jar-http-cookie-hijacking-and-the-exposure-of-private-informa
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   CCS/2025/noise-and-stress-dont-help-with-learning-a-qualitative-study-to-inform-design-of   CCS/2025/noise-and-stress-dont-help-with-learning-a-qualitative-study-to-inform-design-of
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IMC/2025/congestion-patterns-in-a-large-scale-rdma-datacenter   IMC/2025/congestion-patterns-in-a-large-scale-rdma-datacenter
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2025/all-your-database-are-belong-to-us-characterizing-database-ransomware-attacks   NDSS/2025/all-your-database-are-belong-to-us-characterizing-database-ransomware-attacks
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2025/whats-done-is-not-whats-claimed-detecting-and-interpreting-inconsistencies-in-app-behaviors   NDSS/2025/whats-done-is-not-whats-claimed-detecting-and-interpreting-inconsistencies-in-app-behaviors
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2025/who-is-trying-to-access-my-account-exploring-user-perceptions-and-reactions-to-risk-based-authentication-notifications   NDSS/2025/who-is-trying-to-access-my-account-exploring-user-perceptions-and-reactions-to-risk-based-authentication-notifications
       same; its 7 `unsubscri` hits are about turning off security alerts       same; its 7 `unsubscri` hits are about turning off security alerts
Line 1571: Line 1631:
       the notification's wording is the treatment; the channel is incidental       the notification's wording is the treatment; the channel is incidental
   NDSS/2025/attributing-open-source-contributions-is-critical-but-difficult-a-systematic-analysis-of-github-practices-and-their-impact-on-software-supply-chain-security   NDSS/2025/attributing-open-source-contributions-is-critical-but-difficult-a-systematic-analysis-of-github-practices-and-their-impact-on-software-supply-chain-security
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2025/evaluating-llm-based-personal-information-extraction-and-countermeasures   USENIX/2025/evaluating-llm-based-personal-information-extraction-and-countermeasures
       ML-privacy work whose training data happens to be mail       ML-privacy work whose training data happens to be mail
   USENIX/2025/scanned-and-scammed-insecurity-by-obsqrity-measuring-user-susceptibility-and-awa   USENIX/2025/scanned-and-scammed-insecurity-by-obsqrity-measuring-user-susceptibility-and-awa
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2025/private-investigator-extracting-personally-identifiable-information-from-large-l   USENIX/2025/private-investigator-extracting-personally-identifiable-information-from-large-l
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2025/generated-data-with-fake-privacy-hidden-dangers-of-fine-tuning-large-language-mo   USENIX/2025/generated-data-with-fake-privacy-hidden-dangers-of-fine-tuning-large-language-mo
       ML-privacy work whose training data happens to be mail       ML-privacy work whose training data happens to be mail
Line 1587: Line 1647:
       ML-privacy work whose training data happens to be mail       ML-privacy work whose training data happens to be mail
   NDSS/2026/tipso-gan-malicious-network-traffic-detection-using-a-novel-optimized-generative-adversarial-network   NDSS/2026/tipso-gan-malicious-network-traffic-detection-using-a-novel-optimized-generative-adversarial-network
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   PETS/2026/because-i-didnt-touch-these-and-even-dont-know-why-i-should-to-change-these-why   PETS/2026/because-i-didnt-touch-these-and-even-dont-know-why-i-should-to-change-these-why
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2026/trojpix-electromagnetic-covert-channels-via-imperceptible-pixel-modulation   USENIX/2026/trojpix-electromagnetic-covert-channels-via-imperceptible-pixel-modulation
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   WWW/2026/truth-with-a-twist-the-rhetoric-of-persuasion-in-professional-vs-community-autho   WWW/2026/truth-with-a-twist-the-rhetoric-of-persuasion-in-professional-vs-community-autho
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IEEE-SP/2015/vetting-ssl-usage-in-applications-with-sslint   IEEE-SP/2015/vetting-ssl-usage-in-applications-with-sslint
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IEEE-SP/2016/triggerscope-towards-detecting-logic-bombs-in-android-applications   IEEE-SP/2016/triggerscope-towards-detecting-logic-bombs-in-android-applications
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
   IEEE-SP/2018/the-spyware-used-in-intimate-partner-violence   IEEE-SP/2018/the-spyware-used-in-intimate-partner-violence
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IEEE-SP/2017/the-password-reset-mitm-attack   IEEE-SP/2017/the-password-reset-mitm-attack
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2025/a-large-scale-measurement-study-of-the-proxy-protocol-and-its-security-implications   NDSS/2025/a-large-scale-measurement-study-of-the-proxy-protocol-and-its-security-implications
       network-layer work that uses spam as a label source       network-layer work that uses spam as a label source
Line 1607: Line 1667:
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
   PETS/2025/gig-work-at-what-cost-exploring-privacy-risks-of-gig-work-platform-participation   PETS/2025/gig-work-at-what-cost-exploring-privacy-risks-of-gig-work-platform-participation
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   PETS/2025/real-world-deniability-in-messaging   PETS/2025/real-world-deniability-in-messaging
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   USENIX/2025/from-alarms-to-real-bugs-multi-target-multi-step-directed-greybox-fuzzing-for-st   USENIX/2025/from-alarms-to-real-bugs-multi-target-multi-step-directed-greybox-fuzzing-for-st
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
   PETS/2025/sok-web-authentication-and-recovery-in-the-age-of-end-to-end-encryption   PETS/2025/sok-web-authentication-and-recovery-in-the-age-of-end-to-end-encryption
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   PETS/2026/chatbot-confessions-large-scale-analysis-of-private-data-disclosure-in-shared-ai   PETS/2026/chatbot-confessions-large-scale-analysis-of-private-data-disclosure-in-shared-ai
       ML-privacy work whose training data happens to be mail       ML-privacy work whose training data happens to be mail
Line 1621: Line 1681:
       systems security; the message vocabulary is incidental       systems security; the message vocabulary is incidental
   IEEE-SP/2017/under-the-shadow-of-sunshine-understanding-and-detecting-bulletproof-hosting-on   IEEE-SP/2017/under-the-shadow-of-sunshine-understanding-and-detecting-bulletproof-hosting-on
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   NDSS/2025/isolategpt-an-execution-isolation-architecture-for-llm-based-agentic-systems   NDSS/2025/isolategpt-an-execution-isolation-architecture-for-llm-based-agentic-systems
       ML-privacy work whose training data happens to be mail       ML-privacy work whose training data happens to be mail
   IEEE-SP/2022/desperate-times-call-for-desperate-measures-user-concerns-with-mobile-loan-apps   IEEE-SP/2022/desperate-times-call-for-desperate-measures-user-concerns-with-mobile-loan-apps
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IEEE-SP/2022/device-fingerprinting-with-peripheral-timestamps   IEEE-SP/2022/device-fingerprinting-with-peripheral-timestamps
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IEEE-SP/2022/towards-automated-auditing-for-account-and-session-management-flaws-in-single-si   IEEE-SP/2022/towards-automated-auditing-for-account-and-session-management-flaws-in-single-si
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IEEE-SP/2024/dnsbomb-a-new-practical-and-powerful-pulsing-dos-attack-exploiting-dns-queries-a   IEEE-SP/2024/dnsbomb-a-new-practical-and-powerful-pulsing-dos-attack-exploiting-dns-queries-a
       network-layer work that uses spam as a label source       network-layer work that uses spam as a label source
Line 1637: Line 1697:
       measures what campaign websites collect, including addresses; no mail was received or measured       measures what campaign websites collect, including addresses; no mail was received or measured
   IEEE-SP/2023/d-dae-defense-penetrating-model-extraction-attacks   IEEE-SP/2023/d-dae-defense-penetrating-model-extraction-attacks
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IEEE-SP/2022/scraping-sticky-leftovers-app-user-information-left-on-servers-after-account-del   IEEE-SP/2022/scraping-sticky-leftovers-app-user-information-left-on-servers-after-account-del
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IEEE-SP/2023/rulekeeper-gdpr-aware-personal-data-compliance-for-web-frameworks   IEEE-SP/2023/rulekeeper-gdpr-aware-personal-data-compliance-for-web-frameworks
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
   IEEE-SP/2024/erasan-efficient-rust-address-sanitizer   IEEE-SP/2024/erasan-efficient-rust-address-sanitizer
-      read and judged off-topic: the message vocabulary is incidental to the object measured+      not individually annotated: screened at title level, matched no off-topic family rule
  
 --- Rejected outside the pool — 9 --- Rejected outside the pool — 9
Line 1670: Line 1730:
   detection[].phenomenon =~ /spam/i   detection[].phenomenon =~ /spam/i
       (not counted) — free text, ~20% run-to-run stable; and the phenomenon string does not say which channel       (not counted) — free text, ~20% run-to-run stable; and the phenomenon string does not say which channel
-  full text =~ /\bunsubscri|List-Unsubscribe|CAN-?SPAM/i +  full text =~ /\bunsubscrib/i 
-      59 papers — MQTT / SDN protocol verb in of 59; CCPA "opt out of sale" in a further 6; one genuine marketing-mail paper+      56 papers — MQTT / SDN protocol verb in 10 of 56; CCPA "opt out of sale" or an advertising opt-out cookie in 3; turning off security notification in 4the remaining 39 have one passing mention each. Zero papers whose OBJECT is an unsubscription mechanism
   full text =~ /\bspam/i   full text =~ /\bspam/i
       954 papers — a fact about these being security venues, not a population       954 papers — a fact about these being security venues, not a population
Line 1766: Line 1826:
       outside: WWW/2026/towards-token-level-text-anomaly-detection       outside: WWW/2026/towards-token-level-text-anomaly-detection
       outside: IEEE-SP/2022/symbexcel-automated-analysis-and-understanding-of-malicious-excel-4-0-macros       outside: IEEE-SP/2022/symbexcel-automated-analysis-and-understanding-of-malicious-excel-4-0-macros
 +
 +  outside-the-pool mentions across the four probes: 72
 +  DISTINCT papers outside the pool that the probes surfaced: 67
  
   Verdict on the 2026-09-02 run. All four probes' outside-the-pool lists were   Verdict on the 2026-09-02 run. All four probes' outside-the-pool lists were
Line 1792: Line 1855:
       recall probe /spam ?trap|honey ?pot (e-?mail|account)|honey ?token/ (8 hits)       recall probe /spam ?trap|honey ?pot (e-?mail|account)|honey ?token/ (8 hits)
   are-you-sure-you-want-to-contact-us-quantifying-the-leakage-of-pii-via-website-c   are-you-sure-you-want-to-contact-us-quantifying-the-leakage-of-pii-via-website-c
-      recall probe detection[].phenomenon =~ /e-?mail|inbox|.../ — a genuine miss, and the oldest paper in the EID slice: contact forms as the gateway from pseudonym to address, five years before Leaky Forms+      recall probe detection[].phenomenon =~ /e-?mail|inbox|.../ — a genuine miss, and the oldest paper in the EID slice: contact forms as the gateway from pseudonym to address, six years before Leaky Forms
  
 ============================================================================== ==============================================================================
Line 1811: Line 1874:
                       measured. Source: data/extract/README.md and OVERVIEW.md's own note                       measured. Source: data/extract/README.md and OVERVIEW.md's own note
                       that the stability table is not recomputed by the overview script.                       that the stability table is not recomputed by the overview script.
-  12,618              mailing lists subscribed to by Englehardt et al., PoPETs 2018  +  12,618              EMAILS collected by Englehardt et al., PoPETs 2018, from 902 
-                      a per-paper figure, checked by verify_email_figures.mjs+                      distinct senders on 15,700 sites crawled. NOT count of mailing 
 +                      lists — that wording was wrong here and on the page until review 
 +                      on 2026-09-02. All four figures are needles in 
 +                      verify_email_figures.mjs
   56.5%               data/extract/OVERVIEW.md's own artifacts.availability rate, quoted on   56.5%               data/extract/OVERVIEW.md's own artifacts.availability rate, quoted on
                       the page only to say it is NOT comparable with this report's 59.5%                       the page only to say it is NOT comparable with this report's 59.5%
Line 1822: Line 1888:
   Llama-3.1-8B        a model name, not a figure   Llama-3.1-8B        a model name, not a figure
   0.30 / 0.10         Google's spam-rate thresholds, percent   0.30 / 0.10         Google's spam-rate thresholds, percent
 +  November 2025       Google's enforcement escalation to temporary and permanent
 +                      rejections, support.google.com/a/answer/14229414
 +  5,000 / 2025-05-05  Microsoft's high-volume-sender threshold and effective date for
 +                      Outlook.com consumer services (SPF + DKIM + DMARC, no RFC 8058).
 +                      Verified INDIRECTLY: the announcement (Defender for Office 365
 +                      blog post 4399730) is served behind a block that refuses this
 +                      host to curl and to a headless browser
 +  550 5.7.515         the Outlook.com NDR quoted in Microsoft's own Q&A threads
 +  Safari 17.0         the release WebKit names for Link Tracking Protection, described
 +                      there as a PRIVATE BROWSING protection
 +  arXiv:2604.06759    preprint of the IEEE S&P 2026 lead-marketing paper
 +  11573454            its IEEE Xplore document id, from resolving the DOI
 +  arXiv:2508.05276    Agarwal et al., 7726 user reports, 1.35 million reports — a
 +                      PREPRINT, cited as such, used for no figure
 +  arXiv:2606.31790    Altwlkany et al., robocalls, 8.7 million records, 65 countries —
 +                      likewise a preprint
 </file> </file>
  
Line 1833: Line 1915:
   ok   i-never-signed-up-for-this :: 87% ||| 87 %   ok   i-never-signed-up-for-this :: 87% ||| 87 %
   ok   i-never-signed-up-for-this :: 12,618   ok   i-never-signed-up-for-this :: 12,618
 +  ok   i-never-signed-up-for-this :: 902
 +  ok   i-never-signed-up-for-this :: 15,700
 +  ok   i-never-signed-up-for-this :: 38%
 +  ok   i-never-signed-up-for-this :: 32%
   ok   characterizing-pixel-tracking-through-the-lens-of-disposable-email-services :: 24.6% ||| 24.6 %   ok   characterizing-pixel-tracking-through-the-lens-of-disposable-email-services :: 24.6% ||| 24.6 %
   ok   characterizing-pixel-tracking-through-the-lens-of-disposable-email-services :: 573,244   ok   characterizing-pixel-tracking-through-the-lens-of-disposable-email-services :: 573,244
Line 1913: Line 1999:
  
  
-24 papers, 74 figures: 74 present, 0 MISSING+24 papers, 78 figures: 78 present, 0 MISSING
 </file> </file>
  
Line 1933: Line 2019:
 PASS   Gmail image proxy: no IP, no cookies PASS   Gmail image proxy: no IP, no cookies
 PASS   Gmail image proxy does NOT hide the open PASS   Gmail image proxy does NOT hide the open
 +PASS   Google Nov-2025 enforcement escalation
 +PASS   WebKit Link Tracking Protection is a PRIVATE BROWSING feature
 +PASS   arXiv preprint of the lead-marketing paper
 PASS   Yahoo one-click + two days, Q1 2024 PASS   Yahoo one-click + two days, Q1 2024
  
Line 1954: Line 2043:
 | ''node scripts/check_page_numbers.mjs pages/privacy_email_tracking.txt <report+verify+external> '===== Use in Publications =====' '===== What to Report ====='%%'%%'' | OK, after two rounds of fixes | | ''node scripts/check_page_numbers.mjs pages/privacy_email_tracking.txt <report+verify+external> '===== Use in Publications =====' '===== What to Report ====='%%'%%'' | OK, after two rounds of fixes |
 | the same, **whole page** (no markers) | OK, after four rounds of fixes | | the same, **whole page** (no markers) | OK, after four rounds of fixes |
-| ''node scripts/verify_email_figures.mjs''68 figures across 23 papers, 0 missing |+| ''node scripts/verify_email_figures.mjs''78 figures across 24 papers, 0 missing |
 | ''bash scripts/external_checks_email_tracking.sh'' | all checks PASSED after its own RFC 6376 assertion was corrected | | ''bash scripts/external_checks_email_tracking.sh'' | all checks PASSED after its own RFC 6376 assertion was corrected |
  
Line 1970: Line 2059:
 ===== 14. Review pass, 2026-09-02 ===== ===== 14. Review pass, 2026-09-02 =====
  
-PENDING — reviewers have not run yet.+Three focused reviewers ran in parallel against the published text, the scripts and their committed output; the generic pass ran after their findings were applied. Every finding is listed with whether it was **accepted** or **rejected**, because the rejections are the only record of whether a reviewer earns its slot. 
 + 
 +==== 14.1 Figures versus the script (Sonnet) ==== 
 + 
 +^ # ^ Finding ^ Verdict ^ 
 +| 1 | **HIGH. The crawl-configuration, platform and participant figures were computed at 69 papers and never refreshed after {[starov2016_sure]} was added.** Removing exactly that one paper reproduces the page's 26 / 32 / 13-of-26 exactly. The real values are 27 with a ''crawlConfig'', 33 on the ''web'' platform, 14 of 27 stating a consent action — and 29, not 28, ran a crawl. The reviewer also noted the number guard passed anyway, because 26, 32, 13 and 7 all occur elsewhere in a 1,600-line report attached to unrelated facts, and "Twelve" was spelled out so the digit regex never saw it. | **Accepted in full.** All four figures corrected, and — the actual fix — a new ''Study shape'' section in ''report_email_tracking.mjs'' now computes platforms, study types, every ''crawlConfig'' field, participants and ''temporal.mode'' from the population, so these figures can never again be hand-computed. Writing it found a **second bug**: ''crawlConfig.browsers'' is an array, ''isSentinel()'' on an array is false, and the naive check reported 27 of 27 = 100% stated. The real figure is 16 of 27. | 
 +| 2 | **MEDIUM-HIGH. The ''classification.method'' table shows 9 of the 13 values the report computes**, silently dropping ''other'' (7 papers, larger than two rows that are shown), ''dynamic-analysis'', ''graph-analysis'' and ''static-analysis''. The guard cannot see a missing row. | **Accepted.** All 13 rows now published, with the era split folded into the same table and a sentence saying nothing is truncated. | 
 +| 3 | **MEDIUM-HIGH. The instrument table shows 15 of the 19 families**, dropping ''study-apparatus'' (9 papers, larger than nine of the rows shown), ''traffic-capture'', ''crawler-framework'' and ''tracker-blocklist''. | **Accepted.** All 19 published. The omission had also hidden the more interesting reading: only four of the nineteen families are mail- or message-specific, which is now stated. | 
 +| 4 | **MEDIUM. "60 figures across 22 papers" was stale** — the verifier had grown to 68 and the prose on both pages still said 60/22. | **Accepted.** Now 78 across 24, and the count is re-derived from the verifier's own output line on every run. | 
 +| 5 | **Informational. The elided-quote order check is correct but currently unexercised.** A mutant that searches from position 0 each time produces byte-identical output on all 400 real quotes; a synthetic adversarial case distinguishes them. | **Accepted as recorded, no change.** It is a safeguard against a future edit, not a live check, and saying so is more useful than pretending the 400 quotes exercise it. | 
 + 
 +==== 14.2 Citations and quotes (Sonnet) ==== 
 + 
 +^ # ^ Finding ^ Verdict ^ 
 +| 1 | **HIGH. "12,618 mailing lists" is wrong.** {[englehardt2018_email]} crawled 15,700 sites and assembled "12,618 emails from 902 distinct senders"; 12,618 is the corpus size, not the number of lists. The reviewer found it in two places on the page; a later grep found two more — a third occurrence in the residue commentary and the hand-map reason inside ''msg_fold.mjs'', whose output is published on this page, so the wrong figure was in the audit trail as well. | **Accepted, and it is the worst error found.** All four occurrences corrected; the page now gives 15,700 crawled, 902 senders, 12,618 emails, a 38% submission rate of which 32% were list subscriptions, and says explicitly which number means what. All four are now needles in ''verify_email_figures.mjs''. | 
 +| 2 | **HIGH. 53 lines of internal ''bibgen'' QA notes were pasted into the live [[literature:bibliography]] page**, inside the database block, rendering as visible text ("metadata source: venue-page", "cited by ? (OpenAlex)", "check spacing in title"). | **Accepted. This was a real accidental exposure and it was mine.** ''bibgen.mjs'' prints those notes to **stdout**, and the additions file was built by redirecting stdout. Removed from the live page and from the local file; the entry count was checked before and after (816 → 816, zero keys lost) because the first attempt at removing it cut a block and silently dropped {[starov2016_sure]}. | 
 +| 3 | **MEDIUM. The live entry for {[stringhini2012_leveraging]} renders "B@bel" with U+FF20 rather than an ASCII ''@''.** | **Accepted as a finding, rejected as a fix.** The reviewer is right that the glyph is wrong. It is also load-bearing: retested on a clean file at rev 1788333447, the ASCII ''@'' made the entry **and its citation marker** vanish with no warning, references dropping 72 → 71. (An earlier draft of this log said "all five markers". This page carries **one** marker to that key and is the only page on the wiki that cites it; the re-verification pass caught the exaggeration.) ''B{@}bel'', ''B\\@bel'' and ''B&#64;bel'' were all tested too; only the last resolved, and it rendered the entity literally. U+FF20 stays, with the reason in a ''%'' comment above the entry so nobody "fixes" it. | 
 +| 4 | **LOW-MEDIUM. ''sherman2020_going'' renders an author as "Jr., Keith McNamara"** — BibTeX read "Jr." as the family name. | **Accepted.** Corrected to the three-part ''McNamara, Jr., Keith''. | 
 +| 5 | **LOW. Five genuine duplicate papers under two keys each exist in the live bibliography**, three of them sharing a DOI. | **Accepted as a finding, out of scope for this page.** None is cited here. Filed as its own work item rather than fixed in passing, because deduplicating a shared 816-entry bibliography touches every page that cites the losing key. | 
 +| 6 | **LOW. "66 languages" hangs off a three-paper citation group** where it belongs to one of them, two sentences after a "75 languages" that belongs to another. | **Accepted.** Attributed explicitly. | 
 + 
 +The same reviewer independently re-verified the whole run of per-paper figures, the external quotes, and the author lists of the new PETS/USENIX/NDSS entries against each paper's first page, and found no other misattribution. 
 + 
 +==== 14.3 External currency (Sonnet) ==== 
 + 
 +^ # ^ Finding ^ Verdict ^ 
 +| 1 | RFC 2369, 8058 and 6376 as described are all still current; no IETF successor to 2369 or 8058 exists or is in progress; RFC 5321 and 5322 are Draft Standards updated by 7504 and 6854. | **Accepted, no change needed.** | 
 +| 2 | **Gmail's enforcement escalated in November 2025** from spam-foldering to "temporary and permanent rejections". | **Accepted.** Added, quoted verbatim, with a check in the script. It changes what a compliance measurement observes. | 
 +| 3 | **Microsoft is missing and should be there**: since 5 May 2025, ≥5,000 messages a day to Outlook.com consumer services must pass SPF, DKIM and DMARC — and Microsoft does **not** require one-click unsubscribe or cite RFC 8058. | **Accepted, with the verification limit stated on the page.** This is a genuine three-way asymmetry and it matters precisely because the page's unsubscription section is about that mechanism. But the announcement is served behind an Azure Front Door block that refuses this host to both ''curl'' and a headless browser, so the page footnote says the claim rests on Microsoft-hosted Q&A pages that quote and link the announcement, and tells the reader to read the announcement before citing it. | 
 +| 4 | **Apple Link Tracking Protection is missing**: parameters stripped from links in Mail since iOS 17, extended to all of Safari in iOS 26. | **Partly rejected.** The mechanism is added, because it is exactly the thing a reader would mistake for a defence against the address-in-URL leak. The **Mail** and **iOS 26** parts are not: WebKit's own post lists Link Tracking Protection under "protections and defenses added to Private Browsing in **Safari 17.0**" and says nothing about Mail or about iOS 26, and no other Apple primary source could be found for either. The page says so in those words and tells the reader to test. The reviewer's own note said "plus multiple corroborating sources" — which is the listicle bar, not the primary-source bar. | 
 +| 5 | **Add the arXiv preprint and code repository for the IEEE S&P 2026 lead-marketing paper**, since the page tells the reader to read a paywalled paper. | **Accepted.** Both added; the arXiv record's title and all four authors were checked against the IEEE record, and the DOI was resolved. | 
 +| 6 | Two 2026 arXiv preprints (1.35 M operator-side SMS reports; 8.7 M international call records across 65 countries) aim at this page's own open questions. | **Accepted with a label.** Added to Open Questions **as preprints**, used for no figure, and the zero counts on the page left as they were. | 
 +| 7 | **''external_checks_email_tracking.sh'' ended in an unconditional ''exit 0''**, so anything gating on the exit code saw success with checks failing. | **Accepted. A real bug.** Now ''exit "$fails"''. The reviewer found it by mutating a needle and inspecting ''$?'' — which is the right way to review a check script and the reason this slot exists. | 
 +| — | All URLs on both pages return 200 with the content claimed; OpenWPM, libphonenumber, TDLib, SpamAssassin, Snorkel and the Email Privacy Tester are all alive with no renamed successor. | **Accepted, no change.** | 
 + 
 +==== 14.4 What the reviewers did not catch, and what caught it ==== 
 + 
 +Two things were found before the review pass, by the guards rather than by reading, and both are recorded in section 13: four silence base rates copied out of ''OVERVIEW.md'' instead of computed, and the denominator mismatch that copying them concealed. Two more were found by the author while checking a claim the review would probably have accepted: **"on by default since iOS 15"** for Mail Privacy Protection, which Apple documents as a setting the user turns on and for which Apple publishes no take-up figure, and **"Gmail's image proxy has been on since 2013"**, which was recall until the December 2013 announcement was fetched. The provenance page's own assertion that eight below-threshold quotes had been read by hand was written before they had been; reading them is what produced the table in section 8, and checking the sentence after it is what added two papers to the verifier. 
 + 
 +==== 14.5 Author's own pass over the frozen text, after the reviewers ==== 
 + 
 +Four more things were changed after the review, found by reading the page top to bottom rather than by any check: 
 + 
 +  * The intro claimed **"two of the four things you might want to measure have essentially never been measured"**. One slice is zero; the other three all have papers. Rewritten to say what is true: one slice is a measured zero, and a second question — what the mailbox providers' own defences do to a tracking measurement — is unmeasured. 
 +  * **"Four neighbouring clusters are larger than anything on this page"** was false of one of the four: the web-pixel cluster is 7 papers against the SMS slice's 23. Now "three of them are larger than any slice that is". 
 +  * **"This topic has an unusual concentration of ethics problems, and the corpus reflects it"** used a legal-assessment rate as evidence for an ethics claim. The rate (20.0% against 6.9%) is real and stays; the inference does not, and the page now says the number is about legal engagement and that the reporting figures at the foot of the same list say the ethics are not handled well. 
 +  * **"Three of the recent smishing and SMS-spam papers draw on partly overlapping public sources"** was not supported by anything. Two of them mine Twitter ({[tang2022_clues]}, {[agarwal2025_fishing]}); one uses a crowdsourced app and one carrier-side reports. Replaced with the two that actually overlap and the specific comparison nobody has run
 + 
 +The pattern in all four is the same and worth naming: **the number was right and the sentence around it claimed more than the number supports.** Neither the number guard nor a figures reviewer can see that, because they check digits against a script. It is the failure mode the generic review slot exists for, and half of it was still found by re-reading rather than by a reviewer. 
 + 
 +==== 14.6 Generic pass (Fable), and the re-verification that followed ==== 
 + 
 +The generic reviewer's verdict was that the pages are well guarded on their headline figures and that **the provenance page repeatedly failed to practise what it preaches about itself**. That is correct, and the underlying cause is one this project has a memory note about: //a fix survives in the provenance log.// Every accepted review fix had been applied to the content page and almost none had been propagated into the script constants and the prose template that the provenance page republishes. 
 + 
 +^ # ^ Finding ^ Verdict ^ 
 +| 1 | **"183 hand-written verdicts, each with its reason" is overstated, and the residue manufactures evidence of reading.** 57 ''MAP'' entries have an empty reason string, and 171 of the 181 ''OFF'' papers are not in the map at all — they were given a **canned default** reading "read and judged off-topic: the message vocabulary is incidental to the object measured", printed per paper in the published residue. A typo'd ''MAP'' key would have become a silent ''OFF'' still claiming it had been read. | **Accepted, and it is the most serious finding of the whole review.** The default string now says //"not individually annotated: screened at title level, matched no off-topic family rule"//. The report prints the real accounting (183 mapped / 125 with a written reason / 171 defaulted / 46 by family rule / 125 bare) and both pages quote it instead of asserting the stronger claim. All 173 non-''OFF'' papers do carry a hand-written reason; that part was true. | 
 +| 2 | **The error 14.2 called "the worst error found" was still published — inside this page.** "12,618 mailing lists" survived in ''report_email_tracking.mjs''&#39;s section Z and therefore in the embedded output, annotated "checked by verify_email_figures.mjs" — which checks digits, not what the noun claims. Same line: {[starov2016_sure]} described as "five years before Leaky Forms" when 2016→2022 is six. | **Accepted.** Both strings fixed at source in the scripts, re-run, rebuilt. The lesson is the annotation, not the number: a verifier that confirms ''12,618'' appears in the paper says nothing about the sentence around it. | 
 +| 3 | **One provenance page, three values for its own figure check**: §7 said 60 across 22, §13 said 68 across 23, the embedded output said 78 across 24 — and §14.1 had already "accepted" the fix. | **Accepted.** Both prose numbers corrected, and §7 now says where the count comes from so it cannot drift again. | 
 +| 4 | **The published script's own annotations contradict the published output.** ''REJECTED_PROBES'' still said "MQTT in 6 of 59; CCPA a further 6" where section E prints 10 / 3 / 4 of 56, and the header comment said "28 papers social spam against 16 email spam" where the figures are 37 / 15. | **Accepted.** All corrected at source. | 
 +| 5 | **The below-threshold quote rule generalises from a sample that excluded the riskiest cases**: eight quotes read, all from the 40–59% band, then a rule declared about that band — while naming four below 25% as "worth opening" and not opening two of them. | **Accepted.** Both were then read: {[tu2016_security]} at 23% and {[utz2023_comparing]} at 20%, tabulated in §8, both present. The content page now says ten were read, not "every one sampled". | 
 +| 6 | The intro's "two of the four things … have essentially never been measured" has no owner. | **Accepted — and already fixed** in the author's own pass (§14.5) before this review returned. | 
 +| 7 | **"the single biggest coherent slice the pool turned up" (''INFRA'', 31) is contradicted two bullets later** by ''SOC'' at 37. | **Accepted.** Reworded to say the social cluster is larger but is not one topic. | 
 +| 8 | Five smaller self-contradictions: "both CAN-SPAM hits are motivational asides" (one is); the probe table drops the sixth row the report prints, on a page that twice says nothing is truncated; §1 says 56 bib entries and §15 said "56 became 57"; the run log said 145 markers where the source has 147; and the Apple footnote says ''curl'' returns only a shell while the check script's ''curl'' finds the needle. | **All five accepted and fixed.** The Apple one mattered most: ''curl'' does return a noscript copy of the body, so the footnote now tells the reader to trust the browser fetch and not the ''curl'' pass, rather than implying the ''curl'' check is what verified the load-bearing quote. | 
 +| 9 | **The four mail-specific tools are hidden on the provenance page** while the content page tells the student "there is no toolchain for this topic". For someone starting on Monday those four names are the most actionable line on either page. | **Accepted, and the best suggestion of the review.** The Email Privacy Tester, EmailHarvester, the Honey Messages Framework and ''css-inline'' are now named on the content page, together with the CAPTCHA-solver and fake-identity residue, which is what signing up on fifteen thousand sites actually costs. The suggestion to cut the Gmail/Apple point from four appearances to two was **rejected**: the four are a claim, a corpus gap, a study-design consequence and an open question, which are four different uses of the same fact. | 
 + 
 +==== 14.7 Re-verification pass (Sonnet, against the corrected files) ==== 
 + 
 +All eight claimed fixes re-derived independently and **CONFIRMED**: the five Englehardt figures against the paper; the study-shape figures recomputed from ''buildPool()'' (70 / 29 / 27 / 33 / 54 / 12 / 42 and 14-7-5-16-22-22 of 27); both untruncated tables; the verifier count; the non-zero exit; the ''browsers'' array fix at 16 of 27 (and a sweep confirming ''languages'' is the only other array field, and it is not in that table); the bibliography with the QA notes gone and 816 entries intact; and the U+FF20 title rendering with every marker resolving. 
 + 
 +It also found two things nobody else did: 
 + 
 +  * **"73 papers outside the pool" was wrong.** The four probes produce **72** outside-the-pool mentions and **67 distinct papers**, because five papers appear on more than one probe's list. Neither figure is 73. Worse, ''check_page_numbers.mjs'' **passed** it, because its shared ''ALLOW'' map contains a ''73'' written for ''design:platforms'' — a whitelist entry for one page silently blessing a wrong number on another. Fixed on the page (67 distinct, 72 mentions), and the report now prints both counts so neither has to be derived by hand. **The shared ''ALLOW'' map is a cross-page liability and this is the first measured instance of it hiding a real error.** 
 +  * **The "all five citation markers" claim in 14.2 above.** There is one. Corrected in place. 
 + 
 +==== 14.8 A guard that came out of the review ==== 
 + 
 +''scripts/check_attributions.mjs'' already existed, and its own header said "a page that attributes in prose is not covered; report 0 checked rather than assume pass". This page attributes entirely in prose, so it reported **0 checked** — and a fabricated author name shipped anyway: the page said //"Sharevski and Zettlemoyer"// where the authors are Sharevski, Loop, Evans and Ponticello. No other guard could see it. The citekey resolved, the figure beside it was verified against the paper, and the number guard traced the digits. 
 + 
 +A ''--prose'' mode was added: it matches ''Surname'', ''Surname et al.'', ''Surname and Surname'' and ''Surname, Surname and Surname'' immediately before a marker, and requires every capitalised token to appear somewhere in that entry's **full** author string. Two iterations were needed and both are worth recording, because a guard that is wrong in either direction does not get run twice: 
 + 
 +  - A first version scanned every capitalised token in a 90-character window and reported **eight** false positives on a clean page (''XRay'', ''SMTP-dialect'', ''DNSBL'', ''GDPR'', ''URL'', "Older", "Since"). 
 +  - Tightening it to the four name shapes brought it to 52 attributions checked and **one** suspect — //"Kubicek et al."// against ''Kub{\'i}{\v{c}}ek'', because the skeletoniser folded ''\v{c}'' to "vc". A LaTeX-accent stripping step fixed that. 52 checked, 0 suspect. 
 + 
 +Run it in **both** modes on any page it is pointed at. Table mode reports 0 on this page and says so loudly, which is the behaviour that let the error through in the first place.
  
 ===== 15. Run log ===== ===== 15. Run log =====
Line 1977: Line 2149:
 | 2026-09-02 | Whole page, this provenance page, ''msg_fold.mjs'', ''report_email_tracking.mjs'', ''verify_email_figures.mjs'', ''external_checks_email_tracking.sh'', and 56 bibliography entries, written in one sitting against the 5,859-paper corpus | | 2026-09-02 | Whole page, this provenance page, ''msg_fold.mjs'', ''report_email_tracking.mjs'', ''verify_email_figures.mjs'', ''external_checks_email_tracking.sh'', and 56 bibliography entries, written in one sitting against the 5,859-paper corpus |
 | 2026-09-02 | Signal 3 of the pool rule added mid-run after signals 1 and 2 were found to miss the messenger-enumeration papers. Four ''EID'' papers, including the two newest, entered that way | | 2026-09-02 | Signal 3 of the pool rule added mid-run after signals 1 and 2 were found to miss the messenger-enumeration papers. Four ''EID'' papers, including the two newest, entered that way |
-| 2026-09-02 | Recall probes run after the pool was fixed; one paper ({[starov2016_sure]}) added, and 56 bibliography entries became 57 as a result |+| 2026-09-02 | Recall probes run after the pool was fixed; one paper ({[starov2016_sure]}) added, and the bibliography additions grew from 55 entries to 56 as a result |
 | 2026-09-02 | ''external_checks_email_tracking.sh'' failed on its own RFC 6376 assertion (''DRAFT STANDARD'' where the answer is ''INTERNET STANDARD''). The assertion was corrected; the source was right | | 2026-09-02 | ''external_checks_email_tracking.sh'' failed on its own RFC 6376 assertion (''DRAFT STANDARD'' where the answer is ''INTERNET STANDARD''). The assertion was corrected; the source was right |
 | 2026-09-02 | First Apple support URL fetched was the wrong article (Boot Camp). Caught because the fetch answered "I cannot answer this from the provided content" rather than returning something plausible | | 2026-09-02 | First Apple support URL fetched was the wrong article (Boot Camp). Caught because the fetch answered "I cannot answer this from the provided content" rather than returning something plausible |
 | 2026-09-02 | **The bibtex4dw parser splits entries on any literal ASCII ''@'' and silently drops the entry — and every citation marker to it.** {[stringhini2012_leveraging]} is titled //B@bel//. Its five citation markers rendered as nothing at all, with no warning, and the reference was absent from the list; the count of rendered references against distinct markers is what caught it. ''B{@}bel'' and ''B\\@bel'' both still failed; the HTML entity ''B&#64;bel'' resolved the entry but rendered the entity literally. The title now uses **U+FF20 FULLWIDTH COMMERCIAL AT**. Do not "fix" it back to a plain ''@'' without re-checking that the entry still resolves | | 2026-09-02 | **The bibtex4dw parser splits entries on any literal ASCII ''@'' and silently drops the entry — and every citation marker to it.** {[stringhini2012_leveraging]} is titled //B@bel//. Its five citation markers rendered as nothing at all, with no warning, and the reference was absent from the list; the count of rendered references against distinct markers is what caught it. ''B{@}bel'' and ''B\\@bel'' both still failed; the HTML entity ''B&#64;bel'' resolved the entry but rendered the entity literally. The title now uses **U+FF20 FULLWIDTH COMMERCIAL AT**. Do not "fix" it back to a plain ''@'' without re-checking that the entry still resolves |
-| 2026-09-02 | New citekeys need a cache purge. Immediately after publishing, only 30 of the page'145 markers rendered and only 16 references appeared. ''?purge=true'' on [[literature:bibliography]] and then on the content page fixed it. **Always count rendered references against distinct markers after publishing** |+| 2026-09-02 | New citekeys need a cache purge. Immediately after publishing, only 30 of the page's markers rendered — there were 144 at the time, 147 by the end of review — and only 16 references appeared. ''?purge=true'' on [[literature:bibliography]] and then on the content page fixed it. **Always count rendered references against distinct markers after publishing** |
  
 Nothing on this page was carried over from an earlier run. The three ''bench-*'' dossiers and every published figure predating 2026-08-11 were treated as stale by construction and not consulted for numbers. Nothing on this page was carried over from an earlier run. The three ''bench-*'' dossiers and every published figure predating 2026-08-11 were treated as stale by construction and not consulted for numbers.
provenance/privacy/email_tracking.1788333259.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki