| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| provenance:privacy:email_tracking [2026/09/02 07:25] – Add section 14: the full three-reviewer log with accept/reject per finding, including the accidental exposure of bibgen QA notes on the live bibliography page, the 12,618-vs-902 error, the two untruncated tables, the exit-0 bug, and the partly-rejected Ap karel.kubicek.claude | provenance:privacy:email_tracking [2026/09/02 07:41] (current) – Add sections 14.6-14.8: the generic (Fable) pass, the re-verification pass, and the prose-mode attribution guard that came out of it. Fix the provenance page's own three-way disagreement about its figure count, the stale script annotations it republishes, karel.kubicek.claude |
|---|
| ===== 3. The population rule ===== | ===== 3. The population rule ===== |
| |
| **There is no query for this page.** No field in the extraction means "the measured object is a message". The population is a published candidate pool plus one hand-written verdict per paper. Both are in ''scripts/msg_fold.mjs'' and both are printed by the report. | **There is no query for this page.** No field in the extraction means "the measured object is a message". The population is a published candidate pool plus a verdict per paper — but **not** a hand-written verdict for all 354 of them, and an earlier draft of this page and of the content page both said so. The real accounting, which the report now prints: |
| | |
| | ^ ^ Papers ^ |
| | | in the hand map (''MAP'') | **183** | |
| | | ... of which also carry a written reason | 125 | |
| | | ... the rest are ''PHISH'' / ''SOC'' / ''WEBPIXEL'' rows where the verdict is the reason | 58 | |
| | | **not** in the hand map, ''OFF'' by default | **171** | |
| | | ... given a reason by a named ''OFF_FAMILIES'' rule | 46 | |
| | | ... carrying only "not individually annotated" | 125 | |
| | |
| | All 173 non-''OFF'' papers are hand-mapped with a reason. The 125 defaulted papers were screened at **title** level and nothing further was recorded about them individually. Until 2026-09-02 the default reason string read //"read and judged off-topic: the message vocabulary is incidental to the object measured"// and was printed per paper in the published residue — a canned sentence asserting a per-paper reading that did not happen, and one that would have hidden a typo'd ''MAP'' key as a silent ''OFF''. It now says what is true. |
| |
| ==== 3.1 Pool signals ==== | ==== 3.1 Pool signals ==== |
| ''verify_email_figures.mjs'' checks every literal per-paper figure quoted on the content page against that paper's own ''paper.cols.txt'', **not** against ''evidence.quote'' — ''detection[].prevalence'' is a model summary, so a figure can be right in the paper and wrong in the extraction. Whitespace is collapsed on both sides. | ''verify_email_figures.mjs'' checks every literal per-paper figure quoted on the content page against that paper's own ''paper.cols.txt'', **not** against ''evidence.quote'' — ''detection[].prevalence'' is a model summary, so a figure can be right in the paper and wrong in the extraction. Whitespace is collapsed on both sides. |
| |
| **60 figures across 22 papers, all present.** The pass earned its keep three times: | **78 figures across 24 papers, all present.** (It was 60 across 22 when this section was first written and grew twice during review; the count here is now taken from the script's own summary line in §12, not retyped.) The pass earned its keep three times: |
| |
| - **{[agarwal2025_fishing]}: the mathematical-italic //k//.** The paper typesets "27.7𝑘 smishing messages, 19.3𝑘 sender IDs, and 20𝑘 URLs" with U+1D458, not ASCII ''k''. An obvious ASCII needle reported three present figures as missing. The needles now carry the real codepoint and a comment saying why. | - **{[agarwal2025_fishing]}: the mathematical-italic //k//.** The paper typesets "27.7𝑘 smishing messages, 19.3𝑘 sender IDs, and 20𝑘 URLs" with U+1D458, not ASCII ''k''. An obvious ASCII needle reported three present figures as missing. The needles now carry the real codepoint and a comment saying why. |
| | {[li2025_hades]} | 53% | 10 words, //"mail servers of popular ESPs and websites have been included"// | present; the 39,201 (76.88%) is immediately before it | | | {[li2025_hades]} | 53% | 10 words, //"mail servers of popular ESPs and websites have been included"// | present; the 39,201 (76.88%) is immediately before it | |
| |
| The pattern is uniform enough to be worth stating as a rule: **a coverage figure in the 40–60% band on a two-column paper is a column splice, not a fabrication.** A coverage figure below about 25% is worth opening. Four of the 62 are below that — {[tu2016_security]} (23%), {[utz2023_comparing]} (20%), {[reaves2016_sending]} (20%) and {[agarwal2025_fishing]} (18%). Checking that claim is how ''verify_email_figures.mjs'' gained two more papers: {[agarwal2025_fishing]} was already covered, {[reaves2016_sending]} was not and the page quotes three of its figures, so it was added (386,327 messages, 522 containing email addresses, 14 months, over 400 numbers — all present); {[jiang2013_greystar]} was added in the same pass. **The page quotes no figure from the remaining two**, so nothing on it rests on those two quotes. | The pattern is uniform enough to be worth stating as a rule: **a coverage figure in the 40–60% band on a two-column paper is a column splice, not a fabrication.** A coverage figure below about 25% is worth opening, and the rule above was inferred from the middle of the distribution, so the tail was checked separately. **The two lowest were read by hand as well, and both are present:** |
| | |
| | ^ Paper ^ Coverage ^ What the paper says ^ |
| | | {[tu2016_security]} | 23% | //"victim account's activity logs, as shown in Figure 8(d), confirms … that those three attack actions are successful"// — the quote's own ellipsis covers the figure reference, and a column break falls inside it | |
| | | {[utz2023_comparing]} | 20% | //"No HTTPS was rarest, / with 2.85 % of sites"// — verbatim, split across the column boundary | |
| | |
| | Four of the 62 are below 25% — {[tu2016_security]} (23%), {[utz2023_comparing]} (20%), {[reaves2016_sending]} (20%) and {[agarwal2025_fishing]} (18%). Checking that claim is how ''verify_email_figures.mjs'' gained two more papers: {[agarwal2025_fishing]} was already covered, {[reaves2016_sending]} was not and the page quotes three of its figures, so it was added (386,327 messages, 522 containing email addresses, 14 months, over 400 numbers — all present); {[jiang2013_greystar]} was added in the same pass. **The page quotes no figure from the remaining two**, so nothing on it rests on those two quotes. |
| |
| ===== 9. External sources, and how each was verified ===== | ===== 9. External sources, and how each was verified ===== |
| admitted by hand : 3 | admitted by hand : 3 |
| hand-mapped verdicts : 183 | hand-mapped verdicts : 183 |
| | ... of which carry a written reason : 125 (182 MAP entries; the rest are |
| | PHISH / SOC / WEBPIXEL rows where the verdict IS the reason) |
| | NOT in the hand map, verdict OFF by default: 171 |
| | ... given a reason by an OFF_FAMILIES rule: 46 |
| | ... carrying only "not individually annotated": 125 |
| | These were screened at TITLE level. Nothing further was recorded per paper. |
| ON the page : 70 | ON the page : 70 |
| rejected outside the pool : 9 (documented, see section H) | rejected outside the pool : 9 (documented, see section H) |
| --- OFF — 181 papers — homograph or unrelated | --- OFF — 181 papers — homograph or unrelated |
| IMC/2010/estimating-and-sampling-graphs-with-multidimensional-random-walks | IMC/2010/estimating-and-sampling-graphs-with-multidimensional-random-walks |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2010/inference-and-analysis-of-formal-models-of-botnet-command-and-control-protocols | CCS/2010/inference-and-analysis-of-formal-models-of-botnet-command-and-control-protocols |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2010/dissecting-one-click-frauds | CCS/2010/dissecting-one-click-frauds |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2010/building-a-dynamic-reputation-system-for-dns | USENIX/2010/building-a-dynamic-reputation-system-for-dns |
| network-layer work that uses spam as a label source | network-layer work that uses spam as a label source |
| USENIX/2010/searching-the-searchers-with-searchaudit | USENIX/2010/searching-the-searchers-with-searchaudit |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IMC/2011/gq-practical-containment-for-measuring-modern-malware-systems | IMC/2011/gq-practical-containment-for-measuring-modern-malware-systems |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| network-layer work that uses spam as a label source | network-layer work that uses spam as a label source |
| USENIX/2011/dirty-jobs-the-role-of-freelance-labor-in-web-service-abuse | USENIX/2011/dirty-jobs-the-role-of-freelance-labor-in-web-service-abuse |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2011/forensic-triage-for-mobile-phones-with-dec0de | USENIX/2011/forensic-triage-for-mobile-phones-with-dec0de |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2011/measuring-and-analyzing-search-redirection-attacks-in-the-illicit-online-prescri | USENIX/2011/measuring-and-analyzing-search-redirection-attacks-in-the-illicit-online-prescri |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2012/detecting-money-stealing-apps-in-alternative-android-markets | CCS/2012/detecting-money-stealing-apps-in-alternative-android-markets |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| legitimacy of webmail accounts; the unit is the account, not the message | legitimacy of webmail accounts; the unit is the account, not the message |
| CCS/2012/mobile-data-charging-new-attacks-and-countermeasures | CCS/2012/mobile-data-charging-new-attacks-and-countermeasures |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2012/guess-who-s-texting-you-evaluating-the-security-of-smartphone-messaging-applicat | NDSS/2012/guess-who-s-texting-you-evaluating-the-security-of-smartphone-messaging-applicat |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2012/hey-you-get-off-of-my-market-detecting-malicious-apps-in-official-and-alternativ | NDSS/2012/hey-you-get-off-of-my-market-detecting-malicious-apps-in-official-and-alternativ |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2013/cross-origin-pixel-stealing-timing-attacks-using-css-filters | CCS/2013/cross-origin-pixel-stealing-timing-attacks-using-css-filters |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2013/whyper-towards-automating-risk-assessment-of-mobile-applications | USENIX/2013/whyper-towards-automating-risk-assessment-of-mobile-applications |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2014/dialing-back-abuse-on-phone-verified-accounts | CCS/2014/dialing-back-abuse-on-phone-verified-accounts |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2014/are-you-ready-to-lock | CCS/2014/are-you-ready-to-lock |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2013/detecting-passive-content-leaks-and-pollution-in-android-applications | NDSS/2013/detecting-passive-content-leaks-and-pollution-in-android-applications |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| USENIX/2013/jekyll-on-ios-when-benign-apps-become-evil | USENIX/2013/jekyll-on-ios-when-benign-apps-become-evil |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2013/automatically-inferring-the-evolution-of-malicious-activity-on-the-internet | NDSS/2013/automatically-inferring-the-evolution-of-malicious-activity-on-the-internet |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2013/protecting-sensitive-web-content-from-client-side-vulnerabilities-with-cryptons | CCS/2013/protecting-sensitive-web-content-from-client-side-vulnerabilities-with-cryptons |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2013/vetting-undesirable-behaviors-in-android-apps-with-permission-use-analysis | CCS/2013/vetting-undesirable-behaviors-in-android-apps-with-permission-use-analysis |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| CCS/2014/consequences-of-connectivity-characterizing-account-hijacking-on-twitter | CCS/2014/consequences-of-connectivity-characterizing-account-hijacking-on-twitter |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| WWW/2013/two-years-of-short-urls-internet-measurement-security-threats-and-countermeasure | WWW/2013/two-years-of-short-urls-internet-measurement-security-threats-and-countermeasure |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2014/autocog-measuring-the-description-to-permission-fidelity-in-android-applications | CCS/2014/autocog-measuring-the-description-to-permission-fidelity-in-android-applications |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| CCS/2014/real-threats-to-your-data-bills-security-loopholes-and-defenses-in-mobile-data-c | CCS/2014/real-threats-to-your-data-bills-security-loopholes-and-defenses-in-mobile-data-c |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2014/understanding-the-dark-side-of-domain-parking | USENIX/2014/understanding-the-dark-side-of-domain-parking |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IMC/2014/handcrafted-fraud-and-extortion-manual-account-hijacking-in-the-wild | IMC/2014/handcrafted-fraud-and-extortion-manual-account-hijacking-in-the-wild |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2015/leakage-abuse-attacks-against-searchable-encryption | CCS/2015/leakage-abuse-attacks-against-searchable-encryption |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2015/copperdroid-automatic-reconstruction-of-android-malware-behaviors | NDSS/2015/copperdroid-automatic-reconstruction-of-android-malware-behaviors |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| network-layer work that uses spam as a label source | network-layer work that uses spam as a label source |
| NDSS/2015/what-s-in-your-dongle-and-bank-account-mandatory-and-discretionary-protection-of | NDSS/2015/what-s-in-your-dongle-and-bank-account-mandatory-and-discretionary-protection-of |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2015/android-permissions-remystified-a-field-study-on-contextual-integrity | USENIX/2015/android-permissions-remystified-a-field-study-on-contextual-integrity |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| USENIX/2015/cloudy-with-a-chance-of-breach-forecasting-cyber-security-incidents | USENIX/2015/cloudy-with-a-chance-of-breach-forecasting-cyber-security-incidents |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2014/on-the-mismanagement-and-maliciousness-of-networks | NDSS/2014/on-the-mismanagement-and-maliciousness-of-networks |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2014/man-vs-machine-practical-adversarial-detection-of-malicious-crowdsourcing-worker | USENIX/2014/man-vs-machine-practical-adversarial-detection-of-malicious-crowdsourcing-worker |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2014/dspin-detecting-automatically-spun-content-on-the-web | NDSS/2014/dspin-detecting-automatically-spun-content-on-the-web |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IMC/2015/affiliate-crookies-characterizing-affiliate-marketing-abuse | IMC/2015/affiliate-crookies-characterizing-affiliate-marketing-abuse |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IMC/2015/detecting-malicious-activity-with-dns-backscatter | IMC/2015/detecting-malicious-activity-with-dns-backscatter |
| network-layer work that uses spam as a label source | network-layer work that uses spam as a label source |
| IMC/2015/leveraging-internet-background-radiation-for-opportunistic-network-analysis | IMC/2015/leveraging-internet-background-radiation-for-opportunistic-network-analysis |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IMC/2015/the-dark-menace-characterizing-network-based-attacks-in-the-cloud | IMC/2015/the-dark-menace-characterizing-network-based-attacks-in-the-cloud |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2015/evilcohort-detecting-communities-of-malicious-accounts-on-online-services | USENIX/2015/evilcohort-detecting-communities-of-malicious-accounts-on-online-services |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2016/featuresmith-automatically-engineering-features-for-malware-detection-by-mining | CCS/2016/featuresmith-automatically-engineering-features-for-malware-detection-by-mining |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| interconnect (SIM-box) bypass fraud: the object is call routing revenue, not a message delivered to a user | interconnect (SIM-box) bypass fraud: the object is call routing revenue, not a message delivered to a user |
| CCS/2016/predator-proactive-recognition-and-elimination-of-domain-abuse-at-time-of-regist | CCS/2016/predator-proactive-recognition-and-elimination-of-domain-abuse-at-time-of-regist |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IMC/2016/bdrmap-inference-of-borders-between-ip-networks | IMC/2016/bdrmap-inference-of-borders-between-ip-networks |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| PETS/2016/tales-from-the-dark-side-privacy-dark-strategies-and-privacy-dark-patterns | PETS/2016/tales-from-the-dark-side-privacy-dark-strategies-and-privacy-dark-patterns |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2016/the-ever-changing-labyrinth-a-large-scale-analysis-of-wildcard-dns-powered-black | USENIX/2016/the-ever-changing-labyrinth-a-large-scale-analysis-of-wildcard-dns-powered-black |
| network-layer work that uses spam as a label source | network-layer work that uses spam as a label source |
| IEEE-SP/2017/obstacles-to-the-adoption-of-secure-communication-tools | IEEE-SP/2017/obstacles-to-the-adoption-of-secure-communication-tools |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IMC/2017/understanding-the-role-of-registrars-in-dnssec-deployment | IMC/2017/understanding-the-role-of-registrars-in-dnssec-deployment |
| network-layer work that uses spam as a label source | network-layer work that uses spam as a label source |
| pub/sub or SDN "unsubscribe" — protocol verb, not marketing mail | pub/sub or SDN "unsubscribe" — protocol verb, not marketing mail |
| USENIX/2017/characterizing-the-nature-and-dynamics-of-tor-exit-blocking | USENIX/2017/characterizing-the-nature-and-dynamics-of-tor-exit-blocking |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2017/malton-towards-on-device-non-invasive-mobile-malware-analysis-for-art | USENIX/2017/malton-towards-on-device-non-invasive-mobile-malware-analysis-for-art |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| CCS/2018/detecting-attacks-against-robotic-vehicles-a-control-invariant-approach | CCS/2018/detecting-attacks-against-robotic-vehicles-a-control-invariant-approach |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2018/towards-paving-the-way-for-large-scale-windows-malware-analysis-generic-binary-u | CCS/2018/towards-paving-the-way-for-large-scale-windows-malware-analysis-generic-binary-u |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| IEEE-SP/2018/computer-security-and-privacy-for-refugees-in-the-united-states | IEEE-SP/2018/computer-security-and-privacy-for-refugees-in-the-united-states |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IMC/2018/is-the-web-ready-for-ocsp-must-staple | IMC/2018/is-the-web-ready-for-ocsp-must-staple |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| PETS/2018/toward-distribution-estimation-under-local-differential-privacy-with-small-sampl | PETS/2018/toward-distribution-estimation-under-local-differential-privacy-with-small-sampl |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2018/from-patching-delays-to-infection-symptoms-using-risk-profiles-for-an-early-disc | USENIX/2018/from-patching-delays-to-infection-symptoms-using-risk-profiles-for-an-early-disc |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2018/o-single-sign-off-where-art-thou-an-empirical-analysis-of-single-sign-on-account | USENIX/2018/o-single-sign-off-where-art-thou-an-empirical-analysis-of-single-sign-on-account |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2019/a-systematic-framework-to-generate-invariants-for-anomaly-detection-in-industrial-control-systems | NDSS/2019/a-systematic-framework-to-generate-invariants-for-anomaly-detection-in-industrial-control-systems |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2019/cleaning-up-the-internet-of-evil-things-real-world-evidence-on-isp-and-consumer-efforts-to-remove-mirai | NDSS/2019/cleaning-up-the-internet-of-evil-things-real-world-evidence-on-isp-and-consumer-efforts-to-remove-mirai |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2019/privacy-attacks-to-the-4g-and-5g-cellular-paging-protocols-using-side-channel-information | NDSS/2019/privacy-attacks-to-the-4g-and-5g-cellular-paging-protocols-using-side-channel-information |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2019/a-billion-open-interfaces-for-eve-and-mallory-mitm-dos-and-tracking-attacks-on-i | USENIX/2019/a-billion-open-interfaces-for-eve-and-mallory-mitm-dos-and-tracking-attacks-on-i |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2019/policylint-investigating-internal-privacy-policy-contradictions-on-google-play | USENIX/2019/policylint-investigating-internal-privacy-policy-contradictions-on-google-play |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2019/reading-the-tea-leaves-a-comparative-analysis-of-threat-intelligence | USENIX/2019/reading-the-tea-leaves-a-comparative-analysis-of-threat-intelligence |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| WWW/2019/evaluating-login-challenges-as-adefense-against-account-takeover | WWW/2019/evaluating-login-challenges-as-adefense-against-account-takeover |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| WWW/2019/exploring-user-behavior-in-email-re-finding-tasks | WWW/2019/exploring-user-behavior-in-email-re-finding-tasks |
| information retrieval inside a mailbox — no privacy or security measurement | information retrieval inside a mailbox — no privacy or security measurement |
| what a notification should say, not what it discloses to third parties | what a notification should say, not what it discloses to third parties |
| WWW/2019/understanding-the-evolution-of-mobile-app-ecosystems-a-longitudinal-measurement | WWW/2019/understanding-the-evolution-of-mobile-app-ecosystems-a-longitudinal-measurement |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2020/towards-attribution-in-mobile-markets-identifying-developer-account-polymorphism | CCS/2020/towards-attribution-in-mobile-markets-identifying-developer-account-polymorphism |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IEEE-SP/2020/the-many-kinds-of-creepware-used-for-interpersonal-attacks | IEEE-SP/2020/the-many-kinds-of-creepware-used-for-interpersonal-attacks |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IEEE-SP/2020/karonte-detecting-insecure-multi-binary-interactions-in-embedded-firmware | IEEE-SP/2020/karonte-detecting-insecure-multi-binary-interactions-in-embedded-firmware |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| IMC/2020/revisiting-transactional-statistics-of-high-scalability-blockchains | IMC/2020/revisiting-transactional-statistics-of-high-scalability-blockchains |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2020/blag-improving-the-accuracy-of-blacklists | NDSS/2020/blag-improving-the-accuracy-of-blacklists |
| network-layer work that uses spam as a label source | network-layer work that uses spam as a label source |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| NDSS/2020/surfingattack-interactive-hidden-attack-on-voice-assistants-using-ultrasonic-guided-waves | NDSS/2020/surfingattack-interactive-hidden-attack-on-voice-assistants-using-ultrasonic-guided-waves |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| PETS/2020/sok-anatomy-of-data-breaches | PETS/2020/sok-anatomy-of-data-breaches |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2020/security-analysis-of-unified-payments-interface-and-payment-apps-in-india | USENIX/2020/security-analysis-of-unified-payments-interface-and-payment-apps-in-india |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2020/teerex-discovery-and-exploitation-of-memory-corruption-vulnerabilities-in-sgx-en | USENIX/2020/teerex-discovery-and-exploitation-of-memory-corruption-vulnerabilities-in-sgx-en |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| WWW/2020/fairrec-two-sided-fairness-for-personalized-recommendations-in-two-sided-platfor | WWW/2020/fairrec-two-sided-fairness-for-personalized-recommendations-in-two-sided-platfor |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2021/dont-forget-the-stuffing-revisiting-the-security-impact-of-typo-tolerant-passwor | CCS/2021/dont-forget-the-stuffing-revisiting-the-security-impact-of-typo-tolerant-passwor |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2021/alchemist-fusing-application-and-audit-logs-for-precise-attack-provenance-without-instrumentation | NDSS/2021/alchemist-fusing-application-and-audit-logs-for-precise-attack-provenance-without-instrumentation |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2021/from-whois-to-whowas-a-large-scale-measurement-study-of-domain-registration-privacy-under-the-gdpr | NDSS/2021/from-whois-to-whowas-a-large-scale-measurement-study-of-domain-registration-privacy-under-the-gdpr |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| PETS/2021/the-motivated-can-encrypt-even-with-pgp | PETS/2021/the-motivated-can-encrypt-even-with-pgp |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| PETS/2021/we-three-brothers-have-always-known-everything-of-each-other-a-cross-cultural-st | PETS/2021/we-three-brothers-have-always-known-everything-of-each-other-a-cross-cultural-st |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| PETS/2021/defining-privacy-how-users-interpret-technical-terms-in-privacy-policies | PETS/2021/defining-privacy-how-users-interpret-technical-terms-in-privacy-policies |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2021/alpaca-application-layer-protocol-confusion-analyzing-and-mitigating-cracks-in-t | USENIX/2021/alpaca-application-layer-protocol-confusion-analyzing-and-mitigating-cracks-in-t |
| network-layer work that uses spam as a label source | network-layer work that uses spam as a label source |
| USENIX/2021/effect-of-mood-location-trust-and-presence-of-others-on-video-based-social-authe | USENIX/2021/effect-of-mood-location-trust-and-presence-of-others-on-video-based-social-authe |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2021/now-im-a-bit-angry-individuals-awareness-perception-and-responses-to-data-breach | USENIX/2021/now-im-a-bit-angry-individuals-awareness-perception-and-responses-to-data-breach |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2021/privatedrop-practical-privacy-preserving-authentication-for-apple-airdrop | USENIX/2021/privatedrop-practical-privacy-preserving-authentication-for-apple-airdrop |
| contact discovery, but the contribution is a protocol; no measurement of an address space | contact discovery, but the contribution is a protocol; no measurement of an address space |
| USENIX/2021/strategies-and-perceived-risks-of-sending-sensitive-documents | USENIX/2021/strategies-and-perceived-risks-of-sending-sensitive-documents |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2021/the-hijackers-guide-to-the-galaxy-off-path-taking-over-internet-resources | USENIX/2021/the-hijackers-guide-to-the-galaxy-off-path-taking-over-internet-resources |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2021/why-older-adults-dont-use-password-managers | USENIX/2021/why-older-adults-dont-use-password-managers |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| WWW/2021/an-investigation-of-identity-account-inconsistency-in-single-sign-on | WWW/2021/an-investigation-of-identity-account-inconsistency-in-single-sign-on |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| WWW/2021/privacy-policies-over-time-curation-and-analysis-of-a-million-document-dataset | WWW/2021/privacy-policies-over-time-curation-and-analysis-of-a-million-document-dataset |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2022/stolenencoder-stealing-pre-trained-encoders-in-self-supervised-learning | CCS/2022/stolenencoder-stealing-pre-trained-encoders-in-self-supervised-learning |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IMC/2022/stop-drop-and-roa-effectiveness-of-defenses-through-the-lens-of-drop | IMC/2022/stop-drop-and-roa-effectiveness-of-defenses-through-the-lens-of-drop |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2022/auto-draft-238 | NDSS/2022/auto-draft-238 |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| PETS/2022/athena-probabilistic-verification-of-machine-unlearning | PETS/2022/athena-probabilistic-verification-of-machine-unlearning |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| PETS/2022/user-perceptions-of-gmail-s-confidential-mode | PETS/2022/user-perceptions-of-gmail-s-confidential-mode |
| usability of one provider feature | usability of one provider feature |
| network-layer work that uses spam as a label source | network-layer work that uses spam as a label source |
| USENIX/2022/pre-hijacked-accounts-an-empirical-study-of-security-failures-in-user-account-cr | USENIX/2022/pre-hijacked-accounts-an-empirical-study-of-security-failures-in-user-account-cr |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2022/sgxfuzz-efficiently-synthesizing-nested-structures-for-sgx-enclave-fuzzing | USENIX/2022/sgxfuzz-efficiently-synthesizing-nested-structures-for-sgx-enclave-fuzzing |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| CCS/2023/black-ostrich-web-application-scanning-with-string-solvers | CCS/2023/black-ostrich-web-application-scanning-with-string-solvers |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2023/comprehension-from-chaos-towards-informed-consent-for-private-computation | CCS/2023/comprehension-from-chaos-towards-informed-consent-for-private-computation |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2023/fetchbench-systematic-identification-and-characterization-of-proprietary-prefetc | CCS/2023/fetchbench-systematic-identification-and-characterization-of-proprietary-prefetc |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2023/i-still-know-what-you-watched-last-sunday-privacy-of-the-hbbtv-protocol-in-the-european-smart-tv-landscape | NDSS/2023/i-still-know-what-you-watched-last-sunday-privacy-of-the-hbbtv-protocol-in-the-european-smart-tv-landscape |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| PETS/2023/designing-a-location-trace-anonymization-contest | PETS/2023/designing-a-location-trace-anonymization-contest |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2023/weve-disabled-mfa-for-you-an-evaluation-of-the-security-and-usability-of-multi-f | CCS/2023/weve-disabled-mfa-for-you-an-evaluation-of-the-security-and-usability-of-multi-f |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| PETS/2023/no-privacy-among-spies-assessing-the-functionality-and-insecurity-of-consumer-an | PETS/2023/no-privacy-among-spies-assessing-the-functionality-and-insecurity-of-consumer-an |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2023/a-large-scale-measurement-of-website-login-policies | USENIX/2023/a-large-scale-measurement-of-website-login-policies |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2023/a-two-decade-retrospective-analysis-of-a-universitys-vulnerability-to-attacks-ex | USENIX/2023/a-two-decade-retrospective-analysis-of-a-universitys-vulnerability-to-attacks-ex |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2023/security-and-privacy-failures-in-popular-2fa-apps | USENIX/2023/security-and-privacy-failures-in-popular-2fa-apps |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2023/basecomp-a-comparative-analysis-for-integrity-protection-in-cellular-baseband-so | USENIX/2023/basecomp-a-comparative-analysis-for-integrity-protection-in-cellular-baseband-so |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| USENIX/2023/work-from-home-and-covid-19-trajectories-of-endpoint-security-management-in-a-se | USENIX/2023/work-from-home-and-covid-19-trajectories-of-endpoint-security-management-in-a-se |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2023/anatomy-of-a-high-profile-data-breach-dissecting-the-aftermath-of-a-crypto-walle | USENIX/2023/anatomy-of-a-high-profile-data-breach-dissecting-the-aftermath-of-a-crypto-walle |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2023/instructions-unclear-undefined-behaviour-in-cellular-network-specifications | USENIX/2023/instructions-unclear-undefined-behaviour-in-cellular-network-specifications |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| USENIX/2023/notice-the-imposter-a-study-on-user-tag-spoofing-attack-in-mobile-apps | USENIX/2023/notice-the-imposter-a-study-on-user-tag-spoofing-attack-in-mobile-apps |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2023/policycomp-counterpart-comparison-of-privacy-policies-uncovers-overbroad-persona | USENIX/2023/policycomp-counterpart-comparison-of-privacy-policies-uncovers-overbroad-persona |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2024/maginot-line-assessing-a-new-cross-app-threat-to-pii-as-factor-authentication-in-chinese-mobile-apps | NDSS/2024/maginot-line-assessing-a-new-cross-app-threat-to-pii-as-factor-authentication-in-chinese-mobile-apps |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2024/on-precisely-detecting-censorship-circumvention-in-real-world-networks | NDSS/2024/on-precisely-detecting-censorship-circumvention-in-real-world-networks |
| network-layer work that uses spam as a label source | network-layer work that uses spam as a label source |
| WWW/2023/unsupervised-anomaly-detection-on-microservice-traces-through-graph-vae | WWW/2023/unsupervised-anomaly-detection-on-microservice-traces-through-graph-vae |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2024/a-first-look-at-security-and-privacy-risks-in-the-rapidapi-ecosystem | CCS/2024/a-first-look-at-security-and-privacy-risks-in-the-rapidapi-ecosystem |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IEEE-SP/2024/combing-for-credentials-active-pattern-extraction-from-smart-reply | IEEE-SP/2024/combing-for-credentials-active-pattern-extraction-from-smart-reply |
| ML-privacy work whose training data happens to be mail | ML-privacy work whose training data happens to be mail |
| IEEE-SP/2024/deeptheft-stealing-dnn-model-architectures-through-power-side-channel | IEEE-SP/2024/deeptheft-stealing-dnn-model-architectures-through-power-side-channel |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2024/a-mixed-methods-study-on-user-experiences-and-challenges-of-recovery-codes-for-a | USENIX/2024/a-mixed-methods-study-on-user-experiences-and-challenges-of-recovery-codes-for-a |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2024/guardians-of-the-galaxy-content-moderation-in-the-interplanetary-file-system | USENIX/2024/guardians-of-the-galaxy-content-moderation-in-the-interplanetary-file-system |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2024/landscape-exploring-ldap-weaknesses-and-data-leaks-at-internet-scale | USENIX/2024/landscape-exploring-ldap-weaknesses-and-data-leaks-at-internet-scale |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2024/malla-demystifying-real-world-large-language-model-integrated-malicious-services | USENIX/2024/malla-demystifying-real-world-large-language-model-integrated-malicious-services |
| LLM-for-hire services, one output of which is spam; the unit is the service | LLM-for-hire services, one output of which is spam; the unit is the service |
| USENIX/2024/pixel-thief-exploiting-svg-filter-leakage-in-firefox-and-chrome | USENIX/2024/pixel-thief-exploiting-svg-filter-leakage-in-firefox-and-chrome |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2024/simurai-slicing-through-the-complexity-of-sim-card-security-research | USENIX/2024/simurai-slicing-through-the-complexity-of-sim-card-security-research |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| IEEE-SP/2016/the-cracked-cookie-jar-http-cookie-hijacking-and-the-exposure-of-private-informa | IEEE-SP/2016/the-cracked-cookie-jar-http-cookie-hijacking-and-the-exposure-of-private-informa |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| CCS/2025/noise-and-stress-dont-help-with-learning-a-qualitative-study-to-inform-design-of | CCS/2025/noise-and-stress-dont-help-with-learning-a-qualitative-study-to-inform-design-of |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IMC/2025/congestion-patterns-in-a-large-scale-rdma-datacenter | IMC/2025/congestion-patterns-in-a-large-scale-rdma-datacenter |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2025/all-your-database-are-belong-to-us-characterizing-database-ransomware-attacks | NDSS/2025/all-your-database-are-belong-to-us-characterizing-database-ransomware-attacks |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2025/whats-done-is-not-whats-claimed-detecting-and-interpreting-inconsistencies-in-app-behaviors | NDSS/2025/whats-done-is-not-whats-claimed-detecting-and-interpreting-inconsistencies-in-app-behaviors |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2025/who-is-trying-to-access-my-account-exploring-user-perceptions-and-reactions-to-risk-based-authentication-notifications | NDSS/2025/who-is-trying-to-access-my-account-exploring-user-perceptions-and-reactions-to-risk-based-authentication-notifications |
| same; its 7 `unsubscri` hits are about turning off security alerts | same; its 7 `unsubscri` hits are about turning off security alerts |
| the notification's wording is the treatment; the channel is incidental | the notification's wording is the treatment; the channel is incidental |
| NDSS/2025/attributing-open-source-contributions-is-critical-but-difficult-a-systematic-analysis-of-github-practices-and-their-impact-on-software-supply-chain-security | NDSS/2025/attributing-open-source-contributions-is-critical-but-difficult-a-systematic-analysis-of-github-practices-and-their-impact-on-software-supply-chain-security |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2025/evaluating-llm-based-personal-information-extraction-and-countermeasures | USENIX/2025/evaluating-llm-based-personal-information-extraction-and-countermeasures |
| ML-privacy work whose training data happens to be mail | ML-privacy work whose training data happens to be mail |
| USENIX/2025/scanned-and-scammed-insecurity-by-obsqrity-measuring-user-susceptibility-and-awa | USENIX/2025/scanned-and-scammed-insecurity-by-obsqrity-measuring-user-susceptibility-and-awa |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2025/private-investigator-extracting-personally-identifiable-information-from-large-l | USENIX/2025/private-investigator-extracting-personally-identifiable-information-from-large-l |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2025/generated-data-with-fake-privacy-hidden-dangers-of-fine-tuning-large-language-mo | USENIX/2025/generated-data-with-fake-privacy-hidden-dangers-of-fine-tuning-large-language-mo |
| ML-privacy work whose training data happens to be mail | ML-privacy work whose training data happens to be mail |
| ML-privacy work whose training data happens to be mail | ML-privacy work whose training data happens to be mail |
| NDSS/2026/tipso-gan-malicious-network-traffic-detection-using-a-novel-optimized-generative-adversarial-network | NDSS/2026/tipso-gan-malicious-network-traffic-detection-using-a-novel-optimized-generative-adversarial-network |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| PETS/2026/because-i-didnt-touch-these-and-even-dont-know-why-i-should-to-change-these-why | PETS/2026/because-i-didnt-touch-these-and-even-dont-know-why-i-should-to-change-these-why |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2026/trojpix-electromagnetic-covert-channels-via-imperceptible-pixel-modulation | USENIX/2026/trojpix-electromagnetic-covert-channels-via-imperceptible-pixel-modulation |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| WWW/2026/truth-with-a-twist-the-rhetoric-of-persuasion-in-professional-vs-community-autho | WWW/2026/truth-with-a-twist-the-rhetoric-of-persuasion-in-professional-vs-community-autho |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IEEE-SP/2015/vetting-ssl-usage-in-applications-with-sslint | IEEE-SP/2015/vetting-ssl-usage-in-applications-with-sslint |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IEEE-SP/2016/triggerscope-towards-detecting-logic-bombs-in-android-applications | IEEE-SP/2016/triggerscope-towards-detecting-logic-bombs-in-android-applications |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| IEEE-SP/2018/the-spyware-used-in-intimate-partner-violence | IEEE-SP/2018/the-spyware-used-in-intimate-partner-violence |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IEEE-SP/2017/the-password-reset-mitm-attack | IEEE-SP/2017/the-password-reset-mitm-attack |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2025/a-large-scale-measurement-study-of-the-proxy-protocol-and-its-security-implications | NDSS/2025/a-large-scale-measurement-study-of-the-proxy-protocol-and-its-security-implications |
| network-layer work that uses spam as a label source | network-layer work that uses spam as a label source |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| PETS/2025/gig-work-at-what-cost-exploring-privacy-risks-of-gig-work-platform-participation | PETS/2025/gig-work-at-what-cost-exploring-privacy-risks-of-gig-work-platform-participation |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| PETS/2025/real-world-deniability-in-messaging | PETS/2025/real-world-deniability-in-messaging |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| USENIX/2025/from-alarms-to-real-bugs-multi-target-multi-step-directed-greybox-fuzzing-for-st | USENIX/2025/from-alarms-to-real-bugs-multi-target-multi-step-directed-greybox-fuzzing-for-st |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| PETS/2025/sok-web-authentication-and-recovery-in-the-age-of-end-to-end-encryption | PETS/2025/sok-web-authentication-and-recovery-in-the-age-of-end-to-end-encryption |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| PETS/2026/chatbot-confessions-large-scale-analysis-of-private-data-disclosure-in-shared-ai | PETS/2026/chatbot-confessions-large-scale-analysis-of-private-data-disclosure-in-shared-ai |
| ML-privacy work whose training data happens to be mail | ML-privacy work whose training data happens to be mail |
| systems security; the message vocabulary is incidental | systems security; the message vocabulary is incidental |
| IEEE-SP/2017/under-the-shadow-of-sunshine-understanding-and-detecting-bulletproof-hosting-on | IEEE-SP/2017/under-the-shadow-of-sunshine-understanding-and-detecting-bulletproof-hosting-on |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| NDSS/2025/isolategpt-an-execution-isolation-architecture-for-llm-based-agentic-systems | NDSS/2025/isolategpt-an-execution-isolation-architecture-for-llm-based-agentic-systems |
| ML-privacy work whose training data happens to be mail | ML-privacy work whose training data happens to be mail |
| IEEE-SP/2022/desperate-times-call-for-desperate-measures-user-concerns-with-mobile-loan-apps | IEEE-SP/2022/desperate-times-call-for-desperate-measures-user-concerns-with-mobile-loan-apps |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IEEE-SP/2022/device-fingerprinting-with-peripheral-timestamps | IEEE-SP/2022/device-fingerprinting-with-peripheral-timestamps |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IEEE-SP/2022/towards-automated-auditing-for-account-and-session-management-flaws-in-single-si | IEEE-SP/2022/towards-automated-auditing-for-account-and-session-management-flaws-in-single-si |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IEEE-SP/2024/dnsbomb-a-new-practical-and-powerful-pulsing-dos-attack-exploiting-dns-queries-a | IEEE-SP/2024/dnsbomb-a-new-practical-and-powerful-pulsing-dos-attack-exploiting-dns-queries-a |
| network-layer work that uses spam as a label source | network-layer work that uses spam as a label source |
| measures what campaign websites collect, including addresses; no mail was received or measured | measures what campaign websites collect, including addresses; no mail was received or measured |
| IEEE-SP/2023/d-dae-defense-penetrating-model-extraction-attacks | IEEE-SP/2023/d-dae-defense-penetrating-model-extraction-attacks |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IEEE-SP/2022/scraping-sticky-leftovers-app-user-information-left-on-servers-after-account-del | IEEE-SP/2022/scraping-sticky-leftovers-app-user-information-left-on-servers-after-account-del |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IEEE-SP/2023/rulekeeper-gdpr-aware-personal-data-compliance-for-web-frameworks | IEEE-SP/2023/rulekeeper-gdpr-aware-personal-data-compliance-for-web-frameworks |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| IEEE-SP/2024/erasan-efficient-rust-address-sanitizer | IEEE-SP/2024/erasan-efficient-rust-address-sanitizer |
| read and judged off-topic: the message vocabulary is incidental to the object measured | not individually annotated: screened at title level, matched no off-topic family rule |
| |
| --- Rejected outside the pool — 9 | --- Rejected outside the pool — 9 |
| detection[].phenomenon =~ /spam/i | detection[].phenomenon =~ /spam/i |
| (not counted) — free text, ~20% run-to-run stable; and the phenomenon string does not say which channel | (not counted) — free text, ~20% run-to-run stable; and the phenomenon string does not say which channel |
| full text =~ /\bunsubscri|List-Unsubscribe|CAN-?SPAM/i | full text =~ /\bunsubscrib/i |
| 59 papers — MQTT / SDN protocol verb in 6 of 59; CCPA "opt out of sale" in a further 6; one genuine marketing-mail paper | 56 papers — MQTT / SDN protocol verb in 10 of 56; CCPA "opt out of sale" or an advertising opt-out cookie in 3; turning off a security notification in 4; the remaining 39 have one passing mention each. Zero papers whose OBJECT is an unsubscription mechanism |
| full text =~ /\bspam/i | full text =~ /\bspam/i |
| 954 papers — a fact about these being security venues, not a population | 954 papers — a fact about these being security venues, not a population |
| outside: WWW/2026/towards-token-level-text-anomaly-detection | outside: WWW/2026/towards-token-level-text-anomaly-detection |
| outside: IEEE-SP/2022/symbexcel-automated-analysis-and-understanding-of-malicious-excel-4-0-macros | outside: IEEE-SP/2022/symbexcel-automated-analysis-and-understanding-of-malicious-excel-4-0-macros |
| | |
| | outside-the-pool mentions across the four probes: 72 |
| | DISTINCT papers outside the pool that the probes surfaced: 67 |
| |
| Verdict on the 2026-09-02 run. All four probes' outside-the-pool lists were | Verdict on the 2026-09-02 run. All four probes' outside-the-pool lists were |
| recall probe /spam ?trap|honey ?pot (e-?mail|account)|honey ?token/ (8 hits) | recall probe /spam ?trap|honey ?pot (e-?mail|account)|honey ?token/ (8 hits) |
| are-you-sure-you-want-to-contact-us-quantifying-the-leakage-of-pii-via-website-c | are-you-sure-you-want-to-contact-us-quantifying-the-leakage-of-pii-via-website-c |
| recall probe detection[].phenomenon =~ /e-?mail|inbox|.../ — a genuine miss, and the oldest paper in the EID slice: contact forms as the gateway from pseudonym to address, five years before Leaky Forms | recall probe detection[].phenomenon =~ /e-?mail|inbox|.../ — a genuine miss, and the oldest paper in the EID slice: contact forms as the gateway from pseudonym to address, six years before Leaky Forms |
| |
| ============================================================================== | ============================================================================== |
| measured. Source: data/extract/README.md and OVERVIEW.md's own note | measured. Source: data/extract/README.md and OVERVIEW.md's own note |
| that the stability table is not recomputed by the overview script. | that the stability table is not recomputed by the overview script. |
| 12,618 mailing lists subscribed to by Englehardt et al., PoPETs 2018 — | 12,618 EMAILS collected by Englehardt et al., PoPETs 2018, from 902 |
| a per-paper figure, checked by verify_email_figures.mjs | distinct senders on 15,700 sites crawled. NOT a count of mailing |
| | lists — that wording was wrong here and on the page until review |
| | on 2026-09-02. All four figures are needles in |
| | verify_email_figures.mjs |
| 56.5% data/extract/OVERVIEW.md's own artifacts.availability rate, quoted on | 56.5% data/extract/OVERVIEW.md's own artifacts.availability rate, quoted on |
| the page only to say it is NOT comparable with this report's 59.5% | the page only to say it is NOT comparable with this report's 59.5% |
| | ''node scripts/check_page_numbers.mjs pages/privacy_email_tracking.txt <report+verify+external> '===== Use in Publications =====' '===== What to Report ====='%%'%%'' | OK, after two rounds of fixes | | | ''node scripts/check_page_numbers.mjs pages/privacy_email_tracking.txt <report+verify+external> '===== Use in Publications =====' '===== What to Report ====='%%'%%'' | OK, after two rounds of fixes | |
| | the same, **whole page** (no markers) | OK, after four rounds of fixes | | | the same, **whole page** (no markers) | OK, after four rounds of fixes | |
| | ''node scripts/verify_email_figures.mjs'' | 68 figures across 23 papers, 0 missing | | | ''node scripts/verify_email_figures.mjs'' | 78 figures across 24 papers, 0 missing | |
| | ''bash scripts/external_checks_email_tracking.sh'' | all checks PASSED after its own RFC 6376 assertion was corrected | | | ''bash scripts/external_checks_email_tracking.sh'' | all checks PASSED after its own RFC 6376 assertion was corrected | |
| |
| |
| ^ # ^ Finding ^ Verdict ^ | ^ # ^ Finding ^ Verdict ^ |
| | 1 | **HIGH. "12,618 mailing lists" is wrong.** {[englehardt2018_email]} crawled 15,700 sites and assembled "12,618 emails from 902 distinct senders"; 12,618 is the corpus size, not the number of lists. The wrong figure appeared twice, including in Open Questions as //the// reference figure. | **Accepted, and it is the worst error found.** Both occurrences corrected; the page now gives 15,700 crawled, 902 senders, 12,618 emails, a 38% submission rate of which 32% were list subscriptions, and says explicitly which number means what. All four are now needles in ''verify_email_figures.mjs''. | | | 1 | **HIGH. "12,618 mailing lists" is wrong.** {[englehardt2018_email]} crawled 15,700 sites and assembled "12,618 emails from 902 distinct senders"; 12,618 is the corpus size, not the number of lists. The reviewer found it in two places on the page; a later grep found two more — a third occurrence in the residue commentary and the hand-map reason inside ''msg_fold.mjs'', whose output is published on this page, so the wrong figure was in the audit trail as well. | **Accepted, and it is the worst error found.** All four occurrences corrected; the page now gives 15,700 crawled, 902 senders, 12,618 emails, a 38% submission rate of which 32% were list subscriptions, and says explicitly which number means what. All four are now needles in ''verify_email_figures.mjs''. | |
| | 2 | **HIGH. 53 lines of internal ''bibgen'' QA notes were pasted into the live [[literature:bibliography]] page**, inside the database block, rendering as visible text ("metadata source: venue-page", "cited by ? (OpenAlex)", "check spacing in title"). | **Accepted. This was a real accidental exposure and it was mine.** ''bibgen.mjs'' prints those notes to **stdout**, and the additions file was built by redirecting stdout. Removed from the live page and from the local file; the entry count was checked before and after (816 → 816, zero keys lost) because the first attempt at removing it cut a block and silently dropped {[starov2016_sure]}. | | | 2 | **HIGH. 53 lines of internal ''bibgen'' QA notes were pasted into the live [[literature:bibliography]] page**, inside the database block, rendering as visible text ("metadata source: venue-page", "cited by ? (OpenAlex)", "check spacing in title"). | **Accepted. This was a real accidental exposure and it was mine.** ''bibgen.mjs'' prints those notes to **stdout**, and the additions file was built by redirecting stdout. Removed from the live page and from the local file; the entry count was checked before and after (816 → 816, zero keys lost) because the first attempt at removing it cut a block and silently dropped {[starov2016_sure]}. | |
| | 3 | **MEDIUM. The live entry for {[stringhini2012_leveraging]} renders "B@bel" with U+FF20 rather than an ASCII ''@''.** | **Accepted as a finding, rejected as a fix.** The reviewer is right that the glyph is wrong. It is also load-bearing: retested on a clean file at rev 1788333447, the ASCII ''@'' made the entry **and all five citation markers to it** vanish with no warning, references dropping 72 → 71. ''B{@}bel'', ''B\\@bel'' and ''B@bel'' were all tested too; only the last resolved, and it rendered the entity literally. U+FF20 stays, with the reason in a ''%'' comment above the entry so nobody "fixes" it. | | | 3 | **MEDIUM. The live entry for {[stringhini2012_leveraging]} renders "B@bel" with U+FF20 rather than an ASCII ''@''.** | **Accepted as a finding, rejected as a fix.** The reviewer is right that the glyph is wrong. It is also load-bearing: retested on a clean file at rev 1788333447, the ASCII ''@'' made the entry **and its citation marker** vanish with no warning, references dropping 72 → 71. (An earlier draft of this log said "all five markers". This page carries **one** marker to that key and is the only page on the wiki that cites it; the re-verification pass caught the exaggeration.) ''B{@}bel'', ''B\\@bel'' and ''B@bel'' were all tested too; only the last resolved, and it rendered the entity literally. U+FF20 stays, with the reason in a ''%'' comment above the entry so nobody "fixes" it. | |
| | 4 | **LOW-MEDIUM. ''sherman2020_going'' renders an author as "Jr., Keith McNamara"** — BibTeX read "Jr." as the family name. | **Accepted.** Corrected to the three-part ''McNamara, Jr., Keith''. | | | 4 | **LOW-MEDIUM. ''sherman2020_going'' renders an author as "Jr., Keith McNamara"** — BibTeX read "Jr." as the family name. | **Accepted.** Corrected to the three-part ''McNamara, Jr., Keith''. | |
| | 5 | **LOW. Five genuine duplicate papers under two keys each exist in the live bibliography**, three of them sharing a DOI. | **Accepted as a finding, out of scope for this page.** None is cited here. Filed as its own work item rather than fixed in passing, because deduplicating a shared 816-entry bibliography touches every page that cites the losing key. | | | 5 | **LOW. Five genuine duplicate papers under two keys each exist in the live bibliography**, three of them sharing a DOI. | **Accepted as a finding, out of scope for this page.** None is cited here. Filed as its own work item rather than fixed in passing, because deduplicating a shared 816-entry bibliography touches every page that cites the losing key. | |
| |
| Two things were found before the review pass, by the guards rather than by reading, and both are recorded in section 13: four silence base rates copied out of ''OVERVIEW.md'' instead of computed, and the denominator mismatch that copying them concealed. Two more were found by the author while checking a claim the review would probably have accepted: **"on by default since iOS 15"** for Mail Privacy Protection, which Apple documents as a setting the user turns on and for which Apple publishes no take-up figure, and **"Gmail's image proxy has been on since 2013"**, which was recall until the December 2013 announcement was fetched. The provenance page's own assertion that eight below-threshold quotes had been read by hand was written before they had been; reading them is what produced the table in section 8, and checking the sentence after it is what added two papers to the verifier. | Two things were found before the review pass, by the guards rather than by reading, and both are recorded in section 13: four silence base rates copied out of ''OVERVIEW.md'' instead of computed, and the denominator mismatch that copying them concealed. Two more were found by the author while checking a claim the review would probably have accepted: **"on by default since iOS 15"** for Mail Privacy Protection, which Apple documents as a setting the user turns on and for which Apple publishes no take-up figure, and **"Gmail's image proxy has been on since 2013"**, which was recall until the December 2013 announcement was fetched. The provenance page's own assertion that eight below-threshold quotes had been read by hand was written before they had been; reading them is what produced the table in section 8, and checking the sentence after it is what added two papers to the verifier. |
| | |
| | ==== 14.5 Author's own pass over the frozen text, after the reviewers ==== |
| | |
| | Four more things were changed after the review, found by reading the page top to bottom rather than by any check: |
| | |
| | * The intro claimed **"two of the four things you might want to measure have essentially never been measured"**. One slice is zero; the other three all have papers. Rewritten to say what is true: one slice is a measured zero, and a second question — what the mailbox providers' own defences do to a tracking measurement — is unmeasured. |
| | * **"Four neighbouring clusters are larger than anything on this page"** was false of one of the four: the web-pixel cluster is 7 papers against the SMS slice's 23. Now "three of them are larger than any slice that is". |
| | * **"This topic has an unusual concentration of ethics problems, and the corpus reflects it"** used a legal-assessment rate as evidence for an ethics claim. The rate (20.0% against 6.9%) is real and stays; the inference does not, and the page now says the number is about legal engagement and that the reporting figures at the foot of the same list say the ethics are not handled well. |
| | * **"Three of the recent smishing and SMS-spam papers draw on partly overlapping public sources"** was not supported by anything. Two of them mine Twitter ({[tang2022_clues]}, {[agarwal2025_fishing]}); one uses a crowdsourced app and one carrier-side reports. Replaced with the two that actually overlap and the specific comparison nobody has run. |
| | |
| | The pattern in all four is the same and worth naming: **the number was right and the sentence around it claimed more than the number supports.** Neither the number guard nor a figures reviewer can see that, because they check digits against a script. It is the failure mode the generic review slot exists for, and half of it was still found by re-reading rather than by a reviewer. |
| | |
| | ==== 14.6 Generic pass (Fable), and the re-verification that followed ==== |
| | |
| | The generic reviewer's verdict was that the pages are well guarded on their headline figures and that **the provenance page repeatedly failed to practise what it preaches about itself**. That is correct, and the underlying cause is one this project has a memory note about: //a fix survives in the provenance log.// Every accepted review fix had been applied to the content page and almost none had been propagated into the script constants and the prose template that the provenance page republishes. |
| | |
| | ^ # ^ Finding ^ Verdict ^ |
| | | 1 | **"183 hand-written verdicts, each with its reason" is overstated, and the residue manufactures evidence of reading.** 57 ''MAP'' entries have an empty reason string, and 171 of the 181 ''OFF'' papers are not in the map at all — they were given a **canned default** reading "read and judged off-topic: the message vocabulary is incidental to the object measured", printed per paper in the published residue. A typo'd ''MAP'' key would have become a silent ''OFF'' still claiming it had been read. | **Accepted, and it is the most serious finding of the whole review.** The default string now says //"not individually annotated: screened at title level, matched no off-topic family rule"//. The report prints the real accounting (183 mapped / 125 with a written reason / 171 defaulted / 46 by family rule / 125 bare) and both pages quote it instead of asserting the stronger claim. All 173 non-''OFF'' papers do carry a hand-written reason; that part was true. | |
| | | 2 | **The error 14.2 called "the worst error found" was still published — inside this page.** "12,618 mailing lists" survived in ''report_email_tracking.mjs'''s section Z and therefore in the embedded output, annotated "checked by verify_email_figures.mjs" — which checks digits, not what the noun claims. Same line: {[starov2016_sure]} described as "five years before Leaky Forms" when 2016→2022 is six. | **Accepted.** Both strings fixed at source in the scripts, re-run, rebuilt. The lesson is the annotation, not the number: a verifier that confirms ''12,618'' appears in the paper says nothing about the sentence around it. | |
| | | 3 | **One provenance page, three values for its own figure check**: §7 said 60 across 22, §13 said 68 across 23, the embedded output said 78 across 24 — and §14.1 had already "accepted" the fix. | **Accepted.** Both prose numbers corrected, and §7 now says where the count comes from so it cannot drift again. | |
| | | 4 | **The published script's own annotations contradict the published output.** ''REJECTED_PROBES'' still said "MQTT in 6 of 59; CCPA a further 6" where section E prints 10 / 3 / 4 of 56, and the header comment said "28 papers social spam against 16 email spam" where the figures are 37 / 15. | **Accepted.** All corrected at source. | |
| | | 5 | **The below-threshold quote rule generalises from a sample that excluded the riskiest cases**: eight quotes read, all from the 40–59% band, then a rule declared about that band — while naming four below 25% as "worth opening" and not opening two of them. | **Accepted.** Both were then read: {[tu2016_security]} at 23% and {[utz2023_comparing]} at 20%, tabulated in §8, both present. The content page now says ten were read, not "every one sampled". | |
| | | 6 | The intro's "two of the four things … have essentially never been measured" has no owner. | **Accepted — and already fixed** in the author's own pass (§14.5) before this review returned. | |
| | | 7 | **"the single biggest coherent slice the pool turned up" (''INFRA'', 31) is contradicted two bullets later** by ''SOC'' at 37. | **Accepted.** Reworded to say the social cluster is larger but is not one topic. | |
| | | 8 | Five smaller self-contradictions: "both CAN-SPAM hits are motivational asides" (one is); the probe table drops the sixth row the report prints, on a page that twice says nothing is truncated; §1 says 56 bib entries and §15 said "56 became 57"; the run log said 145 markers where the source has 147; and the Apple footnote says ''curl'' returns only a shell while the check script's ''curl'' finds the needle. | **All five accepted and fixed.** The Apple one mattered most: ''curl'' does return a noscript copy of the body, so the footnote now tells the reader to trust the browser fetch and not the ''curl'' pass, rather than implying the ''curl'' check is what verified the load-bearing quote. | |
| | | 9 | **The four mail-specific tools are hidden on the provenance page** while the content page tells the student "there is no toolchain for this topic". For someone starting on Monday those four names are the most actionable line on either page. | **Accepted, and the best suggestion of the review.** The Email Privacy Tester, EmailHarvester, the Honey Messages Framework and ''css-inline'' are now named on the content page, together with the CAPTCHA-solver and fake-identity residue, which is what signing up on fifteen thousand sites actually costs. The suggestion to cut the Gmail/Apple point from four appearances to two was **rejected**: the four are a claim, a corpus gap, a study-design consequence and an open question, which are four different uses of the same fact. | |
| | |
| | ==== 14.7 Re-verification pass (Sonnet, against the corrected files) ==== |
| | |
| | All eight claimed fixes re-derived independently and **CONFIRMED**: the five Englehardt figures against the paper; the study-shape figures recomputed from ''buildPool()'' (70 / 29 / 27 / 33 / 54 / 12 / 42 and 14-7-5-16-22-22 of 27); both untruncated tables; the verifier count; the non-zero exit; the ''browsers'' array fix at 16 of 27 (and a sweep confirming ''languages'' is the only other array field, and it is not in that table); the bibliography with the QA notes gone and 816 entries intact; and the U+FF20 title rendering with every marker resolving. |
| | |
| | It also found two things nobody else did: |
| | |
| | * **"73 papers outside the pool" was wrong.** The four probes produce **72** outside-the-pool mentions and **67 distinct papers**, because five papers appear on more than one probe's list. Neither figure is 73. Worse, ''check_page_numbers.mjs'' **passed** it, because its shared ''ALLOW'' map contains a ''73'' written for ''design:platforms'' — a whitelist entry for one page silently blessing a wrong number on another. Fixed on the page (67 distinct, 72 mentions), and the report now prints both counts so neither has to be derived by hand. **The shared ''ALLOW'' map is a cross-page liability and this is the first measured instance of it hiding a real error.** |
| | * **The "all five citation markers" claim in 14.2 above.** There is one. Corrected in place. |
| | |
| | ==== 14.8 A guard that came out of the review ==== |
| | |
| | ''scripts/check_attributions.mjs'' already existed, and its own header said "a page that attributes in prose is not covered; report 0 checked rather than assume pass". This page attributes entirely in prose, so it reported **0 checked** — and a fabricated author name shipped anyway: the page said //"Sharevski and Zettlemoyer"// where the authors are Sharevski, Loop, Evans and Ponticello. No other guard could see it. The citekey resolved, the figure beside it was verified against the paper, and the number guard traced the digits. |
| | |
| | A ''--prose'' mode was added: it matches ''Surname'', ''Surname et al.'', ''Surname and Surname'' and ''Surname, Surname and Surname'' immediately before a marker, and requires every capitalised token to appear somewhere in that entry's **full** author string. Two iterations were needed and both are worth recording, because a guard that is wrong in either direction does not get run twice: |
| | |
| | - A first version scanned every capitalised token in a 90-character window and reported **eight** false positives on a clean page (''XRay'', ''SMTP-dialect'', ''DNSBL'', ''GDPR'', ''URL'', "Older", "Since"). |
| | - Tightening it to the four name shapes brought it to 52 attributions checked and **one** suspect — //"Kubicek et al."// against ''Kub{\'i}{\v{c}}ek'', because the skeletoniser folded ''\v{c}'' to "vc". A LaTeX-accent stripping step fixed that. 52 checked, 0 suspect. |
| | |
| | Run it in **both** modes on any page it is pointed at. Table mode reports 0 on this page and says so loudly, which is the behaviour that let the error through in the first place. |
| |
| ===== 15. Run log ===== | ===== 15. Run log ===== |
| | 2026-09-02 | Whole page, this provenance page, ''msg_fold.mjs'', ''report_email_tracking.mjs'', ''verify_email_figures.mjs'', ''external_checks_email_tracking.sh'', and 56 bibliography entries, written in one sitting against the 5,859-paper corpus | | | 2026-09-02 | Whole page, this provenance page, ''msg_fold.mjs'', ''report_email_tracking.mjs'', ''verify_email_figures.mjs'', ''external_checks_email_tracking.sh'', and 56 bibliography entries, written in one sitting against the 5,859-paper corpus | |
| | 2026-09-02 | Signal 3 of the pool rule added mid-run after signals 1 and 2 were found to miss the messenger-enumeration papers. Four ''EID'' papers, including the two newest, entered that way | | | 2026-09-02 | Signal 3 of the pool rule added mid-run after signals 1 and 2 were found to miss the messenger-enumeration papers. Four ''EID'' papers, including the two newest, entered that way | |
| | 2026-09-02 | Recall probes run after the pool was fixed; one paper ({[starov2016_sure]}) added, and 56 bibliography entries became 57 as a result | | | 2026-09-02 | Recall probes run after the pool was fixed; one paper ({[starov2016_sure]}) added, and the bibliography additions grew from 55 entries to 56 as a result | |
| | 2026-09-02 | ''external_checks_email_tracking.sh'' failed on its own RFC 6376 assertion (''DRAFT STANDARD'' where the answer is ''INTERNET STANDARD''). The assertion was corrected; the source was right | | | 2026-09-02 | ''external_checks_email_tracking.sh'' failed on its own RFC 6376 assertion (''DRAFT STANDARD'' where the answer is ''INTERNET STANDARD''). The assertion was corrected; the source was right | |
| | 2026-09-02 | First Apple support URL fetched was the wrong article (Boot Camp). Caught because the fetch answered "I cannot answer this from the provided content" rather than returning something plausible | | | 2026-09-02 | First Apple support URL fetched was the wrong article (Boot Camp). Caught because the fetch answered "I cannot answer this from the provided content" rather than returning something plausible | |
| | 2026-09-02 | **The bibtex4dw parser splits entries on any literal ASCII ''@'' and silently drops the entry — and every citation marker to it.** {[stringhini2012_leveraging]} is titled //B@bel//. Its five citation markers rendered as nothing at all, with no warning, and the reference was absent from the list; the count of rendered references against distinct markers is what caught it. ''B{@}bel'' and ''B\\@bel'' both still failed; the HTML entity ''B@bel'' resolved the entry but rendered the entity literally. The title now uses **U+FF20 FULLWIDTH COMMERCIAL AT**. Do not "fix" it back to a plain ''@'' without re-checking that the entry still resolves | | | 2026-09-02 | **The bibtex4dw parser splits entries on any literal ASCII ''@'' and silently drops the entry — and every citation marker to it.** {[stringhini2012_leveraging]} is titled //B@bel//. Its five citation markers rendered as nothing at all, with no warning, and the reference was absent from the list; the count of rendered references against distinct markers is what caught it. ''B{@}bel'' and ''B\\@bel'' both still failed; the HTML entity ''B@bel'' resolved the entry but rendered the entity literally. The title now uses **U+FF20 FULLWIDTH COMMERCIAL AT**. Do not "fix" it back to a plain ''@'' without re-checking that the entry still resolves | |
| | 2026-09-02 | New citekeys need a cache purge. Immediately after publishing, only 30 of the page's 145 markers rendered and only 16 references appeared. ''?purge=true'' on [[literature:bibliography]] and then on the content page fixed it. **Always count rendered references against distinct markers after publishing** | | | 2026-09-02 | New citekeys need a cache purge. Immediately after publishing, only 30 of the page's markers rendered — there were 144 at the time, 147 by the end of review — and only 16 references appeared. ''?purge=true'' on [[literature:bibliography]] and then on the content page fixed it. **Always count rendered references against distinct markers after publishing** | |
| |
| Nothing on this page was carried over from an earlier run. The three ''bench-*'' dossiers and every published figure predating 2026-08-11 were treated as stale by construction and not consulted for numbers. | Nothing on this page was carried over from an earlier run. The three ''bench-*'' dossiers and every published figure predating 2026-08-11 were treated as stale by construction and not consulted for numbers. |