User Tools

Site Tools


privacy:email_tracking

This is an old revision of the document!


Email Tracking and Message-Channel Measurement

You are about to measure something that arrives in a mailbox or on a handset: a tracking pixel in a newsletter, whether an unsubscribe link works, how much spam a filter lets through, what an SMS scam campaign looks like. This page is about the instrument that question needs — an address you control, a way to make mail arrive at it, and a way to render it without becoming part of the measurement — and about which of the field's methods are still current.

It is not a tutorial on SMTP, MIME or what a 1×1 GIF is. It assumes you can read RFC 5322 yourself. What it adds is the part no spec tells you: that the mailbox provider now sits between you and the sender and changes what your instrument can see; that one of the four things you might want to measure here has never been measured in these seven venues, and a second — what those provider defences do to a tracking measurement — has not been measured either; and that the busiest part of this topic stopped being email some years ago.

The one thing to understand before you start: the mailbox provider is now a party to your measurement, and it does not do the same thing to every signal. Gmail proxies remote images, so the sender learns nothing about the recipient's IP, user agent or cookies — but still learns that the message was opened. Google says both parts in the same help article: “Senders can't use image loading to get information about your computer or location”, “Senders can't use the image to set or read cookies in your browser”, and then “Sometimes, senders may know whether you've opened an email that has an image.”1) Apple Mail Privacy Protection, available since iOS 15 (2021), takes the other half too: it “hides your IP address so senders can't link it to your other online activity or determine your exact location” and “prevents senders from seeing if you've opened the email message they sent you”, by fetching remote content in the background on arrival rather than on view.2)

So “we found a tracking pixel in 24.6% of messages” and “the sender learned that 24.6% of messages were opened” are now different claims, and which one you can make depends on which client you rendered in. None of the email-tracking papers in this corpus measures the effect of that change: the newest of them [1Chand, Anish; Nikiforakis, Nick; Vadrevu, Phani (2025): "Doubly Dangerous: Evading Phishing Reporting Systems by Leveraging Email Tracking Techniques", in: Proceedings of the USENIX Security Symposium. (Link)] was published in 2025 and still treats the pixel fetch as the signal. If you are planning this measurement, the client is a treatment you have to vary on purpose, and you have no baseline in this literature to compare against. See What a Pixel Can Still Measure in 2026.

What to Read First

  • Englehardt, Han and Narayanan, I never signed up for this! Privacy implications of email tracking [2Englehardt, Steven; Han, Jeffrey; Narayanan, Arvind (2018): "I never signed up for this! Privacy implications of email tracking", Proceedings on Privacy Enhancing Technologies 2018(1):109-126. (DOI)] (PoPETs 2018) — the method paper for this whole topic. Crawl 15,700 sites, sign up for mail on each with a distinct address, render what arrives in instrumented clients, and look for the address in outbound requests. The corpus that came out is 12,618 emails from 902 distinct senders — a 38% submission success rate, of which 32% were mailing-list subscriptions. Quote the 902, not the 12,618, when you mean senders. Everything since is a variation on it.
  • Hu and Wang, Characterizing Pixel Tracking through the Lens of Disposable Email Services [3Hu, Hang; Peng, Peng; Wang, Gang (2019): "Characterizing Pixel Tracking through the Lens of Disposable Email Services", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] (IEEE S&P 2019) — the other vantage point, and read it for the denominator shock: on a disposable-mail provider, 94.75% of the mail is spam and only 3.63% is registration, so its “24.6% of messages carry tracking” is a statement about a different population than Englehardt's.
  • Kirchner et al., A Black-Box Privacy Analysis of Messaging Service Providers' Chat Message Processing [4Kirchner, Robin; Koch, Simon; Kamangar, Noah; Klein, David; Johns, Martin (2024): "A Black-Box Privacy Analysis of Messaging Service Providers' Chat Message Processing", in: Proceedings on Privacy Enhancing Technologies. (DOI)] (PoPETs 2024) — the same method, carried to 105 messaging platforms with honey messages and honey tokens. The most useful single paper if your channel is not email.
  • Pitsillidis et al., Taster's Choice: A Comparative Analysis of Spam Feeds [5Pitsillidis, Andreas; Kanich, Chris; Voelker, Geoffrey M.; Levchenko, Kirill; Savage, Stefan (2012): "Taster's choice: a comparative analysis of spam feeds", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] (IMC 2012) — old, and still the paper to read before you take a message corpus from anyone. Ten contemporaneous feeds, and 60% of live domains appeared in only one of them.
  • Senol et al., Leaky Forms [6Senol, Asuman; Acar, Gunes; Humbert, Mathias; Zuiderveen Borgesius, Frederik (2022): "Leaky Forms: A Study of Email and Password Exfiltration Before Form Submission", in: 31st USENIX Security Symposium (USENIX Security 22). (Link)] (USENIX Sec 2022) and Starov et al., Are You Sure You Want to Contact Us? [7Starov, Oleksii; Gill, Phillipa; Nikiforakis, Nick (2016): "Are You Sure You Want to Contact Us? Quantifying the Leakage of PII via Website Contact Forms", in: Proceedings on Privacy Enhancing Technologies. (DOI)] (PoPETs 2016) — how the address gets from the user to the sender in the first place, six years apart.
  • Prasad et al., Characterizing Robocalls with Multiple Vantage Points [8Prasad, Sathvik; Nahapetyan, Aleksandr; Reaves, Bradley (2025): "Characterizing Robocalls with Multiple Vantage Points", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] (IEEE S&P 2025) — if your channel is voice or SMS, start here, because it is the paper about vantage points rather than about a single honeypot.

The Boundary: One Page, Not Five

This page was scoped by inventorying the publication corpus (seven venues, 2010–2026, 5,859 extracted papers) for the four slices its brief named, and by counting them before deciding. The result was one page with sections rather than a parent with children, because two of the four slices cannot carry a page and one of them has no papers at all.

Slice Papers Span Decision
Tracking inside a message — a pixel, a CSS load, a link, in mail or in chat 8 2014–2025 A section. Thin, but it is the page's own subject and the papers are the canonical ones
The address itself as the identifier — exfiltration from forms, hashed-email ad matching, alias collision, phone-number enumeration 12 2016–2026 A section. Not in the original brief, and the slice that is growing fastest
Unsubscription, opt-out, List-Unsubscribe 0 No page, no section beyond a gap statement. List-Unsubscribe appears in one paper in 5,859, as a row in a table of signable header fields. See Unsubscription: A Measured Gap
Spam as the object of study (email) 15 2010–2025 A section, dated: ten of the fifteen are 2010–2012 and their methods are not current
SMS, calls and messaging-app abuse 23 2011–2026 A section, and the largest and most current slice on this page. If it grows past ~35 papers it should be split out as privacy:message_spam
The mailbox as the instrument, object elsewhere — honey accounts, canary addresses, notification and DSAR mail 12 2016–2026 A section, because it is where a student most often actually needs a mailbox

Seventy papers, then. A parent plus children would have put one page's worth of material behind five links, one of which has nothing at all to say and three of which have fewer papers than openwpm has for a single tool. The judgement, the counts behind it, and the alternatives considered are on email_tracking.

What is deliberately not here

Four neighbouring clusters were screened out of the same candidate pool and are not on this page. Three of them are larger than any slice that is. Each is named so you do not conclude the corpus is silent about it:

  • Email transport, authentication and encryption deployment — SPF, DKIM, DMARC, DANE, MTA-STS, STARTTLS, S/MIME, sender spoofing, delivery paths. 31 papers, heavily 2018–2026, and the single biggest coherent slice the pool turned up. It is page-sized and it is not tracking; it belongs under security: and does not exist yet. Start from Durumeric et al. [9Durumeric, Zakir; Adrian, David; Mirian, Ariana; Kasten, James; Bursztein, Elie; Lidzborski, Nicolas; Thomas, Kurt; Eranti, Vijay; Bailey, Michael D.; Halderman, J. Alex (2015): "Neither Snow Nor Rain Nor MITM...: An Empirical Analysis of Email Delivery Security", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] and Shen et al. [10Shen, Kaiwen; Wang, Chuhan; Guo, Minglei; Zheng, Xiaofeng; Lu, Chaoyi; Liu, Baojun; Zhao, Yuxuan; Hao, Shuang; Duan, Haixin; Pan, Qingfeng; Yang, Min (2021): "Weak Links in Authentication Chains: A Large-scale Analysis of Email Sender Spoofing Attacks", in: Proceedings of the USENIX Security Symposium. (Link)] if you need it now.
  • Phishing and its interventions, with email as the vector — 28 papers. Phishing owns this. The line drawn here: a paper about what a message discloses to a third party is on this page; a paper about whether a user falls for a message is on that one.
  • Social-platform, review, forum, SEO and ad-click spam37 papers, more than any slice on this page. “Spam” in these venues usually means Twitter accounts or product reviews, not mail. It is a platform question (Platforms), not a message-channel one.
  • The tracking pixel on a web page7 papers, including the Meta Pixel work and the invisible-pixel filter-list studies. Requests owns it. Two things that share a name and share almost no method: a web pixel runs in a browser with JavaScript, cookies and a referrer; an email pixel runs in a mail client with none of those.

Pick the Instrument Before the Question

Every measurement on this page needs an address, and how you got the address decides your population, your ethics section, and what a reviewer will ask. This is the choice to make first.

Instrument What you get What it costs Papers
Subscribe with your own addresses — one per site, from a domain you control Full control of the client, the render, the network view; you know exactly which site got which address, so a leak is attributable Slow and manual-ish; the population is whatever you could sign up to, so it is a convenience sample of sites that have a newsletter; and you have created accounts under a pseudonym, which is an ethics-section problem [2Englehardt, Steven; Han, Jeffrey; Narayanan, Arvind (2018): "I never signed up for this! Privacy implications of email tracking", Proceedings on Privacy Enhancing Technologies 2018(1):109-126. (DOI)] [11Kubíček, Karel; Merane, Jakob; Cotrini, Carlos; Stremitzer, Alexander; Bechtold, Stefan; Basin, David (2022): "Checking Websites' GDPR Consent Compliance for Marketing Emails", Proceedings on Privacy Enhancing Technologies 2022(2). (DOI)] [12Martin, Eric Burton Samuel; Shirazi, Hossein; Ray, Indrakshi (2023): "Poster: Towards a Dataset for the Discrimination between Warranted and Unwarranted Emails", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)]
A disposable-mail provider's stream Volume, and mail nobody would have signed you up for; catches spam and account-management mail You did not choose the population and cannot describe it; 94.75% spam [3Hu, Hang; Peng, Peng; Wang, Gang (2019): "Characterizing Pixel Tracking through the Lens of Disposable Email Services", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]; and you are reading other people's mail, which is why that paper's ethics section is worth copying [3Hu, Hang; Peng, Peng; Wang, Gang (2019): "Characterizing Pixel Tracking through the Lens of Disposable Email Services", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]
A provider-side dataset (an ESP, a carrier, a mail operator) Scale nothing else reaches, and the only way to see delivery and filtering decisions Not reproducible by anyone without the same partner; and the paper cannot usually say enough about the population for you to bound it [13Dasgupta, Anirban; Punera, Kunal; Rao, Justin M.; Wang, Xuanhui (2012): "Impact of Spam Exposure on User Engagement", in: Proceedings of the USENIX Security Symposium. (Link)] [14Huh, Jun Ho; Shin, Hyejin; Ahn, Sunwoo; Yi, Hayoon; Cho, Joonho; Kim, Taewoo; Lim, Minchae; Choi, Nuel (2025): "Preventing Artificially Inflated SMS Attacks through Large-Scale Traffic Inspection", in: Proceedings of the USENIX Security Symposium. (Link)]
Seeded accounts at real providers, as an A/B treatment Measures the provider's behaviour — its filter, its warnings, its proxy — which is often the actual question Provider anti-abuse will fight you; account age and reputation confound everything; N is small [15Iqbal, Hassan; Khan, Usman Mahmood; Khan, Hassan Ali; Shahzad, Muhammad (2022): "Left or Right: A Peek into the Political Biases in Email Spam Filtering Algorithms During US Election 2020", in: Proceedings of the ACM Web Conference. (DOI)] [16Rao, Sumanth; Liu, Enze; Ho, Grant; Voelker, Geoffrey M.; Savage, Stefan (2024): "Unfiltered: Measuring Cloud-based Email Filtering Bypasses", in: Proceedings of the ACM Web Conference. (DOI)] [17Lécuyer, Mathias; Ducoffe, Guillaume; Lan, Francis; Papancea, Andrei; Petsios, Theofilos; Spahn, Riley; Chaintreau, Augustin; Geambasu, Roxana (2014): "XRay: Enhancing the Web’s Transparency with Differential Correlation", in: Proceedings of the USENIX Security Symposium. (Link)]
Honey addresses / honeytokens given to a third party and then watched Attributes misuse to a specific recipient, which no crawl can do Detects only what someone chooses to send; a negative is not evidence of no misuse [18Farooqi, Shehroze; Musa, Maaz; Shafiq, Zubair; Zaffar, Fareed (2020): "CanaryTrap: Detecting Data Misuse by Third-Party Apps on Online Social Networks", in: Proceedings on Privacy Enhancing Technologies. (DOI)] [19DeBlasio, Joe; Savage, Stefan; Voelker, Geoffrey M.; Snoeren, Alex C. (2017): "Tripwire: inferring internet site compromise", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] [4Kirchner, Robin; Koch, Simon; Kamangar, Noah; Klein, David; Johns, Martin (2024): "A Black-Box Privacy Analysis of Messaging Service Providers' Chat Message Processing", in: Proceedings on Privacy Enhancing Technologies. (DOI)]
A telephony or SMS honeypot (numbers you own, calls you receive) The only unsolicited-message population for voice and SMS that a researcher can actually own Number blocks have history and age effects — older blocks receive significantly more calls [20Gupta, Payas; Srinivasan, Bharat; Balasubramaniyan, Vijay; Ahamad, Mustaque (2015): "Phoneypot: Data-driven Understanding of Telephony Threats", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]; and a honeypot receives, it cannot sample [20Gupta, Payas; Srinivasan, Bharat; Balasubramaniyan, Vijay; Ahamad, Mustaque (2015): "Phoneypot: Data-driven Understanding of Telephony Threats", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] [21Prasad, Sathvik; Bouma-Sims, Elijah; Mylappan, Athishay Kiran; Reaves, Bradley (2020): "Who's Calling? Characterizing Robocalls through Audio and Metadata Analysis", in: Proceedings of the USENIX Security Symposium. (Link)] [8Prasad, Sathvik; Nahapetyan, Aleksandr; Reaves, Bradley (2025): "Characterizing Robocalls with Multiple Vantage Points", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]
Crowdsourced public reports (forum posts, complaint services, tweeted screenshots) Reach and language coverage no honeypot gets; 66 languages in one recent smishing study Reporting bias in every direction, and the message often arrives as a screenshot rather than as text, so you need OCR or a vision model before you have data [22Murynets, Ilona; Jover, Roger Piqueras (2012): "Crime scene investigation: SMS spam data analysis", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] [23Tang, Siyuan; Mi, Xianghang; Li, Ying; Wang, XiaoFeng; Chen, Kai (2022): "Clues in Tweets: Twitter-Guided Discovery and Analysis of SMS Spam", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] [24Agarwal, Sharad; Papasavva, Antonis; Suarez-Tangil, Guillermo; Vasek, Marie (2025): "Fishing for Smishing: Understanding SMS Phishing Infrastructure and Strategies by Mining Public User Reports", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]

These are not interchangeable and the papers that used different ones do not disagree with each other. Englehardt's “85% of emails contain third-party content” and Hu and Wang's “24.6% contain tracking links” are both correct about different mail.

Methods, and Which Ones Are Current

A ranking of what the 2010–2026 literature did is a fact about the literature, not advice. Every row below is dated and its status stated as of 2026-09-02. The corpus reaches 2026 but its 2025–2026 venue-years are provisional — CCS and IMC 2026 have not been held, and IEEE S&P and TheWebConf 2026 are incompletely selected — so current rows were checked against material outside the corpus as well.

Era Method Representative work Status in 2026
2014– Seed the mailbox, observe the output. Put content in an account you control, measure what the provider does downstream XRay on Gmail ads [17Lécuyer, Mathias; Ducoffe, Guillaume; Lan, Francis; Papancea, Andrei; Petsios, Theofilos; Spahn, Riley; Chaintreau, Augustin; Geambasu, Roxana (2014): "XRay: Enhancing the Web’s Transparency with Differential Correlation", in: Proceedings of the USENIX Security Symposium. (Link)] (80–90% precision and recall) Current, and the most under-used design here. It is the only way to measure the provider rather than the sender
2018– Subscribe-and-render: distinct address per site, render in an instrumented client, match the address against outbound traffic in every encoding you can think of [2Englehardt, Steven; Han, Jeffrey; Narayanan, Arvind (2018): "I never signed up for this! Privacy implications of email tracking", Proceedings on Privacy Enhancing Technologies 2018(1):109-126. (DOI)] (31 hash and encoding functions tested), [11Kubíček, Karel; Merane, Jakob; Cotrini, Carlos; Stremitzer, Alexander; Bechtold, Stefan; Basin, David (2022): "Checking Websites' GDPR Consent Compliance for Marketing Emails", Proceedings on Privacy Enhancing Technologies 2022(2). (DOI)] Current, and the default. Copy the encoding-enumeration step, not just the design
2019– Disposable-provider vantage: read a stream you did not solicit [3Hu, Hang; Peng, Peng; Wang, Gang (2019): "Characterizing Pixel Tracking through the Lens of Disposable Email Services", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] Current for volume, but answers a different question and needs its own ethics argument
2024– Honey messages and honey tokens across channels: put a unique URL in a message and watch who fetches it [4Kirchner, Robin; Koch, Simon; Kamangar, Noah; Klein, David; Johns, Martin (2024): "A Black-Box Privacy Analysis of Messaging Service Providers' Chat Message Processing", in: Proceedings on Privacy Enhancing Technologies. (DOI)] — 34% of 105 messaging services fetched the URL server-side; one fetched a page 30 days after the chat was closed Current, and the method to copy if your channel is not email. It also generalises the “is anyone reading this” question beyond tracking
2025– CSS-only fingerprinting, because mail clients run no JavaScript [25Trampert, Leon; Weber, Daniel; Gerlach, Lukas; Rossow, Christian; Schwarz, Michael (2025): "Cascading Spy Sheets: Exploiting the Complexity of Modern CSS for Email and Browser Fingerprinting", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] — 1,152 of 1,176 browser–OS combinations distinguished (97.95%) Current. The reason “mail clients don't execute JS” is not a privacy property
2015–2021 Value-length, expiry and pairwise-similarity heuristics for deciding whether a string in a message is an identifier not represented on this page; see Cookies Superseded wherever an alternative exists. In mail the alternative is easy: you planted the address, so you can search for it under every encoding rather than guessing what looks like an ID
2010–2012 Spam-feed domain extraction and value-chain tracing [26Levchenko, Kirill; Pitsillidis, Andreas; Chachra, Neha; Enright, Brandon; Félegyházi, Márk; Grier, Chris; Halvorson, Tristan; Kanich, Chris; Kreibich, Christian; Liu, He; McCoy, Damon; Weaver, Nicholas; Paxson, Vern; Voelker, Geoffrey M.; Savage, Stefan (2011): "Click Trajectories: End-to-End Analysis of the Spam Value Chain", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] [27McCoy, Damon; Pitsillidis, Andreas; Jordan, Grant; Weaver, Nicholas; Kreibich, Christian; Krebs, Brian; Voelker, Geoffrey M.; Savage, Stefan; Levchenko, Kirill (2012): "PharmaLeaks: Understanding the Business of Online Pharmaceutical Affiliate Programs", in: Proceedings of the USENIX Security Symposium. (Link)] [28Kanich, Chris; Weaver, Nicholas; McCoy, Damon; Halvorson, Tristan; Kreibich, Christian; Levchenko, Kirill; Paxson, Vern; Voelker, Geoffrey M.; Savage, Stefan (2011): "Show Me the Money: Characterizing Spam-advertised Revenue", in: Proceedings of the USENIX Security Symposium. (Link)] [29Stringhini, Gianluca; Holz, Thorsten; Stone-Gross, Brett; Kruegel, Christopher; Vigna, Giovanni (2011): "BOTMAGNIFIER: Locating Spambots on the Internet", in: Proceedings of the USENIX Security Symposium. (Link)] Historical. Ten of the fifteen email-spam papers here are from these three years, the affiliate programmes they traced are gone, and nothing has replaced the method. Read [5Pitsillidis, Andreas; Kanich, Chris; Voelker, Geoffrey M.; Levchenko, Kirill; Savage, Stefan (2012): "Taster's choice: a comparative analysis of spam feeds", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] for the feed-bias lesson, which has not aged
2022– Black-box audit of the filter, with seeded accounts as the unit [15Iqbal, Hassan; Khan, Usman Mahmood; Khan, Hassan Ali; Shahzad, Muhammad (2022): "Left or Right: A Peek into the Political Biases in Email Spam Filtering Algorithms During US Election 2020", in: Proceedings of the ACM Web Conference. (DOI)] — Gmail marked 67.6% of right-affiliated and 8.2% of left-affiliated campaign mail as spam; [16Rao, Sumanth; Liu, Enze; Ho, Grant; Voelker, Geoffrey M.; Savage, Stefan (2024): "Unfiltered: Measuring Cloud-based Email Filtering Bypasses", in: Proceedings of the ACM Web Conference. (DOI)] — 80% of 1,577 evaluated configurations allowed mail to bypass the cloud filter entirely Current, and the live email-spam method. The object moved from the spammer to the filter
2025– Manipulating the label source: treat the blocklist as an attack surface [30Li, Ruixuan; Lu, Chaoyi; Liu, Baojun; Zhang, Yunyi; Hong, Geng; Duan, Haixin; Lin, Yanzhong; Pan, Qingfeng; Yang, Min; Shao, Jun (2025): "HADES Attack: Understanding and Evaluating Manipulation Risks of Email Blocklists", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] — 39,201 (76.88%) of surveyed outgoing mail servers had been listed on at least one blocklist Current. If your ground truth is a DNSBL, this is the paper that says why that is a choice
2025– LLM classification of message content [31Hao, Wei; Tran, Van; Rideout, Vincent; Wang, Zixi; Dasbach-Prisk, AnMei; Afifi, M. H.; Yang, Junfeng; Katz-Bassett, Ethan; Ho, Grant; Cidon, Asaf (2025): "Do Spammers Dream of Electric Sheep? Characterizing the Prevalence of LLM-Generated Malicious Emails", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] (LLM-generated malicious mail, Llama-3.1-8B as the evaluator); [24Agarwal, Sharad; Papasavva, Antonis; Suarez-Tangil, Guillermo; Vasek, Marie (2025): "Fishing for Smishing: Understanding SMS Phishing Infrastructure and Strategies by Mining Public User Reports", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] (GPT-4o to extract sender IDs and URLs from 64,284 screenshots) Emerging, and this is where it emerged. Exactly 2 of the 61 on-page papers with a classification tuple carry an llm method, and both are 2025. Both use the model for extraction or scoring, not as the classifier of record
2015– Telephony honeypot [20Gupta, Payas; Srinivasan, Bharat; Balasubramaniyan, Vijay; Ahamad, Mustaque (2015): "Phoneypot: Data-driven Understanding of Telephony Threats", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] (1,297,517 calls from 252,621 sources), [21Prasad, Sathvik; Bouma-Sims, Elijah; Mylappan, Athishay Kiran; Reaves, Bradley (2020): "Who's Calling? Characterizing Robocalls through Audio and Metadata Analysis", in: Proceedings of the USENIX Security Symposium. (Link)] (1,481,201 calls), [32Prasad, Sathvik; Dunlap, Trevor; Ross, Alexander; Reaves, Bradley (2023): "Diving into Robocall Content with SnorCall", in: Proceedings of the USENIX Security Symposium. (Link)] (26,791 campaigns from 232,723 calls, weak supervision over transcripts) Current, and the best-instrumented sub-field on this page. The 2023 paper's use of Snorkel-style weak supervision instead of hand labels is the transferable idea
2025– Multiple honeypot vantage points, compared [8Prasad, Sathvik; Nahapetyan, Aleksandr; Reaves, Bradley (2025): "Characterizing Robocalls with Multiple Vantage Points", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] — WavLM audio clustering assigned 76.55% of benchmark calls to campaigns; 362 campaigns used ringless-voicemail injection Current, and the state of the art for this slice. One honeypot is one vantage point, and the paper measures how much that matters
2013– Carrier- or aggregator-side traffic [33Jiang, Nan; Jin, Yu; Skudlark, Ann; Zhang, Zhi-Li (2013): "Greystar: Fast and Accurate Detection of SMS Spam Numbers in Large Cellular Networks Using Gray Phone Space", in: Proceedings of the USENIX Security Symposium. (Link)] (over 34K spam numbers in five months, median detection 1.2 h), [14Huh, Jun Ho; Shin, Hyejin; Ahn, Sunwoo; Yi, Hayoon; Cho, Joonho; Kim, Taewoo; Lim, Minchae; Choi, Nuel (2025): "Preventing Artificially Inflated SMS Attacks through Large-Scale Traffic Inspection", in: Proceedings of the USENIX Security Symposium. (Link)] (SMS pumping at an aggregator) Current where you have the partner, and unavailable otherwise
2012– Crowdsourced complaint mining [22Murynets, Ilona; Jover, Roger Piqueras (2012): "Crime scene investigation: SMS spam data analysis", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], [34Li, Zhenhua; Wang, Weiwei; Wilson, Christo; Chen, Jian; Qian, Chen; Jung, Taeho; Zhang, Lan; Liu, Kebin; Li, Xiangyang; Liu, Yunhao (2017): "FBS-Radar: Uncovering Fake Base Stations at Scale in the Wild", in: Proceedings of the Network and Distributed System Security Symposium. (Link)], [23Tang, Siyuan; Mi, Xianghang; Li, Ying; Wang, XiaoFeng; Chen, Kai (2022): "Clues in Tweets: Twitter-Guided Discovery and Analysis of SMS Spam", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] (21,918 unique messages in 75 languages), [24Agarwal, Sharad; Papasavva, Antonis; Suarez-Tangil, Guillermo; Vasek, Marie (2025): "Fishing for Smishing: Understanding SMS Phishing Infrastructure and Strategies by Mining Public User Reports", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] Current and cheap, and now the standard way to get non-English message data

What the corpus cannot tell you

  • Nothing here measures what a mailbox provider's own defences do to a tracking measurement. Gmail's image proxy has been on since December 20133) and Apple's Mail Privacy Protection since iOS 15 in 2021; no paper on this page reports its results per client, or holds the client fixed and varies the provider. This is the single largest methodological gap on the page and it is squarely a measurement question.
  • No paper here measures unsubscription. See below; this is not a corpus-coverage caveat, it is a zero.
  • The lead-marketing pipeline is measured in halves. [35Venkatadri, Giridhari; Sapiezynski, Piotr; Redmiles, Elissa M.; Mislove, Alan; Goga, Oana; Mazurek, Michelle L.; Gummadi, Krishna P. (2019): "Auditing Offline Data Brokers via Facebook's Advertising Platform", in: Proceedings of the ACM Web Conference. (DOI)] audits offline brokers through an ad platform; [36Kempen, Elina van; Bagayatkar, Isita; Frolikov, Pavel; Georgiou, Chloe; Tsudik, Gene (2026): "Consumer Beware! Exploring Data Brokers' CCPA Compliance", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] sends CCPA requests to brokers; nothing in the extraction joins “who collected the address” to “what arrived in the mailbox”. The paper that does exist — Vekaria, Demir, Kollnig and Shafiq, Understanding Data Collection, Brokerage, and Spam in the Lead Marketing Ecosystem, IEEE S&P 2026, doi 10.1109/sp63933.2026.00162is in the bibliographic index and absent from the extraction, because IEEE S&P 2026 has 252 index records and only 58 with an abstract, selection screens on abstracts, and 194 of that venue-year were therefore never screened. A free preprint and the authors' code and data are available, so “read it” is not blocked by the paywall.4) It is not in any figure on this page.
  • Nobody has repeated the “your feed is not the population” comparison since 2012 [5Pitsillidis, Andreas; Kanich, Chris; Voelker, Geoffrey M.; Levchenko, Kirill; Savage, Stefan (2012): "Taster's choice: a comparative analysis of spam feeds", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], on any channel. It was true of spam feeds then. Whether it is true of the public smishing forums that the 2022–2025 papers all draw on is unknown, and three of them draw on overlapping sources.

What a Pixel Can Still Measure in 2026

The measurable signals, and what each provider and client leaves of them. Fill this table in for your own setup before you crawl; the rows are what a reviewer will ask about.

Signal the sender wants Direct render (your own client, images on) Gmail web Apple Mail with Protect Mail Activity
That the message was opened yes yes — Google says so explicitly5) no — content is prefetched on arrival, not on view
Recipient IP address, hence coarse location yes no — served from Google's proxy no — “hides your IP address”
User agent, hence client and OS yes no no
Cookies set or read by the image request yes no no
Which links were clicked yes yes — clicks leave the mail client yes
Address echoed in a URL (the leak [2Englehardt, Steven; Han, Jeffrey; Narayanan, Arvind (2018): "I never signed up for this! Privacy implications of email tracking", Proceedings on Privacy Enhancing Technologies 2018(1):109-126. (DOI)] measured) yes yes — proxying the fetch does not remove the token from the URL yes, if a link is clicked; the prefetch also fetches it. Apple's Link Tracking Protection does not change this — see below

A note on Apple's Link Tracking Protection, because it looks like it belongs in that last row and does not. Safari 17 / iOS 17 (2023) added Link Tracking Protection, which strips a curated list of known tracking parameters from the query string and fragment of a URL when the user navigates between sites. WebKit describes it as a Safari Private Browsing protection and specifies what it covers: “The specific parts of the URL covered are query parameters and the fragment”, against “known tracking” parameters, so that third-party scripts on the destination cannot read them.6) That is a different mechanism from the one this table's last row is about: it removes named ad-network parameters such as gclid, not an arbitrary hashed recipient address embedded in a path or query, which is what [2Englehardt, Steven; Han, Jeffrey; Narayanan, Arvind (2018): "I never signed up for this! Privacy implications of email tracking", Proceedings on Privacy Enhancing Technologies 2018(1):109-126. (DOI)] and [3Hu, Hang; Peng, Peng; Wang, Gang (2019): "Characterizing Pixel Tracking through the Lens of Disposable Email Services", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] measured. If your identifier happens to be a known ad-network parameter it may be stripped; if it is your own hash it will not be. Test rather than assume, in the exact client you are reporting on.

Two consequences for a study design. First, “tracker present in the message” and “tracker learned something” have come apart, and a prevalence figure has to say which it is. Second, the client is now a treatment: render the same corpus in two or three clients and report per client. No paper on this page does this, so if you do it you are producing the baseline, not comparing against one. The same logic as Browser protection, one layer up the stack.

The Address as the Identifier

The other half of the topic, and the half that is growing. An email address or phone number is a cross-context identifier that survives everything third-party cookies do not: it is stable, the user hands it over voluntarily, and it can be hashed and matched into an advertising platform.

  • How the address escapes the page. Starov et al. [7Starov, Oleksii; Gill, Phillipa; Nikiforakis, Nick (2016): "Are You Sure You Want to Contact Us? Quantifying the Leakage of PII via Website Contact Forms", in: Proceedings on Privacy Enhancing Technologies. (DOI)] measured contact forms as the point where a pseudonymous visitor becomes identified. Six years later Senol et al. [6Senol, Asuman; Acar, Gunes; Humbert, Mathias; Zuiderveen Borgesius, Frederik (2022): "Leaky Forms: A Study of Email and Password Exfiltration Before Form Submission", in: 31st USENIX Security Symposium (USENIX Security 22). (Link)] measured leakage before submit: the address was exfiltrated on 1,844 EU and 2,950 US websites, and 41 of the receiving domains were on none of four blocklists. Acar et al. [37Acar, Gunes; Englehardt, Steven; Narayanan, Arvind (2020): "No boundaries: data exfiltration by third parties embedded on web pages", Proceedings on Privacy Enhancing Technologies 2020(4):220-238. (DOI)] found the same behaviour in session-replay and login-manager scripts, and Kieserman et al. [38Kieserman, Julia B.; Andreou, Athanasios; Geeng, Chris; Lauinger, Tobias; McCoy, Damon (2025): "Tracker Installations Are Not Created Equal: Understanding Tracker Configuration of Form Data Collection", in: Proceedings on Privacy Enhancing Technologies, pp. 679-695. (DOI)] showed that whether it happens depends on how a given tracker installation is configured, not on which tracker it is.
  • What the address is worth to an ad platform. Venkatadri et al. measured this twice: [39Venkatadri, Giridhari; Andreou, Athanasios; Liu, Yabing; Mislove, Alan; Gummadi, Krishna P.; Loiseau, Patrick; Goga, Oana (2018): "Privacy Risks with Facebook's PII-Based Targeting: Auditing a Data Broker's Advertising Interface", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] showed PII-based targeting could be inverted — 18 of 20 volunteers who visited a page were correctly identified from a tracking-pixel audience, and phone numbers were inferred from addresses — and [40Venkatadri, Giridhari; Lucherini, Elena; Sapiezynski, Piotr; Mislove, Alan (2019): "Investigating sources of PII used in Facebook’s targeted advertising", in: Proceedings on Privacy Enhancing Technologies. (DOI)] tracked where the platform's PII comes from, including that a phone number added purely for two-factor authentication became advertising-targetable after 22 days. [35Venkatadri, Giridhari; Sapiezynski, Piotr; Redmiles, Elissa M.; Mislove, Alan; Goga, Oana; Mazurek, Michelle L.; Gummadi, Krishna P. (2019): "Auditing Offline Data Brokers via Facebook's Advertising Platform", in: Proceedings of the ACM Web Conference. (DOI)] then measured the offline-broker side: over 90% of targetable US Facebook identities carried at least one broker attribute, and coverage was strongly unequal — 64.3% of US counties reached 80% coverage against 13.8% of the poorest decile.
  • The address is not one identifier. Wu et al. [41Wu, Mengying; Hong, Geng; Chen, Jiatao; Liu, Baojun; Liu, Mingxuan; Yang, Min (2026): "One Email, Many Faces: A Deep Dive into Identity Confusion in Email Aliases", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] measured alias handling — dots, plus-addressing, sub-addressing — and found providers disagree about whether two strings are the same account. If your method matches on the address, this is your false-positive and false-negative source, and it is provider-dependent.
  • The phone number is worse, because it is enumerable. Hagen et al. [42Hagen, Christoph; Weinert, Christian; Sendner, Christoph; Dmitrienko, Alexandra; Schneider, Thomas (2021): "All the Numbers are US: Large-scale Abuse of Contact Discovery in Mobile Messengers", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] abused contact discovery to enumerate messenger users: 10% of US mobile numbers for WhatsApp and 100% for Signal, finding 5.0 million users among 46.2 million checked numbers. Gegenhuber et al. [43Gegenhuber, Gabriel K.; Frenzel, Philipp E.; Günther, Maximilian; Ullrich, Johanna; Judmayer, Aljosha (2026): "Hey there! You are using WhatsApp: Enumerating Three Billion Accounts for Security and Privacy", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] carried it to 3.5 billion accounts. Kang and Lee [44Kang, Junkyu; Lee, Soyoung; Kwon, Yonghwi; Son, Sooel (2026): "Connecting the Dots: An Investigative Study on Linking Private User Data Across Messaging Apps", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] joined contact discovery, nearby-user search and SSO linking across apps, and Niksirat et al. [45Niksirat, Kavous Salehzadeh; Velykoivanenko, Lev; Mätzler, Samuel; Mulders, Stephan; Tamò-Larrieux, Aurelia; Boldi, Marc-Olivier; Humbert, Mathias; Huguenin, Kévin (2025): "Addressing the Address Books' (Interdependent) Privacy Issues", in: Proceedings of the USENIX Security Symposium. (Link)] measured what an uploaded address book discloses about people who never installed the app. If your study uploads contacts or enumerates numbers, read all four before writing the ethics section — and note that these papers are also the reason the same trick supplies lead-generation lists.

Unsubscription: A Measured Gap

This is the slice the brief for this page asked about, and the honest answer is that these seven venues have not studied it. The probes below ran over full text for all 5,859 papers; the script is on the provenance page.

Probe over paper.cols.txt, all 5,859 papers Papers
List-Unsubscribe (the header field, RFC 2369) 1
one-click unsubscribe / List-Unsubscribe-Post / RFC 8058 0
RFC 2369 cited by number 0
CAN-SPAM 2
the word unsubscrib* anywhere 56

The one List-Unsubscribe hit names the header in a table of RFC 6376 signable header fields; it measures nothing about it.7) Both CAN-SPAM hits are motivational asides. Of the 56 unsubscrib* papers, ten are using the MQTT, pub/sub or SDN protocol verb, three mean the CCPA “do not sell” opt-out or an advertising opt-out cookie, four mean turning off a security notification, and the remaining 39 have a single passing mention in a paper about something else.

The nearest thing in the corpus is one sub-check inside one paper. Kubicek et al. [11Kubíček, Karel; Merane, Jakob; Cotrini, Carlos; Stremitzer, Alexander; Bechtold, Stefan; Basin, David (2022): "Checking Websites' GDPR Consent Compliance for Marketing Emails", Proceedings on Privacy Enhancing Technologies 2022(2). (DOI)] report that 16% of websites sending marketing mail provided neither an unsubscribe method nor a legal notice, alongside 17.3% of such websites having at least one potential consent violation at the registration form, 21.9% having one somewhere, 59% sending a double opt-in confirmation first, and 2.3% mailing the user's own password back in plaintext.

That is availability, not function. Nothing in these venues measures whether an unsubscribe link works, how long it takes to take effect, or whether the one-click mechanism that Google and Yahoo have required of bulk senders since 1 February 2024 is actually implemented. The requirement is specific and testable: bulk senders — more than 5,000 messages a day to Gmail addresses — must supply List-Unsubscribe and List-Unsubscribe-Post, citing RFC 2369 and RFC 8058, and must “process and honor unsubscribe requests within 48 hours”.8) Google also publishes the spam-rate thresholds it enforces: keep reported spam below 0.10%, and never reach 0.30%. And the consequence of non-compliance changed in November 2025, which changes what a measurement would observe: Google's own FAQ says “Starting November 2025, Gmail is ramping up its enforcement on non-compliant traffic. Messages that fail to meet the email sender requirements will experience disruptions, including temporary and permanent rejections.”9) Before that, non-compliance meant the spam folder; now it can mean an SMTP rejection, so a compliance measurement should be looking for bounces as well as for foldering.

Microsoft is the asymmetry worth knowing about. Since 5 May 2025, Outlook.com, Hotmail and Live require domains sending 5,000 or more messages a day to their consumer services to pass SPF, DKIM and DMARC, with non-compliant mail routed to Junk and rejection announced for later. It does not require one-click unsubscribe or cite RFC 8058 at all. So the three largest consumer mailbox providers do not agree on this mechanism: Google and Yahoo mandate RFC 8058, Microsoft mandates authentication only. Anyone measuring one-click-unsubscribe deployment has to say which providers' senders are in the population, because two of the three create the incentive and the third does not.10)

  • Nobody has measured one-click unsubscribe compliance in these seven venues. Header presence, mechanism correctness, whether the 48-hour deadline is met, and whether unsubscribing changes anything are four separate measurable claims and there are zero papers on any of them.
  • Nobody has measured what happens after you unsubscribe. The address you used is still an identifier; whether it stops being mailed, gets sold on, or reappears under a different sender is unstudied here.
  • This gap is being worked on. Do not fill this section from a preprint or from work in progress; it is a real hole in the published record and should be shown as one until something is published. If you are planning this study, the useful thing this page can tell you is that you have no baseline and no comparable prior denominator.

Spam as the Object of Study

Fifteen papers, and the shape matters more than the count: ten are from 2010–2012, then a single 2014 paper, then nothing until 2022. The old cluster traced the spam value chain end to end — purchases, affiliate programmes, payment processors [26Levchenko, Kirill; Pitsillidis, Andreas; Chachra, Neha; Enright, Brandon; Félegyházi, Márk; Grier, Chris; Halvorson, Tristan; Kanich, Chris; Kreibich, Christian; Liu, He; McCoy, Damon; Weaver, Nicholas; Paxson, Vern; Voelker, Geoffrey M.; Savage, Stefan (2011): "Click Trajectories: End-to-End Analysis of the Spam Value Chain", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] [27McCoy, Damon; Pitsillidis, Andreas; Jordan, Grant; Weaver, Nicholas; Kreibich, Christian; Krebs, Brian; Voelker, Geoffrey M.; Savage, Stefan; Levchenko, Kirill (2012): "PharmaLeaks: Understanding the Business of Online Pharmaceutical Affiliate Programs", in: Proceedings of the USENIX Security Symposium. (Link)] [28Kanich, Chris; Weaver, Nicholas; McCoy, Damon; Halvorson, Tristan; Kreibich, Christian; Levchenko, Kirill; Paxson, Vern; Voelker, Geoffrey M.; Savage, Stefan (2011): "Show Me the Money: Characterizing Spam-advertised Revenue", in: Proceedings of the USENIX Security Symposium. (Link)], botnet populations inferred from mail transactions [29Stringhini, Gianluca; Holz, Thorsten; Stone-Gross, Brett; Kruegel, Christopher; Vigna, Giovanni (2011): "BOTMAGNIFIER: Locating Spambots on the Internet", in: Proceedings of the USENIX Security Symposium. (Link)], SMTP-dialect fingerprinting of spamming software [46Stringhini, Gianluca; Egele, Manuel; Zarras, Apostolis; Holz, Thorsten; Kruegel, Christopher; Vigna, Giovanni (2012): "B@bel: Leveraging Email Delivery for Spam Mitigation", in: Proceedings of the USENIX Security Symposium. (Link)], the delivery technique itself [47Qian, Zhiyun; Mao, Zhuoqing Morley; Xie, Yinglian; Yu, Fang (2010): "Investigation of Triangular Spamming: A Stealthy and Efficient Spamming Technique", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)], real-time URL spam filtering trained on both mail and tweet spam [48Thomas, Kurt; Grier, Chris; Ma, Justin; Paxson, Vern; Song, Dawn (2011): "Design and Evaluation of a Real-Time URL Spam Filtering Service", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)], and a randomised measurement of what inbox spam does to a mail provider's users [13Dasgupta, Anirban; Punera, Kunal; Rao, Justin M.; Wang, Xuanhui (2012): "Impact of Spam Exposure on User Engagement", in: Proceedings of the USENIX Security Symposium. (Link)]. It is excellent work about an ecosystem that no longer exists in that form, and no method in it is one you should start from.

Two things from that era have not aged. Pitsillidis et al. [5Pitsillidis, Andreas; Kanich, Chris; Voelker, Geoffrey M.; Levchenko, Kirill; Savage, Stefan (2012): "Taster's choice: a comparative analysis of spam feeds", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] compared ten spam feeds and found 60% of live domains exclusive to a single feed, which is still the reason to describe your message source rather than name it. And Park et al. [49Park, Youngsam; Jones, Jackie; McCoy, Damon; Shi, Elaine; Jakobsson, Markus (2014): "Scambaiter: Understanding Targeted Nigerian Scams on Craigslist", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] showed how to solicit a scam population on purpose — bait adverts, automated reply engine, 13,215 first responses, 9.6 scam attempts per advert — which is the design the honeypot papers on the SMS side later rediscovered.

What is current is the audit of the filter, not of the spammer. Iqbal et al. [15Iqbal, Hassan; Khan, Usman Mahmood; Khan, Hassan Ali; Shahzad, Muhammad (2022): "Left or Right: A Peek into the Political Biases in Email Spam Filtering Algorithms During US Election 2020", in: Proceedings of the ACM Web Conference. (DOI)] ran seeded accounts across three providers through the 2020 US election and measured a political skew in spam classification — Gmail marked 67.6% of right-affiliated and 8.2% of left-affiliated campaign mail as spam, Outlook the reverse direction at 75.4% and 95.8% — and showed the effect survived propensity-score matching on content and metadata, and that five user interactions moved it. Rao et al. [16Rao, Sumanth; Liu, Enze; Ho, Grant; Voelker, Geoffrey M.; Savage, Stefan (2024): "Unfiltered: Measuring Cloud-based Email Filtering Bypasses", in: Proceedings of the ACM Web Conference. (DOI)] measured bypasses of cloud filtering: of 1,577 evaluated configurations, 80% let mail reach the recipient without passing the filter at all. Li et al. [30Li, Ruixuan; Lu, Chaoyi; Liu, Baojun; Zhang, Yunyi; Hong, Geng; Duan, Haixin; Lin, Yanzhong; Pan, Qingfeng; Yang, Min; Shao, Jun (2025): "HADES Attack: Understanding and Evaluating Manipulation Risks of Email Blocklists", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] treated the blocklist as an attack surface. Zhang et al. [50Zhang, Jiahe; Chen, Jianjun; Wang, Qi; Zhang, Hangyu; Wang, Chuhan; Zhuge, Jianwei; Duan, Haixin (2024): "Inbox Invasion: Exploiting MIME Ambiguities to Evade Email Attachment Detectors", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] measured MIME-parsing disagreement between filters and clients: 180 of 237 samples bypassed detection in at least one of 102 vulnerable product–client combinations (75.95%). And Hao et al. [31Hao, Wei; Tran, Van; Rideout, Vincent; Wang, Zixi; Dasbach-Prisk, AnMei; Afifi, M. H.; Yang, Junfeng; Katz-Bassett, Ethan; Ho, Grant; Cidon, Asaf (2025): "Do Spammers Dream of Electric Sheep? Characterizing the Prevalence of LLM-Generated Malicious Emails", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] is the first paper here to measure the prevalence of LLM-generated malicious mail.

If you are designing a spam measurement now: your object is a classifier you do not control, your instrument is seeded accounts, and your hardest problem is account reputation, not message collection.

Beyond Email: SMS, Calls and Messaging Apps

Twenty-three papers, six of them in 2025–2026, and this is the live slice of the topic. It is on a privacy: page because it shares the instrument — an address or number you control, and messages that arrive at it unbidden — not because it shares a platform. Almost none of it involves a browser.

  • Honeypots are the backbone. Gupta et al. [20Gupta, Payas; Srinivasan, Bharat; Balasubramaniyan, Vijay; Ahamad, Mustaque (2015): "Phoneypot: Data-driven Understanding of Telephony Threats", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] established the design and its confounder in the same paper: older number blocks receive significantly more calls (t-test, p = 0.005), so a honeypot's numbers have history. Prasad et al. [21Prasad, Sathvik; Bouma-Sims, Elijah; Mylappan, Athishay Kiran; Reaves, Bradley (2020): "Who's Calling? Characterizing Robocalls through Audio and Metadata Analysis", in: Proceedings of the USENIX Security Symposium. (Link)] scaled it to 1,481,201 calls with audio and metadata; [32Prasad, Sathvik; Dunlap, Trevor; Ross, Alexander; Reaves, Bradley (2023): "Diving into Robocall Content with SnorCall", in: Proceedings of the USENIX Security Symposium. (Link)] added weak supervision over transcripts to get 26,791 campaigns from 232,723 calls without hand-labelling everything; [8Prasad, Sathvik; Nahapetyan, Aleksandr; Reaves, Bradley (2025): "Characterizing Robocalls with Multiple Vantage Points", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] compared vantage points and found that a 90% audio-similarity threshold identifies campaigns common to several honeypots, which is the first evidence on this page about how much one vantage point misses.
  • Crowdsourced reports are how you get scale and languages. Murynets and Piqueras Jover [22Murynets, Ilona; Jover, Roger Piqueras (2012): "Crime scene investigation: SMS spam data analysis", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] used a carrier reporting service; Li et al. [34Li, Zhenhua; Wang, Weiwei; Wilson, Christo; Chen, Jian; Qian, Chen; Jung, Taeho; Zhang, Lan; Liu, Kebin; Li, Xiangyang; Liu, Yunhao (2017): "FBS-Radar: Uncovering Fake Base Stations at Scale in the Wild", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] crowdsourced fake-base-station detection; Tang et al. [23Tang, Siyuan; Mi, Xianghang; Li, Ying; Wang, XiaoFeng; Chen, Kai (2022): "Clues in Tweets: Twitter-Guided Discovery and Analysis of SMS Spam", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] mined tweeted screenshots for 21,918 unique SMS messages in 75 languages, recovering the text with OCR at 90% exact-match accuracy on a 1,000-image sample; Agarwal et al. [51Agarwal, Sharad; Harvey, Emma; Vasek, Marie (2024): "Poster: A Comprehensive Categorization of SMS Scams", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] built a sector taxonomy from firewall-side data — delivery impersonation was the largest category at 830.2k recipients — and [24Agarwal, Sharad; Papasavva, Antonis; Suarez-Tangil, Guillermo; Vasek, Marie (2025): "Fishing for Smishing: Understanding SMS Phishing Infrastructure and Strategies by Mining Public User Reports", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] pushed the forum approach to five sources — 220,585 posts and 64,284 image attachments, from which an LLM extracted 27,718 unique messages, 19,314 sender IDs and 20,060 URLs, in that paper across 66 languages. The screenshot is the data format in this slice, so budget for OCR or a vision model.
  • Carrier and aggregator vantage points see what nothing else does. Jiang et al. [33Jiang, Nan; Jin, Yu; Skudlark, Ann; Zhang, Zhi-Li (2013): "Greystar: Fast and Accurate Detection of SMS Spam Numbers in Large Cellular Networks Using Gray Phone Space", in: Proceedings of the USENIX Security Symposium. (Link)] found over 34K spam numbers in five months from grey-number traffic with a 1.2-hour median detection time; Huh et al. [14Huh, Jun Ho; Shin, Hyejin; Ahn, Sunwoo; Yi, Hayoon; Cho, Joonho; Kim, Taewoo; Lim, Minchae; Choi, Nuel (2025): "Preventing Artificially Inflated SMS Attacks through Large-Scale Traffic Inspection", in: Proceedings of the USENIX Security Symposium. (Link)] measured SMS pumping at an aggregator. Both are unreproducible without the partner, and both say so.
  • The delivery path is itself measurable and leaky. Bitsikas et al. [52Bitsikas, Evangelos; Schnitzler, Theodor; Pöpper, Christina; Ranganathan, Aanjhan (2023): "Freaky Leaky SMS: Extracting User Locations by Analyzing SMS Timings", in: Proceedings of the USENIX Security Symposium. (Link)] inferred receiver location from SMS delivery-report timing at up to 96% accuracy across countries; Schnitzler et al. [53Schnitzler, Theodor; Kohls, Katharina; Bitsikas, Evangelos; Pöpper, Christina (2023): "Hope of Delivery: Extracting User Locations From Mobile Instant Messengers", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] did the messenger equivalent with delivery receipts, over 80% for three locations inside one city. Reaves et al. [54Reaves, Bradley; Scaife, Nolen; Tian, Dave; Blue, Logan; Traynor, Patrick; Butler, Kevin R. B. (2016): "Sending Out an SMS: Characterizing the Security of the SMS Ecosystem with Public Gateways", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] used public SMS gateways as a vantage point onto the whole ecosystem — 386,327 messages over 14 months, including 522 that contained email addresses. Mulliner et al. [55Mulliner, Collin; Golde, Nico; Seifert, Jean-Pierre (2011): "SMS of Death: From Analyzing to Attacking Mobile Phones on a Large Scale", in: Proceedings of the USENIX Security Symposium. (Link)] delivered malformed SMS to real handsets and found faults in feature phones from six manufacturers, two of which could not be restored afterwards; Tu et al. [56Tu, Guan-Hua; Li, Chi-Yu; Peng, Chunyi; Li, Yuanjie; Lu, Songwu (2016): "New Security Threats Caused by IMS-based SMS Service in 4G LTE Networks", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] and Lei et al. [57Lei, Zeyu; Nan, Yuhong; Fratantonio, Yanick; Bianchi, Antonio (2021): "On the Insecurity of SMS One-Time Password Messages against Local Attackers in Modern Mobile Devices", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] measured the path and the handset; Wang et al. [58Wang, Qi; Chen, Jianjun; Yang, Jingcheng; Zhang, Jiahe; Yang, Yaru; Duan, Haixin (2026): "SIPConfusion: Exploiting SIP Semantic Ambiguities for Caller ID and SMS Spoofing", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] found caller-ID and SMS spoofing through SIP ambiguity in 47 of 54 server–user-agent combinations.
  • Injection at the radio layer is a Chinese-market phenomenon with real numbers. Zhang et al. [59Zhang, Yiming; Liu, Baojun; Lu, Chaoyi; Li, Zhou; Duan, Haixin; Hao, Shuang; Liu, Mingxuan; Liu, Ying; Wang, Dong; Li, Qiang (2020): "Lies in the Air: Characterizing Fake-base-station Spam Ecosystem in China", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] characterised fake-base-station SMS spam: 279,017 message logs over 97 days, 7,884 campaigns, and illegal businesses accounting for over 75% of the messages.
  • Messaging apps. Edu et al. [60Edu, Jide S.; Mulligan, Cliona; Pierazzi, Fabio; Polakis, Jason; Suarez-Tangil, Guillermo; Such, Jose M. (2022): "Exploring the security and privacy risks of chatbots in messaging services", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] measured third-party chatbots in messaging channels — 8,521 of 15,525 valid chatbots (54.86%) requested administrator permissions and 95.67% had no privacy policy. Kirchner et al. [4Kirchner, Robin; Koch, Simon; Kamangar, Noah; Klein, David; Johns, Martin (2024): "A Black-Box Privacy Analysis of Messaging Service Providers' Chat Message Processing", in: Proceedings on Privacy Enhancing Technologies. (DOI)] is the honey-message study. Together they are the reason “email tracking” is the wrong frame for the next study in this area.
  • User-facing work exists and is small. Sherman et al. [61Sherman, Imani N.; Bowers, Jasmine D.; McNamara, Jr., Keith; Gilbert, Juan E.; Ruiz, Jaime; Traynor, Patrick (2020): "Are You Going to Answer That? Measuring User Responses to Anti-Robocall Application Indicators", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] measured responses to anti-robocall indicators (answered calls fell 43% with an “Avail-Spam” warning); Sharevski and Zettlemoyer [62Sharevski, Filipo; Loop, Jennifer Vander; Evans, Bill; Ponticello, Alexander (2025): "(Blind) Users Really Do Heed Aural Telephone Scam Warnings", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] did aural scam warnings with blind participants; Agarwal et al. [63Agarwal, Sharad; Harvey, Emma; Mariconti, Enrico; Suarez-Tangil, Guillermo; Vasek, Marie (2025): "'Hey mum, I dropped my phone down the toilet': Investigating Hi Mum and Dad SMS Scams in the United Kingdom", in: Proceedings of the USENIX Security Symposium. (Link)] characterised UK “Hi Mum” impersonation scams — 582 mule accounts, over £577k requested in 13 weeks, and a 14-day median lifetime for the originating sender IDs.

The Mailbox as an Instrument

Twelve papers on this page do not measure messages at all: they needed an address, and the mail is plumbing or a channel. This is where most students actually meet the topic, and the methods are worth stealing.

  • Honey mailboxes as a compromise detector. DeBlasio et al. [19DeBlasio, Joe; Savage, Stefan; Voelker, Geoffrey M.; Snoeren, Alex C. (2017): "Tripwire: inferring internet site compromise", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] registered a unique account, with a unique address and password, at each of ~2,302 sites, and treated a login to the mailbox as evidence that the site had leaked — 19 compromises found. Onaolapo et al. [64Onaolapo, Jeremiah; Mariconti, Enrico; Stringhini, Gianluca (2016): "What Happens After You Are Pwnd: Understanding the Use of Leaked Webmail Credentials in the Wild", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] went the other way and leaked instrumented webmail credentials on purpose: 326 unique accesses, 147 messages opened, 845 sent, 132 of the accesses via Tor.
  • Honeytokens against a third party. Farooqi et al. [18Farooqi, Shehroze; Musa, Maaz; Shafiq, Zubair; Zaffar, Fareed (2020): "CanaryTrap: Detecting Data Misuse by Third-Party Apps on Online Social Networks", in: Proceedings on Privacy Enhancing Technologies. (DOI)] gave a distinct honeytoken address to each social-platform app and watched: 422 unrecognised emails tied to 20 apps, 76 outright malicious. This is the cheapest way to attribute misuse to a named recipient, and its limitation is structural — silence is not evidence of no misuse.
  • Mail as the notification channel, where the response rate is the result. Stock et al. [65Stock, Ben; Pellegrino, Giancarlo; Rossow, Christian; Johns, Martin; Backes, Michael (2016): "Hey, You Have a Problem: On the Feasibility of Large-Scale Web Vulnerability Notification", in: Proceedings of the USENIX Security Symposium. (Link)] measured the brutal part first: only 2,064 of 35,832 transmitted vulnerability reports (5.8%) were received at all. Maass et al. [66Maass, Max; Stöver, Alina; Pridöhl, Henning; Bretthauer, Sebastian; Herrmann, Dominik; Hollick, Matthias; Spiecker, Indra (2021): "Effective Notification Campaigns on the Web: A Matter of Trust, Framing, and Support", in: Proceedings of the USENIX Security Symposium. (Link)] varied framing and got 56.6% remediation against 9.2% in control; Utz et al. [67Utz, Christine; Michels, Matthias; Degeling, Martin; Marnau, Ninja; Stock, Ben (2023): "Comparing Large-Scale Privacy and Security Notifications", in: Proceedings on Privacy Enhancing Technologies. (DOI)] compared campaigns; Sasaki et al. [68Sasaki, Takayuki; Inazawa, Tomoya; Yamaguchi, Youhei; Parkin, Simon; Eeten, Michel van; Yoshioka, Katsunari; Matsumoto, Tsutomu (2025): "Am I Infected? Lessons from Operating a Large-Scale IoT Security Diagnostic Service", in: Proceedings of the USENIX Security Symposium. (Link)] ran a diagnostic service. If notification is your method, Notifying websites is the page you want, and this section is only about the delivery problem.
  • Mail as the legal channel. Subject access and erasure requests are sent and answered by mail: Di Martino et al. [69Martino, Mariano Di; Meers, Isaac; Quax, Peter; Andries, Ken; Lamotte, Wim (2022): "Revisiting Identification Issues in GDPR ‘Right Of Access’ Policies: A Technical and Longitudinal Analysis", in: Proceedings on Privacy Enhancing Technologies. (DOI)], Rupp et al. [70Rupp, Eduard; Syrmoudis, Emmanuel; Grossklags, Jens (2022): "Leave No Data Behind – Empirical Insights into Data Erasure from Online Services", in: Proceedings on Privacy Enhancing Technologies. (DOI)] (24 of 90 services, 27%, failed to erase completely), and van Kempen et al. [36Kempen, Elina van; Bagayatkar, Isita; Frolikov, Pavel; Georgiou, Chloe; Tsudik, Gene (2026): "Consumer Beware! Exploring Data Brokers' CCPA Compliance", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] on CCPA — 57% of data brokers responded at all, and email was the identity-verification datum requested by 444 of 454 brokers.
  • Mail as registration plumbing. Drakonakis et al. [71Drakonakis, Kostas; Ioannidis, Sotiris; Polakis, Jason (2020): "The Cookie Hunter: Automated Black-box Auditing for Web Authentication and Authorization Flaws", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] and Oh et al. [72Oh, ChangSeok; Kanich, Chris; McCoy, Damon; Pearce, Paul (2022): "Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] both need a mailbox per account. Registration is the page for the mechanics of getting one; this page is about what the mailbox then lets you measure.

Ethics

Papers on this topic engage with law about three times as often as the corpus average: 14 of the 70 assess compliance with a law (20.0%), against 402 of 5,859 (6.9%) corpus-wide — GDPR in 9 of the 14, then the ePrivacy Directive in 2, and one each of the TRACED Act, the CCPA, the Digital Markets Act, the Gramm-Leach-Bliley Act, HIPAA, Chinese law, German competition law, the German UWG and TMG, US federal law, and “the law of the state where the honeypot is operated”. That is a fact about legal engagement, not proof that the ethics here are handled well — the reporting figures at the bottom of this list say otherwise. Read the relevant papers' ethics sections rather than this list, and see Ethics for the general apparatus. What is specific to this topic:

  • You are creating accounts under identities that are not yours in order to subscribe. Say so, say how many, and say what you did with the accounts afterwards.
  • You may be reading other people's mail. The disposable-provider design [3Hu, Hang; Peng, Peng; Wang, Gang (2019): "Characterizing Pixel Tracking through the Lens of Disposable Email Services", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] is the clearest case, and that paper found 1,399 credit-card numbers, 926 SSNs and 89,329 account-management messages in the stream. Its ethics section is the model.
  • A honeypot receives messages meant for someone. A telephony honeypot's numbers are recycled real numbers [20Gupta, Payas; Srinivasan, Bharat; Balasubramaniyan, Vijay; Ahamad, Mustaque (2015): "Phoneypot: Data-driven Understanding of Telephony Threats", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]; a leaked-credential study invites attackers into an account [64Onaolapo, Jeremiah; Mariconti, Enrico; Stringhini, Gianluca (2016): "What Happens After You Are Pwnd: Understanding the Use of Leaked Webmail Credentials in the Wild", in: Proceedings of the ACM Internet Measurement Conference. (DOI)].
  • Enumerating an address space touches everyone in it. [42Hagen, Christoph; Weinert, Christian; Sendner, Christoph; Dmitrienko, Alexandra; Schneider, Thomas (2021): "All the Numbers are US: Large-scale Abuse of Contact Discovery in Mobile Messengers", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] and [43Gegenhuber, Gabriel K.; Frenzel, Philipp E.; Günther, Maximilian; Ullrich, Johanna; Judmayer, Aljosha (2026): "Hey there! You are using WhatsApp: Enumerating Three Billion Accounts for Security and Privacy", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] both assess this against the GDPR, and [43Gegenhuber, Gabriel K.; Frenzel, Philipp E.; Günther, Maximilian; Ullrich, Johanna; Judmayer, Aljosha (2026): "Hey there! You are using WhatsApp: Enumerating Three Billion Accounts for Security and Privacy", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] additionally against the DMA. If you enumerate, you will be asked about rate limits, data minimisation and disclosure.
  • Scam-baiting means conversing with criminals and possibly wasting a third party's time [49Park, Youngsam; Jones, Jackie; McCoy, Damon; Shi, Elaine; Jakobsson, Markus (2014): "Scambaiter: Understanding Targeted Nigerian Scams on Craigslist", in: Proceedings of the Network and Distributed System Security Symposium. (Link)].
  • Reporting rates are low. ethics.reviewOutcome carries a stated value for 36 of the 66 on-page papers with an ethics object (54.5%), against 1,728 of 4,472 (38.6%) for empirical papers corpus-wide on the same denominator rule. Better than the field, and still only half. 11)

Use in Publications

Every figure in this section is a count of papers in the 7-venue, 2010–2026, 5,859-paper extraction, with its denominator in the same sentence. The script is report_email_tracking.mjs; the population rule and the hand map are msg_fold.mjs; the query log, the unedited output and the full residue are on email_tracking.

The population is a hand map, not a query

No field in the extraction means “the measured object is a message”. classification.target does have an email-message value, and it is noisy in a way an enum disguises: it fires on 54 papers, of which 12 are outside this page's candidate pool entirely and include The Matter of Heartbleed, The State of the SameSite and Traveling the Silk Road. So the population here is a published candidate pool plus an explicit verdict per paper.

The pool is a union of three signals — the title names a message channel; the full text is dense in message-channel vocabulary; the full text names a messaging-app address space (“contact discovery”, “address book”). It contains 354 papers, 183 of which carry a hand-written verdict, and 70 of which are on this page.

Verdict Papers Where it goes
Message tracking (T) 8 this page
The address as identifier (EID) 12 this page
Marketing mail and opt-out (U) 0 this page, as a gap
Email spam (SE) 15 this page
SMS / calls / messaging (SM) 23 this page
Mailbox as instrument (INST) 12 this page
Email transport and authentication (INFRA) 31 no page exists yet
Phishing (PHISH) 28 Phishing
Social-platform, review and SEO spam (SOC) 37 Platforms
Web-page pixel (WEBPIXEL) 7 Requests
Read and off-topic 181

Four rejected probes, with the count each produced, are listed on the provenance page — including /\bspam/i over full text (954 papers, which is a fact about these being security venues) and /\be-?mail/i (2,502, i.e. every paper that mentions contacting an author).

Where the papers are

Of the 70: USENIX Security 19, NDSS 13, PoPETs 11, IEEE S&P 9, IMC 8, CCS 7, TheWebConf 3. 29 ran an automated web crawl and 27 have a recorded crawl configuration; other-online-service is the platform for 54 of them and web for 33. Twelve recruited human participants, and manual-audit is the most common study type at 42 of 70.

Per year and slice. 2025 is thin at the edges and 2026 is provisional — treat the last two columns as a lower bound, not a trend:

Slice Total ≤2014 2015–2021 2022–2024 2025–2026 (prov.)
Message tracking 8 1 2 3 2
Address as identifier 12 0 6 1 5
Email spam 15 10 0 3 2
SMS / calls / messaging 23 3 8 6 6
Mailbox as instrument 12 0 6 4 2

Read the rows, not the totals. Email spam is a 2010–2012 literature with a different 2022–2025 aftershock; the address-as-identifier slice did not exist before 2016 and half of it is 2025–2026; SMS and calls have been continuous since 2011 and are the only slice with papers in every bucket.

How they classify

Of the 70, 61 have a classification tuple. Shares are of those 61 and do not sum to 100% because the field is multi-valued.

Method (enum) Papers Share of 61 2010–2019 (n=24) 2020–2026 (n=37)
heuristic-rules 36 59.0% 58.3% 59.5%
manual-labelling 23 37.7% 25.0% 45.9%
third-party-service 15 24.6% 25.0% 24.3%
supervised-ml 13 21.3% 12.5% 27.0%
curated-database 12 19.7% 29.2% 13.5%
regex-or-signature 12 19.7% 16.7% 21.6%
other 7 11.5% 12.5% 10.8%
unsupervised-ml 5 8.2% 8.3% 8.1%
blocklist 4 6.6% 8.3% 5.4%
dynamic-analysis 3 4.9% 4.2% 5.4%
graph-analysis 2 3.3% 0.0% 5.4%
static-analysis 2 3.3% 0.0% 5.4%
llm 2 3.3% 0.0% 5.4%

All thirteen values the enum takes on this population are listed; nothing is truncated.

Two things to take from this. Hand labelling went up, not down — 25.0% to 45.9% — which is what you would expect of a topic whose data is increasingly screenshots and transcripts in many languages. And llm is 2 papers, both 2025, so anyone claiming LLM classification is the current practice here is claiming it from two data points in the corpus's thinnest years.

Instruments, folded

tools[].name is free text and agrees with an independent extraction run on roughly 20% of exact strings, so it is folded into families tagged by the question the instrument answers. Denominator: the 70 on-page papers.

Instrument family Papers Share of 70 Distinct strings What it answers
browser-automation 18 25.7% 18 driving a browser or app
network-intelligence 16 22.9% 19 geolocating or attributing an address (ip_classification)
ml-model 16 22.9% 32 a classifier or clustering algorithm
telephony-rig 13 18.6% 21 placing, receiving or fingerprinting calls and SMS
browser 13 18.6% 11 the browser itself
maliciousness-oracle 13 18.6% 13 a maliciousness verdict from a vendor
study-apparatus 9 12.9% 10 recruiting participants or analysing their answers
mail-plumbing 7 10.0% 16 sending, receiving or parsing mail
language-tooling 7 10.0% 4 translating non-English text (multilingual_support)
spam-blocklist 6 8.6% 7 a spam or abuse blocklist as ground truth
mail-filter 6 8.6% 5 a spam filter run by the authors as a labeller
traffic-capture 6 8.6% 8 recording traffic (traffic_files)
mail-endpoint 5 7.1% 8 a mailbox, provider or mail-specific detector as the measurement point
number-parsing 5 7.1% 5 parsing and validating phone numbers
llm-or-weak-supervision 5 7.1% 11 an LLM or a weak-supervision labeller
crawler-framework 4 5.7% 4 a web-measurement framework (crawler)
tracker-blocklist 4 5.7% 7 a tracking blocklist (filter_lists)
audio-pipeline 4 5.7% 7 turning call audio into text or fingerprints
messenger-client 4 5.7% 6 driving a messaging app programmatically

All nineteen families the fold produces on this population are listed; nothing is truncated. Only four of them are mail- or message-specific — the rest is the ordinary apparatus of a web or mobile measurement, which is the point: there is no toolchain for this topic.

The fold is not cosmetic. Spamhaus appears under four spellings on this small a population (Spamhaus, Spamhaus blacklist, Spamhaus.org, Spamhaus passive DNS API), libphonenumber under four, and MaxMind under four. Counting exact strings would have put every one of them below the reporting threshold. 159 distinct strings matched no family and are printed in full on the provenance page rather than dropped; they are almost all one-off infrastructure (Faker, Fake Name Generator, 2Captcha, DeCaptcher — which between them tell you something about what signing up on 15,700 sites [2Englehardt, Steven; Han, Jeffrey; Narayanan, Arvind (2018): "I never signed up for this! Privacy implications of email tracking", Proceedings on Privacy Enhancing Technologies 2018(1):109-126. (DOI)] actually involves).

Where these papers go quiet

Share of the on-page papers that gave a real value rather than a not-stated-class sentinel, with the whole-corpus base rate beside it. A slice can look bad and be average, so the base-rate column is the one to read.

Field On-page papers Whole corpus
classification.validation 58/61 = 95.1% 3,977/4,439 = 89.6%
classification.groundTruthSource 44/61 = 72.1% 3,234/4,439 = 72.9%
classification.taxonomy 43/61 = 70.5% 2,854/4,439 = 64.3%
population.n 69/70 = 98.6% 5,496/5,712 = 96.2%
population.listVersion 33/70 = 47.1% 2,455/5,712 = 43.0%
humanAnnotation.annotatorCount 18/51 = 35.3% 1,288/3,318 = 38.8%
humanAnnotation.agreementMetric 4/51 = 7.8% 512/3,318 = 15.4%
artifacts.availability 36/67 = 53.7% 2,890/4,854 = 59.5%

Both columns are computed by the same script with the same rule: the denominator is papers for which the field's family fired, and a paper whose artifacts object is null is excluded rather than counted as silent. corpus and the dataset's own overview use the full empirical population as the artifact denominator and therefore report 56.5% where this table reports 59.5%; the two are not comparable, and a base-rate column that is not comparable to the column beside it is worse than none.

One row is genuinely worse than the field: inter-annotator agreement, at 4 of 51 papers that coded data by hand, against 15.4% corpus-wide. On a topic where the unit of analysis is often a message that a human read and judged — is this marketing or servicing mail, is this scam or legitimate SMS, what sector is this smishing campaign — that is the reporting gap to fix in your own paper. Of the crawling subset, only 14 of 27 with a recorded crawl configuration state a consent action, 7 of 27 state whether the crawl was stateful and 5 of 27 state headless or headful; both of the first two matter here, because a registration form behind a consent banner is a form you may not have filled in.

Methodology and limitations of these figures

  • How they were produced. One structured record per paper, extracted from full text, each tuple carrying a verbatim evidence quote and its section. report_email_tracking.mjs prints every number here with its denominator; msg_fold.mjs holds the population rule, the per-paper verdicts and the instrument fold with its residue; verify_email_figures.mjs re-checks every literal per-paper figure on this page against the paper's own text; external_checks_email_tracking.sh re-fetches every external fact. All four, their unedited output, and the full residue are on email_tracking. Corpus-level caveats: corpus.
  • The population is a judgement. 183 hand-written verdicts, each with its reason, published on the provenance page so you can disagree with individual papers rather than with the total.
  • A paper counts once, never once per tuple. Shares do not sum to 100%.
  • Sentinels are never counted as answers. The zero in the unsubscription row is a measured zero, not a missing value.
  • Every per-paper figure was checked against the paper, not against the extraction. 78 figures across 24 papers, all present. That pass earned its keep twice: the smishing paper's figures are typeset with the mathematical-italic k rather than an ASCII k, so an obvious needle reported a real figure as missing; and the extraction's summary of Venkatadri et al. [39Venkatadri, Giridhari; Andreou, Athanasios; Liu, Yabing; Mislove, Alan; Gummadi, Krishna P.; Loiseau, Patrick; Goga, Oana (2018): "Privacy Risks with Facebook's PII-Based Targeting: Auditing a Data Broker's Advertising Interface", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] reads “18 of 20 visitors inferred”, where the paper says “18 of the volunteers who did visit the webpage” out of 20 — correct, but a paraphrase, and this page quotes the paper.
  • Quotes. 400 detection evidence quotes attach to the 70 papers: 220 located verbatim, 118 after whitespace and punctuation normalisation, 62 below the matching threshold. Below-threshold is not “unsupported” — every one sampled by hand was present, spliced by two-column reading order. All 62 are listed on the provenance page.
  • Recall. Four schema-side probes that were not used to build the pool were run against it afterwards. They surfaced 73 papers outside it; all were read at title level, the twelve from the email-message enum were opened, and exactly one was a genuine miss — Starov et al. [7Starov, Oleksii; Gill, Phillipa; Nikiforakis, Nick (2016): "Are You Sure You Want to Contact Us? Quantifying the Leakage of PII via Website Contact Forms", in: Proceedings on Privacy Enhancing Technologies. (DOI)], now the oldest paper in the address-as-identifier slice. Read that as the measured cost of a title-and-density pool: about one paper in seventy, in the direction of missing the oldest work in a slice.
  • Venue coverage. Seven venues. EuroS&P, ACSAC, RAID, AsiaCCS, WPES, CHI and SOUPS are absent, and so is the CEAS/anti-spam literature entirely — for a topic with this much of its history in dedicated anti-spam and telephony venues, every count here is a lower bound and a weaker one than on most pages of this wiki.
  • Stability. classification.method agrees with an independent extraction run on 58% of papers and free-text names on about 20% of exact strings, which is why methods are given as rankings and enum fields as percentages. Those figures were measured on the previous, 4,322-paper run and have not been re-measured.

What to Report

  1. How you got the addresses, and how many. “We subscribed to newsletters” is not a population; “one address per site on a domain we control, across N sites drawn from X” is.
  2. The client you rendered in, by name and version — and whether its remote-content and privacy settings were default. A prevalence figure without this is not comparable with anything, given What a Pixel Can Still Measure in 2026.
  3. The provider you received at. Gmail, Apple, self-hosted and disposable are four different measurement instruments.
  4. Every encoding you searched the address under. Englehardt et al. [2Englehardt, Steven; Han, Jeffrey; Narayanan, Arvind (2018): "I never signed up for this! Privacy implications of email tracking", Proceedings on Privacy Enhancing Technologies 2018(1):109-126. (DOI)] tested 31 hash and encoding functions and two-function combinations; Hu and Wang [3Hu, Hang; Peng, Peng; Wang, Gang (2019): "Characterizing Pixel Tracking through the Lens of Disposable Email Services", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] found MD5 alone accounted for 91.7% of single-layer obfuscation. Report the list, and report which ones fired.
  5. Whether you distinguish “tracker present” from “tracker learned something”, and which one your headline number is.
  6. The label source, and that it is a choice. If your ground truth is a DNSBL or a provider's spam verdict, say so and cite [30Li, Ruixuan; Lu, Chaoyi; Liu, Baojun; Zhang, Yunyi; Hong, Geng; Duan, Haixin; Lin, Yanzhong; Pan, Qingfeng; Yang, Min; Shao, Jun (2025): "HADES Attack: Understanding and Evaluating Manipulation Risks of Email Blocklists", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] on what that means.
  7. Inter-annotator agreement, if a human judged messages. Four papers in fifty-one do; be the fifth.
  8. Your consent action and statefulness at the registration form, if you subscribed by crawling. Half the crawling papers here do not say.
  9. The ethics of the address: whose mail you read, whose number you enumerated, what you did with the accounts afterwards, and which law you assessed yourself against.
  10. For SMS and calls: the number block and its age. Older blocks receive more traffic [20Gupta, Payas; Srinivasan, Bharat; Balasubramaniyan, Vijay; Ahamad, Mustaque (2015): "Phoneypot: Data-driven Understanding of Telephony Threats", in: Proceedings of the Network and Distributed System Security Symposium. (Link)], so an unqualified rate is not comparable.
  11. For crowdsourced message data: the sources, the dates, and the overlap. [23Tang, Siyuan; Mi, Xianghang; Li, Ying; Wang, XiaoFeng; Chen, Kai (2022): "Clues in Tweets: Twitter-Guided Discovery and Analysis of SMS Spam", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] and [24Agarwal, Sharad; Papasavva, Antonis; Suarez-Tangil, Guillermo; Vasek, Marie (2025): "Fishing for Smishing: Understanding SMS Phishing Infrastructure and Strategies by Mining Public User Reports", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] both mine Twitter for reported SMS spam, three years apart, and neither says how much of the other's data it would have seen. Name your sources and their date ranges so the next paper can.

Open Questions

  • No published measurement of one-click unsubscribe. Zero papers in these seven venues; a specific, dated, testable requirement in force since 1 February 2024. This is the largest hole on the page. See Unsubscription: A Measured Gap before starting, and note that work is under way.
  • No paper measures a tracking-pixel prevalence per mail client. Gmail's proxy and Apple's Mail Privacy Protection change what the sender learns, in different ways, and every email-tracking figure in this corpus predates the question.
  • No successor to Englehardt et al. [2Englehardt, Steven; Han, Jeffrey; Narayanan, Arvind (2018): "I never signed up for this! Privacy implications of email tracking", Proceedings on Privacy Enhancing Technologies 2018(1):109-126. (DOI)] at comparable scale. Its 12,618 emails from 902 senders are still the reference figures for third-party content in mail, eight years on, and they predate both Mail Privacy Protection and the 2024 bulk-sender rules.
  • The lead-marketing pipeline is unjoined in the extraction. Who collects the address, who brokers it, and what arrives are three separate literatures here. The IEEE S&P 2026 paper that joins them is missing from this corpus by construction (see What the corpus cannot tell you) — read it before you propose the study.
  • Two 2026 preprints are about to land on this page's gaps and are not published yet. Agarwal, Suarez-Tangil and Vasek, “An Overview of 7726 User Reports” (arXiv:2508.05276), works from 1.35 million operator-side SMS reports; Altwlkany et al., “Robocalls: A Worldwide or US-only Problem?” (arXiv:2606.31790), from 8.7 million international call records across 65 countries. Neither is peer-reviewed at a venue as of 2026-09-02, so the zero counts on this page still stand. They are aimed at two things this corpus does not have: the cross-source bias question in the next bullet, and international coverage of robocalls, which nothing in these seven venues addresses at all.12)
  • Source bias in message corpora has not been re-measured since 2012. [5Pitsillidis, Andreas; Kanich, Chris; Voelker, Geoffrey M.; Levchenko, Kirill; Savage, Stefan (2012): "Taster's choice: a comparative analysis of spam feeds", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] did it for spam feeds and found 60% of live domains exclusive to a single feed. Nobody has repeated it on any channel. The two recent papers that mine Twitter for reported SMS spam — [23Tang, Siyuan; Mi, Xianghang; Li, Ying; Wang, XiaoFeng; Chen, Kai (2022): "Clues in Tweets: Twitter-Guided Discovery and Analysis of SMS Spam", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] and [24Agarwal, Sharad; Papasavva, Antonis; Suarez-Tangil, Guillermo; Vasek, Marie (2025): "Fishing for Smishing: Understanding SMS Phishing Infrastructure and Strategies by Mining Public User Reports", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] — are the obvious pair to compare, and neither reports the overlap with the other; nor has anyone compared a crowdsourced source against a carrier-side one ([22Murynets, Ilona; Jover, Roger Piqueras (2012): "Crime scene investigation: SMS spam data analysis", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], [14Huh, Jun Ho; Shin, Hyejin; Ahn, Sunwoo; Yi, Hayoon; Cho, Joonho; Kim, Taewoo; Lim, Minchae; Choi, Nuel (2025): "Preventing Artificially Inflated SMS Attacks through Large-Scale Traffic Inspection", in: Proceedings of the USENIX Security Symposium. (Link)]) on the same period.
  • Nobody has measured the messaging-app channel at email's scale. [4Kirchner, Robin; Koch, Simon; Kamangar, Noah; Klein, David; Johns, Martin (2024): "A Black-Box Privacy Analysis of Messaging Service Providers' Chat Message Processing", in: Proceedings on Privacy Enhancing Technologies. (DOI)] covers 105 platforms with honey messages; there is no equivalent of the newsletter-subscription crawl for chat, and no prevalence figure for tracking in machine-generated chat messages.
  • Inter-annotator agreement is essentially unreported in this slice (4 of 51). Since the field is moving toward LLM-assisted labelling of message content [31Hao, Wei; Tran, Van; Rideout, Vincent; Wang, Zixi; Dasbach-Prisk, AnMei; Afifi, M. H.; Yang, Junfeng; Katz-Bassett, Ethan; Ho, Grant; Cidon, Asaf (2025): "Do Spammers Dream of Electric Sheep? Characterizing the Prevalence of LLM-Generated Malicious Emails", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] [24Agarwal, Sharad; Papasavva, Antonis; Suarez-Tangil, Guillermo; Vasek, Marie (2025): "Fishing for Smishing: Understanding SMS Phishing Infrastructure and Strategies by Mining Public User Reports", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], the human baseline it would be validated against does not exist.
  • security:email_authentication does not exist. 31 papers in this corpus measure SPF, DKIM, DMARC, DANE, MTA-STS, STARTTLS or S/MIME deployment, and no page on this wiki covers them. It is the clearest missing page adjacent to this one.
  • Registration — getting an account, which is upstream of getting mail. This page is what the mailbox then lets you measure.
  • Requests — the web-page pixel, filter lists, and request classification. The other “pixel”, with a different method.
  • Consent — the banner in front of the registration form you are about to fill in.
  • Browser protection — the same argument one layer down: the instrument is blocking part of what you are trying to measure.
  • Fingerprinting — where [25Trampert, Leon; Weber, Daniel; Gerlach, Lukas; Rossow, Christian; Schwarz, Michael (2025): "Cascading Spy Sheets: Exploiting the Complexity of Modern CSS for Email and Browser Fingerprinting", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] also belongs, since CSS-only fingerprinting is not email-specific.
  • Phishing — email as an attack vector, and the 28 papers this page hands over.
  • Notifying websites — mail as a disclosure channel, and the response rates.
  • Ethics — the general apparatus behind this page's ethics section.
  • Mobile and app measurement — for the handset end of the SMS and messaging-app slice.
  • Multilingual support — because message data in this slice is multilingual by default.
  • Interrater agreement — the metric 47 of 51 papers here did not report.
  • email_tracking — every query, the scripts, their unedited output, the folds and the full residue.

References

[1]
Chand, Anish; Nikiforakis, Nick; Vadrevu, Phani (2025): "Doubly Dangerous: Evading Phishing Reporting Systems by Leveraging Email Tracking Techniques", in: Proceedings of the USENIX Security Symposium. (Link)
[2]
Englehardt, Steven; Han, Jeffrey; Narayanan, Arvind (2018): "I never signed up for this! Privacy implications of email tracking", Proceedings on Privacy Enhancing Technologies 2018(1):109-126. (DOI)
[3]
Hu, Hang; Peng, Peng; Wang, Gang (2019): "Characterizing Pixel Tracking through the Lens of Disposable Email Services", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[4]
Kirchner, Robin; Koch, Simon; Kamangar, Noah; Klein, David; Johns, Martin (2024): "A Black-Box Privacy Analysis of Messaging Service Providers' Chat Message Processing", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[5]
Pitsillidis, Andreas; Kanich, Chris; Voelker, Geoffrey M.; Levchenko, Kirill; Savage, Stefan (2012): "Taster's choice: a comparative analysis of spam feeds", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[6]
Senol, Asuman; Acar, Gunes; Humbert, Mathias; Zuiderveen Borgesius, Frederik (2022): "Leaky Forms: A Study of Email and Password Exfiltration Before Form Submission", in: 31st USENIX Security Symposium (USENIX Security 22). (Link)
[7]
Starov, Oleksii; Gill, Phillipa; Nikiforakis, Nick (2016): "Are You Sure You Want to Contact Us? Quantifying the Leakage of PII via Website Contact Forms", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[8]
Prasad, Sathvik; Nahapetyan, Aleksandr; Reaves, Bradley (2025): "Characterizing Robocalls with Multiple Vantage Points", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[9]
Durumeric, Zakir; Adrian, David; Mirian, Ariana; Kasten, James; Bursztein, Elie; Lidzborski, Nicolas; Thomas, Kurt; Eranti, Vijay; Bailey, Michael D.; Halderman, J. Alex (2015): "Neither Snow Nor Rain Nor MITM...: An Empirical Analysis of Email Delivery Security", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[10]
Shen, Kaiwen; Wang, Chuhan; Guo, Minglei; Zheng, Xiaofeng; Lu, Chaoyi; Liu, Baojun; Zhao, Yuxuan; Hao, Shuang; Duan, Haixin; Pan, Qingfeng; Yang, Min (2021): "Weak Links in Authentication Chains: A Large-scale Analysis of Email Sender Spoofing Attacks", in: Proceedings of the USENIX Security Symposium. (Link)
[11]
Kubíček, Karel; Merane, Jakob; Cotrini, Carlos; Stremitzer, Alexander; Bechtold, Stefan; Basin, David (2022): "Checking Websites' GDPR Consent Compliance for Marketing Emails", Proceedings on Privacy Enhancing Technologies 2022(2). (DOI)
[12]
Martin, Eric Burton Samuel; Shirazi, Hossein; Ray, Indrakshi (2023): "Poster: Towards a Dataset for the Discrimination between Warranted and Unwarranted Emails", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[13]
Dasgupta, Anirban; Punera, Kunal; Rao, Justin M.; Wang, Xuanhui (2012): "Impact of Spam Exposure on User Engagement", in: Proceedings of the USENIX Security Symposium. (Link)
[14]
Huh, Jun Ho; Shin, Hyejin; Ahn, Sunwoo; Yi, Hayoon; Cho, Joonho; Kim, Taewoo; Lim, Minchae; Choi, Nuel (2025): "Preventing Artificially Inflated SMS Attacks through Large-Scale Traffic Inspection", in: Proceedings of the USENIX Security Symposium. (Link)
[15]
Iqbal, Hassan; Khan, Usman Mahmood; Khan, Hassan Ali; Shahzad, Muhammad (2022): "Left or Right: A Peek into the Political Biases in Email Spam Filtering Algorithms During US Election 2020", in: Proceedings of the ACM Web Conference. (DOI)
[16]
Rao, Sumanth; Liu, Enze; Ho, Grant; Voelker, Geoffrey M.; Savage, Stefan (2024): "Unfiltered: Measuring Cloud-based Email Filtering Bypasses", in: Proceedings of the ACM Web Conference. (DOI)
[17]
Lécuyer, Mathias; Ducoffe, Guillaume; Lan, Francis; Papancea, Andrei; Petsios, Theofilos; Spahn, Riley; Chaintreau, Augustin; Geambasu, Roxana (2014): "XRay: Enhancing the Web’s Transparency with Differential Correlation", in: Proceedings of the USENIX Security Symposium. (Link)
[18]
Farooqi, Shehroze; Musa, Maaz; Shafiq, Zubair; Zaffar, Fareed (2020): "CanaryTrap: Detecting Data Misuse by Third-Party Apps on Online Social Networks", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[19]
DeBlasio, Joe; Savage, Stefan; Voelker, Geoffrey M.; Snoeren, Alex C. (2017): "Tripwire: inferring internet site compromise", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[20]
Gupta, Payas; Srinivasan, Bharat; Balasubramaniyan, Vijay; Ahamad, Mustaque (2015): "Phoneypot: Data-driven Understanding of Telephony Threats", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[21]
Prasad, Sathvik; Bouma-Sims, Elijah; Mylappan, Athishay Kiran; Reaves, Bradley (2020): "Who's Calling? Characterizing Robocalls through Audio and Metadata Analysis", in: Proceedings of the USENIX Security Symposium. (Link)
[22]
Murynets, Ilona; Jover, Roger Piqueras (2012): "Crime scene investigation: SMS spam data analysis", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[23]
Tang, Siyuan; Mi, Xianghang; Li, Ying; Wang, XiaoFeng; Chen, Kai (2022): "Clues in Tweets: Twitter-Guided Discovery and Analysis of SMS Spam", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[24]
Agarwal, Sharad; Papasavva, Antonis; Suarez-Tangil, Guillermo; Vasek, Marie (2025): "Fishing for Smishing: Understanding SMS Phishing Infrastructure and Strategies by Mining Public User Reports", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[25]
Trampert, Leon; Weber, Daniel; Gerlach, Lukas; Rossow, Christian; Schwarz, Michael (2025): "Cascading Spy Sheets: Exploiting the Complexity of Modern CSS for Email and Browser Fingerprinting", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[26]
Levchenko, Kirill; Pitsillidis, Andreas; Chachra, Neha; Enright, Brandon; Félegyházi, Márk; Grier, Chris; Halvorson, Tristan; Kanich, Chris; Kreibich, Christian; Liu, He; McCoy, Damon; Weaver, Nicholas; Paxson, Vern; Voelker, Geoffrey M.; Savage, Stefan (2011): "Click Trajectories: End-to-End Analysis of the Spam Value Chain", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[27]
McCoy, Damon; Pitsillidis, Andreas; Jordan, Grant; Weaver, Nicholas; Kreibich, Christian; Krebs, Brian; Voelker, Geoffrey M.; Savage, Stefan; Levchenko, Kirill (2012): "PharmaLeaks: Understanding the Business of Online Pharmaceutical Affiliate Programs", in: Proceedings of the USENIX Security Symposium. (Link)
[28]
Kanich, Chris; Weaver, Nicholas; McCoy, Damon; Halvorson, Tristan; Kreibich, Christian; Levchenko, Kirill; Paxson, Vern; Voelker, Geoffrey M.; Savage, Stefan (2011): "Show Me the Money: Characterizing Spam-advertised Revenue", in: Proceedings of the USENIX Security Symposium. (Link)
[29]
Stringhini, Gianluca; Holz, Thorsten; Stone-Gross, Brett; Kruegel, Christopher; Vigna, Giovanni (2011): "BOTMAGNIFIER: Locating Spambots on the Internet", in: Proceedings of the USENIX Security Symposium. (Link)
[30]
Li, Ruixuan; Lu, Chaoyi; Liu, Baojun; Zhang, Yunyi; Hong, Geng; Duan, Haixin; Lin, Yanzhong; Pan, Qingfeng; Yang, Min; Shao, Jun (2025): "HADES Attack: Understanding and Evaluating Manipulation Risks of Email Blocklists", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[31]
Hao, Wei; Tran, Van; Rideout, Vincent; Wang, Zixi; Dasbach-Prisk, AnMei; Afifi, M. H.; Yang, Junfeng; Katz-Bassett, Ethan; Ho, Grant; Cidon, Asaf (2025): "Do Spammers Dream of Electric Sheep? Characterizing the Prevalence of LLM-Generated Malicious Emails", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[32]
Prasad, Sathvik; Dunlap, Trevor; Ross, Alexander; Reaves, Bradley (2023): "Diving into Robocall Content with SnorCall", in: Proceedings of the USENIX Security Symposium. (Link)
[33]
Jiang, Nan; Jin, Yu; Skudlark, Ann; Zhang, Zhi-Li (2013): "Greystar: Fast and Accurate Detection of SMS Spam Numbers in Large Cellular Networks Using Gray Phone Space", in: Proceedings of the USENIX Security Symposium. (Link)
[34]
Li, Zhenhua; Wang, Weiwei; Wilson, Christo; Chen, Jian; Qian, Chen; Jung, Taeho; Zhang, Lan; Liu, Kebin; Li, Xiangyang; Liu, Yunhao (2017): "FBS-Radar: Uncovering Fake Base Stations at Scale in the Wild", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[35]
Venkatadri, Giridhari; Sapiezynski, Piotr; Redmiles, Elissa M.; Mislove, Alan; Goga, Oana; Mazurek, Michelle L.; Gummadi, Krishna P. (2019): "Auditing Offline Data Brokers via Facebook's Advertising Platform", in: Proceedings of the ACM Web Conference. (DOI)
[36]
Kempen, Elina van; Bagayatkar, Isita; Frolikov, Pavel; Georgiou, Chloe; Tsudik, Gene (2026): "Consumer Beware! Exploring Data Brokers' CCPA Compliance", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[37]
Acar, Gunes; Englehardt, Steven; Narayanan, Arvind (2020): "No boundaries: data exfiltration by third parties embedded on web pages", Proceedings on Privacy Enhancing Technologies 2020(4):220-238. (DOI)
[38]
Kieserman, Julia B.; Andreou, Athanasios; Geeng, Chris; Lauinger, Tobias; McCoy, Damon (2025): "Tracker Installations Are Not Created Equal: Understanding Tracker Configuration of Form Data Collection", in: Proceedings on Privacy Enhancing Technologies, pp. 679-695. (DOI)
[39]
Venkatadri, Giridhari; Andreou, Athanasios; Liu, Yabing; Mislove, Alan; Gummadi, Krishna P.; Loiseau, Patrick; Goga, Oana (2018): "Privacy Risks with Facebook's PII-Based Targeting: Auditing a Data Broker's Advertising Interface", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[40]
Venkatadri, Giridhari; Lucherini, Elena; Sapiezynski, Piotr; Mislove, Alan (2019): "Investigating sources of PII used in Facebook’s targeted advertising", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[41]
Wu, Mengying; Hong, Geng; Chen, Jiatao; Liu, Baojun; Liu, Mingxuan; Yang, Min (2026): "One Email, Many Faces: A Deep Dive into Identity Confusion in Email Aliases", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[42]
Hagen, Christoph; Weinert, Christian; Sendner, Christoph; Dmitrienko, Alexandra; Schneider, Thomas (2021): "All the Numbers are US: Large-scale Abuse of Contact Discovery in Mobile Messengers", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[43]
Gegenhuber, Gabriel K.; Frenzel, Philipp E.; Günther, Maximilian; Ullrich, Johanna; Judmayer, Aljosha (2026): "Hey there! You are using WhatsApp: Enumerating Three Billion Accounts for Security and Privacy", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[44]
Kang, Junkyu; Lee, Soyoung; Kwon, Yonghwi; Son, Sooel (2026): "Connecting the Dots: An Investigative Study on Linking Private User Data Across Messaging Apps", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[45]
Niksirat, Kavous Salehzadeh; Velykoivanenko, Lev; Mätzler, Samuel; Mulders, Stephan; Tamò-Larrieux, Aurelia; Boldi, Marc-Olivier; Humbert, Mathias; Huguenin, Kévin (2025): "Addressing the Address Books' (Interdependent) Privacy Issues", in: Proceedings of the USENIX Security Symposium. (Link)
[46]
Stringhini, Gianluca; Egele, Manuel; Zarras, Apostolis; Holz, Thorsten; Kruegel, Christopher; Vigna, Giovanni (2012): "B@bel: Leveraging Email Delivery for Spam Mitigation", in: Proceedings of the USENIX Security Symposium. (Link)
[47]
Qian, Zhiyun; Mao, Zhuoqing Morley; Xie, Yinglian; Yu, Fang (2010): "Investigation of Triangular Spamming: A Stealthy and Efficient Spamming Technique", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[48]
Thomas, Kurt; Grier, Chris; Ma, Justin; Paxson, Vern; Song, Dawn (2011): "Design and Evaluation of a Real-Time URL Spam Filtering Service", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[49]
Park, Youngsam; Jones, Jackie; McCoy, Damon; Shi, Elaine; Jakobsson, Markus (2014): "Scambaiter: Understanding Targeted Nigerian Scams on Craigslist", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[50]
Zhang, Jiahe; Chen, Jianjun; Wang, Qi; Zhang, Hangyu; Wang, Chuhan; Zhuge, Jianwei; Duan, Haixin (2024): "Inbox Invasion: Exploiting MIME Ambiguities to Evade Email Attachment Detectors", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[51]
Agarwal, Sharad; Harvey, Emma; Vasek, Marie (2024): "Poster: A Comprehensive Categorization of SMS Scams", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[52]
Bitsikas, Evangelos; Schnitzler, Theodor; Pöpper, Christina; Ranganathan, Aanjhan (2023): "Freaky Leaky SMS: Extracting User Locations by Analyzing SMS Timings", in: Proceedings of the USENIX Security Symposium. (Link)
[53]
Schnitzler, Theodor; Kohls, Katharina; Bitsikas, Evangelos; Pöpper, Christina (2023): "Hope of Delivery: Extracting User Locations From Mobile Instant Messengers", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[54]
Reaves, Bradley; Scaife, Nolen; Tian, Dave; Blue, Logan; Traynor, Patrick; Butler, Kevin R. B. (2016): "Sending Out an SMS: Characterizing the Security of the SMS Ecosystem with Public Gateways", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[55]
Mulliner, Collin; Golde, Nico; Seifert, Jean-Pierre (2011): "SMS of Death: From Analyzing to Attacking Mobile Phones on a Large Scale", in: Proceedings of the USENIX Security Symposium. (Link)
[56]
Tu, Guan-Hua; Li, Chi-Yu; Peng, Chunyi; Li, Yuanjie; Lu, Songwu (2016): "New Security Threats Caused by IMS-based SMS Service in 4G LTE Networks", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[57]
Lei, Zeyu; Nan, Yuhong; Fratantonio, Yanick; Bianchi, Antonio (2021): "On the Insecurity of SMS One-Time Password Messages against Local Attackers in Modern Mobile Devices", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[58]
Wang, Qi; Chen, Jianjun; Yang, Jingcheng; Zhang, Jiahe; Yang, Yaru; Duan, Haixin (2026): "SIPConfusion: Exploiting SIP Semantic Ambiguities for Caller ID and SMS Spoofing", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[59]
Zhang, Yiming; Liu, Baojun; Lu, Chaoyi; Li, Zhou; Duan, Haixin; Hao, Shuang; Liu, Mingxuan; Liu, Ying; Wang, Dong; Li, Qiang (2020): "Lies in the Air: Characterizing Fake-base-station Spam Ecosystem in China", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[60]
Edu, Jide S.; Mulligan, Cliona; Pierazzi, Fabio; Polakis, Jason; Suarez-Tangil, Guillermo; Such, Jose M. (2022): "Exploring the security and privacy risks of chatbots in messaging services", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[61]
Sherman, Imani N.; Bowers, Jasmine D.; McNamara, Jr., Keith; Gilbert, Juan E.; Ruiz, Jaime; Traynor, Patrick (2020): "Are You Going to Answer That? Measuring User Responses to Anti-Robocall Application Indicators", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[62]
Sharevski, Filipo; Loop, Jennifer Vander; Evans, Bill; Ponticello, Alexander (2025): "(Blind) Users Really Do Heed Aural Telephone Scam Warnings", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[63]
Agarwal, Sharad; Harvey, Emma; Mariconti, Enrico; Suarez-Tangil, Guillermo; Vasek, Marie (2025): "'Hey mum, I dropped my phone down the toilet': Investigating Hi Mum and Dad SMS Scams in the United Kingdom", in: Proceedings of the USENIX Security Symposium. (Link)
[64]
Onaolapo, Jeremiah; Mariconti, Enrico; Stringhini, Gianluca (2016): "What Happens After You Are Pwnd: Understanding the Use of Leaked Webmail Credentials in the Wild", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[65]
Stock, Ben; Pellegrino, Giancarlo; Rossow, Christian; Johns, Martin; Backes, Michael (2016): "Hey, You Have a Problem: On the Feasibility of Large-Scale Web Vulnerability Notification", in: Proceedings of the USENIX Security Symposium. (Link)
[66]
Maass, Max; Stöver, Alina; Pridöhl, Henning; Bretthauer, Sebastian; Herrmann, Dominik; Hollick, Matthias; Spiecker, Indra (2021): "Effective Notification Campaigns on the Web: A Matter of Trust, Framing, and Support", in: Proceedings of the USENIX Security Symposium. (Link)
[67]
Utz, Christine; Michels, Matthias; Degeling, Martin; Marnau, Ninja; Stock, Ben (2023): "Comparing Large-Scale Privacy and Security Notifications", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[68]
Sasaki, Takayuki; Inazawa, Tomoya; Yamaguchi, Youhei; Parkin, Simon; Eeten, Michel van; Yoshioka, Katsunari; Matsumoto, Tsutomu (2025): "Am I Infected? Lessons from Operating a Large-Scale IoT Security Diagnostic Service", in: Proceedings of the USENIX Security Symposium. (Link)
[69]
Martino, Mariano Di; Meers, Isaac; Quax, Peter; Andries, Ken; Lamotte, Wim (2022): "Revisiting Identification Issues in GDPR ‘Right Of Access’ Policies: A Technical and Longitudinal Analysis", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[70]
Rupp, Eduard; Syrmoudis, Emmanuel; Grossklags, Jens (2022): "Leave No Data Behind – Empirical Insights into Data Erasure from Online Services", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[71]
Drakonakis, Kostas; Ioannidis, Sotiris; Polakis, Jason (2020): "The Cookie Hunter: Automated Black-box Auditing for Web Authentication and Authorization Flaws", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[72]
Oh, ChangSeok; Kanich, Chris; McCoy, Damon; Pearce, Paul (2022): "Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
1)
Google, "Turn images on or off in Gmail", section “Learn how Gmail helps make images safe”. Fetched 2026-09-02.
2)
Apple, "Use Mail Privacy Protection on iPhone". Fetched 2026-09-02 with a real browser; the page is client-rendered and curl alone returns a shell. iOS 15 is the oldest version the guide's own version selector offers, which is the date evidence. Apple documents it as a setting the user turns on and publishes no take-up figure, so do not assume it is on and do not assume it is off — from the sender side you cannot tell, and that is itself the measurement problem.
3)
Gmail Blog, "Images Now Showing", 12 December 2013: “Instead of serving images directly from their original external host servers, Gmail will now serve all images through Google's own secure proxy servers”, rolling out on desktop that day and to the mobile apps in early 2014. Fetched 2026-09-02.
4)
arxiv.org/abs/2604.06759, submitted 8 April 2026, title and all four authors matching the IEEE record; code and data at github.com/Yash-Vekaria/lead-marketing-spam. The DOI resolves to ieeexplore.ieee.org/document/11573454. Checked 2026-09-02.
5)
See the footnote in the box at the top of this page.
6)
WebKit, "Private Browsing 2.0", the “Link Tracking Protection” section; fetched 2026-09-02. It lists the protection under “protections and defenses added to Private Browsing in Safari 17.0”. Widely repeated claims that the same stripping applies to links opened from Mail, and that iOS 26 extended it to all of Safari, are not in this or any other Apple primary source this page's author could locate; they are therefore not asserted here. If your method depends on it, test it.
7)
Chen et al., A Large-scale and Longitudinal Measurement Study of DKIM Deployment, USENIX Security 2022. The header appears in its Class-2 field list. Read by hand 2026-09-02.
8)
Google, "Email sender guidelines" and "Email subscription guidelines for senders", fetched 2026-09-02. Yahoo announced the same requirements first, on 3 October 2023: “we will require senders to support one-click unsubscribe and honor our users requests within two days” — Marcel Becker, "More Secure, Less Spam".
9)
Google, "Email sender guidelines FAQ", fetched 2026-09-02.
10)
Microsoft, “Strengthening Email Ecosystem: Outlook's New Requirements for High-Volume Senders”, Microsoft Defender for Office 365 blog, 4399730. Verified only indirectly. The announcement URL and Microsoft's own Q&A threads on learn.microsoft.com that quote it are served behind a block that refuses this host, both to curl and to a real headless browser (Azure Front Door: “The request is blocked”), so the wording above rests on Microsoft-hosted Q&A pages that quote and link the announcement rather than on the announcement itself, plus the 550 5.7.515 Access denied NDR text those threads reproduce. Checked 2026-09-02. Read the announcement before you cite this in a paper.
11)
The dataset's own overview reports 33.8% for this field, over all 5,118 empirical papers, counting the 646 with no ethics object at all as not-stated. That is a defensible denominator and not the one used here; mixing the two would flatter this slice by five points.
12)
Both checked on arXiv 2026-09-02. Cited as preprints, not as results, and deliberately not used for any figure on this page.
You could leave a comment if you were logged in.
privacy/email_tracking.1788334273.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki