User Tools

Site Tools


privacy:email_tracking

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
privacy:email_tracking [2026/09/02 07:31] – Self-review pass: correct three overstatements — 'two of four never measured' (one is zero, one is unmeasured), 'four clusters larger than anything here' (three of four), and 'unusual concentration of ethics problems' (the number is about legal engagement karel.kubicek.claudeprivacy:email_tracking [2026/09/02 07:40] (current) – Generic + re-verification review fixes: state the real hand-map accounting (183 mapped / 125 with a written reason / 171 defaulted at title level) instead of claiming 183 hand-written reasons; correct '73 papers outside the pool' to 67 distinct / 72 menti karel.kubicek.claude
Line 6: Line 6:
  
 <WRAP important> <WRAP important>
-**The one thing to understand before you start: the mailbox provider is now a party to your measurement, and it does not do the same thing to every signal.** Gmail proxies remote images, so the sender learns nothing about the recipient's IP, user agent or cookies — but still learns that the message was opened. Google says both parts in the same help article: "Senders can't use image loading to get information about your computer or location", "Senders can't use the image to set or read cookies in your browser", and then "Sometimes, senders may know whether you've opened an email that has an image."((Google, [[https://support.google.com/mail/answer/145919|"Turn images on or off in Gmail"]], section "Learn how Gmail helps make images safe". Fetched 2026-09-02.)) Apple Mail Privacy Protection, available since iOS 15 (2021), takes the other half too: it "hides your IP address so senders can't link it to your other online activity or determine your exact location" and "prevents senders from seeing if you've opened the email message they sent you", by fetching remote content in the background on arrival rather than on view.((Apple, [[https://support.apple.com/guide/iphone/use-mail-privacy-protection-iphf084865c7/ios|"Use Mail Privacy Protection on iPhone"]]. Fetched 2026-09-02 with a real browserthe page is client-rendered and ''curl'' alone returns a shell. iOS 15 is the oldest version the guide's own version selector offers, which is the date evidence. Apple documents it as a setting the user turns on and publishes no take-up figure, so **do not assume it is on and do not assume it is off** — from the sender side you cannot tell, and that is itself the measurement problem.))+**The one thing to understand before you start: the mailbox provider is now a party to your measurement, and it does not do the same thing to every signal.** Gmail proxies remote images, so the sender learns nothing about the recipient's IP, user agent or cookies — but still learns that the message was opened. Google says both parts in the same help article: "Senders can't use image loading to get information about your computer or location", "Senders can't use the image to set or read cookies in your browser", and then "Sometimes, senders may know whether you've opened an email that has an image."((Google, [[https://support.google.com/mail/answer/145919|"Turn images on or off in Gmail"]], section "Learn how Gmail helps make images safe". Fetched 2026-09-02.)) Apple Mail Privacy Protection, available since iOS 15 (2021), takes the other half too: it "hides your IP address so senders can't link it to your other online activity or determine your exact location" and "prevents senders from seeing if you've opened the email message they sent you", by fetching remote content in the background on arrival rather than on view.((Apple, [[https://support.apple.com/guide/iphone/use-mail-privacy-protection-iphf084865c7/ios|"Use Mail Privacy Protection on iPhone"]]. Fetched 2026-09-02. The quoted sentences were read out of a real headless browser, because the guide is client-rendered''curl'' returns a navigation shell **plus** a noscript copy of the body, which is why the check in ''external_checks_email_tracking.sh'' also finds the needle. Trust the browser fetch, not the ''curl'' pass. iOS 15 is the oldest version the guide's own version selector offers, which is the date evidence. Apple documents it as a setting the user turns on and publishes no take-up figure, so **do not assume it is on and do not assume it is off** — from the sender side you cannot tell, and that is itself the measurement problem.))
  
 So "we found a tracking pixel in 24.6% of messages" and "the sender learned that 24.6% of messages were opened" are now different claims, and which one you can make depends on **which client you rendered in**. None of the email-tracking papers in this corpus measures the effect of that change: the newest of them {[chand2025_doubly]} was published in 2025 and still treats the pixel fetch as the signal. If you are planning this measurement, the client is a treatment you have to vary on purpose, and you have no baseline in this literature to compare against. See [[#What a Pixel Can Still Measure in 2026]]. So "we found a tracking pixel in 24.6% of messages" and "the sender learned that 24.6% of messages were opened" are now different claims, and which one you can make depends on **which client you rendered in**. None of the email-tracking papers in this corpus measures the effect of that change: the newest of them {[chand2025_doubly]} was published in 2025 and still treats the pixel fetch as the signal. If you are planning this measurement, the client is a treatment you have to vary on purpose, and you have no baseline in this literature to compare against. See [[#What a Pixel Can Still Measure in 2026]].
Line 38: Line 38:
 Four neighbouring clusters were screened out of the same candidate pool and are **not** on this page. Three of them are larger than any slice that is. Each is named so you do not conclude the corpus is silent about it: Four neighbouring clusters were screened out of the same candidate pool and are **not** on this page. Three of them are larger than any slice that is. Each is named so you do not conclude the corpus is silent about it:
  
-  * **Email transport, authentication and encryption deployment** — SPF, DKIM, DMARC, DANE, MTA-STS, STARTTLS, S/MIME, sender spoofing, delivery paths. **31 papers**, heavily 2018–2026, and the single biggest coherent slice the pool turned up. It is page-sized and it is not tracking; it belongs under ''security:'' and does not exist yet. Start from Durumeric et al. {[durumeric2015_neither]} and Shen et al. {[shen2021_weak]} if you need it now.+  * **Email transport, authentication and encryption deployment** — SPF, DKIM, DMARC, DANE, MTA-STS, STARTTLS, S/MIME, sender spoofing, delivery paths. **31 papers**, heavily 2018–2026. The social-spam cluster below is larger, but it is not one topic; this is, and it is page-sized. It is also not tracking; it belongs under ''security:'' and does not exist yet. Start from Durumeric et al. {[durumeric2015_neither]} and Shen et al. {[shen2021_weak]} if you need it now.
   * **Phishing and its interventions**, with email as the vector — **28 papers**. [[Security:Phishing]] owns this. The line drawn here: a paper about what a message //discloses to a third party// is on this page; a paper about whether a user //falls for// a message is on that one.   * **Phishing and its interventions**, with email as the vector — **28 papers**. [[Security:Phishing]] owns this. The line drawn here: a paper about what a message //discloses to a third party// is on this page; a paper about whether a user //falls for// a message is on that one.
   * **Social-platform, review, forum, SEO and ad-click spam** — **37 papers**, more than any slice on this page. "Spam" in these venues usually means Twitter accounts or product reviews, not mail. It is a platform question ([[Design:Platforms]]), not a message-channel one.   * **Social-platform, review, forum, SEO and ad-click spam** — **37 papers**, more than any slice on this page. "Spam" in these venues usually means Twitter accounts or product reviews, not mail. It is a platform question ([[Design:Platforms]]), not a message-channel one.
Line 122: Line 122:
 | CAN-SPAM | **2** | | CAN-SPAM | **2** |
 | the word %%unsubscrib*%% anywhere | 56 | | the word %%unsubscrib*%% anywhere | 56 |
 +| %%opt out of the mailing / e-mail / newsletter / marketing / list%% | **2** |
  
-The one ''List-Unsubscribe'' hit names the header in a table of RFC 6376 signable header fields; it measures nothing about it.((Chen et al., //A Large-scale and Longitudinal Measurement Study of DKIM Deployment//, USENIX Security 2022. The header appears in its Class-2 field list. Read by hand 2026-09-02.)) Both CAN-SPAM hits are motivational asides. Of the 56 ''unsubscrib*'' papers, **ten** are using the MQTT, pub/sub or SDN protocol verb, **three** mean the CCPA "do not sell" opt-out or an advertising opt-out cookie, **four** mean turning off a security notification, and the remaining 39 have a single passing mention in a paper about something else.+The one ''List-Unsubscribe'' hit names the header in a table of RFC 6376 signable header fields; it measures nothing about it.((Chen et al., //A Large-scale and Longitudinal Measurement Study of DKIM Deployment//, USENIX Security 2022. The header appears in its Class-2 field list. Read by hand 2026-09-02.)) Of the two CAN-SPAM hits, one is a motivational aside in an introduction and the other is a paper that extracts opt-out statements from privacy-policy //text// — neither measures a mechanism. Of the 56 ''unsubscrib*'' papers, **ten** are using the MQTT, pub/sub or SDN protocol verb, **three** mean the CCPA "do not sell" opt-out or an advertising opt-out cookie, **four** mean turning off a security notification, and the remaining 39 have a single passing mention in a paper about something else.
  
 **The nearest thing in the corpus is one sub-check inside one paper.** Kubicek et al. {[kubicek2022_emails]} report that **16%** of websites sending marketing mail provided neither an unsubscribe method nor a legal notice, alongside **17.3%** of such websites having at least one potential consent violation at the registration form, **21.9%** having one somewhere, **59%** sending a double opt-in confirmation first, and **2.3%** mailing the user's own password back in plaintext. **The nearest thing in the corpus is one sub-check inside one paper.** Kubicek et al. {[kubicek2022_emails]} report that **16%** of websites sending marketing mail provided neither an unsubscribe method nor a legal notice, alongside **17.3%** of such websites having at least one potential consent violation at the registration form, **21.9%** having one somewhere, **59%** sending a double opt-in confirmation first, and **2.3%** mailing the user's own password back in plaintext.
Line 157: Line 158:
   * **Injection at the radio layer is a Chinese-market phenomenon with real numbers.** Zhang et al. {[zhang2020_lies]} characterised fake-base-station SMS spam: 279,017 message logs over 97 days, 7,884 campaigns, and illegal businesses accounting for over 75% of the messages.   * **Injection at the radio layer is a Chinese-market phenomenon with real numbers.** Zhang et al. {[zhang2020_lies]} characterised fake-base-station SMS spam: 279,017 message logs over 97 days, 7,884 campaigns, and illegal businesses accounting for over 75% of the messages.
   * **Messaging apps.** Edu et al. {[edu2022_exploring]} measured third-party chatbots in messaging channels — 8,521 of 15,525 valid chatbots (54.86%) requested administrator permissions and 95.67% had no privacy policy. Kirchner et al. {[kirchner2024_black]} is the honey-message study. Together they are the reason "email tracking" is the wrong frame for the next study in this area.   * **Messaging apps.** Edu et al. {[edu2022_exploring]} measured third-party chatbots in messaging channels — 8,521 of 15,525 valid chatbots (54.86%) requested administrator permissions and 95.67% had no privacy policy. Kirchner et al. {[kirchner2024_black]} is the honey-message study. Together they are the reason "email tracking" is the wrong frame for the next study in this area.
-  * **User-facing work exists and is small.** Sherman et al. {[sherman2020_going]} measured responses to anti-robocall indicators (answered calls fell 43% with an "Avail-Spam" warning); Sharevski and Zettlemoyer {[sharevski2025_blind]} did aural scam warnings with blind participants; Agarwal et al. {[agarwal2025_dropped]} characterised UK "Hi Mum" impersonation scams — 582 mule accounts, over £577k requested in 13 weeks, and a 14-day median lifetime for the originating sender IDs.+  * **User-facing work exists and is small.** Sherman et al. {[sherman2020_going]} measured responses to anti-robocall indicators (answered calls fell 43% with an "Avail-Spam" warning); Sharevski et al. {[sharevski2025_blind]} did aural scam warnings with blind participants; Agarwal et al. {[agarwal2025_dropped]} characterised UK "Hi Mum" impersonation scams — 582 mule accounts, over £577k requested in 13 weeks, and a 14-day median lifetime for the originating sender IDs.
  
 ===== The Mailbox as an Instrument ===== ===== The Mailbox as an Instrument =====
Line 188: Line 189:
 No field in the extraction means "the measured object is a message". ''classification.target'' does have an ''email-message'' value, and it is **noisy in a way an enum disguises**: it fires on 54 papers, of which 12 are outside this page's candidate pool entirely and include //The Matter of Heartbleed//, //The State of the SameSite// and //Traveling the Silk Road//. So the population here is a published candidate pool plus an explicit verdict per paper. No field in the extraction means "the measured object is a message". ''classification.target'' does have an ''email-message'' value, and it is **noisy in a way an enum disguises**: it fires on 54 papers, of which 12 are outside this page's candidate pool entirely and include //The Matter of Heartbleed//, //The State of the SameSite// and //Traveling the Silk Road//. So the population here is a published candidate pool plus an explicit verdict per paper.
  
-The pool is a union of three signals — the title names a message channel; the full text is dense in message-channel vocabulary; the full text names a messaging-app address space ("contact discovery", "address book"). It contains **354** papers**183** of which carry a hand-written verdict, and **70** of which are on this page.+The pool is a union of three signals — the title names a message channel; the full text is dense in message-channel vocabulary; the full text names a messaging-app address space ("contact discovery", "address book"). It contains **354** papers**183** carry a hand-written verdict; of those, **125** also carry a written reason, and the rest are the phishing, social-spam and web-pixel rows where the verdict //is// the reason. The other **171** are ''OFF'' by default: **46** matched a named off-topic family rule, and **125** carry only "not individually annotated" — they were screened at title level and nothing further was recorded about them. **70** are on this page. Do not read the 181-paper ''OFF'' row below as 181 individual judgements; it is 10 of them plus 171 title-level screens.
  
 ^ Verdict ^ Papers ^ Where it goes ^ ^ Verdict ^ Papers ^ Where it goes ^
Line 269: Line 270:
  
 All nineteen families the fold produces on this population are listed; nothing is truncated. Only four of them are mail- or message-specific — the rest is the ordinary apparatus of a web or mobile measurement, which is the point: **there is no toolchain for this topic.** All nineteen families the fold produces on this population are listed; nothing is truncated. Only four of them are mail- or message-specific — the rest is the ordinary apparatus of a web or mobile measurement, which is the point: **there is no toolchain for this topic.**
 +
 +**Four tools in this whole population are mail- or message-specific**, and they are the most directly reusable thing on this page: the **Email Privacy Tester** (''emailprivacytester.com'', alive as of 2026-09-02 — sends a message to you and reports which tracking techniques your client executed), **EmailHarvester**, the **Honey Messages Framework** from {[kirchner2024_black]}, and **css-inline** (needed because mail clients want inlined CSS, which is also why {[trampert2025_cascading]} works). Everything else in the table is the ordinary apparatus of a web or mobile measurement. Alongside them, the residue contains ''Faker'', ''Fake Name Generator'', ''This person does not exist'', ''2Captcha'' and ''DeCaptcher'' — which is what signing up on fifteen thousand sites actually requires, and worth budgeting for.
  
 The fold is not cosmetic. **Spamhaus appears under four spellings** on this small a population (''Spamhaus'', ''Spamhaus blacklist'', ''Spamhaus.org'', ''Spamhaus passive DNS API''), **libphonenumber under four**, and MaxMind under four. Counting exact strings would have put every one of them below the reporting threshold. **159 distinct strings** matched no family and are printed in full on the provenance page rather than dropped; they are almost all one-off infrastructure (''Faker'', ''Fake Name Generator'', ''2Captcha'', ''DeCaptcher'' — which between them tell you something about what signing up on 15,700 sites {[englehardt2018_email]} actually involves). The fold is not cosmetic. **Spamhaus appears under four spellings** on this small a population (''Spamhaus'', ''Spamhaus blacklist'', ''Spamhaus.org'', ''Spamhaus passive DNS API''), **libphonenumber under four**, and MaxMind under four. Counting exact strings would have put every one of them below the reporting threshold. **159 distinct strings** matched no family and are printed in full on the provenance page rather than dropped; they are almost all one-off infrastructure (''Faker'', ''Fake Name Generator'', ''2Captcha'', ''DeCaptcher'' — which between them tell you something about what signing up on 15,700 sites {[englehardt2018_email]} actually involves).
Line 297: Line 300:
   * **Sentinels are never counted as answers.** The zero in the unsubscription row is a measured zero, not a missing value.   * **Sentinels are never counted as answers.** The zero in the unsubscription row is a measured zero, not a missing value.
   * **Every per-paper figure was checked against the paper, not against the extraction.** 78 figures across 24 papers, all present. That pass earned its keep twice: the smishing paper's figures are typeset with the mathematical-italic //k// rather than an ASCII ''k'', so an obvious needle reported a real figure as missing; and the extraction's summary of Venkatadri et al. {[venkatadri2018_privacy]} reads "18 of 20 visitors inferred", where the paper says "18 of the volunteers who did visit the webpage" out of 20 — correct, but a paraphrase, and this page quotes the paper.   * **Every per-paper figure was checked against the paper, not against the extraction.** 78 figures across 24 papers, all present. That pass earned its keep twice: the smishing paper's figures are typeset with the mathematical-italic //k// rather than an ASCII ''k'', so an obvious needle reported a real figure as missing; and the extraction's summary of Venkatadri et al. {[venkatadri2018_privacy]} reads "18 of 20 visitors inferred", where the paper says "18 of the volunteers who did visit the webpage" out of 20 — correct, but a paraphrase, and this page quotes the paper.
-  * **Quotes.** 400 ''detection'' evidence quotes attach to the 70 papers: 220 located verbatim, 118 after whitespace and punctuation normalisation, 62 below the matching threshold. Below-threshold is **not** "unsupported" — every one sampled by hand was present, spliced by two-column reading order. All 62 are listed on the provenance page. +  * **Quotes.** 400 ''detection'' evidence quotes attach to the 70 papers: 220 located verbatim, 118 after whitespace and punctuation normalisation, 62 below the matching threshold. Below-threshold is **not** "unsupported". **Ten** were read by hand and all ten were present, spliced by two-column reading order — the eight in the 40–60% band that the provenance page tabulates, plus the two lowest-coverage ones (23% and 20%) read afterwards because a rule inferred from the middle of a distribution is not a rule about its tail. All 62 are listed on the provenance page. 
-  * **Recall.** Four schema-side probes that were //not// used to build the pool were run against it afterwards. They surfaced 73 papers outside it; all were read at title level, the twelve from the ''email-message'' enum were opened, and exactly **one** was a genuine miss — Starov et al. {[starov2016_sure]}, now the oldest paper in the address-as-identifier slice. Read that as the measured cost of a title-and-density pool: about one paper in seventy, in the direction of missing the oldest work in a slice.+  * **Recall.** Four schema-side probes that were //not// used to build the pool were run against it afterwards. Between them they produced **72** outside-the-pool mentions, which are **67 distinct papers** (five appear on more than one probe's list, so the four columns double-count). All 67 were read at title level, the twelve from the ''email-message'' enum were opened, and exactly **one** was a genuine miss — Starov et al. {[starov2016_sure]}, now the oldest paper in the address-as-identifier slice. Read that as the measured cost of a title-and-density pool: one miss in 67 papers re-screened, in the direction of missing the oldest work in a slice.
   * **Venue coverage.** Seven venues. EuroS&P, ACSAC, RAID, AsiaCCS, WPES, CHI and SOUPS are absent, and so is the CEAS/anti-spam literature entirely — for a topic with this much of its history in dedicated anti-spam and telephony venues, every count here is a **lower bound and a weaker one than on most pages of this wiki**.   * **Venue coverage.** Seven venues. EuroS&P, ACSAC, RAID, AsiaCCS, WPES, CHI and SOUPS are absent, and so is the CEAS/anti-spam literature entirely — for a topic with this much of its history in dedicated anti-spam and telephony venues, every count here is a **lower bound and a weaker one than on most pages of this wiki**.
   * **Stability.** ''classification.method'' agrees with an independent extraction run on 58% of papers and free-text names on about 20% of exact strings, which is why methods are given as rankings and enum fields as percentages. Those figures were measured on the previous, 4,322-paper run and have not been re-measured.   * **Stability.** ''classification.method'' agrees with an independent extraction run on 58% of papers and free-text names on about 20% of exact strings, which is why methods are given as rankings and enum fields as percentages. Those figures were measured on the previous, 4,322-paper run and have not been re-measured.
privacy/email_tracking.txt · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki