User Tools

Site Tools


privacy:server_side_tracking

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
privacy:server_side_tracking [2026/08/21 06:38] – Opening box: the study statuses changed with the CCS 2026 finding — one preprint, one accepted-not-yet-presented. Authored by Claude karel.kubicek.claudeprivacy:server_side_tracking [2026/08/21 08:28] (current) – [Open Questions] karelkubicek
Line 6: Line 6:
  
 <WRAP important> <WRAP important>
-**The one thing to understand before you start: the four published prevalence estimates differ by a factor of a hundred, and almost none of the difference is real.** They range from **0.38%** of visited sites to **38%**, and the spread is produced by what each study defined as SST, what it crawled, and — most of all — whether the crawler did anything on the page. SST fires on //events// (''page_view'', ''scroll'', ''click''); a landing-page-only crawl that never interacts sees the fewest of them. Of the four studies, one is an unrefereed preprint and another is accepted but not yet presented. **Until CCS 2026 is held there is exactly one peer-reviewed SST detection method in the seven venues of this site's [[Literature:Corpus|publication corpus]]** — Fouad et al. {[fouad2024_devil]}, PETS 2024 — and it cannot be re-run, because it requires a crawl made before August 2020. A second, Mertens et al. {[mertens2026_gtm]}, is **accepted at CCS 2026** and is therefore absent from every corpus figure on this page purely because that venue-year has not happened yet. So do not open with "prior work finds SST on X% of sites". Say which definition, which population, and which interaction depth, or the number means nothing.+**The one thing to understand before you start: the four published prevalence estimates differ by a factor of a hundred, and almost none of the difference is real.** They range from **0.38%** of visited sites to **38%**, and the spread is produced by what each study defined as SST, what it crawled, and — plausibly most of all, though nobody has isolated it — whether the crawler did anything on the page. SST fires on //events// (''page_view'', ''scroll'', ''click''); a landing-page-only crawl that never interacts sees the fewest of them. Of the four studies, one is an unrefereed preprint and another is accepted but not yet presented. **Until CCS 2026 is held there is exactly one peer-reviewed SST detection method in the seven venues of this site's [[Literature:Corpus|publication corpus]]** — Fouad et al. {[fouad2024_devil]}, PETS 2024 — and it cannot be re-run, because it requires a crawl made before August 2020. A second, Mertens et al. {[mertens2026_gtm]}, is **accepted at CCS 2026** and is therefore absent from every corpus figure on this page purely because that venue-year has not happened yet. So do not open with "prior work finds SST on X% of sites". Say which definition, which population, and which interaction depth, or the number means nothing.
 </WRAP> </WRAP>
  
Line 13: Line 13:
   * **The Devil is in the Details** {[fouad2024_devil]}, PETS 2024 — the first and still the only peer-reviewed measurement of SST in the wild, plus a legal analysis with a co-author who is a legal scholar. Read it for the detection logic and for the auditing argument, not for the prevalence figure.   * **The Devil is in the Details** {[fouad2024_devil]}, PETS 2024 — the first and still the only peer-reviewed measurement of SST in the wild, plus a legal analysis with a co-author who is a legal scholar. Read it for the detection logic and for the auditing argument, not for the prevalence figure.
   * **Client-side and Server-side Tracking on Meta** {[elfraihi2024_meta]}, PETS 2024 — the complementary experiment: rather than detecting SST from the outside, the authors //deployed// the Meta Pixel and the Conversions API themselves and drove 2,400 recruited users through them. It is the only study that can say how well SST actually works, and the answer is "well enough, and less accurately": CAPI matched **34%–51%** of visitors to Meta profiles against the Pixel's **42%–61%**, but **fewer than 65%** of CAPI's matches were the right person against the Pixel's 100%.   * **Client-side and Server-side Tracking on Meta** {[elfraihi2024_meta]}, PETS 2024 — the complementary experiment: rather than detecting SST from the outside, the authors //deployed// the Meta Pixel and the Conversions API themselves and drove 2,400 recruited users through them. It is the only study that can say how well SST actually works, and the answer is "well enough, and less accurately": CAPI matched **34%–51%** of visitors to Meta profiles against the Pixel's **42%–61%**, but **fewer than 65%** of CAPI's matches were the right person against the Pixel's 100%.
-  * **SST-Guard** {[jazlan2026_sstguard]} — an **arXiv preprint of 30 April 2026, not peer-reviewed** (checked 2026-08-21: v1 only, no ''journal_ref''). It was shown as a **poster at PETS 2026** on 22 July 2026, which is a venue appearance but not a review: the call for posters says proposals are "lightly reviewed for relevance to PETS and adherence to formatting guidelines" and "Posters will not be peer-reviewed".((''petsymposium.org/2026/cfposters.php'' and ''.../accepted-posters.php'', both fetched 2026-08-21.))Read it because it is the current state of the detection art and because its artefacts are public; read the audit in [[#Auditing SST-Guard against its own released artefacts]] before you reuse either.+  * **SST-Guard** {[jazlan2026_sstguard]} — an **arXiv preprint of 30 April 2026, not peer-reviewed** (checked 2026-08-21: v1 only, no ''journal_ref''). It was shown as a **poster at PETS 2026** on 22 July 2026, which is a venue appearance but not a review: the call for posters says proposals are "lightly reviewed for relevance to PETS and adherence to formatting guidelines" and "Posters will not be peer-reviewed".((''petsymposium.org/2026/cfposters.php'' and ''.../accepted-posters.php'', both fetched 2026-08-21.)) Read it because it is the current state of the detection art and because its artefacts are public; read the audit in [[#Auditing SST-Guard against its own released artefacts]] before you reuse either.
   * **SoK: Advances and Open Problems in Web Tracking** {[vekaria2025_soktracking]} for where SST sits in the wider tracking literature, and **The Bitter Pill** {[moti2025_bitterpill]} (DPM 2025) for what SST looks like in one high-monetisation vertical measured by hand.   * **SoK: Advances and Open Problems in Web Tracking** {[vekaria2025_soktracking]} for where SST sits in the wider tracking literature, and **The Bitter Pill** {[moti2025_bitterpill]} (DPM 2025) for what SST looks like in one high-monetisation vertical measured by hand.
  
Line 21: Line 21:
  
 ^ Study ^ Status ^ Population, vantage, interaction ^ Definition detected ^ Result ^ ^ Study ^ Status ^ Population, vantage, interaction ^ Definition detected ^ Result ^
-| Fouad et al. {[fouad2024_devil]} | **Peer-reviewed**, PETS 2024 | Alexa top 10,000 → **7,367** sites visited in all crawls; EU vantage; OpenWPM + Firefox; **stateless**; **no banner interaction**; **landing page only** | //First-party// SST: a subdomain that (1) appeared only after Aug 2020, (2) resolves to a different WHOIS organisation, (3) sets or receives an ID cookie, and (4) receives parameters or cookies that two or more now-absent trackers used to receive | **28 sites (0.38%)** deploy SST. **389 sites (5.28%)** carry a cloaked tracker |+| Fouad et al. {[fouad2024_devil]} | **Peer-reviewed**, PETS 2024 | Alexa top 10,000 → **7,367** sites visited in all crawls; EU vantage; OpenWPM + Firefox; **stateless**; **no banner interaction**; **landing page only** | //First-party// SST: a subdomain that (1) appeared only after Aug 2020, (2) resolves to a different WHOIS organisation, (3) sets or receives an ID cookie, and (4) receives parameters or cookies that two or more now-absent trackers used to receive | **28 sites (0.38%)** deploy SST. **389 sites (5.28%)** carry a //cloaked tracker// — conditions (1)–(3) without (4), i.e. a third party hidden behind a first-party subdomain that handles an ID cookie but with no evidence that anything moved |
 | Moti et al. {[moti2025_bitterpill]} | Peer-reviewed workshop (DPM 2025), **outside these seven venues** | **50** EU pharmacy sites, five countries; **manual shopping for a pregnancy test**; accept and reject conditions | Requests carrying any of **36** parameters common to GTM-initiated requests, confirmed by WHOIS organisation mismatch on the first-party subdomain | **19 of 50 (38%)**; 12 of the 19 SST servers hosted by Google | | Moti et al. {[moti2025_bitterpill]} | Peer-reviewed workshop (DPM 2025), **outside these seven venues** | **50** EU pharmacy sites, five countries; **manual shopping for a pregnancy test**; accept and reject conditions | Requests carrying any of **36** parameters common to GTM-initiated requests, confirmed by WHOIS organisation mismatch on the first-party subdomain | **19 of 50 (38%)**; 12 of the 19 SST servers hosted by Google |
 | Mertens et al. {[mertens2026_gtm]} | **Accepted at ACM CCS 2026**, to appear; the only text available today is the HAL preprint of 2026-02-06 | **80,000** popular sites worldwide | Server-side GTM identified by **template signatures** derived from the GTM template marketplace | sGTM on **~3,000 sites (~3.8%)**; 398 server-side Tag instances | | Mertens et al. {[mertens2026_gtm]} | **Accepted at ACM CCS 2026**, to appear; the only text available today is the HAL preprint of 2026-02-06 | **80,000** popular sites worldwide | Server-side GTM identified by **template signatures** derived from the GTM template marketplace | sGTM on **~3,000 sites (~3.8%)**; 398 server-side Tag instances |
Line 28: Line 28:
 Three things follow, and they are the reusable part of this table. Three things follow, and they are the reusable part of this table.
  
-  * **The ordering of the results is the ordering of the interaction depth.** Landing page, no clicks → 0.38%. Scroll and click and one subpage → 4.9%. A human shopping for a product → 38%. SST reports events; if your crawler generates no events past ''page_view'', you are measuring the floor. This is the single largest methodological lever on this page and nobody has isolated it — see [[#Open Questions]].+  * **The ordering of the results is the ordering of the interaction depth.** Landing page, no clicks → 0.38%. Scroll and click and one subpage → 4.9%. A human shopping for a product → 38%. SST reports events; if your crawler generates no events past ''page_view'', you are probably measuring the floor. **This is an ordering argument, not a measurement, and it is confounded**: Moti et al.'s 38% is also the loosest definition, the latest year, and 50 hand-picked sites in the single most tracking-heavy vertical on the list. Interaction depth is the most plausible driver and by far the cheapest to test — see [[#Open Questions]], where it is the first entry.
   * **Fouad et al. and SST-Guard are not measuring the same object.** Fouad requires evidence that trackers //moved// — a shift visible only against a pre-2020 baseline — and so reports a deliberate, conservative lower bound: "we only detect a subset of the servers partaking in SST". SST-Guard requires only that a Google Analytics artefact reaches a non-Google endpoint. **0.38% and 4.9% are consistent with each other.**   * **Fouad et al. and SST-Guard are not measuring the same object.** Fouad requires evidence that trackers //moved// — a shift visible only against a pre-2020 baseline — and so reports a deliberate, conservative lower bound: "we only detect a subset of the servers partaking in SST". SST-Guard requires only that a Google Analytics artefact reaches a non-Google endpoint. **0.38% and 4.9% are consistent with each other.**
   * **Denominator discipline.** SST-Guard's own headline, 4.21%, divides by the 150,000-name list rather than by the 128,222 domains it actually classified. Over the crawled population the figure is 4.92%. Use the population you measured.   * **Denominator discipline.** SST-Guard's own headline, 4.21%, divides by the 150,000-name list rather than by the 128,222 domains it actually classified. Over the crawled population the figure is 4.92%. Use the population you measured.
Line 42: Line 42:
 | 2022–2024 | **Temporal shift plus organisation cloaking**: diff a pre-SST crawl against a post-SST crawl, keep subdomains that are new, WHOIS-mismatched, ID-bearing, and that inherited parameters from vanished trackers | {[fouad2024_devil]} | **Historical, and not reproducible.** It is the only peer-reviewed SST detector in these seven venues and its logic is still the clearest published statement of what SST //is//. But step one needs a crawl predating August 2020. **You cannot run this method today** unless you happen to own a 2020 crawl of your population. Cite it for the definition; do not plan a study around it | | 2022–2024 | **Temporal shift plus organisation cloaking**: diff a pre-SST crawl against a post-SST crawl, keep subdomains that are new, WHOIS-mismatched, ID-bearing, and that inherited parameters from vanished trackers | {[fouad2024_devil]} | **Historical, and not reproducible.** It is the only peer-reviewed SST detector in these seven venues and its logic is still the clearest published statement of what SST //is//. But step one needs a crawl predating August 2020. **You cannot run this method today** unless you happen to own a 2020 crawl of your population. Cite it for the definition; do not plan a study around it |
 | 2025 | **Parameter-intersection heuristic**, history-free: take the parameter names common to GTM-initiated requests, look for them on first-party endpoints, confirm with WHOIS | {[moti2025_bitterpill]} | **Current, cheap, and the one to start with** if you need an SST flag as a side quantity in a larger study. It needs only request logs plus DNS and WHOIS. Its cost is measured: SST-Guard replicated it with the 36 parameters obtained from those authors and recovered **129 of 403** ground-truth domains, **32%** accuracy, against SST-Guard's 401 and 99.5% | | 2025 | **Parameter-intersection heuristic**, history-free: take the parameter names common to GTM-initiated requests, look for them on first-party endpoints, confirm with WHOIS | {[moti2025_bitterpill]} | **Current, cheap, and the one to start with** if you need an SST flag as a side quantity in a larger study. It needs only request logs plus DNS and WHOIS. Its cost is measured: SST-Guard replicated it with the 36 parameters obtained from those authors and recovered **129 of 403** ground-truth domains, **32%** accuracy, against SST-Guard's 401 and 99.5% |
-| 2026 | **Vendor template signatures**: enumerate the sGTM template marketplace and fingerprint each template's client-side artefacts | {[mertens2026_gtm]} | **Current, and about to become the only peer-reviewed runnable method** — it is accepted at CCS 2026. Narrow by construction, though: it covers sGTM only, so it misses ''gtag.js'' deployments and any custom container, and it can only recognise templates that are on the marketplace. SST-Guard declines to compare against it, on the stated grounds of low dataset overlap and no available validation | +| 2026 | **Vendor template signatures**: enumerate the sGTM template marketplace and fingerprint each template's client-side artefacts | {[mertens2026_gtm]} | **Current, and about to become the only runnable method published at a main venue** — it is accepted at CCS 2026 (Moti et al.'s is peer-reviewed too, but at a workshop). Narrow by construction, though: it covers sGTM only, so it misses ''gtag.js'' deployments and any custom container, and it can only recognise templates that are on the marketplace. SST-Guard declines to compare against it, on the stated grounds of low dataset overlap and no available validation | 
-| 2026 | **Value templates across three browser modalities** — regexes matched against the //values// in request parameters, first-party cookies and ''window'' variables, one logistic regression per modality plus a meta-classifier | {[jazlan2026_sstguard]} | **The state of the art, and not peer-reviewed.** It is the right idea: the reporting endpoint is attacker-controlled but the semantics of what GA collects are not, because Google cannot rename ''dataLayer'' without breaking millions of sites. Its measured numbers are the best published. **Its released artefacts do not reproduce its own figures** — see the next section before you build on it |+| 2026 | **Value templates across three browser modalities** — regexes matched against the //values// in request parameters, first-party cookies and ''window'' variables, one logistic regression per modality plus a meta-classifier | {[jazlan2026_sstguard]} | **The state of the art, and not peer-reviewed.** It is the right idea: the reporting endpoint is attacker-controlled but the semantics of what GA collects are not, because Google cannot rename ''dataLayer'' without breaking millions of sites. Its measured numbers are the best published. **Its released //extension// does not reproduce its released //feature data//, and two of its headline splits cannot be re-derived from the release at all** — see the next section before you build on it |
 | — | **Taint tracking from DOM event to network request** | nobody | **The open problem**, and SST-Guard names it as such: "the more a tracker decouples client-side collection from server-side reporting, the less our approach has to match against". It is also the only approach that would generalise past Google | | — | **Taint tracking from DOM event to network request** | nobody | **The open problem**, and SST-Guard names it as such: "the more a tracker decouples client-side collection from server-side reporting, the less our approach has to match against". It is also the only approach that would generalise past Google |
  
Line 50: Line 50:
   * **Matching the third-party hostname, with or without a filter list, as a definition of tracking.** This is the whole point of SST. Under SST the tracking request //is// first-party.   * **Matching the third-party hostname, with or without a filter list, as a definition of tracking.** This is the whole point of SST. Under SST the tracking request //is// first-party.
   * **Requiring a CNAME record.** 73.51% of subdomain-based sGA uses A/AAAA {[jazlan2026_sstguard]}. A CNAME-only detector reports roughly a quarter of the deployments and reports it as the total.   * **Requiring a CNAME record.** 73.51% of subdomain-based sGA uses A/AAAA {[jazlan2026_sstguard]}. A CNAME-only detector reports roughly a quarter of the deployments and reports it as the total.
-  * **Assuming SST lives on a distinct subdomain.** **18.4%** of SST-Guard's detections route through a path on the main domain (''example.com/collect''), which no subdomain-based or DNS-based method can see — SST-Guard says so plainly and excludes them from its own network analysis. And this is the direction of travel: **Google's current documentation recommends same-origin serving first**, ahead of the subdomain option.((''developers.google.com/tag-platform/tag-manager/server-side/custom-domain'', read 2026-08-21, page's own "Last updated 2025-06-27": "Same-origin serving is a best practice that lets you leverage the security and durability benefits of server-set cookies" and "To unlock the benefits of a first-party context, such as more durable cookies, your tagging server and your website have to run on the same domain." The setup guide adds: "Make sure to host your tagging server in the same origin (best practice) or as a subdomain of your current website" (''.../manual-setup-guide'', "Last updated 2026-05-08"). These pages render client-side, so ''curl'' returns a shell; the quotes were read from the rendered page and only the HTTP 200 is checked by the provenance script.)) Design for path-based deployments now, not later. +  * **Assuming SST lives on a distinct subdomain.** **18.4%** of SST-Guard's detections route through a path on the main domain (''example.com/collect''), which no subdomain-based or DNS-based method can see — SST-Guard says so plainly and excludes them from its own network analysis. That 18.4% is the paper's own figure and **is not re-derivable from its release**: three defensible readings of the released columns put path-based deployments anywhere from 4% to 46% (see the audit below). Treat it as an order of magnitude, not a planning constant. And this is the direction of travel: **Google's current documentation recommends same-origin serving first**, ahead of the subdomain option.((''developers.google.com/tag-platform/tag-manager/server-side/custom-domain'', read 2026-08-21, page's own "Last updated 2025-06-27": "Same-origin serving is a best practice that lets you leverage the security and durability benefits of server-set cookies" and "To unlock the benefits of a first-party context, such as more durable cookies, your tagging server and your website have to run on the same domain." The setup guide adds: "Make sure to host your tagging server in the same origin (best practice) or as a subdomain of your current website" (''.../manual-setup-guide'', "Last updated 2026-05-08"). These pages render client-side, so ''curl'' returns a shell; the quotes were read from the rendered page and only the HTTP 200 is checked by the provenance script.)) Design for path-based deployments now, not later. 
-  * **"Third-party cookies are going away, so SST is about to matter."** They are not. Google announced on **2025-04-22** that it would keep third-party cookie choice in Chrome and would not ship the planned prompt.((Anthony Chavez, VP Privacy Sandbox, "Privacy Sandbox: Next steps", ''privacysandbox.google.com/blog/privacy-sandbox-next-steps'', 2025-04-22: "we've made the decision to maintain our current approach to offering users third-party cookie choice in Chrome, and will not be rolling out a new standalone prompt for third-party cookies." Checked 2026-08-21.)) Fouad et al. wrote in 2024 that "the end of third-party cookies planned for 2025 is having severe ramifications", and that premise is now false — **while SST adoption grew anyway**. Filter-list evasion and signal loss to browser defences, not the cookie deprecation, are the live motivation. Do not reproduce the 2024 framing. +  * **"Third-party cookies are going away, so SST is about to matter."** They are not. Google announced on **2025-04-22** that it would keep third-party cookie choice in Chrome and would not ship the planned prompt.((Anthony Chavez, VP Privacy Sandbox, "Privacy Sandbox: Next steps", ''privacysandbox.google.com/blog/privacy-sandbox-next-steps'', 2025-04-22: "we've made the decision to maintain our current approach to offering users third-party cookie choice in Chrome, and will not be rolling out a new standalone prompt for third-party cookies." Checked 2026-08-21.)) Fouad et al. wrote in 2024 that "the end of third-party cookies planned for 2025 is having severe ramifications", and that premise is now false. SST did not go away with it — **later studies, under broader definitions, find far more of it than the 2024 one did** — but see [[#Open Questions]]: there is no time series, so nobody can say it //grew//. Filter-list evasion and signal loss to browser defences, not the cookie deprecation, are the live motivation. Do not reproduce the 2024 framing. 
-  * **"Filter lists cannot see SST."** Too strong as of today. EasyPrivacy has evolved specific rules for it, and SST-Guard measures them blocking **93.50%** of the 40,199 sGA requests it found. Verified live on 2026-08-21 (EasyPrivacy ''Version: 202608210608'', commit ''0ad9e733''): the list carries ''?v=2&tid=G-$~third-party'' — note the ''$~third-party'' modifier, which exists precisely to catch first-party tracking — plus ''&sst.sw_exp='' and ''&sst.gcsub='' for server-side container parameters, and per-site rules including ''||mstm.motorsport.com^'' and nine hosts beginning ''sst''. What lists cannot see is the **customised** tail: 166 domains (2.62%) where nothing was blocked, and 223 domains sending base64-encoded payloads.+  * **"Filter lists cannot see SST."** Too strong as of today. EasyPrivacy has evolved specific rules for it, and SST-Guard measures them blocking **93.50%** of the 40,199 sGA requests it reports finding (40,198 in the released file — see the audit). Verified live on 2026-08-21 (EasyPrivacy ''Version: 202608210644'', commit ''94b83d3b'' — the version the provenance page's committed script output shows; it had already moved three times in the half hour before this sentence was written, which is the point): the list carries ''?v=2&tid=G-$~third-party'' — note the ''$~third-party'' modifier, which exists precisely to catch first-party tracking — plus ''&sst.sw_exp='' and ''&sst.gcsub='' for server-side container parameters, and per-site rules including ''||mstm.motorsport.com^'' and nine hosts beginning ''sst''. What lists cannot see is the **customised** tail: 166 domains (2.62%) where nothing was blocked, and 223 domains sending base64-encoded payloads.
  
 ===== Auditing SST-Guard against its own released artefacts ===== ===== Auditing SST-Guard against its own released artefacts =====
  
-The SST-Guard repository((''github.com/jazlan01/sst-guard'', cloned 2026-08-21 at commit ''9e013d4'' of 2026-04-29 — the only commit.)) publishes the ground truth, the classifier output on the ground truth, the 40k-row request-level output of the 150K crawl, the list of 6,314 detected domains, and a packaged Chrome extension. **It does not publish the pipeline that produced the paper's numbers.** The extension is the only code, and it ships minified.+The SST-Guard repository((''github.com/jazlan01/sst-guard'', cloned 2026-08-21 at commit ''9e013d4'' of 2026-04-30 (UTC) — the only commit.)) publishes the ground truth, the classifier output on the ground truth, the 40k-row request-level output of the 150K crawl, the list of 6,314 detected domains, and a packaged Chrome extension. **It does not publish the pipeline that produced the paper's numbers.** The extension is the only code, and it ships minified.
  
-The detection heuristics that decide whether a request is sGA are therefore in the bundle, not in the paper, so they were extracted from it and exercised. Two scripts do this and their unedited output is on [[provenance:privacy:server_side_tracking|the provenance page]]. **This is a source audit, not a reproduction attempt**: no crawl was re-run.+The detection heuristics that decide whether a request is sGA are therefore in the bundle, not in the paper, so they were extracted from it and exercised. Three scripts do this and their unedited output is on [[provenance:privacy:server_side_tracking|the provenance page]]. **This is a source audit, not a reproduction attempt**: no crawl was re-run.
  
 ==== Which signals a reader can actually reproduce ==== ==== Which signals a reader can actually reproduce ====
Line 64: Line 64:
 ^ Signal ^ Reproducible from a crawl you can run? ^ ^ Signal ^ Reproducible from a crawl you can run? ^
 | **''window'' variables** — ''dataLayer'' event shapes, ''gaGlobal[hid]'', ''gaGlobal[from_cookie]'', ''google_tag_data'' container ID | **Yes, and this is the modality to build on.** It is the paper's own strongest argument and it holds: Google cannot rename ''dataLayer'' or ''google_tag_data'' without breaking every publisher integration built on them. Note this modality needs a ''window'' snapshot, which most crawlers do not take | | **''window'' variables** — ''dataLayer'' event shapes, ''gaGlobal[hid]'', ''gaGlobal[from_cookie]'', ''google_tag_data'' container ID | **Yes, and this is the modality to build on.** It is the paper's own strongest argument and it holds: Google cannot rename ''dataLayer'' or ''google_tag_data'' without breaking every publisher integration built on them. Note this modality needs a ''window'' snapshot, which most crawlers do not take |
-| **''_ga'' / ''_ga_X'' / ''_gid'' cookie patterns**, and the network ''cid'' | **Yes — until 2027-01-15.** All five shipped cookie templates and the ''cid'' template hard-code the literal ''17'' prefix of a ten-digit Unix timestamp (''/^GA1\.[123](-2)?\.[0-9]{6,10}\.17[0-9]{8,13}$/''). That prefix covers 2023-11-14 to 2027-01-15 and nothing after it. Verified by running the shipped regexes against synthetic values: ''t=1799999999'' matches, ''t=1800000000'' does not |+| **''_ga'' / ''_ga_X'' / ''_gid'' cookie patterns**, and the network ''cid'' | **Yes — until 2027-01-15.** All five shipped cookie templates and the ''cid'' template hard-code the literal ''17'' prefix of a ten-digit Unix timestamp (''%%/^GA1\.[123](-2)?\.[0-9]{6,10}\.17[0-9]{8,13}$/%%''). That prefix covers 2023-11-14 to 2027-01-15 and nothing after it. Verified by running the shipped regexes against synthetic values: ''t=1799999999'' matches, ''t=1800000000'' does not |
 | **''window'' ''chrome_version'', ''platform_version'', ''architecture'', ''bitness''** | **No.** These are literals: ''%%/"144\.0\.7559\.97"/%%'', ''%%/"26\.2\.0"/%%'', ''%%/"arm"/%%'', ''%%/"64"/%%''. They fire only for one Chrome build on one Apple-Silicon machine and match nothing on your crawler | | **''window'' ''chrome_version'', ''platform_version'', ''architecture'', ''bitness''** | **No.** These are literals: ''%%/"144\.0\.7559\.97"/%%'', ''%%/"26\.2\.0"/%%'', ''%%/"arm"/%%'', ''%%/"64"/%%''. They fire only for one Chrome build on one Apple-Silicon machine and match nothing on your crawler |
 | **Network ''uap'', ''uapv'', ''uaa''** | **No.** ''/Linux/'', ''/5\.15\.0/'', ''/x86/'' — that is the paper's own crawl host (Ubuntu 22.04, Linux 5.15.0). Note these pin a //different// machine from the ''window'' templates above, so the shipped extension carries two mutually inconsistent environment fingerprints | | **Network ''uap'', ''uapv'', ''uaa''** | **No.** ''/Linux/'', ''/5\.15\.0/'', ''/x86/'' — that is the paper's own crawl host (Ubuntu 22.04, Linux 5.15.0). Note these pin a //different// machine from the ''window'' templates above, so the shipped extension carries two mutually inconsistent environment fingerprints |
-| **Network ''dl'' (document location)** | **No — the shipped regex is broken.** ''/https:\/\/[\^\s&#]+/'' is a character class containing an //escaped literal caret//, not a negation, so it matches ''https://'' followed only by ''^'', whitespace, ''&'' or ''#''. It matches no ordinary URL. It is 1 in **98.4%** of the authors' own published rows and 0.0% under the released extension |+| **Network ''dl'' (document location)** | **No — the shipped regex is broken.** ''%%/https:\/\/[\^\s&#]+/%%'' is a character class containing an //escaped literal caret//, not a negation, so it matches ''https://'' followed only by ''%%^%%'', whitespace, ''&'' or ''#''. It matches no ordinary URL. It is 1 in **98.4%** of the authors' own published rows and 0.0% under the released extension |
 | **Network ''uafvl''** | **No.** 99.5% in the published data, 0.0% under the shipped regex | | **Network ''uafvl''** | **No.** 99.5% in the published data, 0.0% under the shipped regex |
 | **Network ''sid'' ''/\d{10}/'', ''_p'' ''/\d{13}/'', ''uab'' ''/64/'', ''tfd'' ''/\d{3,4}/'', ''_eu'', ''ul''** | **Yes, but they carry no information.** ''1234567890'' sets ''sid'' and ''tfd''; ''x86_64'' sets ''uab''; ''de'' sets ''_eu'' and ''ul''. Any HTTP status code matches ''tfd''. These are consistent with the paper's own 73.36% request-level precision on its training labels | | **Network ''sid'' ''/\d{10}/'', ''_p'' ''/\d{13}/'', ''uab'' ''/64/'', ''tfd'' ''/\d{3,4}/'', ''_eu'', ''ul''** | **Yes, but they carry no information.** ''1234567890'' sets ''sid'' and ''tfd''; ''x86_64'' sets ''uab''; ''de'' sets ''_eu'' and ''ul''. Any HTTP status code matches ''tfd''. These are consistent with the paper's own 73.36% request-level precision on its training labels |
Line 74: Line 74:
 ==== Replaying the shipped extractor over the published rows ==== ==== Replaying the shipped extractor over the published rows ====
  
-The extension's extractor tests each template against the query parameter of the //same name//. Replaying it over the 40,198 published request URLs and diffing against the published feature columns gives **79.3%** per-cell agreement. A variant that tests each template against //every// parameter value gives **83.1%**, and reaches **100%** for 12 of the 23 features. The residue is exactly the rows above''dl'' and ''uafvl'' (shipped regex dead), ''gcs'' / ''tcfd'' / ''ep.user_agent'' (published column dead), and ''gtm'' at 91.5%.+The extension's extractor tests each template against the query parameter of the //same name//. Replaying it over the 40,198 published request URLs and diffing against the published feature columns gives **79.3%** per-cell agreement. A variant that tests each template against //every// parameter value gives **83.1%**, and reaches **100%** for **15** of the 23 features — 14 of them exactly, the fifteenth (''uapv'') at 99.9975%, which the report rounds to 100.0%. The residue is the rows above — ''dl'' and ''uafvl'' (shipped regex dead), ''gcs'' / ''tcfd'' / ''ep.user_agent'' (published column dead), ''gtm'' at 91.5% — plus two near-misses, ''_eu'' at 99.8% and ''_gid'' at 98.9%.
  
 **Read that as: the offline pipeline is not the shipped extension, and the offline pipeline is not released.** That is a normal state of affairs for a preprint artefact and it is not evidence against the paper's conclusions — the ''window''-variable insight in particular is unaffected. It does mean you cannot reuse SST-Guard as a black box and report its accuracy as yours. **Read that as: the offline pipeline is not the shipped extension, and the offline pipeline is not released.** That is a normal state of affairs for a preprint artefact and it is not evidence against the paper's conclusions — the ''window''-variable insight in particular is unaffected. It does mean you cannot reuse SST-Guard as a black box and report its accuracy as yours.
Line 86: Line 86:
 | ''/g/collect'' on 71.94% of sGA endpoints, ''*/collect'' on 79.97% | **Close, not exact: 72.71% and 82.23%** of the 6,314 domains | | ''/g/collect'' on 71.94% of sGA endpoints, ''*/collect'' on 79.97% | **Close, not exact: 72.71% and 82.23%** of the 6,314 domains |
 | 81.59% subdomain-based / 18.4% path-based | **Not reproducible, because the rule is not stated.** Three defensible readings of the released columns give **53.67%**, **85.92%** and **95.91%** | | 81.59% subdomain-based / 18.4% path-based | **Not reproducible, because the rule is not stated.** Three defensible readings of the released columns give **53.67%**, **85.92%** and **95.91%** |
-| "21.05% [of subdomain deployments] use CNAME cloaking" (abstract) | **Inconsistent with the body**, which says 1,365 of 5,152 subdomain deployments, **26.49%**. 21.05% is 1,329 divided by all 6,314 detected domains, not by the subdomain ones |+| "21.05% [of subdomain deployments] use CNAME cloaking" (§1 and §7.3) | **Inconsistent with §6.2**, which says 1,365 of 5,152 subdomain deployments, **26.49%**. 21.05% is 1,329 divided by all 6,314 detected domains, not by the subdomain ones |
 | README names ''ground-truth-labels.csv'' | No such file; the repository ships ''sst-guard-output.csv'' | | README names ''ground-truth-labels.csv'' | No such file; the repository ships ''sst-guard-output.csv'' |
  
Line 96: Line 96:
  
   * **[[Privacy:Requests|Request classification]] loses its most reliable label source.** The current best practice for deciding what a cookie or a request is for is the //provenance// of the thing that created it — a cookie set by a resource a filter list blocked as advertising is an advertising cookie. Under SST there is no blocked advertising resource: there is one first-party endpoint that received data destined for an unknown number of platforms. SST-Guard's payload analysis makes this concrete — requests to sGA endpoints carry ''ep.fb_event_name'' and ''ep.event_id'', i.e. Meta's Conversions API riding in the same container.   * **[[Privacy:Requests|Request classification]] loses its most reliable label source.** The current best practice for deciding what a cookie or a request is for is the //provenance// of the thing that created it — a cookie set by a resource a filter list blocked as advertising is an advertising cookie. Under SST there is no blocked advertising resource: there is one first-party endpoint that received data destined for an unknown number of platforms. SST-Guard's payload analysis makes this concrete — requests to sGA endpoints carry ''ep.fb_event_name'' and ''ep.event_id'', i.e. Meta's Conversions API riding in the same container.
-  * **[[Privacy:Cookies|Cookie classification]] mislabels by construction.** A ''_ga'' cookie set by ''analytics.example.com'' looks like a first-party analytics cookie and will be classified as one. Fouad et al. looked their SST cookies up in Cookiepedia and **35%** came back as Targeting/Advertising. Worse, they found **119** cloaked trackers receiving ID cookies set by **91 distinct** other third-party domains — a Same-Origin Policy bypass, because a first-party subdomain can read cookies its operator never set.+  * **[[Privacy:Cookies|Cookie classification]] mislabels by construction.** A ''_ga'' cookie set by ''analytics.example.com'' looks like a first-party analytics cookie and will be classified as one. Fouad et al. looked up their SST cookies in Cookiepedia and **35%** came back as Targeting/Advertising. Worse, they found **119** cloaked trackers receiving ID cookies set by **91 distinct** other third-party domains — a Same-Origin Policy bypass, because a first-party subdomain can read cookies its operator never set.
   * **Consent measurement gets an artefact you must declare.** The two main studies sit at opposite extremes and it changes what their numbers mean. Fouad et al. "did not interact with cookie banners on the visited [websites]", so every flow they report is **pre-consent** — which is what makes their legal analysis possible. SST-Guard drove Consent-O-Matic to **accept everything**, which maximises what fires and says nothing about lawfulness. Neither is wrong; reporting a prevalence without saying which one you did is.   * **Consent measurement gets an artefact you must declare.** The two main studies sit at opposite extremes and it changes what their numbers mean. Fouad et al. "did not interact with cookie banners on the visited [websites]", so every flow they report is **pre-consent** — which is what makes their legal analysis possible. SST-Guard drove Consent-O-Matic to **accept everything**, which maximises what fires and says nothing about lawfulness. Neither is wrong; reporting a prevalence without saying which one you did is.
   * **Legal and compliance work loses its evidence.** Fouad et al., with a legal scholar, argue SST infringes both the GDPR and the ePrivacy Directive, and the operative reason is evidentiary rather than doctrinal: an auditor cannot see which third parties received the data, cannot attribute a purpose to a cookie whose setter is hidden, and a first-party cookie set by a first-party subdomain invites the (wrong) inference that it is strictly necessary and consent-exempt. See [[Practices:Legal enforcement]] for who to report to.   * **Legal and compliance work loses its evidence.** Fouad et al., with a legal scholar, argue SST infringes both the GDPR and the ePrivacy Directive, and the operative reason is evidentiary rather than doctrinal: an auditor cannot see which third parties received the data, cannot attribute a purpose to a cookie whose setter is hidden, and a first-party cookie set by a first-party subdomain invites the (wrong) inference that it is strictly necessary and consent-exempt. See [[Practices:Legal enforcement]] for who to report to.
Line 117: Line 117:
   - **Your population and your denominator.** The list, its version, how many names you attempted and how many you successfully classified. Not the list length.   - **Your population and your denominator.** The list, its version, how many names you attempted and how many you successfully classified. Not the list length.
   - **Interaction depth, consent action, statefulness, vantage.** Without all four the prevalence is uninterpretable. These are the four fields that the corpus shows most often go unreported for crawls in general (see [[Literature:Corpus]]).   - **Interaction depth, consent action, statefulness, vantage.** Without all four the prevalence is uninterpretable. These are the four fields that the corpus shows most often go unreported for crawls in general (see [[Literature:Corpus]]).
-  - **Whether you counted path-based deployments**, and if not, that your figure excludes roughly a fifth of them.+  - **Whether you counted path-based deployments**, and if not, that your figure excludes roughly a fifth of them //by the only published count//, which is itself not reproducible.
   - **Whether you required a CNAME.** If yes, your figure is roughly a quarter of the DNS-visible deployments.   - **Whether you required a CNAME.** If yes, your figure is roughly a quarter of the DNS-visible deployments.
   - **Which tracking platform.** "SST" measured through Google Analytics artefacts is not "SST". Meta CAPI, TikTok Events API and any Measurement-Protocol deployment leave different artefacts or none.   - **Which tracking platform.** "SST" measured through Google Analytics artefacts is not "SST". Meta CAPI, TikTok Events API and any Measurement-Protocol deployment leave different artefacts or none.
Line 127: Line 127:
 Two independent passes over the [[Literature:Corpus|publication corpus]] — 7 venues (CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P), 2010–2026. Every figure below names its own population; the script and its unedited output are on [[provenance:privacy:server_side_tracking|the provenance page]]. Two independent passes over the [[Literature:Corpus|publication corpus]] — 7 venues (CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P), 2010–2026. Every figure below names its own population; the script and its unedited output are on [[provenance:privacy:server_side_tracking|the provenance page]].
  
-**Pass A — the structured extraction.** Of the **5,859** papers with an extraction record, **two** name server-side tracking in their title, and both are PETS 2024: {[fouad2024_devil]} and {[elfraihi2024_meta]}. The extraction schema has no field for SST, and a deliberately wide keyword sweep of ''detection[]'' tuples for "cname" or "server-side" returns **63** papers, of which only **7** concern web tracking at all — the other **56** are unrelated server-side security work (server-side malware, phishing cloaking, TCP injection, CDN CNAME chains) and are printed in full on the provenance page. **The dataset cannot answer this page's question on its own; the full text had to be probed.**+**Pass A — the structured extraction.** Of the **5,859** papers with an extraction record, **two** name server-side tracking in their title, and both are PETS 2024: {[fouad2024_devil]} and {[elfraihi2024_meta]}. The extraction schema has no field for SST, and a deliberately wide keyword sweep of ''detection[]'' tuples for "cname" or "server-side" returns **63** papers, of which only **7** concern web tracking by the predicate — and on inspection only **5**, of which **2** actually study SST — the other **56** are unrelated server-side security work (server-side malware, phishing cloaking, TCP injection, CDN CNAME chains) and are printed in full on the provenance page. **The dataset cannot answer this page's question on its own; the full text had to be probed.**
  
 **Pass B — full-text probe.** Denominator: **5,869** papers with a readable ''paper.cols.txt''. End-of-line hyphenation was joined and whitespace collapsed before matching, because ''server-\nside tracking'' otherwise silently fails. **Pass B — full-text probe.** Denominator: **5,869** papers with a readable ''paper.cols.txt''. End-of-line hyphenation was joined and whitespace collapsed before matching, because ''server-\nside tracking'' otherwise silently fails.
Line 135: Line 135:
 | Google's product names (''sGTM'', ''sGA'', server-side Tag Manager/Analytics) | 2 | 0.0% | | Google's product names (''sGTM'', ''sGA'', server-side Tag Manager/Analytics) | 2 | 0.0% |
 | ''Conversions API'' / ''CAPI'' / ''Events API'' | 16 | 0.3% | | ''Conversions API'' / ''CAPI'' / ''Events API'' | 16 | 0.3% |
 +| the same, narrowed to ''%%/Conversions? API/%%'' — the product name, capitalised | **4** | 0.1% |
 | any description of moving collection to the server side | 18 | 0.3% | | any description of moving collection to the server side | 18 | 0.3% |
 | ''CNAME'' cloaking / tracking / redirection | 46 | 0.8% | | ''CNAME'' cloaking / tracking / redirection | 46 | 0.8% |
Line 140: Line 141:
 | ''Measurement Protocol'' | 7 | 0.1% | | ''Measurement Protocol'' | 7 | 0.1% |
  
-**This is one of the quietest topics on this site.** Eleven papers in sixteen years of seven venues use the term, and only one of the eleven detects SST in the wild. By contrast the neighbouring technique, CNAME cloaking, has 46 — and the 105 papers that say ''CNAME'' without saying anything about cloaking are printed in full on the provenance page rather than silently droppedbecause most of them are DNS, CDN and TLS work with no bearing on tracking.+**Read those two API rows together: they are the clearest illustration on this page of how much a probe's width decides its answer.** Widening from ''%%/Conversions? API/%%'' to ''%%/conversions? api|\bCAPI\b|events api/i%%'' takes the count from **4** to **16**, and all twelve extra papers are noise — GitHub'and Android's //Events API//, Microsoft's CryptoAPI, C++ type-conversion APIs, a voice-assistant paper, and a hyphenated "social capi-talists" whose trailing hyphen falls at a column break and so counts as a word boundary. All twelve are listed on the provenance page. **Nothing on this page rests on the wide count**, and the four narrow hits are the Meta paper {[elfraihi2024_meta]} plus three 2026 papers.
  
-**And the trend is not "growing interest".** Papers matching the term, the products or the APIs, by year:+**This is one of the quietest topics on this site.** Eleven papers in seventeen years of seven venues use the term, and only one of the eleven detects SST in the wild. By contrast the neighbouring technique, CNAME cloaking, has 46 — and the 105 papers that say ''CNAME'' without saying anything about cloaking are printed in full on the provenance page rather than silently dropped, because most of them are DNS, CDN and TLS work with no bearing on tracking. 
 + 
 +**And the trend is not "growing interest" — nor is it mostly signal.** Papers matching the narrow term, Google's product names, or the //wide// API probe, by year. Because the wide probe is in the union, the pre-2024 rows are dominated by the false positives just described, so read this as an upper bound on attention rather than a measure of it:
  
 ^ Year ^ Papers ^ Of scanned ^ ^ Year ^ Papers ^ Of scanned ^ ^ Year ^ Papers ^ Of scanned ^ ^ Year ^ Papers ^ Of scanned ^
Line 152: Line 155:
 2010–2018 contributes three hits in total and all three are false positives: a 2010 paper describing an extension that warns about "server-side tracking using web bugs" — the phrase, but not the architecture, which did not exist until 2020 — and two matches on ''CAPI'' meaning Microsoft's CryptoAPI and a hyphenated "social capi-talists". **2026 is provisional and must not be read as a complete year** — CCS 2026 and IMC 2026 have not been held and two other venue-years are incompletely selected. 2010–2018 contributes three hits in total and all three are false positives: a 2010 paper describing an extension that warns about "server-side tracking using web bugs" — the phrase, but not the architecture, which did not exist until 2020 — and two matches on ''CAPI'' meaning Microsoft's CryptoAPI and a hyphenated "social capi-talists". **2026 is provisional and must not be read as a complete year** — CCS 2026 and IMC 2026 have not been held and two other venue-years are incompletely selected.
  
-**What the 2025–2026 papers actually do with SST is cite it as the reason their own method fails.** None measures it. The clearest instance is CookieGuard {[bahrami2025_cookieguard]} (IMC 2025), whose own conclusion reads: "Emerging practices like server-side tracking bypass client-side defenses, including our own CookieGuard […], by proxying exfiltration through seemingly first-party endpoints […]." (the elisions are a section cross-reference and three citation markers) Others — the ''gclid'' study at PETS 2026, the WebView study {[weerasekara2025_webviews]} at PETS 2025, ''PiiXel'' at CCS 2025the localhost web-to-app study at USENIX Security 2026 — cite {[elfraihi2024_meta]} for the proposition that Meta's CAPI bypasses browser defences and move on. **If you are planning an SST measurement, that is your gap: the field has agreed SST defeats client-side defences and has published almost nothing measuring it.**+**What the 2025–2026 papers actually do with SST is cite it as the reason their own method fails.** None measures it. The clearest instance is CookieGuard {[bahrami2025_cookieguard]} (IMC 2025), whose own conclusion reads: "Emerging practices like server-side tracking bypass client-side defenses, including our own CookieGuard […], by proxying exfiltration through seemingly first-party endpoints […]." (the elisions are a section cross-reference and three citation markers) Two 2026 papers cite {[elfraihi2024_meta]} for the proposition that Meta's Conversions API bypasses browser defences and then move on: Clicking into Exposure {[dao2026_gclid]} (PETS 2026)which quotes its 34%–51% matching figure by name, and Bridges to Self {[vlummens2026_bridges]} (USENIX Security 2026)which cites it in a grouped reference. ''PIIxel Leaks'' {[bekos2025_piixel]} (CCS 2025) carries both {[fouad2024_devil]} and {[dao2021_cname]} in its reference list without measuring either. And the closest thing to a measurement in the 2025–2026 slice is not about SST at all: Tracking Without Borders {[weerasekara2025_webviews]} finds CNAME tracking — the neighbouring technique — inside //mobile WebViews//, attributed to Ensighten and Adobe Experience Cloud. **If you are planning an SST measurement, that is your gap: the field has agreed SST defeats client-side defences and has published almost nothing measuring it.**
  
 **Legal assessment.** Of the **402** corpus papers that assess compliance with a law, {[fouad2024_devil]} is the only one that assesses SST, against both the GDPR and the ePrivacy Directive, finding violations of each. **Legal assessment.** Of the **402** corpus papers that assess compliance with a law, {[fouad2024_devil]} is the only one that assesses SST, against both the GDPR and the ePrivacy Directive, finding violations of each.
Line 158: Line 161:
 ===== Open Questions ===== ===== Open Questions =====
  
-<wrap todo>+<WRAP todo>
   * **Nobody has isolated the effect of interaction depth on measured SST prevalence.** It is the largest apparent driver of the 0.38%–38% spread and it is a clean, cheap experiment: one population, one vantage, four interaction depths. This page would most like this done.   * **Nobody has isolated the effect of interaction depth on measured SST prevalence.** It is the largest apparent driver of the 0.38%–38% spread and it is a clean, cheap experiment: one population, one vantage, four interaction depths. This page would most like this done.
-  * **As of today, no peer-reviewed SST detector exists that can be run.** The one refereed method needs a 2020 baseline crawl; of the three runnable methods, one is peer-reviewed but at a workshop outside these venues {[moti2025_bitterpill]}, one is a preprint {[jazlan2026_sstguard]}, and one is accepted at CCS 2026 {[mertens2026_gtm]} but not yet presented or published. **That last one closes this gap when CCS 2026 is held** — check before repeating this sentence.+  * **No SST detector published at a main security or measurement venue can be run today.** The one that was — Fouad et al. at PETS 2024 — needs a 2020 baseline crawl. Of the three runnable methods, one //is// peer-reviewed but at a workshop outside these seven venues {[moti2025_bitterpill]}, one is a preprint {[jazlan2026_sstguard]}, and one is accepted at CCS 2026 {[mertens2026_gtm]} but not yet presented. **That last one closes this gap when CCS 2026 is held** — check before repeating this sentence. An earlier draft of it read "no peer-reviewed SST detector exists that can be run", which was simply wrong about Moti et al.
   * **Nothing detects SST for any platform other than Google.** SST-Guard says so explicitly and names the reason: Meta, TikTok, Snapchat and Reddit ship no debugging tool equivalent to Google Tag Assistant, so there is no ground-truth source to bootstrap from. Meta CAPI is the obvious next target — {[elfraihi2024_meta]} shows it works — and it is undetected in the wild.   * **Nothing detects SST for any platform other than Google.** SST-Guard says so explicitly and names the reason: Meta, TikTok, Snapchat and Reddit ship no debugging tool equivalent to Google Tag Assistant, so there is no ground-truth source to bootstrap from. Meta CAPI is the obvious next target — {[elfraihi2024_meta]} shows it works — and it is undetected in the wild.
-  * **Path-based and same-origin deployments are unmeasured by anyone.** They are ~18% of detections today, they are invisible to every DNS-based method, and Google's documentation now recommends them first. Any method that finds them is new.+  * **Path-based and same-origin deployments are unmeasured by anyone.** They are ~18% of detections by the one published count, and between 4% and 46% depending on how you read its release, they are invisible to every DNS-based method, and Google's documentation now recommends them first. Any method that finds them is new.
   * **Measurement-Protocol deployments leave no client-side artefact at all** and are excluded from the scope of every study above. Whether they are a rounding error or the real story is unknown.   * **Measurement-Protocol deployments leave no client-side artefact at all** and are excluded from the scope of every study above. Whether they are a rounding error or the real story is unknown.
   * **SST-Guard's own suggestion — taint tracking from DOM event to network request — is unattempted for SST.** The taint-tracking browsers exist (see [[Privacy:JavaScript]]); nobody has pointed one at this.   * **SST-Guard's own suggestion — taint tracking from DOM event to network request — is unattempted for SST.** The taint-tracking browsers exist (see [[Privacy:JavaScript]]); nobody has pointed one at this.
Line 168: Line 171:
   * **Longitudinal growth is unmeasured.** Fouad et al. named it as future work in 2024 and it has not been done. There is no published SST time series, so "SST is growing" is currently an assertion.   * **Longitudinal growth is unmeasured.** Fouad et al. named it as future work in 2024 and it has not been done. There is no published SST time series, so "SST is growing" is currently an assertion.
   * **Venue coverage is itself a limitation here.** A reading list built from this corpus alone would contain one paper: of the four studies in the main table, one is a preprint, one is a DPM workshop paper outside these venues, and one is accepted at CCS 2026 but not yet held. The state of a literature and the state of a corpus of it are not the same thing.   * **Venue coverage is itself a limitation here.** A reading list built from this corpus alone would contain one paper: of the four studies in the main table, one is a preprint, one is a DPM workshop paper outside these venues, and one is accepted at CCS 2026 but not yet held. The state of a literature and the state of a corpus of it are not the same thing.
-</wrap>+</WRAP>
  
 ===== Methodology and Limitations of These Figures ===== ===== Methodology and Limitations of These Figures =====
Line 174: Line 177:
   * **Every claim about "the literature" is a claim about seven venues, and "absent" does not mean "outside".** EuroS&P, ACSAC, RAID, AsiaCCS, WPES, CHI and SOUPS are genuinely outside the corpus, so Mertens et al.'s earlier EuroS&P 2025 work {[mertens2025_tag]} and Moti et al.'s DPM 2025 workshop paper {[moti2025_bitterpill]} are invisible to every count here. **Mertens et al. 2026 {[mertens2026_gtm]} is a different case: it is accepted at CCS, one of the seven, and is missing only because CCS 2026 has not been held.** It is the cleanest available illustration of why a 2026 figure from this corpus is a floor and not a measurement.   * **Every claim about "the literature" is a claim about seven venues, and "absent" does not mean "outside".** EuroS&P, ACSAC, RAID, AsiaCCS, WPES, CHI and SOUPS are genuinely outside the corpus, so Mertens et al.'s earlier EuroS&P 2025 work {[mertens2025_tag]} and Moti et al.'s DPM 2025 workshop paper {[moti2025_bitterpill]} are invisible to every count here. **Mertens et al. 2026 {[mertens2026_gtm]} is a different case: it is accepted at CCS, one of the seven, and is missing only because CCS 2026 has not been held.** It is the cleanest available illustration of why a 2026 figure from this corpus is a floor and not a measurement.
   * **The corpus figures count papers, never tuples**, and the full-text probe counts a paper once however many times it uses a term.   * **The corpus figures count papers, never tuples**, and the full-text probe counts a paper once however many times it uses a term.
-  * **Probe width decides the number.** The narrow probe (''server-side tracking'') returns 11 papers; adding product names and API names returns 26; the widest phrasing probe adds 18 moreAll widths are printed above and on the provenance page precisely so the choice is visible. No claim on this page rests on a probe alone: the eleven narrow hits were read individually.+  * **Probe width decides the number.** The narrow probe (''server-side tracking'') returns **11** papers; its union with the product names and the wide API probe is **25**adding the widest phrasing probe brings it to **41** (that probe finds 18, of which 16 are new)These are unions, never sums, and the script prints them as such. The most relevant widths are in the table above and **all nine are on the provenance page**, including two the table omits (''first_party_proxy'', 64 papers; ''tag_manager_any'', 48). No claim here rests on a probe alone: the eleven narrow hits were read individually, and so were the three pre-2019 hits and the twelve wide-only API hits. **The 2019–2026 API-probe hits were not individually read**, which is why the year table is labelled an upper bound.
   * **All 25 figures and quotes taken from the two corpus papers were checked against the papers' own full text**, in three renderings. All 25 located. Six required short fragments rather than whole sentences, because the two-column ''.cols'' repair interleaves the columns mid-sentence — which is a property of the corpus, not of the papers.   * **All 25 figures and quotes taken from the two corpus papers were checked against the papers' own full text**, in three renderings. All 25 located. Six required short fragments rather than whole sentences, because the two-column ''.cols'' repair interleaves the columns mid-sentence — which is a property of the corpus, not of the papers.
   * **The SST-Guard audit is a source audit.** No crawl was re-run and no claim here is that the paper's measurements are wrong. What is shown is that its released artefacts do not reproduce its released numbers, and which of its signals a reader can and cannot reuse.   * **The SST-Guard audit is a source audit.** No crawl was re-run and no claim here is that the paper's measurements are wrong. What is shown is that its released artefacts do not reproduce its released numbers, and which of its signals a reader can and cannot reuse.
privacy/server_side_tracking.1787294280.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki