User Tools

Site Tools


privacy:cookie_syncing

Measuring Cookie and ID Syncing

Cookie syncing (also cookie matching, ID syncing, user matching) is the mechanism by which two parties that each hold their own identifier for the same browser learn that the two identifiers belong to the same person. The same-origin policy stops tracker.com from reading advertiser.com's cookie, so instead tracker.com sends the browser to advertiser.com with its own identifier in the URL — a redirect, an image, an iframe — and both sides write the pair into a match table. Everything afterwards happens server to server, invisibly.

This page is about detecting that exchange from a measurement, not about the ad-tech business it serves. It is the mechanism that turns per-domain identifiers into a joinable graph, and it is the reason a per-party site-prevalence figure understates what that party can see. It is also, from the crawler's side, the tracking behaviour most sensitive to how the crawl was set up. Run a Safari- or Firefox-default browser and you measure near-zero third-party syncing by construction, not because it stopped. Run stateless and you see only first-contact syncing and none of the accumulated graph — a real measurement, but of a different thing.

The load-bearing decision on this page is not “which detector” — it is “which string counts as an identifier”. Every family of method below, including the graph and machine-learning ones, ultimately asks whether a value that came out of client-side storage reappeared in a request to a different party. That question is answered by a heuristic with thresholds, those thresholds have changed materially since 2014, and the error they introduce has now been measured: Calzavara et al. [1Calzavara, Stefano; Casarin, Samuele; Squarcina, Marco; Maffei, Matteo (2026): "From Syntactic Matching to Taint Tracking and Back: A Comparative Study of Web Tracking Detection Techniques", in: Proceedings on Privacy Enhancing Technologies. (Link)] (PoPETs 2026) estimate that 16%–19% of the tracking requests exposed by classic syntactic matching are false positives, rising to 27%–30% for requests that only syntactic matching finds, and that syntactic matching misses 7,021 requests that JavaScript taint tracking finds — 17% of the two methods' union, and about 30% of what taint tracking sees on its own. A decade earlier Bashir et al. [2Bashir, Muhammad Ahmad; Arshad, Sajjad; Robertson, William; Wilson, Christo (2016): "Tracing information flows between ad exchanges using retargeted ads", in: 25th USENIX Security Symposium (USENIX Security 16), pp. 481-496. (Link)] showed the false-negative side independently: the string heuristics of the day missed 31% of the ad-exchange partners that were demonstrably sharing data. (Partners, not pairs: the same paper finds that some pairs are detectable in one direction and not the other.) So do not report a syncing prevalence without reporting the identifier heuristic and its thresholds — and see The identifier heuristic is the measurement for what to use in 2026 rather than what 2016 used.

What to Read First

  • Online Tracking: A 1-million-site Measurement and Analysis [3Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)], CCS 2016 — the reference implementation. Section 4 states the ID-cookie criteria that most later work either copies or explicitly modifies, Section 5.6 measures syncing, and Section 3 explains why the crawl has to be stateful. Its tool, OpenWPM, is still the default instrument.
  • Cookie Synchronization: Everything You Always Wanted to Know But Were Afraid to Ask [4Papadopoulos, Panagiotis; Kourtellis, Nicolas; Markatos, Evangelos P. (2019): "Cookie Synchronization: Everything You Always Wanted to Know But Were Afraid to Ask", in: Proceedings of the ACM Web Conference. (DOI)], TheWebConf 2019 — the only paper in this corpus whose whole subject is the phenomenon, and one of only three that measure it on real users' traffic rather than on a crawl. Read it for the CONRAD algorithm — its rule-based detector for identifier sharing in passive traffic, plus a machine-learning fallback for encrypted identifiers — and for what syncing does to a user over a year, not for a site-level prevalence figure (it does not produce one).
  • Tracing Information Flows Between Ad Exchanges Using Retargeted Ads [2Bashir, Muhammad Ahmad; Arshad, Sajjad; Robertson, William; Wilson, Christo (2016): "Tracing information flows between ad exchanges using retargeted ads", in: 25th USENIX Security Symposium (USENIX Security 16), pp. 481-496. (Link)], USENIX Security 2016 — the one paper that does not look for identifiers in URLs at all, and therefore the only external check on everything that does. It infers sharing from the semantics of which retargeted ad gets served, which works even when the identifiers are encrypted.
  • From Syntactic Matching to Taint Tracking and Back [1Calzavara, Stefano; Casarin, Samuele; Squarcina, Marco; Maffei, Matteo (2026): "From Syntactic Matching to Taint Tracking and Back: A Comparative Study of Web Tracking Detection Techniques", in: Proceedings on Privacy Enhancing Technologies. (Link)], PoPETs 2026 — the current methodological state of the art on the detection question itself. Read this before you implement anything; it is a systematisation of every identifier heuristic in the literature plus a measurement of what each gets wrong.
  • Measuring UID Smuggling in the Wild [5Randall, Audrey; Snyder, Peter; Ukani, Alisha; Snoeren, Alex C.; Voelker, Geoffrey M.; Savage, Stefan; Schulman, Aaron (2022): "Measuring UID smuggling in the wild", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], IMC 2022 — read for the crawler design (four synchronised crawlers instead of two) and because it measures the phenomenon syncing turns into once third-party cookies are unavailable.

Pick the Unit Before You Pick the Method

Published syncing figures look wildly inconsistent because they are not measuring the same object. Nothing on this page is comparable across rows of this table; each number is correct for its own unit.

Unit A published figure in that unit Its denominator, stated
Third parties that sync with at least one other party 45 of the top 50, 85 of the top 100, 157 of the top 200, 460 of the top 1,000 [3Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] Third parties ranked by prominence in a stateful crawl of the top 100,000 sites, January 2016
Sites on which a sync is observed 67.96% of visited domains show first-to-third-party cookie syncing [6Fouad, Imane; Bielova, Nataliia; Legout, Arnaud; Sarafijanovic-Djukic, Natasa (2020): "Missed by Filter Lists: Detecting Unknown Third-Party Trackers with Invisible Pixels", in: Proceedings on Privacy Enhancing Technologies, pp. 499-518. (DOI)] A stateful crawl of the Alexa top 10,000, from France, February 2019
Sites, conditioned on the consent action 24.03% (No Action) / 26.20% (Reject All) / 29.61% (Accept All) carry third-party ID synchronisation [7Papadogiannakis, Emmanouil; Papadopoulos, Panagiotis; Kourtellis, Nicolas; Markatos, Evangelos P. (2021): "User Tracking in the Post-cookie Era: How Websites Bypass GDPR Consent to Track Users", in: Proceedings of the ACM Web Conference. (DOI)] 27,180 sites — those with a CMP and no error in all three consent runs — out of the Tranco top 850K crawled
Syncs per identifier 3.51 (No Action) / 3.91 (Reject All) / 4.86 (Accept All) third parties learn a given third-party ID [7Papadogiannakis, Emmanouil; Papadopoulos, Panagiotis; Kourtellis, Nicolas; Markatos, Evangelos P. (2021): "User Tracking in the Post-cookie Era: How Websites Bypass GDPR Consent to Track Users", in: Proceedings of the ACM Web Conference. (DOI)] The same 27,180 sites
Users exposed 97% of regular web users, median user ID leaked to 3.5 domains, tracking domains up by a factor of 6.75 [4Papadopoulos, Panagiotis; Kourtellis, Nicolas; Markatos, Evangelos P. (2019): "Cookie Synchronization: Everything You Always Wanted to Know But Were Afraid to Ask", in: Proceedings of the ACM Web Conference. (DOI)] 850 real mobile users, one year of passive traffic — not a crawl and not a site sample
Request chains more than half of request chains participated in cookie syncing in most crawl configurations [8Iqbal, Umar; Wolfe, Charlie; Nguyen, Charles; Englehardt, Steven; Shafiq, Zubair (2022): "Khaleesi: Breaker of Advertising and Tracking Request Chains", in: 31st USENIX Security Symposium (USENIX Security 22), pp. 2911-2928. USENIX Association, Boston, MA. (Link)] Request chains, not sites: chains observed in the paper's own crawls of a top list under several cookie-blocking configurations
Sites in one vertical 2,867 porn sites, covering 58% of the top-100 most popular ones [9Vallina, Pelayo; Feal, Álvaro; Gamba, Julien; Vallina-Rodriguez, Narseo; Anta, Antonio Fernández (2019): "Tales from the Porn: A Comprehensive Privacy Analysis of the Web Porn Ecosystem", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] A vertical-specific population, not a general top-list
Cookies 76 of 2,545 unique “intractable” cookies (3%) were synchronised at least once [10Rasaii, Ali; Dao, Ha; Feldmann, Anja; Javid, Mohammadmahdi; Gasser, Oliver; Gosain, Devashish (2025): "Intractable Cookie Crumbs: Unveiling the Nexus of Stateful Banner Interaction and Tracking Cookies", in: Proceedings on Privacy Enhancing Technologies, pp. 429-445. (DOI)] Intractable cookies from one crawl run: set on a site where the banner was accepted, then sent by a different site to the tracker before that site's own banner was touched. Cookies, not sites, and a population defined by an interaction sequence
Domain pairs / organisations 1,190 second-level domains involved, 44% of them advertising-or-tracking [11Weerasekara, Nipuna; Moreno, José Miguel; Matic, Srdjan; Reardon, Joel; Tapiador, Juan; Vallina-Rodríguez, Narseo (2025): "Tracking Without Borders: Studying the Role of WebViews in Bridging Mobile and Web Tracking", Proceedings on Privacy Enhancing Technologies 2025(4). (DOI)] Mobile WebView traffic, not desktop web

Two consequences. First, “X% of sites do cookie syncing” is meaningless without the crawl's statefulness, consent action, and vantage — the same population moves by 5.6 percentage points across the three consent actions in [7Papadogiannakis, Emmanouil; Papadopoulos, Panagiotis; Kourtellis, Nicolas; Markatos, Evangelos P. (2021): "User Tracking in the Post-cookie Era: How Websites Bypass GDPR Consent to Track Users", in: Proceedings of the ACM Web Conference. (DOI)] alone. Second, the directional unit (A sent its ID to B) and the pair unit (A and B are synced) differ by roughly a factor of two, and papers are not consistent about which they report; [3Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] says explicitly that its count “includes both events where it is a referer and where it is a receiver”.

Methods, and Which Ones Are Current

A ranking of what the literature did is a fact about the literature, not advice about what to do now. Statuses are as of 2026-08-26. Each superseded judgement below rests on a named, dated source that supersedes the thing it retires, not on a corpus count, and the browser and vendor landscape was checked against vendor documentation — see cookie_syncing. The current labels also rest on judgement about the corpus's 2025–2026 venue-years, which are provisional, so read them as arguments rather than as measurements.

Thirty papers in the corpus measure identifier sharing between parties (see Use in Publications for how that set was built). Twenty-nine of them field their own detector, in six families; the thirtieth reuses another paper's syncing labels rather than detecting anything, and is the last row.

Family What it does First / most recent in corpus Papers Status in 2026
Syntactic matching Collect client-side storage, decide which values are identifiers, then look for those values (and their encodings) in URLs, paths, referrers and POST bodies going to another party 2014 → 2026 15 — e.g. [12Acar, Gunes; Eubank, Christian; Englehardt, Steven; Juarez, Marc; Narayanan, Arvind; Díaz, Claudia (2014): "The Web Never Forgets: Persistent Tracking Mechanisms in the Wild", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], [3Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)], [6Fouad, Imane; Bielova, Nataliia; Legout, Arnaud; Sarafijanovic-Djukic, Natasa (2020): "Missed by Filter Lists: Detecting Unknown Third-Party Trackers with Invisible Pixels", in: Proceedings on Privacy Enhancing Technologies, pp. 499-518. (DOI)], [13Di Tizio, Giorgio; Massacci, Fabio (2021): "A Calculus of Tracking: Theory and Practice", in: Proceedings on Privacy Enhancing Technologies. (DOI)], [5Randall, Audrey; Snyder, Peter; Ukani, Alisha; Snoeren, Alex C.; Voelker, Geoffrey M.; Savage, Stefan; Schulman, Aaron (2022): "Measuring UID smuggling in the wild", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] Still the default, and now the best-characterised. Cheap, works on any HTTP log, no browser modification. Its error is no longer unknown: 16%–19% of its own detections are false positives, and it misses about 7,000 requests that taint tracking finds — 17% of the union of the two [1Calzavara, Stefano; Casarin, Samuele; Squarcina, Marco; Maffei, Matteo (2026): "From Syntactic Matching to Taint Tracking and Back: A Comparative Study of Web Tracking Detection Techniques", in: Proceedings on Privacy Enhancing Technologies. (Link)]. The two rates have different denominators and are not two halves of one figure. Use it, but use the 2026 refinements below, not the 2016 thresholds
Request-chain / redirect-chain analysis Reason about the sequence of requests rather than any single one: who redirected to whom, with what carried along 2022 → 2023 4 — [14Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)], [15Musa, Maaz Bin; Nithyanand, Rishab (2022): "ATOM: Ad-network Tomography", in: Proceedings on Privacy Enhancing Technologies. (DOI)], [8Iqbal, Umar; Wolfe, Charlie; Nguyen, Charles; Englehardt, Steven; Shafiq, Zubair (2022): "Khaleesi: Breaker of Advertising and Tracking Request Chains", in: 31st USENIX Security Symposium (USENIX Security 22), pp. 2911-2928. USENIX Association, Boston, MA. (Link)], [16Iqbal, Umar; Bahrami, Pouneh Nikkhah; Trimananda, Rahmadi; Cui, Hao; Gamero-Garrido, Alexander; Dubois, Daniel J.; Choffnes, David R.; Markopoulou, Athina; Roesner, Franziska; Shafiq, Zubair (2023): "Tracking, Profiling, and Ad Targeting in the Alexa Echo Smart Speaker Ecosystem", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] Current, and the right frame for what syncing became. Chains also capture bounce tracking and link decoration, which pure value-matching misses. [8Iqbal, Umar; Wolfe, Charlie; Nguyen, Charles; Englehardt, Steven; Shafiq, Zubair (2022): "Khaleesi: Breaker of Advertising and Tracking Request Chains", in: 31st USENIX Security Symposium (USENIX Security 22), pp. 2911-2928. USENIX Association, Boston, MA. (Link)] is the reference; it also releases a classifier
Graph + machine learning Build a graph of the page load (requests, scripts, storage, DOM) and learn which nodes are advertising-or-tracking 2022 → 2024 3 — [17Siby, Sandra; Iqbal, Umar; Englehardt, Steven; Shafiq, Zubair; Troncoso, Carmela (2022): "WebGraph: Capturing Advertising and Tracking Information Flows for Robust Blocking", in: 31st USENIX Security Symposium (USENIX Security 22), pp. 2875-2892. USENIX Association, Boston, MA. (Link)], [18Munir, Shaoor; Siby, Sandra; Iqbal, Umar; Englehardt, Steven; Shafiq, Zubair; Troncoso, Carmela (2023): "CookieGraph: Understanding and Detecting First-Party Tracking Cookies", pp. 3490–3504. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)], [19Munir, Shaoor; Lee, Patrick; Iqbal, Umar; Shafiq, Zubair; Siby, Sandra (2024): "PURL: Safe and Effective Sanitization of Link Decoration", in: 33rd USENIX Security Symposium (USENIX Security 24), pp. 4103-4120. USENIX Association, Philadelphia, PA. (Link)] Current for blocking, indirect for measuring syncing. [17Siby, Sandra; Iqbal, Umar; Englehardt, Steven; Shafiq, Zubair; Troncoso, Carmela (2022): "WebGraph: Capturing Advertising and Tracking Information Flows for Robust Blocking", in: 31st USENIX Security Symposium (USENIX Security 22), pp. 2875-2892. USENIX Association, Boston, MA. (Link)], [18Munir, Shaoor; Siby, Sandra; Iqbal, Umar; Englehardt, Steven; Shafiq, Zubair; Troncoso, Carmela (2023): "CookieGraph: Understanding and Detecting First-Party Tracking Cookies", pp. 3490–3504. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] and [19Munir, Shaoor; Lee, Patrick; Iqbal, Umar; Shafiq, Zubair; Siby, Sandra (2024): "PURL: Safe and Effective Sanitization of Link Decoration", in: 33rd USENIX Security Symposium (USENIX Security 24), pp. 4103-4120. USENIX Association, Philadelphia, PA. (Link)] detect the tracking behaviour that syncing is part of, and PURL's decoration graph is the closest thing to a purpose-built successor detector. But they are trained on filter-list labels, so they inherit those labels' blind spots — see Ground truth, and the circularity
Ad-semantics inference Ignore the wire format; infer that two exchanges shared data because a retargeted ad, or a bid, could not otherwise have been served 2016 → 2022 3 — [2Bashir, Muhammad Ahmad; Arshad, Sajjad; Robertson, William; Wilson, Christo (2016): "Tracing information flows between ad exchanges using retargeted ads", in: 25th USENIX Security Symposium (USENIX Security 16), pp. 481-496. (Link)], [20Cook, John; Nithyanand, Rishab; Shafiq, Zubair (2020): "Inferring Tracker-Advertiser Relationships in the Online Advertising Ecosystem using Header Bidding", in: Proceedings on Privacy Enhancing Technologies. (DOI)], [21Oh, ChangSeok; Kanich, Chris; McCoy, Damon; Pearce, Paul (2022): "Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] Underused and still the only independent check. It is the only family that sees server-to-server sharing and encrypted identifiers. Expensive: [2Bashir, Muhammad Ahmad; Arshad, Sajjad; Robertson, William; Wilson, Christo (2016): "Tracing information flows between ad exchanges using retargeted ads", in: 25th USENIX Security Symposium (USENIX Security 16), pp. 481-496. (Link)] trained 90 personas and collected 35,448 inclusion chains; [20Cook, John; Nithyanand, Rishab; Shafiq, Zubair (2020): "Inferring Tracker-Advertiser Relationships in the Online Advertising Ecosystem using Header Bidding", in: Proceedings on Privacy Enhancing Technologies. (DOI)] needs header-bidding bid streams. No paper in these venues has repeated it since 2022 — see Open Questions
Passive traffic analysis Apply the same identifier logic to real users' HTTP logs instead of a crawl 2017 → 2019 3 — incl. [22Papadopoulos, Panagiotis; Kourtellis, Nicolas; Markatos, Evangelos P. (2018): "The Cost of Digital Advertisement: Comparing User and Advertiser Views", in: Proceedings of the ACM Web Conference. (DOI)], [4Papadopoulos, Panagiotis; Kourtellis, Nicolas; Markatos, Evangelos P. (2019): "Cookie Synchronization: Everything You Always Wanted to Know But Were Afraid to Ask", in: Proceedings of the ACM Web Conference. (DOI)] Historical in these venues, for access reasons, not methodological ones. It answers questions a crawl cannot (“how many users are affected, how fast”), and no paper in this corpus has done it since 2019
JavaScript taint tracking Instrument the engine so a value read from document.cookie or localStorage carries a taint into every derived string, and report when a tainted string reaches the network 2026 → 2026 1 — [1Calzavara, Stefano; Casarin, Samuele; Squarcina, Marco; Maffei, Matteo (2026): "From Syntactic Matching to Taint Tracking and Back: A Comparative Study of Web Tracking Detection Techniques", in: Proceedings on Privacy Enhancing Technologies. (Link)] Newly evaluated, and explicitly not a drop-in replacement. [1Calzavara, Stefano; Casarin, Samuele; Squarcina, Marco; Maffei, Matteo (2026): "From Syntactic Matching to Taint Tracking and Back: A Comparative Study of Web Tracking Detection Techniques", in: Proceedings on Privacy Enhancing Technologies. (Link)] finds Foxhound-based taint tracking has far fewer false positives (4%–7%) but misses a great deal: 17,496 requests, 43% of the 40,605-request union and 52% of the 33,584 that syntactic matching finds, are exposed by syntactic matching alone. Taint propagation is limited to string operations and does not model the full JavaScript semantics. The paper's conclusion is to run both
(Reuses another paper's labels) Takes an existing published list of syncing domains and asks a question of it, rather than detecting syncing 2021 → 2021 1 — [23Iqbal, Umar; Englehardt, Steven; Shafiq, Zubair (2021): "Fingerprinting the Fingerprinters: Learning to Detect Browser Fingerprinting Behaviors", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] Not a detection method, and listed only so the seven rows sum to 30. It is nonetheless the cheapest way to get a syncing variable into a study about something else: [23Iqbal, Umar; Englehardt, Steven; Shafiq, Zubair (2021): "Fingerprinting the Fingerprinters: Learning to Detect Browser Fingerprinting Behaviors", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] uses [6Fouad, Imane; Bielova, Nataliia; Legout, Arnaud; Sarafijanovic-Djukic, Natasa (2020): "Missed by Filter Lists: Detecting Unknown Third-Party Trackers with Invisible Pixels", in: Proceedings on Privacy Enhancing Technologies, pp. 499-518. (DOI)]'s list to find that 17.28% of fingerprinting vendors also participate in cookie syncing

What is genuinely superseded

  • A fixed length window on the cookie value as the identifier test. 8 ≤ length ≤ 100 [3Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] and length > 10 [4Papadopoulos, Panagiotis; Kourtellis, Nicolas; Markatos, Evangelos P. (2019): "Cookie Synchronization: Everything You Always Wanted to Know But Were Afraid to Ask", in: Proceedings of the ACM Web Conference. (DOI)] were reasonable when they were written and are weak now: they admit constant strings such as a publisher's own domain name, which is exactly the false-positive class Calzavara et al. traced by hand — matches on bat.bing.net's p parameter (the publisher domain), tags.creativecdn.com's path constant, analytics.twitter.com's landing-page title. Replace or supplement the length test with an entropy / guessability test: zxcvbn guesses ≥ 10^9 is the criterion used by [24Sanchez-Rola, Iskander; Dell'Amico, Matteo; Kotzias, Platon; Balzarotti, Davide; Bilge, Leyla; Vervier, Pierre-Antoine; Santos, Igor (2019): "Can I Opt Out Yet? GDPR and the Global Illusion of Cookie Control", pp. 340–351. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)], [14Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] and adopted by [1Calzavara, Stefano; Casarin, Samuele; Squarcina, Marco; Maffei, Matteo (2026): "From Syntactic Matching to Taint Tracking and Back: A Comparative Study of Web Tracking Detection Techniques", in: Proceedings on Privacy Enhancing Technologies. (Link)].
  • A long expiry as a necessary condition. [12Acar, Gunes; Eubank, Christian; Englehardt, Steven; Juarez, Marc; Narayanan, Arvind; Díaz, Claudia (2014): "The Web Never Forgets: Persistent Tracking Mechanisms in the Wild", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] required > 30 days, [25Englehardt, Steven; Reisman, Dillon; Eubank, Christian; Zimmerman, Peter; Mayer, Jonathan R.; Narayanan, Arvind; Felten, Edward W. (2015): "Cookies That Give You Away: The Surveillance Implications of Web Tracking", in: Proceedings of the ACM Web Conference. (DOI)] and [3Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] > 90 days. [5Randall, Audrey; Snyder, Peter; Ukani, Alisha; Snoeren, Alex C.; Voelker, Geoffrey M.; Savage, Stefan; Schulman, Aaron (2022): "Measuring UID smuggling in the wild", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] drops the lifetime condition entirely and says why: under partitioned storage and capped cookie lifetimes, tracking identifiers are increasingly short-lived, so a lifetime filter now excludes the interesting cases. Do not filter on expiry unless you can justify it for your population.
  • Two synchronised profiles as the way to tell an identifier from a constant. The classic design runs two machines and keeps values that differ between them [3Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)]. [5Randall, Audrey; Snyder, Peter; Ukani, Alisha; Snoeren, Alex C.; Voelker, Geoffrey M.; Savage, Stefan; Schulman, Aaron (2022): "Measuring UID smuggling in the wild", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] shows this discards a large number of genuine identifiers — those that appear on only one crawler, and those that cannot be distinguished from session IDs — and uses four crawlers instead: three different user profiles plus one that repeats the first profile's visit, so a value can be tested for “varies across users” and “stable for the same user” separately. If you can afford four browsers, run four.
  • Assuming the identifier travels in plaintext. DoubleClick was already encrypting synced identifiers by 2016 [2Bashir, Muhammad Ahmad; Arshad, Sajjad; Robertson, William; Wilson, Christo (2016): "Tracing information flows between ad exchanges using retargeted ads", in: 25th USENIX Security Symposium (USENIX Security 16), pp. 481-496. (Link)], and [4Papadopoulos, Panagiotis; Kourtellis, Nicolas; Markatos, Evangelos P. (2019): "Cookie Synchronization: Everything You Always Wanted to Know But Were Afraid to Ask", in: Proceedings of the ACM Web Conference. (DOI)] added a machine-learning “cookie-less” detector for exactly this reason. Plaintext-only matching is a lower bound and should be labelled as one.
  • “Third-party cookies are about to disappear, so this is about to stop mattering.” They are not. Google announced on 2025-04-22 that it would keep third-party-cookie choice in Chrome and not ship the planned prompt,1) and on 2025-10-17 it announced the retirement of ten Privacy Sandbox technologies, among them Attribution Reporting, Topics and Protected Audience, while stating that “Chrome will maintain our current approach to offering users third-party cookie choice in Chrome”.2) The framing to avoid is “post-cookie”; the framing that holds is “Chrome keeps them, Safari and Firefox do not” — which makes syncing a browser-conditional phenomenon, not a disappearing one. Note that the retirement is a live process rather than a completed one, and that Google's blog posts give no Chrome milestone for it while Chrome's own engineering channel does: as of 2026-08-26 the Privacy Sandbox feature-status page lists Protected Audience, Topics, Attribution Reporting, Private Aggregation, Shared Storage and Related Website Sets as “Intent to deprecate and remove filed”, and the Chrome Platform Status entry for Protected Audience gives a removal milestone of Chrome 153 with status “Proposed”.3)

What the corpus cannot tell you

Syncing is measured in venues this site does not index. The AsiaCCS and EuroS&P line of work is entirely absent — Urban et al.'s GDPR-and-data-sharing measurements are cited by corpus papers but are not in the corpus — and so is most of the ad-tech-economics literature. Any count on this page is a count over CCS, IMC, NDSS, PoPETs, USENIX Security, TheWebConf and IEEE S&P, 2010–2026. See Corpus.

The Identifier Heuristic Is the Measurement

Calzavara et al. [1Calzavara, Stefano; Casarin, Samuele; Squarcina, Marco; Maffei, Matteo (2026): "From Syntactic Matching to Taint Tracking and Back: A Comparative Study of Web Tracking Detection Techniques", in: Proceedings on Privacy Enhancing Technologies. (Link)] tabulate the identifier-detection heuristics used across the literature. This is the single most useful table for anyone about to write one, so it is reproduced here with the original papers attributed. “RO” is Ratcliff–Obershelp string similarity between the values two independent clients received; “Guesses” is the zxcvbn cracking-cost estimate.

Work Year Storage lifetime Tokenised Length Randomness Extra conditions
Roesner et al. [26Roesner, Franziska; Kohno, Tadayoshi; Wetherall, David (2012): "Detecting and Defending Against Third-Party Tracking on the Web", in: 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI 12), pp. 155-168. USENIX Association, San Jose, CA. (Link)] 2012 > session no unique across unrelated visits
Acar et al. [12Acar, Gunes; Eubank, Christian; Englehardt, Steven; Juarez, Marc; Narayanan, Arvind; Díaz, Claudia (2014): "The Web Never Forgets: Persistent Tracking Mechanisms in the Wild", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] 2014 > 30 days yes RO < 33% same across related visits; same length across unrelated visits
Englehardt et al. [25Englehardt, Steven; Reisman, Dillon; Eubank, Christian; Zimmerman, Peter; Mayer, Jonathan R.; Narayanan, Arvind; Felten, Edward W. (2015): "Cookies That Give You Away: The Surveillance Implications of Web Tracking", in: Proceedings of the ACM Web Conference. (DOI)] 2015 > 90 days no RO < 55% same across related visits; same length across unrelated visits; unique across unrelated visits
Englehardt & Narayanan [3Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] 2016 > 90 days yes 8–100 RO < 66% same across related visits; unique across unrelated visits
Papadopoulos et al. [4Papadopoulos, Panagiotis; Kourtellis, Nicolas; Markatos, Evangelos P. (2019): "Cookie Synchronization: Everything You Always Wanted to Know But Were Afraid to Ask", in: Proceedings of the ACM Web Conference. (DOI)] 2019 > session yes ≥ 104)
Sánchez-Rola et al. [24Sanchez-Rola, Iskander; Dell'Amico, Matteo; Kotzias, Platon; Balzarotti, Davide; Bilge, Leyla; Vervier, Pierre-Antoine; Santos, Igor (2019): "Can I Opt Out Yet? GDPR and the Global Illusion of Cookie Control", pp. 340–351. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] 2019 no guesses ≥ 109
Fouad et al. [6Fouad, Imane; Bielova, Nataliia; Legout, Arnaud; Sarafijanovic-Djukic, Natasa (2020): "Missed by Filter Lists: Detecting Unknown Third-Party Trackers with Invisible Pixels", in: Proceedings on Privacy Enhancing Technologies, pp. 499-518. (DOI)] 2020 yes
Chen et al. [27Chen, Quan; Ilia, Panagiotis; Polychronakis, Michalis; Kapravelos, Alexandros (2021): "Cookie Swap Party: Abusing First-Party Cookies for Web Tracking", in: Proceedings of the ACM Web Conference. (DOI)] 2021 > session no ≥ 8 RO < 66% length after URL-decoding; RO after removing timestamps and common subsequences > 2
Sánchez-Rola et al. [14Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] 20225) no guesses ≥ 109
Nikkhah Bahrami et al. [28Nikkhah Bahrami, Pouneh; Fass, Aurore; Shafiq, Zubair (2025): "CookieGuard: Characterizing and Isolating the First-Party Cookie Jar", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] 2025 yes ≥ 8
Calzavara et al. [1Calzavara, Stefano; Casarin, Samuele; Squarcina, Marco; Maffei, Matteo (2026): "From Syntactic Matching to Taint Tracking and Back: A Comparative Study of Web Tracking Detection Techniques", in: Proceedings on Privacy Enhancing Technologies. (Link)] 2026 > session no6) ≥ 8 RO < 66% and guesses ≥ 109 length after URL-decoding; RO after removing timestamps and common subsequences > 2

The paper's own summary of that table is the sentence to quote in your related-work section: “there is no consensus on the most effective heuristic due to the absence of a solid ground truth.”

Three practical instructions follow from their evaluation.

  1. Match after decoding, not only after encoding. Most prior work generated encoded forms of the identifier (Base64, MD5, SHA-1, SHA-256) and looked for those. That misses the case where the identifier is inside a structure that is then encoded — base64('{“uid”: 1234}') contains no encoding of 1234. Parse and decode the request too, up to a bounded depth (three layers is the convention).
  2. Confirm each match with a canary. Their validation technique is cheap and reusable: overwrite the storage value with a distinctive canary, revisit the page, and check whether the matching request now carries the canary. If it does, the flow is real; if the request still carries the original value, the match was spurious. Applying this filter removed 24% of tracking requests and 36% of distinct trackers from their own results, and manual inspection confirmed the removed ones were not tracking.
  3. If you can run both matching and taint tracking, do. The two disagree on 60% of the union of 40,605 requests: 17,496 (43% of the union, 52% of syntactic matching's own 33,584) found by syntactic matching alone, 7,021 (17% of the union) by taint tracking alone. Neither is a superset.

Crawl Configuration That Decides Whether You See It At All

Syncing is the tracking behaviour most sensitive to crawler setup. Get any of these wrong and the measurement floor, not the phenomenon, is what you report.

  • Stateful, and with a seed profile. [3Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] is explicit: for cookie syncing, statefulness “is essential”, because the sync graph of an accumulated identity is what you are trying to reconstruct. They also solve the parallelism problem in a way worth copying — build one seed profile by visiting the top 10,000 sites serially, then load that profile into every parallel browser instance. Their measured justification: such a profile “will have communicated with 76% of all third-party domains present on more than 5 of the top 100,000 sites”. The residual bias is stated too: third parties absent from the seed hand out a different identifier in each parallel instance and so appear to sync with themselves — and no paper in these venues has measured how large that inflation is, which is the open question on the statefulness page.
  • But “stateful sees more” is not a law. Zeber et al. [29Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] give the mechanism for the opposite: “cookie syncing is not necessary for users who have already had their cookies synced, whereas a stateless crawler browser instance with a fresh profile would be a clear target for cookie syncing”. A fresh profile over-triggers first-contact syncing; an aged profile is what you need for the graph. Decide which of the two you are measuring and say so — the full comparison is on the statefulness page. That six of the corpus's 25 crawling papers ran both conditions is the right instinct.
  • The browser decides the answer. Safari's ITP “by default blocks all third-party cookies. There are no exceptions to this blocking”,7) and Firefox has partitioned cookies by top-level site for all users since Firefox 103.8) A default-configured Firefox or Safari crawl measures near-zero third-party syncing by construction. If your instrument is OpenWPM on Firefox, check what protections are on before you interpret a low number; [18Munir, Shaoor; Siby, Sandra; Iqbal, Umar; Englehardt, Steven; Shafiq, Zubair; Troncoso, Carmela (2023): "CookieGraph: Understanding and Detecting First-Party Tracking Cookies", pp. 3490–3504. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] turns Firefox's additional protections off deliberately and says so. Conversely, running with third-party cookies blocked is a legitimate experimental condition — [12Acar, Gunes; Eubank, Christian; Englehardt, Steven; Juarez, Marc; Narayanan, Arvind; Díaz, Claudia (2014): "The Web Never Forgets: Persistent Tracking Mechanisms in the Wild", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] used it and found synced IDs and parties fell “by nearly a factor of two”.
  • Consent action changes the number, and not in the direction you expect. [7Papadogiannakis, Emmanouil; Papadopoulos, Panagiotis; Kourtellis, Nicolas; Markatos, Evangelos P. (2021): "User Tracking in the Post-cookie Era: How Websites Bypass GDPR Consent to Track Users", in: Proceedings of the ACM Web Conference. (DOI)]: sites carrying third-party ID synchronisation went 24.03% → 26.20% → 29.61% for No Action → Reject All → Accept All, and syncs per third-party ID 3.51 → 3.91 → 4.86. Rejecting produced more syncing than not interacting at all. Record the consent action, and prefer measuring more than one. See Consent and Interaction.
  • Interaction depth and timing. Syncing is triggered by the ad stack, which runs after the auction. Prebid.js — the dominant open-source header-bidding wrapper — documents its user-sync defaults as syncDelay 3000 ms after the auction ends, syncsPerBidder 5, image syncs enabled and iframe syncs disabled by default.9) A crawler that closes the page 3 seconds after load can miss the syncs entirely, and any per-adapter count is capped at 5 by the publisher's own configuration rather than by the adapter's appetite. In the corpus's measuring set, 11 of the 25 that crawled visited the landing page only.
  • Vantage. Which exchanges bid, and therefore which sync, depends on where the browser appears to be, and EU vantage points additionally bring a consent banner into the path. Of the 30 measuring papers, 19 state any vantage location at all; the United States accounts for 13 of those.
  • Logging. You need request URLs and referrers and POST bodies, plus Set-Cookie and the resulting cookie jar, plus redirect chains with their initiators. HAR alone is usually not enough — see Traffic files.

Where the Phenomenon Went

Third-party-cookie syncing is a Chrome-and-blocklist-permitting behaviour. Everywhere it is blocked, the same function is served by mechanisms that a syncing detector will not see, and these are where the measurable action is now.

  • Link decoration and UID smuggling. The identifier moves into the URL of a top-level navigation, so no third-party cookie is needed. [5Randall, Audrey; Snyder, Peter; Ukani, Alisha; Snoeren, Alex C.; Voelker, Geoffrey M.; Savage, Stefan; Schulman, Aaron (2022): "Measuring UID smuggling in the wild", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] measures it directly: UID smuggling on 8.11% of the unique URL paths its crawler took (Table 2 puts that population at 10,814), against 2.7% of the navigation paths that were bounce tracking without a UID transfer, and it names 27 “dedicated smugglers” — redirectors with no purpose in the path except carrying an identifier.10) [19Munir, Shaoor; Lee, Patrick; Iqbal, Umar; Shafiq, Zubair; Siby, Sandra (2024): "PURL: Safe and Effective Sanitization of Link Decoration", in: 33rd USENIX Security Symposium (USENIX Security 24), pp. 4103-4120. USENIX Association, Philadelphia, PA. (Link)] finds tracking link decorations on 73.02% of tested sites, averaging 10.75 per site. The 2026 entry in this line is [30Dao, Ha; Shinde, Abhishek; Athar, Sana; Gosain, Devashish (2026): "Clicking into Exposure: Uncovering Privacy Risks of Google Click Identifier in YouTube Ads", Proceedings on Privacy Enhancing Technologies 2026(2):92-107. (DOI)], which follows one specific click identifier: all 568 YouTube ad interactions it observed carried a gclid in a URL path, 64 of 76 advertisers stored it as a first-party cookie (41 of 74 even after the banner was rejected), and 133 distinct third-party domains received gclid values. That last figure is the syncing question in link-decoration clothing: an identifier the advertiser did not set, reaching parties that did not set it either. Detection, tooling and the parameter-list landscape are on Link Decoration and Tracking Parameters — do not re-derive them here.
  • Bounce tracking. A redirect through the tracker's own domain so that its cookie becomes first-party for one hop. WebKit classifies it explicitly: ITP “counts the number of unique such redirects” and “will count it as a bounce even if the redirect is delayed by landing on a webpage and triggering a navigation a couple of seconds later”, and it “caps the expiry of cookies created in JavaScript on the landing webpage to 24 hours” when it detects link decoration.11) [8Iqbal, Umar; Wolfe, Charlie; Nguyen, Charles; Englehardt, Steven; Shafiq, Zubair (2022): "Khaleesi: Breaker of Advertising and Tracking Request Chains", in: 31st USENIX Security Symposium (USENIX Security 22), pp. 2911-2928. USENIX Association, Boston, MA. (Link)] detects it with a chain heuristic requiring a top-level navigation, third-party cookies and a return navigation. Chrome ships mitigations for it too, and — importantly for anyone reading the Privacy Sandbox retirement as “Chrome gave up” — they are on the keep list: a site that a navigation redirected through, and that the user has not interacted with in 45 days, has its storage deleted, for users who block third-party cookies.12)
  • First-party identifier sharing. The identifier is set as a first-party cookie — often by a third-party script — and then sent onward. [6Fouad, Imane; Bielova, Nataliia; Legout, Arnaud; Sarafijanovic-Djukic, Natasa (2020): "Missed by Filter Lists: Detecting Unknown Third-Party Trackers with Invisible Pixels", in: Proceedings on Privacy Enhancing Technologies, pp. 499-518. (DOI)] calls this first-to-third-party syncing and finds it on 67.96% of the domains it visited; [14Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] names the practice cookie ghostwriting, where an entity creates a cookie in another party's name, and measures the resulting graph over 138M cookie-creation events from 6.2M pages on 1M sites. [18Munir, Shaoor; Siby, Sandra; Iqbal, Umar; Englehardt, Steven; Shafiq, Zubair; Troncoso, Carmela (2023): "CookieGraph: Understanding and Detecting First-Party Tracking Cookies", pp. 3490–3504. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] and [28Nikkhah Bahrami, Pouneh; Fass, Aurore; Shafiq, Zubair (2025): "CookieGuard: Characterizing and Isolating the First-Party Cookie Jar", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] are the detection follow-ups. This is the family most likely to be what you actually need to measure in 2026.
  • Server-side. If the exchange happens between two servers, no client-side method sees it at all. That is the subject of Server side tracking, and it is the reason the ad-semantics family [2Bashir, Muhammad Ahmad; Arshad, Sajjad; Robertson, William; Wilson, Christo (2016): "Tracing information flows between ad exchanges using retargeted ads", in: 25th USENIX Security Symposium (USENIX Security 16), pp. 481-496. (Link)] remains the only complete check.
  • Deterministic, non-cookie identifiers. Hashed-email identity frameworks are the industry's declared replacement. UID2 describes itself as “a framework that enables deterministic identity for advertising opportunities on the open internet”.13) These are joined at the identity-provider rather than in the browser, so a syncing crawl sees a token being set, not an exchange. No paper in this corpus measures their deployment, which is a gap, not a finding.

A note on the vendor mechanism, for the related-work paragraph. Google's Authorized Buyers documentation still describes the classic flow in the vendor's own words: Cookie Matching “lets you match your cookie … with a corresponding bidder-specific Google User ID”, implemented as a redirect that hands the bidder a google_gid — “an unpadded web-safe base64-encoded string” — and a google_cver, with the bidder answering with a 1×1 pixel; the match is stored in a match table, optionally Google-hosted.14) Citing the vendor documentation rather than a secondary description is worth the extra paragraph: it is the ground truth for what parameter names to look for.

Use in Publications

All figures below come from the publication corpus — seven venues, 2010–2026, 5,869 papers with readable full text. The full query log, the report script and its unedited output are on cookie_syncing.

How many papers touch it at all

The phenomenon has no field in the extraction schema, so reach is measured from the text. Denominator: 5,869 papers with full text.

Wording Papers Share
cookie sync / cookie synchronisation 86 1.5%
cookie matching 36 0.6%
ID / UID syncing 10 0.2%
pixel / tag syncing 2 0.0%
any of the above, at least once 106 1.8%
any of the above, five or more times 31 0.5%

The gap between 106 and 31 is the point: most mentions are one line of related work. The regex is not the population — see Methodology and limitations of these figures.

By venue

Denominator: that venue's own papers with full text.

Venue Papers ≥1 mention ≥5 mentions
PoPETs 510 28 (5.5%) 9 (1.8%)
IMC 637 18 (2.8%) 4 (0.6%)
TheWebConf 843 20 (2.4%) 6 (0.7%)
NDSS 701 11 (1.6%) 0 (0.0%)
USENIX Security 1,410 15 (1.1%) 6 (0.4%)
CCS 989 8 (0.8%) 4 (0.4%)
IEEE S&P 779 6 (0.8%) 2 (0.3%)

PoPETs is where this literature lives. Its ≥1-mention rate (5.5%) is about twice IMC's, the next venue, and about seven times CCS's; at the ≥5 threshold it holds 9 of the 31 papers. If you are choosing a venue, that is the signal. No NDSS paper in the corpus mentions syncing five times or more, despite NDSS having the fourth-highest ≥1-mention rate in the table.

By year

Papers mentioning any syncing wording, per year, against that year's corpus size. This measures attention in the literature, not how common syncing was in any year — no row here is a prevalence figure. 2025 is thin at the edges and 2026 is provisional — CCS 2026 and IMC 2026 have not been held, and IEEE S&P 2026 and TheWebConf 2026 are incompletely selected — so the last two rows are not evidence of decline.

Year Corpus papers ≥1 mention ≥5 mentions
2014 165 1 1
2015 190 2 0
2016 182 5 3
2017 232 6 2
2018 254 5 1
2019 402 9 3
2020 402 9 4
2021 380 10 3
2022 546 17 8
2023 720 14 1
2024 701 10 3
2025 (thin) 770 13 2
2026 (provisional) 415 3 0

2010–2013 produced two mentions in 510 papers and no paper engaging with it. 2022 is the peak year of attention on both columns, which is also when the request-chain and graph families arrive. That is a coincidence worth noticing, not a cause.

The 30 measuring papers

Built by hand from the 44 candidates the automated sweep produced: 30 measure identifier sharing between distinct parties, 5 cite it without measuring it, and 9 use the phrase for something else entirely (a rendering side-channel that “synchronises cookies” across browsers, a cookie-match predicate in a formal browser model, opt-out cookie name matching, and a redefinition of “cookie syncing” to mean cross-site syncing enabled by a browser bug). Unlabelled residue: zero — every candidate carries a hand label.

Detection family Papers Share of 30
Syntactic matching 15 50.0%
Request-chain analysis 4 13.3%
Ad-semantics inference 3 10.0%
Graph + machine learning 3 10.0%
Passive traffic analysis 3 10.0%
JavaScript taint tracking 1 3.3%
Reuses another paper's labels 1 3.3%

Crawl configuration of these papers

Of the 30, 25 ran an automated crawl (crawlConfig fired). Sentinels are shown but are never counted as answers.

Field Value Papers Share of 25
statefulness stateful 10 40.0%
both stateful and stateless 6 24.0%
stateless 3 12.0%
not stated 6 24.0%
consentAction no interaction 9 36.0%
accept and reject 4 16.0%
not stated 12 48.0%
interactionDepth landing page only 11 44.0%
landing plus subpages 7 28.0%
single target page 4 16.0%
deep crawl 1 4.0%
not stated 2 8.0%
headless headless 3 12.0%
headful 2 8.0%
not stated 20 80.0%

Just under a quarter of the papers measuring the one behaviour that most needs a stateful crawl do not say whether their crawl was stateful, and just under half do not say what they did about consent banners. Those are the two fields this page asks you to report. See Stateful stateless.

Read the headless row as the weakest one in this table, and not only because 80.0% of the papers do not state a value. The extraction stores one evidence quote for the whole crawlConfig object, so a spot-check can confirm at most one of its fields per paper; of five quotes read by hand for this page, none supported the headless value and only one supported statefulness. The one headless value that could be traced is also the ambiguous kind: [3Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] is recorded as headless because the paper describes launching measurement instances in a “headless” container “by using the pyvirtualdisplay library” to drive Xvfb — which is a headful browser on a virtual framebuffer, and behaves differently from a genuinely headless one under bot detection. The extraction is faithful to the paper's own word; the paper's own word is loose. If headless-ness matters to your argument, read the papers rather than this row.

Artifacts

Of the 30, 13 (43.3%) state a public artifact and 10 (33.3%) mention none. What is actually released and reusable:

Artifact Paper What you get
OpenWPM [3Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] The crawler, actively maintained (last push 2026-08-25). The ID-cookie logic is not in it and has to be re-implemented from the paper. The citp/OpenWPM URL printed in the paper now redirects here
CrumbCruncher [5Randall, Audrey; Snyder, Peter; Ukani, Alisha; Snoeren, Alex C.; Voelker, Geoffrey M.; Savage, Stefan; Schulman, Aaron (2022): "Measuring UID smuggling in the wild", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] The four-crawler synchronised design, Puppeteer-based
Khaleesi [8Iqbal, Umar; Wolfe, Charlie; Nguyen, Charles; Englehardt, Steven; Shafiq, Zubair (2022): "Khaleesi: Breaker of Advertising and Tracking Request Chains", in: 31st USENIX Security Symposium (USENIX Security 22), pp. 2911-2928. USENIX Association, Boston, MA. (Link)] Request-chain classifier
CookieGraph [18Munir, Shaoor; Siby, Sandra; Iqbal, Umar; Englehardt, Steven; Shafiq, Zubair; Troncoso, Carmela (2023): "CookieGraph: Understanding and Detecting First-Party Tracking Cookies", pp. 3490–3504. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] First-party tracking-cookie classifier
PURL [19Munir, Shaoor; Lee, Patrick; Iqbal, Umar; Shafiq, Zubair; Siby, Sandra (2024): "PURL: Safe and Effective Sanitization of Link Decoration", in: 33rd USENIX Security Symposium (USENIX Security 24), pp. 4103-4120. USENIX Association, Philadelphia, PA. (Link)] Link-decoration classifier and sanitiser. The purl-sanitizer organisation in the paper no longer exists; the URL redirects here, last push 2024-08-22
CookieGuard [28Nikkhah Bahrami, Pouneh; Fass, Aurore; Shafiq, Zubair (2025): "CookieGuard: Characterizing and Isolating the First-Party Cookie Jar", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] First-party cookie-jar isolation
Web-Tracking-Detection [1Calzavara, Stefano; Casarin, Samuele; Squarcina, Marco; Maffei, Matteo (2026): "From Syntactic Matching to Taint Tracking and Back: A Comparative Study of Web Tracking Detection Techniques", in: Proceedings on Privacy Enhancing Technologies. (Link)] Syntactic-matching and taint-tracking pipelines and the comparison data. The paper cites purl.org/tracking-detection-paper, which redirects here; last push 2026-08-03
Retargeting dataset [2Bashir, Muhammad Ahmad; Arshad, Sajjad; Robertson, William; Wilson, Christo (2016): "Tracing information flows between ad exchanges using retargeted ads", in: 25th USENIX Security Symposium (USENIX Security 16), pp. 481-496. (Link)] 7K labelled targeted and retargeted ads, inclusion chains, full HTTP traces — the only external ground truth released by any of the 30 papers in this corpus, and it is from 2016
doi.org/10.17617/3.H5T0W4 [30Dao, Ha; Shinde, Abhishek; Athar, Sana; Gosain, Devashish (2026): "Clicking into Exposure: Uncovering Privacy Risks of Google Click Identifier in YouTube Ads", Proceedings on Privacy Enhancing Technologies 2026(2):92-107. (DOI)] The gclid crawl data, in a repository with a DOI rather than on GitHub

Nine of the thirteen public artifacts are listed; the other four are conference-artifact or project pages whose reuse value is narrower than a tool or a dataset. The full list of links the extraction found is in the report output on cookie_syncing.

How syncing papers classify requests

Crossing the measuring set with the extraction's classification family: 20 of the 30 carry a tuple whose target is web-request (corpus-wide, 262 of the 4,439 papers that classified anything target web-request). Of those 20, 15 use a blocklist to decide which of the parties involved counts as a tracker: 12 name EasyList and/or EasyPrivacy and 4 name Disconnect. That is a second, independent dependency on filter lists layered on top of the identifier heuristic, and it inherits everything those lists miss. Only 9 of the 20 report any validation of that classification stronger than a sentinel.

What to Report

If a reviewer is to accept a syncing figure, the paper has to answer all of these. The corpus can only speak to three of them, and on those three the record is poor: 24.0% of the papers do not state statefulness, 48.0% do not state a consent action, and 80.0% do not state whether the browser was headless.

  1. The unit and its denominator. Sites? Directional flows? Unordered domain pairs? Users? Requests? Chains? And of what population — see the table in Pick the Unit Before You Pick the Method.
  2. The identifier heuristic, in full. Lifetime condition, tokenisation, length bound, randomness test and its threshold, and how many independent profiles you compared. Reproduce the row you would occupy in the table in The identifier heuristic is the measurement.
  3. Which request components you searched: query parameters, path segments, Referer, request body, response headers. Papers differ here and rarely say so.
  4. Which encodings and decodings you applied, and to what depth.
  5. How you decided two domains are different parties, and with which entity list at which version. Without this, amazon.comamazonaws.com is a sync.
  6. Statefulness, seed profile, parallelism. If parallel, say how the seed was built and acknowledge the self-sync artefact [3Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)].
  7. Consent action, and ideally more than one.
  8. Browser and its tracking protections, explicitly. “Firefox” is not a configuration.
  9. Vantage, and whether it is in the EEA.
  10. How long you stayed on the page after load, given the 3-second Prebid default.
  11. Whether your figure is a lower bound because of encryption, server-side exchange, or a blocked configuration. It almost always is.
  12. Any validation. A canary re-visit [1Calzavara, Stefano; Casarin, Samuele; Squarcina, Marco; Maffei, Matteo (2026): "From Syntactic Matching to Taint Tracking and Back: A Comparative Study of Web Tracking Detection Techniques", in: Proceedings on Privacy Enhancing Technologies. (Link)] is cheap; a manual sample is the minimum.

Open Questions

  • No paper in these seven venues has repeated the ad-semantics check since 2022. [2Bashir, Muhammad Ahmad; Arshad, Sajjad; Robertson, William; Wilson, Christo (2016): "Tracing information flows between ad exchanges using retargeted ads", in: 25th USENIX Security Symposium (USENIX Security 16), pp. 481-496. (Link)] is the only independent measurement of what identifier heuristics miss, it is from 2016, and its 31%-missed figure is cited as if it were current. Header bidding [20Cook, John; Nithyanand, Rishab; Shafiq, Zubair (2020): "Inferring Tracker-Advertiser Relationships in the Online Advertising Ecosystem using Header Bidding", in: Proceedings on Privacy Enhancing Technologies. (DOI)] and the retargeting design are both still runnable. This is the highest-value replication on this page.
  • There is no time series. Every prevalence figure here comes from a different population, crawler and year, so nothing in this corpus can say whether syncing grew, shrank, or moved. A single stateful crawl of a fixed population, repeated quarterly with a fixed heuristic, would be the first.
  • The identifier heuristics have never been compared head-to-head on one dataset. [1Calzavara, Stefano; Casarin, Samuele; Squarcina, Marco; Maffei, Matteo (2026): "From Syntactic Matching to Taint Tracking and Back: A Comparative Study of Web Tracking Detection Techniques", in: Proceedings on Privacy Enhancing Technologies. (Link)] tabulates ten prior heuristics and then implements an eleventh — a “representative” union of them — explicitly declining to evaluate the ten individually. Running all ten over one crawl and reporting the spread would tell the field how much of its published variance is heuristic choice rather than measurement.
  • No paper in these venues measures the deterministic-ID frameworks (UID2, EUID, and their competitors) in the wild, even though they are the industry's stated replacement for syncing.
  • Machine-learning detectors have not been evaluated against taint tracking. [1Calzavara, Stefano; Casarin, Samuele; Squarcina, Marco; Maffei, Matteo (2026): "From Syntactic Matching to Taint Tracking and Back: A Comparative Study of Web Tracking Detection Techniques", in: Proceedings on Privacy Enhancing Technologies. (Link)] names this as future work: its comparison covers syntactic matching and taint tracking but not [17Siby, Sandra; Iqbal, Umar; Englehardt, Steven; Shafiq, Zubair; Troncoso, Carmela (2022): "WebGraph: Capturing Advertising and Tracking Information Flows for Robust Blocking", in: 31st USENIX Security Symposium (USENIX Security 22), pp. 2875-2892. USENIX Association, Boston, MA. (Link)], [18Munir, Shaoor; Siby, Sandra; Iqbal, Umar; Englehardt, Steven; Shafiq, Zubair; Troncoso, Carmela (2023): "CookieGraph: Understanding and Detecting First-Party Tracking Cookies", pp. 3490–3504. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] or the behaviour-based detectors.
  • Passive measurement has not been repeated since 2019. [4Papadopoulos, Panagiotis; Kourtellis, Nicolas; Markatos, Evangelos P. (2019): "Cookie Synchronization: Everything You Always Wanted to Know But Were Afraid to Ask", in: Proceedings of the ACM Web Conference. (DOI)]'s 850-user, year-long dataset is the only evidence in this corpus about what syncing does to a person over time, and it predates every browser countermeasure discussed above.

Methodology and Limitations of These Figures

The full query log, the report script, its unedited output, the hand labels and the quote checks are on cookie_syncing. Corpus-wide caveats — how the corpus was built, which venue-years are provisional, how stable each extracted field is — are on Corpus.

The limitations specific to this page:

  • The population is defined by a regex over full text, then corrected by hand. The regex is a filter for reading effort, not a definition. It under-recalls badly at the low end: [1Calzavara, Stefano; Casarin, Samuele; Squarcina, Marco; Maffei, Matteo (2026): "From Syntactic Matching to Taint Tracking and Back: A Comparative Study of Web Tracking Detection Techniques", in: Proceedings on Privacy Enhancing Technologies. (Link)], the most methodologically central paper here, uses the phrase twice and would have been excluded by any threshold. Three papers were added to the set by hand for exactly this reason, and a fourth to correct a boundary inconsistency rather than to expand the set. There is no way to know how many others were missed, and the honest reading of “106 papers mention it” is “at least 106”.
  • The 30-paper “measures it” set is one person's judgement, recorded as an explicit label per paper in the report script so that a disagreement can be located. Nine candidates were excluded because the phrase means something else in them; that call is defensible but not unique.
  • The boundary of that set is genuinely fuzzy, and it moved during writing. It is “measures an identifier being conveyed to a party that did not set it”, which includes both a two-way sync and a one-way hand-off. Two corpus papers meet that description and are not in the 30, because they were found after the set was fixed and adding papers as one stumbles on them is how a hand-curated set stops being reproducible: [31Bekos, Paschalis; Papadopoulos, Panagiotis; Markatos, Evangelos P.; Kourtellis, Nicolas (2023): "The Hitchhiker's Guide to Facebook Web Tracking with Invisible Pixels and Click IDs", in: Proceedings of the ACM Web Conference. (DOI)] measures a median website passing identifiers to 6.2 third parties, and [32Dimova, Yana; Acar, Gunes; Olejnik, Lukasz; Joosen, Wouter; Van Goethem, Tom (2021): "The CNAME of the game: Large-scale analysis of DNS-based tracking evasion", Proceedings on Privacy Enhancing Technologies 2021:394–412. (DOI) (Link)] finds 1,899 cookie leaks in request URLs on 1,295 distinct sites. Both are named here rather than quietly omitted; a reader who counts them gets 32, and the shape of every table above is unchanged by two papers. One paper was added late — the gclid study — but to correct a boundary that had been applied two different ways, not to expand the set; the difference is argued on cookie_syncing. Drawn strictly — only papers measuring a genuine two-way match between two parties' identifiers — the set would be closer to twenty.
  • The detection-family assignment is coarse. Papers combine families — [19Munir, Shaoor; Lee, Patrick; Iqbal, Umar; Shafiq, Zubair; Siby, Sandra (2024): "PURL: Safe and Effective Sanitization of Link Decoration", in: 33rd USENIX Security Symposium (USENIX Security 24), pp. 4103-4120. USENIX Association, Philadelphia, PA. (Link)] builds a graph and does syntactic value matching — and each was assigned the family that does the identifier-sharing work. Read the family table as a ranking, not as a partition.
  • Every prevalence figure quoted from a paper carries that paper's denominator, and they are not comparable. This is stated once in Pick the Unit Before You Pick the Method and is worth repeating: none of the percentages on this page can be averaged, ordered, or plotted against each other.
  • crawlConfig is an object with one shared evidence quote, so the crawl-configuration table cannot be quote-checked field by field. Of five quotes read by hand, four do not touch the field the row reports. See the box beside that table, and the detail in the provenance page.
  • Corpus-level counts come from the .cols rendering of the PDFs. Column reading order was repaired but not perfectly; a phrase split across a column break can still be missed even after whitespace collapsing and de-hyphenation.
  • The extraction fields used here (crawlConfig.statefulness, .consentAction, .interactionDepth) are among the most stable in the schema — 98%, 93% and 97% paper-level agreement between two independent extraction runs — but that stability was measured on the previous corpus and has not been re-measured.

References

[1]
Calzavara, Stefano; Casarin, Samuele; Squarcina, Marco; Maffei, Matteo (2026): "From Syntactic Matching to Taint Tracking and Back: A Comparative Study of Web Tracking Detection Techniques", in: Proceedings on Privacy Enhancing Technologies. (Link)
[2]
Bashir, Muhammad Ahmad; Arshad, Sajjad; Robertson, William; Wilson, Christo (2016): "Tracing information flows between ad exchanges using retargeted ads", in: 25th USENIX Security Symposium (USENIX Security 16), pp. 481-496. (Link)
[3]
Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
[4]
Papadopoulos, Panagiotis; Kourtellis, Nicolas; Markatos, Evangelos P. (2019): "Cookie Synchronization: Everything You Always Wanted to Know But Were Afraid to Ask", in: Proceedings of the ACM Web Conference. (DOI)
[5]
Randall, Audrey; Snyder, Peter; Ukani, Alisha; Snoeren, Alex C.; Voelker, Geoffrey M.; Savage, Stefan; Schulman, Aaron (2022): "Measuring UID smuggling in the wild", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[6]
Fouad, Imane; Bielova, Nataliia; Legout, Arnaud; Sarafijanovic-Djukic, Natasa (2020): "Missed by Filter Lists: Detecting Unknown Third-Party Trackers with Invisible Pixels", in: Proceedings on Privacy Enhancing Technologies, pp. 499-518. (DOI)
[7]
Papadogiannakis, Emmanouil; Papadopoulos, Panagiotis; Kourtellis, Nicolas; Markatos, Evangelos P. (2021): "User Tracking in the Post-cookie Era: How Websites Bypass GDPR Consent to Track Users", in: Proceedings of the ACM Web Conference. (DOI)
[8]
Iqbal, Umar; Wolfe, Charlie; Nguyen, Charles; Englehardt, Steven; Shafiq, Zubair (2022): "Khaleesi: Breaker of Advertising and Tracking Request Chains", in: 31st USENIX Security Symposium (USENIX Security 22), pp. 2911-2928. USENIX Association, Boston, MA. (Link)
[9]
Vallina, Pelayo; Feal, Álvaro; Gamba, Julien; Vallina-Rodriguez, Narseo; Anta, Antonio Fernández (2019): "Tales from the Porn: A Comprehensive Privacy Analysis of the Web Porn Ecosystem", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[10]
Rasaii, Ali; Dao, Ha; Feldmann, Anja; Javid, Mohammadmahdi; Gasser, Oliver; Gosain, Devashish (2025): "Intractable Cookie Crumbs: Unveiling the Nexus of Stateful Banner Interaction and Tracking Cookies", in: Proceedings on Privacy Enhancing Technologies, pp. 429-445. (DOI)
[11]
Weerasekara, Nipuna; Moreno, José Miguel; Matic, Srdjan; Reardon, Joel; Tapiador, Juan; Vallina-Rodríguez, Narseo (2025): "Tracking Without Borders: Studying the Role of WebViews in Bridging Mobile and Web Tracking", Proceedings on Privacy Enhancing Technologies 2025(4). (DOI)
[12]
Acar, Gunes; Eubank, Christian; Englehardt, Steven; Juarez, Marc; Narayanan, Arvind; Díaz, Claudia (2014): "The Web Never Forgets: Persistent Tracking Mechanisms in the Wild", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[13]
Di Tizio, Giorgio; Massacci, Fabio (2021): "A Calculus of Tracking: Theory and Practice", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[14]
Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[15]
Musa, Maaz Bin; Nithyanand, Rishab (2022): "ATOM: Ad-network Tomography", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[16]
Iqbal, Umar; Bahrami, Pouneh Nikkhah; Trimananda, Rahmadi; Cui, Hao; Gamero-Garrido, Alexander; Dubois, Daniel J.; Choffnes, David R.; Markopoulou, Athina; Roesner, Franziska; Shafiq, Zubair (2023): "Tracking, Profiling, and Ad Targeting in the Alexa Echo Smart Speaker Ecosystem", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[17]
Siby, Sandra; Iqbal, Umar; Englehardt, Steven; Shafiq, Zubair; Troncoso, Carmela (2022): "WebGraph: Capturing Advertising and Tracking Information Flows for Robust Blocking", in: 31st USENIX Security Symposium (USENIX Security 22), pp. 2875-2892. USENIX Association, Boston, MA. (Link)
[18]
Munir, Shaoor; Siby, Sandra; Iqbal, Umar; Englehardt, Steven; Shafiq, Zubair; Troncoso, Carmela (2023): "CookieGraph: Understanding and Detecting First-Party Tracking Cookies", pp. 3490–3504. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
[19]
Munir, Shaoor; Lee, Patrick; Iqbal, Umar; Shafiq, Zubair; Siby, Sandra (2024): "PURL: Safe and Effective Sanitization of Link Decoration", in: 33rd USENIX Security Symposium (USENIX Security 24), pp. 4103-4120. USENIX Association, Philadelphia, PA. (Link)
[20]
Cook, John; Nithyanand, Rishab; Shafiq, Zubair (2020): "Inferring Tracker-Advertiser Relationships in the Online Advertising Ecosystem using Header Bidding", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[21]
Oh, ChangSeok; Kanich, Chris; McCoy, Damon; Pearce, Paul (2022): "Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[22]
Papadopoulos, Panagiotis; Kourtellis, Nicolas; Markatos, Evangelos P. (2018): "The Cost of Digital Advertisement: Comparing User and Advertiser Views", in: Proceedings of the ACM Web Conference. (DOI)
[23]
Iqbal, Umar; Englehardt, Steven; Shafiq, Zubair (2021): "Fingerprinting the Fingerprinters: Learning to Detect Browser Fingerprinting Behaviors", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[24]
Sanchez-Rola, Iskander; Dell'Amico, Matteo; Kotzias, Platon; Balzarotti, Davide; Bilge, Leyla; Vervier, Pierre-Antoine; Santos, Igor (2019): "Can I Opt Out Yet? GDPR and the Global Illusion of Cookie Control", pp. 340–351. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
[25]
Englehardt, Steven; Reisman, Dillon; Eubank, Christian; Zimmerman, Peter; Mayer, Jonathan R.; Narayanan, Arvind; Felten, Edward W. (2015): "Cookies That Give You Away: The Surveillance Implications of Web Tracking", in: Proceedings of the ACM Web Conference. (DOI)
[26]
Roesner, Franziska; Kohno, Tadayoshi; Wetherall, David (2012): "Detecting and Defending Against Third-Party Tracking on the Web", in: 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI 12), pp. 155-168. USENIX Association, San Jose, CA. (Link)
[27]
Chen, Quan; Ilia, Panagiotis; Polychronakis, Michalis; Kapravelos, Alexandros (2021): "Cookie Swap Party: Abusing First-Party Cookies for Web Tracking", in: Proceedings of the ACM Web Conference. (DOI)
[28]
Nikkhah Bahrami, Pouneh; Fass, Aurore; Shafiq, Zubair (2025): "CookieGuard: Characterizing and Isolating the First-Party Cookie Jar", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[29]
Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
[30]
Dao, Ha; Shinde, Abhishek; Athar, Sana; Gosain, Devashish (2026): "Clicking into Exposure: Uncovering Privacy Risks of Google Click Identifier in YouTube Ads", Proceedings on Privacy Enhancing Technologies 2026(2):92-107. (DOI)
[31]
Bekos, Paschalis; Papadopoulos, Panagiotis; Markatos, Evangelos P.; Kourtellis, Nicolas (2023): "The Hitchhiker's Guide to Facebook Web Tracking with Invisible Pixels and Click IDs", in: Proceedings of the ACM Web Conference. (DOI)
[32]
Dimova, Yana; Acar, Gunes; Olejnik, Lukasz; Joosen, Wouter; Van Goethem, Tom (2021): "The CNAME of the game: Large-scale analysis of DNS-based tracking evasion", Proceedings on Privacy Enhancing Technologies 2021:394–412. (DOI) (Link)
1)
Anthony Chavez, “Next steps for Privacy Sandbox and tracking protections in Chrome”, privacysandbox.google.com/blog/privacy-sandbox-next-steps, 2025-04-22: “we've made the decision to maintain our current approach to offering users third-party cookie choice in Chrome, and will not be rolling out a new standalone prompt for third-party cookies.” Fetched 2026-08-26.
2)
privacysandbox.google.com/blog/update-on-plans-for-privacy-sandbox-technologies, 2025-10-17, fetched 2026-08-26. The full list, verbatim: “Attribution Reporting API (Chrome and Android), IP Protection, On-Device Personalization, Private Aggregation (including Shared Storage), Protected Audience (Chrome and Android), Protected App Signals, Related Website Sets (including requestStorageAccessFor and Related Website Partition), SelectURL, SDK Runtime and Topics (Chrome and Android).”
3)
privacysandbox.google.com/overview/status and chromestatus.com/api/v0/features/6552486106234880, both fetched 2026-08-26. If you need a date rather than a milestone, read Chrome Platform Status, not the blog.
4)
The source table says ≥ 10; the original paper says strictly greater — “strings with specific length (> 10 characters)”. The difference is one character and is the source table's, not this page's.
5)
Dated by conference: it appeared at the 43rd IEEE S&P, May 2022, and the corpus files it under 2022. The PDF's own running header and copyright line read 2021, and this site's bibliography entry follows the IEEE Xplore record in saying 2021, so the reference list below will show a different year from this cell.
6)
The Parse column of the source table is blank for its own row, which is what this cell reports. The paper's §3.3 nonetheless describes slicing on non-alphanumeric characters as one of its supported transformations, so the blank looks like an omission in the printed table rather than a design choice. Checked against the PDF's word coordinates, not the text layer, because the column is marked only by a tick glyph.
7) , 11)
webkit.org/tracking-prevention/, fetched 2026-08-26.
8)
developer.mozilla.org/en-US/docs/Web/Privacy/Guides/State_Partitioning, fetched 2026-08-26: “Dynamic Partitioning: Enabled by default for all users since Firefox 103.”
9)
docs.prebid.org/dev-docs/publisher-api-reference/setConfig.html, fetched 2026-08-26.
10)
The paper's Table 2 also reports 850 unique URL paths with UID smuggling, which is 7.86% of 10,814 rather than 8.11%. The two figures are not reconciled in the paper; the page quotes the stated percentage and gives the table count so a reader can see the gap.
12)
privacysandbox.google.com/protections/bounce-tracking-mitigations, fetched 2026-08-26. This is conditional on the user blocking third-party cookies, so it does not fire in a default Chrome crawl.
13)
unifiedid.com/docs/intro, fetched 2026-08-26.
14)
developers.google.com/authorized-buyers/rtb/cookie-guide, fetched 2026-08-26. The same page states that the “Cookie Match Assist feature will be deprecated starting on October 28th, 2025”.
You could leave a comment if you were logged in.
privacy/cookie_syncing.txt · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki