User Tools

Site Tools


design:blocking_and_geodifference

This is an old revision of the document!


Blocking and Geodifference

Your crawl finished, the error rate looks fine, and some of the pages in your dataset are not the pages you think you measured. Somebody decided your client was not allowed to see them: a state filtering the path, a CDN refusing your country, a site that shows nothing to European visitors, a platform that will not ship to your region. The failure is quiet — a 200 with an interstitial, a resolution to the wrong address, a connection that closes after the ClientHello — and it lands in your data as this site has no trackers or this site has no consent banner.

This page is about the phenomenon and about how a blocking claim is made: what you need in the experiment before you can say a client could not reach something, and what the field has measured about how often that claim is wrong. It is not an HTTP tutorial: the status codes are on MDN and in RFC 9110, and how DNS resolution works is DNS.

If your question is… Then read
whether some clients cannot reach content, and how to show it this page
where to put your vantage point, and how to verify you got it Crawling location
whether a DNS answer you received was tampered with DNS, which owns the resolution-layer methods
whether you were blocked for looking like a crawler Crawler detection
whether an existing censorship dataset can answer your question Existing datasets, then the instrument section below
how a consent banner changes with the visitor's region Consent and Design:Crawling location § Jurisdiction
what a filter list blocks in your own browser Filter lists

The short version. A blocking claim needs a treatment vantage (inside the condition), a control vantage (outside it), and a written definition of “blocked” that survives an HTTP 200 carrying a block page. Of the 63 papers in the population below, 32 use the phrase control vantage or a direct equivalent, 47 put the word control anywhere near a vantage word, and 16 never do either — so a quarter of the literature making a blocking claim never writes down the comparison its claim rests on.

Three reasons your client did not get the page

The symptom is one thing — you did not get the content — and the causes need separating before anything else, because each has a different control and a different owner on this wiki.

Cause Keyed on Typical evidence Where it is covered
Network interference: a filter on the path resets, drops, poisons or throttles the connection where your packets go, and what is in them injected RST, forged DNS answer, timeout that depends on the SNI or on a keyword this page; resolution layer in DNS
Server-side refusal: the origin, its CDN or the platform declines to serve your region your IP's country or network, or the jurisdiction it implies 403, 451, a block page, a “not available in your region” body, a missing feature this page
Bot management: you were served a challenge because you looked automated your client's behaviour and fingerprint CAPTCHA, JS challenge, 403 after n requests Crawler detection

The first two are this page's subject. The third is a different literature with its own page, and the cheapest way to tell them apart is the one that also underpins every claim here: re-request from a second vantage point. If the same URL from a different country succeeds with the same client, it is where you were. If it fails from everywhere with your client and succeeds with another, it is you.

A block page is a page

Nothing forces a blocker to use an error status, and plenty of them do not. Ablove et al. probing 10,093 domains from Cuban and control vantage points found 546 domains geoblocking somewhere in the stack, of which 395 (72.3%) did so with a block page in an HTTP(S) response rather than a network-level failure — and 32 of those served the block page with a 200 OK [1Ablove, Anna; Chandrashekaran, Shreyas; Le, Hieu; Raman, Ram Sundara; Ramesh, Reethika; Oppenheimer, Harry; Ensafi, Roya (2024): "Digital Discrimination of Users in Sanctioned States: The Case of the Cuba Embargo", in: Proceedings of the USENIX Security Symposium. (Link)]. In the same dataset 88% (480 of 546) of geoblocked domains “do not serve informative notice of why they are blocked”, so the body does not necessarily tell you either.

This is why a crawl-side “success” rate is not a coverage rate. Roth et al., crawling 10,000 sites from several countries to compare security headers, had to build error-and-block-page detection into their deduplication step, on the grounds that “not every URL returns the same content on each load, in particular in the presence of errors or block pages” [2Roth, Sebastian; Calzavara, Stefano; Wilhelm, Moritz; Rabitti, Alvise; Stock, Ben (2022): "The Security Lottery: Measuring Client-Side Web Security Inconsistencies", in: Proceedings of the USENIX Security Symposium, pp. 2047-2064. (Link)] — the block page was a nuisance variable in a study that was not about blocking at all. If your pipeline treats HTTP 200 as data, block pages are in your dataset right now.

Three cheap habits catch most of it: record the status code, final URL, body length and body hash for every fetch; keep the first 2 KB of the body for any response you later count as content; and re-fetch a stratified sample from a second country in the same crawl window. The third one is what turns “we had a 3% error rate” into a measurement.

How a blocking claim is made

The two-vantage minimum

“Blocked in X” is a comparative claim, and it needs two measurements that differ in one thing. The canonical construction is in the two platform papers. ICLab probes from commercial VPN exits and volunteer machines and compares what it sees against a control node it operates: to detect DNS manipulation it “compares them with responses to matching DNS queries from our control node”, and its architecture diagram labels one component “Control vantage” [3Niaki, Arian Akhavan; Cho, Shinyoung; Weinberg, Zachary; Hoang, Nguyen Phong; Razaghpanah, Abbas; Christin, Nicolas; Gill, Phillipa (2020): "ICLab: A Global, Longitudinal Internet Censorship Measurement Platform", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]. Censored Planet takes the comparison remote, combining “four remote measurement techniques (Augur, Satellite/Iris, Quack, and Hyperquack)” into “synchronized measurements on 6 different Internet protocols (IP, DNS, HTTP, HTTPS, Echo and Discard)”, so the treatment vantage does not have to be a machine you own inside the country [4Raman, Ram Sundara; Shenoy, Prerana; Kohls, Katharina; Ensafi, Roya (2020): "Censored Planet: An Internet-wide, Longitudinal Censorship Observatory", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)].

Two designs recur, and the choice is not cosmetic:

  • Direct, from inside the condition — a VPS, a VPN exit, a volunteer device, a residential proxy. You see what a client there sees, including the block page. You also expose a real host to whatever the censor does next, which is an ethics question (Ethics) and sometimes a legal one.
  • Remote, by side channel — Augur uses “TCP/IP side channels to measure reachability between two Internet locations without directly controlling a measurement vantage point at either location”, by watching a shared IP-ID counter [5Pearce, Paul; Ensafi, Roya; Li, Frank; Feamster, Nick; Paxson, Vern (2017): "Augur: Internet-Wide Detection of Connectivity Disruptions", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]; Quack sends HTTP and TLS payloads through Internet-wide echo servers to detect application-layer keyword filtering, observing echo servers in “184 countries with 4458 unique ASes” [6VanderSloot, Benjamin; McDonald, Allison; Scott, Will; Halderman, J. Alex; Ensafi, Roya (2018): "Quack: Scalable Remote Measurement of Application-Layer Censorship", in: Proceedings of the USENIX Security Symposium. (Link)]; Hyperquack extends the idea to ordinary web servers, and FilterMap — the framework built on Hyperquack and Quack, with Censys responses used to check its signatures — produced 90 block-page clusters naming a vendor or a deploying actor across 103 countries, each unique block page manually verified to avoid false positives [7Raman, Ram Sundara; Stoll, Adrian; Dalek, Jakub; Ramesh, Reethika; Scott, Will; Ensafi, Roya (2020): "Measuring the Deployment of Network Censorship Filters at Global Scale", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]. The 2025 refinement drops the requirement that anything answer at all: Nourin et al. trigger interference against unresponsive addresses, reaching 8.6% of 15,158,447 IPv4 /24s for HTTP and 36.9% of the IPv6 /48s they could probe [8Nourin, Sadia; Rye, Erik C.; Bock, Kevin; Hoang, Nguyen Phong; Levin, Dave (2025): "Is Nobody There? Good! Globally Measuring Connection Tampering Without Responsive Endhosts", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)].

Remote methods buy scale and lose ground truth: you are inferring from a side channel, and the paper has to argue that what it saw is the censor rather than the reflector. Direct methods buy interpretability and lose coverage. Papers that can afford it do both.

Your control is a measurement, not an assumption. Two results in the population make this concrete. Hoang et al. compared the sets of domains that trigger China's HTTPS filter from each side of the border and found “about 1K domains that only trigger the HTTPS filter to inject RST packets when probed from inside the country” [9Hoang, Nguyen Phong; Dalek, Jakub; Crete-Nishihata, Masashi; Christin, Nicolas; Yegneswaran, Vinod; Polychronakis, Michalis; Feamster, Nick (2024): "GFWeb: Measuring the Great Firewall's Web Censorship at Scale", in: Proceedings of the USENIX Security Symposium. (Link)] — the direction of the probe changes the answer. Wu et al. (2025) found that Henan's provincial firewall blocked “only traffic going out of Henan”, not inbound traffic, and in one experiment blocked 4,196,532 domains, which the paper notes is “more than five times the 741,542 domains ever blocked by the GFW” in the same measurement [10Wu, Mingshi; Zohaib, Ali; Durumeric, Zakir; Houmansadr, Amir; Wustrow, Eric (2025): "A Wall Behind A Wall: Emerging Regional Censorship in China", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]. A single vantage point cannot see either fact.

And one vantage point is not one path. Bhaskar and Pearce showed that varying only the source port and the source IP within a subnet changes which censorship you observe — the finding design:dns states in full — and their follow-up measured the same effect globally as an artefact of ECMP routing rather than of policy [11Bhaskar, Abhishek; Pearce, Paul (2022): "Many Roads Lead To Rome: How Packet Headers Influence DNS Censorship Measurement", in: Proceedings of the USENIX Security Symposium. (Link)] [12Bhaskar, Abhishek; Pearce, Paul (2024): "Understanding Routing-Induced Censorship Changes Globally", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)]. Report your source addresses and your port-selection policy, and vary them.

What "blocked" has to be operationalised as

The population's papers check several signals per paper — the median is 4 of the 11 below — and the table is a menu, dated. Each column is a full-text probe over the 63 papers, so it is an upper bound on use (a sentence in related work counts) and a lower bound on the idea (a paper can compare against a control without using the word); the exact patterns are in the report on blocking_and_geodifference.

Signal Papers of 63 Share First Last ≤2019 2020–24 2025–26
status code or HTTP error 41 65.1% 2014 2026 13 22 6
injected RST or connection reset 35 55.6% 2011 2026 9 20 6
block-page fingerprint or keyword 33 52.4% 2013 2026 10 18 5
comparison against a control vantage 32 50.8% 2015 2025 8 19 5
DNS answer consistency 22 34.9% 2014 2026 4 14 4
page similarity or length outlier 14 22.2% 2014 2026 5 7 2
manual or in-browser validation 13 20.6% 2015 2024 7 6 0
repeated probe or retry 11 17.5% 2013 2025 4 5 2
supervised classifier or clustering 11 17.5% 2014 2026 2 7 2
TLS certificate verification 3 4.8% 2023 2024 0 3 0
LLM or transformer model 3 4.8% 2024 2026 0 1 2

Which of these is current, and which has been retired by a measurement. This is the part a 2018 methods section will get wrong:

  • A consistency check over the DNS answer alone is superseded. It was the state of the art in 2017: Iris identified “41,778 responses (0.31%) as manipulated, spread across 58 countries” out of 13,594,683 by comparing answers for consistency and independent verifiability [13Pearce, Paul; Jones, Ben; Li, Frank; Ensafi, Roya; Feamster, Nick; Weaver, Nick; Paxson, Vern (2017): "Global Measurement of DNS Manipulation", in: Proceedings of the USENIX Security Symposium. (Link)]. In 2023 Tsai et al. measured what that design costs — “a staggering number of 72.45% DNS resolutions that are tagged as 'manipulation' by consistency-based heuristics are false positives” — and replaced it by connecting to the returned address and checking the TLS certificate, which separates a censor from a CDN [14Tsai, Elisa; Kumar, Deepak; Sundara Raman, Ram; Li, Gavin; Eiger, Yael; Ensafi, Roya (2023): "CERTainty: Detecting DNS Manipulation at Scale using TLS Certificates", Proceedings on Privacy Enhancing Technologies 2023(3):122-137. (DOI)]. The same paper reports the other direction, 9.70% of true manipulations missed, and the awkward corollary for anyone relying on block pages: 82.39% of the invalid certificates it found “come without a blockpage”.
  • A page-length or size threshold alone was never good, and its numbers are published. Jones et al. built the original block-page detector and reported that a 30% page-size difference gives a 95% true-positive and 1.37% false-positive rate against a labelled corpus, while term-frequency clustering reached F1 = 0.98 against page-length clustering's 0.64 [15Jones, Ben; Lee, Tzu-Wen; Feamster, Nick; Gill, Phillipa (2014): "Automated Detection and Fingerprinting of Censorship Block Pages", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]. Four years later McDonald et al. measured the length heuristic's recall in the wild at 58.3% [16McDonald, Allison; Bernhard, Matthew; Valenta, Luke; VanderSloot, Benjamin; Scott, Will; Sullivan, Nick; Halderman, J. Alex; Ensafi, Roya (2018): "403 Forbidden: A Global View of CDN Geoblocking", in: Proceedings of the ACM Internet Measurement Conference. (DOI)].
  • Block-page fingerprint corpora are still current, and they are maintained as artefacts, not re-derived per paper: FilterMap's 90 clusters [7Raman, Ram Sundara; Stoll, Adrian; Dalek, Jakub; Ramesh, Reethika; Scott, Will; Ensafi, Roya (2020): "Measuring the Deployment of Network Censorship Filters at Global Scale", in: Proceedings of the Network and Distributed System Security Symposium. (Link)], ICLab's 48 previously undetected block-page signatures from 13 countries [3Niaki, Arian Akhavan; Cho, Shinyoung; Weinberg, Zachary; Hoang, Nguyen Phong; Razaghpanah, Abbas; Christin, Nicolas; Gill, Phillipa (2020): "ICLab: A Global, Longitudinal Internet Censorship Measurement Platform", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)], CERTainty's 226 new fingerprints [14Tsai, Elisa; Kumar, Deepak; Sundara Raman, Ram; Li, Gavin; Eiger, Yael; Ensafi, Roya (2023): "CERTainty: Detecting DNS Manipulation at Scale using TLS Certificates", Proceedings on Privacy Enhancing Technologies 2023(3):122-137. (DOI)], and Censored Planet's dns_blockpage_fingerprint repository (last commit 2025-05-01).
  • Fingerprinting the blocking device is the live front. Wallbleed found a buffer over-read in the Great Firewall's DNS injector that leaked “up to 125 bytes” of its memory [17Fan, Shencha; Sippe, Jackson; San, Sakamoto; Sheffey, Jade; Fifield, David; Houmansadr, Amir; Wedwards, Elson; Wustrow, Eric (2025): "Wallbleed: A Memory Disclosure Vulnerability in the Great Firewall of China", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]; Wu et al. identified Henan's firewall by a 10-byte RST payload and a fixed TTL [10Wu, Mingshi; Zohaib, Ali; Durumeric, Zakir; Houmansadr, Amir; Wustrow, Eric (2025): "A Wall Behind A Wall: Emerging Regional Censorship in China", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]; and Xue et al. fingerprint DPI devices by the ambiguities in how they parse a request [18Xue, Diwen; Huremagic, Armin; Wang, Wayne; Raman, Ram Sundara; Ensafi, Roya (2025): "Fingerprinting Deep Packet Inspection Devices by their Ambiguities", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)]. If you want to know who blocked you rather than whether, this is where to start.
  • LLMs have arrived in this literature, but not as the blocking classifier. All three papers the probe finds use them for something else: Tang et al. expand topics with a large language model and BERTopic to generate probe lists [19Tang, Jenny; Alvarez, Léo; Brar, Arjun; Hoang, Nguyen Phong; Christin, Nicolas (2024): "Automatic Generation of Web Censorship Probe Lists", Proceedings on Privacy Enhancing Technologies 2024(4):44-60. (DOI)]; Ablove et al. study Chinese LLM services as the censor [20Ablove, Anna; Chandrashekaran, Shreyas; Qiang, Xiao; Ensafi, Roya (2026): "Characterizing the Implementation of Censorship Policies in Chinese LLM Services", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]; and Lipphardt et al. classify refusals in 700k model responses with a fine-tuned DeBERTa cross-checked against ChatGPT and Gemini as judges, with two annotators hand-checking 12k classifications [21Lipphardt, Friedemann; Ali, Moonis; Banzer, Martin; Feldmann, Anja; Gosain, Devashish (2026): "There is No War in Ba Sing Se: A Global Analysis of Content Moderation in Large Language Models", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]. A widened probe over all 5,859 papers — every LLM/GPT/BERT term within 200 characters of a blocking term — returns 11 papers, 3 of them in the population, and none of the three decides whether a page is a block page. Treat “we classified block pages with an LLM” as an unoccupied slot rather than as current practice.

The false-positive rates the field has already measured

The strongest reason to publish your definition of “blocked” is that everyone who has checked theirs by hand has found it optimistic. These are the published error rates, each against its own ground truth:

Paper What was validated Error found
McDonald et al., IMC 2018 [16McDonald, Allison; Bernhard, Matthew; Valenta, Luke; VanderSloot, Benjamin; Scott, Will; Sullivan, Nick; Halderman, J. Alex; Ensafi, Roya (2018): "403 Forbidden: A Global View of CDN Geoblocking", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] 1,068 automated geoblocking detections in the Alexa Top 1M, manually inspected 286 (27%) false positives
Tsai et al., PoPETs 2023 [14Tsai, Elisa; Kumar, Deepak; Sundara Raman, Ram; Li, Gavin; Eiger, Yael; Ensafi, Roya (2023): "CERTainty: Detecting DNS Manipulation at Scale using TLS Certificates", Proceedings on Privacy Enhancing Technologies 2023(3):122-137. (DOI)] consistency-based DNS manipulation verdicts 72.45% false positives, 9.70% missed
Jones et al., IMC 2014 [15Jones, Ben; Lee, Tzu-Wen; Feamster, Nick; Gill, Phillipa (2014): "Automated Detection and Fingerprinting of Censorship Block Pages", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] block-page detection against a labelled corpus 95% TPR at 1.37% FPR
Knockel et al., PoPETs 2026 [22Knockel, Jeffrey; Dałek, Jakub; Aljizawi, Noura; Ahmed, Mohamed; Meletti, Levi; Lau, Justin (2026): "Banned Books: Analysis of Censorship on Amazon.com", Proceedings on Privacy Enhancing Technologies 2026(3):200-214. (DOI)] Amazon shipping-restriction messages, sampled and reviewed 47% of “temporarily out of stock” and 11% of “cannot be shipped” were false positives
McDonald et al., IMC 2018 [16McDonald, Allison; Bernhard, Matthew; Valenta, Luke; VanderSloot, Benjamin; Scott, Will; Sullivan, Nick; Halderman, J. Alex; Ensafi, Roya (2018): "403 Forbidden: A Global View of CDN Geoblocking", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] page-length outlier heuristic against hand-identified block pages recall 58.3%

Two lessons a reviewer will apply to your paper. First, hand-validate a sample and publish the rate — every row above exists because somebody did. Second, watch the base rate: McDonald et al. found a median of 3 inaccessible domains per country in the Alexa Top 10K, with a maximum of 71 in Syria [16McDonald, Allison; Bernhard, Matthew; Valenta, Luke; VanderSloot, Benjamin; Scott, Will; Sullivan, Nick; Halderman, J. Alex; Ensafi, Roya (2018): "403 Forbidden: A Global View of CDN Geoblocking", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]. When the true positives number in the tens, a percent or two of false positives is the whole result, which is the same arithmetic design:dns applies to Iris's 0.31%.

Blocking is not stable, so one probe is not a measurement

Residual censorship, retries and route changes all mean the second probe disagrees with the first. GFWeb observed traffic dropping continue “for up to 350 seconds for TCP packets that share the same three-tuple” after a trigger [9Hoang, Nguyen Phong; Dalek, Jakub; Crete-Nishihata, Masashi; Christin, Nicolas; Yegneswaran, Vinod; Polychronakis, Michalis; Feamster, Nick (2024): "GFWeb: Measuring the Great Firewall's Web Censorship at Scale", in: Proceedings of the USENIX Security Symposium. (Link)], while Wu et al. (2025) found Henan's firewall performed no residual censorship, so the same three-tuple worked immediately [10Wu, Mingshi; Zohaib, Ali; Durumeric, Zakir; Houmansadr, Amir; Wustrow, Eric (2025): "A Wall Behind A Wall: Emerging Regional Censorship in China", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]. McDonald et al. required agreement across 23 samples at an 80% threshold before calling a domain-country pair blocked [16McDonald, Allison; Bernhard, Matthew; Valenta, Luke; VanderSloot, Benjamin; Scott, Will; Sullivan, Nick; Halderman, J. Alex; Ensafi, Roya (2018): "403 Forbidden: A Global View of CDN Geoblocking", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]. Ablove et al. measured Chinese LLM services five times per query and found output blocking was mostly inconsistent — 29 of 349 output-blocked queries were blocked in all five samples — while input blocking was near-deterministic [20Ablove, Anna; Chandrashekaran, Shreyas; Qiang, Xiao; Ensafi, Roya (2026): "Characterizing the Implementation of Censorship Policies in Chinese LLM Services", in: Proceedings of the Network and Distributed System Security Symposium. (Link)].

Report the number of repeats, the agreement rule, and the interval between probes. A single-shot measurement of a stateful censor measures the censor's state.

Blocking follows the protocol

A blocking measurement is pinned to a protocol, and the protocol moves. Quack found that switching the probe from HTTP to TLS took Iran's blocked-domain count in its own experiment from 25 to 374 [6VanderSloot, Benjamin; McDonald, Allison; Scott, Will; Halderman, J. Alex; Ensafi, Roya (2018): "Quack: Scalable Remote Measurement of Application-Layer Censorship", in: Proceedings of the USENIX Security Symposium. (Link)]; Elmenhorst et al. sent paired HTTPS/TCP and HTTP/3/QUIC requests and found censors that handled one and not the other [23Elmenhorst, Kathrin; Schütz, Bertram; Aschenbruck, Nils; Basso, Simone (2021): "Web censorship measurements of HTTP/3 over QUIC", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]; and Wu et al. (2023) showed China blocking fully encrypted traffic on the basis of entropy and byte distributions rather than on any name at all [24Wu, Mingshi; Sippe, Jackson; Sivakumar, Danesh; Burg, Jack; Anderson, Peter; Wang, Xiaokang; Bock, Kevin; Houmansadr, Amir; Levin, Dave; Wustrow, Eric (2023): "How the Great Firewall of China Detects and Blocks Fully Encrypted Traffic", in: Proceedings of the USENIX Security Symposium. (Link)].

Encrypted DNS is the clearest case, because it is both a target and a workaround. Lu et al.'s 2019 deployment measurement found that “Over 99% global users can normally access large DNS-over-Encryption servers” [25Lu, Chaoyi; Liu, Baojun; Li, Zhou; Hao, Shuang; Duan, Hai-Xin; Zhang, Mingming; Leng, Chunying; Liu, Ying; Zhang, Zaifeng; Wu, Jianping (2019): "An End-to-End, Large-Scale Measurement of DNS-over-Encryption: How Far Have We Come?", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]; five years later Li et al. measured the same question from 5,031 vantage points in 102 countries, against control nodes, and found 592K of 10M DoE-over-IPv4 queries (5.92%) blocked [26Li, Ruixuan; Liu, Baojun; Lu, Chaoyi; Duan, Haixin; Shao, Jun (2024): "A Worldwide View on the Reachability of Encrypted DNS Services", in: Proceedings of the ACM Web Conference. (DOI)]. Their decision tree has seven blocking types — pre-resolve, ping, TCP, TLS, QUIC version negotiation, QUIC and response — and the largest is the first hop: “62.83% of DoEv4 services are inaccessible due to Ping blocking”. That is worth pausing on, because it is this page's central point in miniature: the headline depends on where in the stack you decided to call something blocked, and a taxonomy with seven layers finds most of its blocking at the layer a cruder instrument never reaches. If your study depends on a protocol being reachable, that is a claim to measure, not assume.

Blocking also follows the client. A Tor exit or a datacenter address is refused by sites that serve everyone else: Khattak et al. measured differential treatment of Tor users [27Khattak, Sheharbano; Fifield, David; Afroz, Sadia; Javed, Mobin; Sundaresan, Srikanth; McCoy, Damon; Paxson, Vern; Murdoch, Steven J. (2016): "Do You See What I See? Differential Treatment of Anonymous Users", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] — cited here from outside the corpus, because NDSS 2016 is one of the venue-years the extraction is missing — and Singh et al. (2017), which is in the population, characterised website-side Tor exit blocking directly [28Singh, Rachee; Nithyanand, Rishab; Afroz, Sadia; Pearce, Paul; Tschantz, Michael Carl; Gill, Phillipa; Paxson, Vern (2017): "Characterizing the Nature and Dynamics of Tor Exit Blocking", in: 26th USENIX Security Symposium (USENIX Security 17), pp. 325-341. USENIX Association. (Link)]. If your vantage point is unusual, part of what you measure is how the web treats that vantage point — see Crawling location.

Who is doing the blocking

Each paper in the population was hand-labelled with who the blocking decision belongs to, one label per paper, and the split is produced by the report rather than tallied in prose. Read it as an index: the fastest way into the method is the study of the censor you care about.

Who is blocking Papers of 63 Share Years
many countries at once (an observatory, a global scan, a method paper) 23 36.5% 2013–2026
China 13 20.6% 2015–2026
website operators and CDNs 4 6.3% 2018–2025
Russia 4 6.3% 2020–2023
a platform — IPFS, LLM services, a retailer 4 6.3% 2024–2026
Cuba 2 3.2% 2015–2024
website operators, law-driven (GDPR wall, CCPA geofence) 2 3.2% 2022–2024
app stores and developers 2 3.2% 2022–2025
Egypt and Libya 1 1.6% 2011
Pakistan 1 1.6% 2014
Syria 1 1.6% 2014
website operators refusing Tor exits 1 1.6% 2017
India 1 1.6% 2018
Kazakhstan 1 1.6% 2020
Turkmenistan 1 1.6% 2023
protective-DNS resolver operators 1 1.6% 2024
Iran 1 1.6% 2025

Three things follow. China is a fifth of the literature on its ownDNS injection [29Hoang, Nguyen Phong; Niaki, Arian Akhavan; Dalek, Jakub; Knockel, Jeffrey; Lin, Pellaeon; Marczak, Bill; Crete-Nishihata, Masashi; Gill, Phillipa; Polychronakis, Michalis (2021): "How Great is the Great Firewall? Measuring China's DNS Censorship", in: Proceedings of the USENIX Security Symposium. (Link)], HTTP(S) filtering at scale [9Hoang, Nguyen Phong; Dalek, Jakub; Crete-Nishihata, Masashi; Christin, Nicolas; Yegneswaran, Vinod; Polychronakis, Michalis; Feamster, Nick (2024): "GFWeb: Measuring the Great Firewall's Web Censorship at Scale", in: Proceedings of the USENIX Security Symposium. (Link)], keyword filtering [30Rambert, Raymond; Weinberg, Zachary; Barradas, Diogo; Christin, Nicolas (2021): "Chinese Wall or Swiss Cheese? Keyword filtering in the Great Firewall of China", in: Proceedings of the ACM Web Conference. (DOI)], fully encrypted traffic [24Wu, Mingshi; Sippe, Jackson; Sivakumar, Danesh; Burg, Jack; Anderson, Peter; Wang, Xiaokang; Bock, Kevin; Houmansadr, Amir; Levin, Dave; Wustrow, Eric (2023): "How the Great Firewall of China Detects and Blocks Fully Encrypted Traffic", in: Proceedings of the USENIX Security Symposium. (Link)], SNI-based QUIC censorship, a provincial second firewall [10Wu, Mingshi; Zohaib, Ali; Durumeric, Zakir; Houmansadr, Amir; Wustrow, Eric (2025): "A Wall Behind A Wall: Emerging Regional Censorship in China", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)], a memory-disclosure bug in the injector [17Fan, Shencha; Sippe, Jackson; San, Sakamoto; Sheffey, Jade; Fifield, David; Houmansadr, Amir; Wedwards, Elson; Wustrow, Eric (2025): "Wallbleed: A Memory Disclosure Vulnerability in the Great Firewall of China", in: Proceedings of the Network and Distributed System Security Symposium. (Link)], and now LLM services [20Ablove, Anna; Chandrashekaran, Shreyas; Qiang, Xiao; Ensafi, Roya (2026): "Characterizing the Implementation of Censorship Policies in Chinese LLM Services", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]. If you are reading one censor to learn the method, read that one: its papers are the most methodologically explicit in the corpus, and the GFW's behaviour has been reverse-engineered in more detail than any other censor's.

Everything else is thin. Iran, India, Pakistan, Syria, Turkmenistan and Kazakhstan have one paper each in these seven venues, all with a different mechanism — DNS poisoning and blockpage injection [31Tai, Jonas; Sengottuvelavan, Karthik Nishanth; Whiting, Peter; Hoang, Nguyen Phong (2025): "IRBlock: A Large-Scale Measurement Study of the Great Firewall of Iran", in: Proceedings of the USENIX Security Symposium. (Link)], per-ISP URL filters across nine ISPs [32Yadav, Tarun Kumar; Sinha, Akshat; Gosain, Devashish; Sharma, Piyush Kumar; Chakravarty, Sambuddho (2018): "Where The Light Gets In: Analyzing Web Censorship Mechanisms in India", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], leaked Blue Coat proxy logs [33Chaabane, Abdelberi; Chen, Terence; Cunche, Mathieu; De Cristofaro, Emiliano; Friedman, Arik; Kaafar, Mohamed Ali (2014): "Censorship in the Wild: Analyzing Internet Filtering in Syria", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], remotely reverse-engineered blocking rules [34Nourin, Sadia; Tran, Van; Jiang, Xi; Bock, Kevin; Feamster, Nick; Hoang, Nguyen Phong; Levin, Dave (2023): "Measuring and Evading Turkmenistan's Internet Censorship: A Case Study in Large-Scale Measurements of a Low-Penetration Country", in: Proceedings of the ACM Web Conference. (DOI)], HTTPS interception rather than blocking [35Raman, Ram Sundara; Evdokimov, Leonid; Wustrow, Eric; Halderman, J. Alex; Ensafi, Roya (2020): "Investigating Large Scale HTTPS Interception in Kazakhstan", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]. The mechanism does not port, so “we ran the standard censorship measurement” is not a method; the country study you are extending will have told you what to look for, and a different country may not do that at all.

And the largest single row is nobody in particular — 23 of 63 measure many countries at once, which is what an observatory or a global scan produces. That is the shape of the field: a handful of deep single-censor studies, a broad layer of global instruments, and very little in between. The leaked-logs case is the only one in the population with ground truth [33Chaabane, Abdelberi; Chen, Terence; Cunche, Mathieu; De Cristofaro, Emiliano; Friedman, Arik; Kaafar, Mohamed Ali (2014): "Censorship in the Wild: Analyzing Internet Filtering in Syria", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]: everywhere else the censor's rule set is inferred, which is why the false-positive rates above matter so much.

Geoblocking and geodifference

Geoblocking is the version an ordinary web measurement meets, and it is not rare in absolute terms even though it is rare per site. McDonald et al. found 4.4% of Alexa Top Million domains used their CDN's geoblocking feature in at least one country, with Syria (71), Iran (67), Sudan (66) and Cuba (66) leading the per-country counts in the Top 10K [16McDonald, Allison; Bernhard, Matthew; Valenta, Luke; VanderSloot, Benjamin; Scott, Will; Sullivan, Nick; Halderman, J. Alex; Ensafi, Roya (2018): "403 Forbidden: A Global View of CDN Geoblocking", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]. The pattern is sanctions-shaped, and the follow-up work says so directly: Ablove et al. decompose Cuban geoblocking by layer — 37 domains at DNS (6.8%), 97 at the TCP handshake (17.7%), 23 at TLS (4.2%), 24 at the HTTP GET (4.4%), 395 by block page (72.3%) [1Ablove, Anna; Chandrashekaran, Shreyas; Le, Hieu; Raman, Ram Sundara; Ramesh, Reethika; Oppenheimer, Harry; Ensafi, Roya (2024): "Digital Discrimination of Users in Sanctioned States: The Case of the Cuba Embargo", in: Proceedings of the USENIX Security Symposium. (Link)]. Blocking a country happens at whichever layer is convenient. A measurement that only reads HTTP bodies sees the largest share of it — 72.3% here — and misses the quarter that never gets that far, and one that only counts connection failures sees the opposite quarter.

Geodifference is the larger phenomenon and it does not need anybody to be blocked. The site loads, and it is a different site:

  • Kumar et al. collected Google Play from 26 countries and found 3,672 of 5,385 apps geoblocked in at least one country, 2,419 (44.9%) blocked by the developer rather than by a government, 61 subject to government takedown requests, and — the part that should worry a privacy measurement — 596 apps differing in their binaries, 127 in requested permissions, 118 in embedded ad trackers and 103 in their privacy policies across countries [36Kumar, Renuka; Virkud, Apurva; Sundara Raman, Ram; Prakash, Atul; Ensafi, Roya (2022): "A Large-scale Investigation into Geodifferences in Mobile Apps", in: Proceedings of the USENIX Security Symposium. (Link)]. Guo et al. reach the same conclusion from the code side in 2025, differencing call paths across country-specific APKs [37Guo, Jiawei; Nong, Yu; Lin, Zhiqiang; Cai, Haipeng (2025): "Code Speaks Louder: Exploring Security and Privacy Relevant Regional Variations in Mobile Applications", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)].
  • Roth et al. crawled 10,000 sites under varying countries, user agents, network paths and languages and found inconsistent client-side security headers on 321 sites [2Roth, Sebastian; Calzavara, Stefano; Wilhelm, Moritz; Rabitti, Alvise; Stock, Ben (2022): "The Security Lottery: Measuring Client-Side Web Security Inconsistencies", in: Proceedings of the USENIX Security Symposium, pp. 2047-2064. (Link)].
  • Singh et al. (2025) had volunteers in 23 countries load regional and government sites, which is the tracking-flow version of the same question [38Singh, Sachin Kumar; Ricci, Robert; Gamero-Garrido, Alexander (2025): "Where in the World Are My Trackers? Mapping Web Tracking Flow Across Diverse Geographic Regions", in: Proceedings of the ACM Internet Measurement Conference. (DOI)].
  • Ramesh et al., measuring the network response to the 2022 invasion of Ukraine, found geoblocking running in the other direction as well: 136 Russian government domains (25.09%) blocked access from every tested country outside Russia, and a further 112 (20.66%) were unreachable from anywhere except Russia and Kazakhstan [39Ramesh, Reethika; Raman, Ram Sundara; Virkud, Apurva; Dirksen, Alexandra; Huremagic, Armin; Fifield, David; Rodenburg, Dirk; Hynes, Rod; Madory, Doug; Ensafi, Roya (2023): "Network Responses to Russia's Invasion of Ukraine in 2022: A Cautionary Tale for Internet Freedom", in: Proceedings of the USENIX Security Symposium. (Link)].
  • Knockel et al. tested product availability by region on one retailer and found 17,842 products restricted from shipment to at least one world region, including 8,965 of 796,081 books (1.1%) restricted to at least one of Saudi Arabia, the UAE, Qatar or Yemen [22Knockel, Jeffrey; Dałek, Jakub; Aljizawi, Noura; Ahmed, Mohamed; Meletti, Levi; Lau, Justin (2026): "Banned Books: Analysis of Censorship on Amazon.com", Proceedings on Privacy Enhancing Technologies 2026(3):200-214. (DOI)].

For a measurement that is about something else, the consequence is a sampling one: if your seed list is global and your vantage point is not, geodifference is a confound in every per-site number you report, and Crawling location is the page that argues the choice.

GDPR walls: the case that bites an ordinary privacy crawl

After May 2018 some publishers responded to the GDPR by refusing European visitors outright. This is the blocking most likely to hit a privacy crawl — it is aimed at exactly the vantage point a consent study needs — and it is the thinnest slice of this literature.

One paper in the population measures a GDPR wall, and one more measures the closest analogue — a law-driven geofence under a different law. Both do it incidentally to another question:

  • Take et al., auditing privacy rights on 20 people-search websites from the US, report that “nine of the sites attempt to block EU IPs” and that the researchers used a VPN to reach them [40Take, Kejsi; Young, Jordyn; Bhalerao, Rasika; Gallagher, Kevin; Forte, Andrea; McCoy, Damon; Greenstadt, Rachel (2024): "What to Expect When You’re Accessing: An Exploration of User Privacy Rights in People Search Websites", in: Proceedings on Privacy Enhancing Technologies. (DOI)].
  • Van Nortwick and Wilson, measuring CCPA compliance — California, not the GDPR — treat geofencing of the opt-out link as one of the behaviours to measure rather than as noise [41Van Nortwick, Maggie; Wilson, Christo (2022): "Setting the Bar Low: Are Websites Complying With the Minimum Requirements of the CCPA?", in: Proceedings on Privacy Enhancing Technologies. (DOI)]. It is the method transposed to another jurisdiction, not a second GDPR datapoint.

Three more papers hit the wall and said so without measuring it: Dimova et al. excluded from scope every site that “did not explicitly block EU-users”, i.e. scoped their population around the wall [42Dimova, Yana; Van Goethem, Tom; Joosen, Wouter (2023): "Everybody's Looking for SSOmething: A large-scale evaluation on the privacy of OAuth authentication on the web", Proceedings on Privacy Enhancing Technologies 2023(4). (DOI)]; Shezan et al. treat “unless they specifically block EU traffic” as the default assumption for GDPR applicability when deciding which WordPress plugins are in scope [43Shezan, Faysal Hossain; Su, Zihao; Kang, Mingqing; Phair, Nicholas; Thomas, Patrick William; van Dam, Michelangelo; Cao, Yinzhi; Tian, Yuan (2023): "CHKPLUG: Checking GDPR Compliance of WordPress Plugins via Cross-language Code Property Graph", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]; and Haghighi et al., coding developer discussions on Reddit, record practitioners proposing to block EU users as a compliance strategy [44Haghighi, Sara; LaChance, Clark; Pourghasemi Fatideh, Ali; Breaux, Travis; Ghanavati, Sepideh (2026): "The Role of Online Forums in Developer Understanding of Privacy Law - A Reddit Case Study", in: Proceedings on Privacy Enhancing Technologies. (DOI)].

Nobody in these seven venues has measured how common GDPR-driven blocking is on any general population of websites. The two measurements above are on 20 hand-picked people-search sites and on the CCPA opt-out link; neither is a prevalence estimate for the web, or even for news publishers. A tight full-text probe for the phenomenon returns 9 papers of 5,859 and a deliberately loose one — the tight patterns plus three wider ones — returns 24; reading all 24 turns up nothing else, and an independent re-probe during review (a GDPR term within a sentence of a blocking or geofencing term and a Europe term, whitespace-collapsed, over every full text) found four hits and no new paper. So the honest state of knowledge is: the mechanism is documented, its prevalence is not, and the best public figures remain journalistic 1). If you are looking for a well-defined, cheap, publishable measurement, this is one, and the method is on this page.

Two things to know before running it. HTTP 451 exists and is barely used: RFC 7725 “An HTTP Status Code to Report Legal Obstacles” has been a Proposed Standard since February 2016 and is not obsoleted,2) but ICLab observed only “23 unique websites that return status 451, from vantages in 21 countries” [3Niaki, Arian Akhavan; Cho, Shinyoung; Weinberg, Zachary; Hoang, Nguyen Phong; Razaghpanah, Abbas; Christin, Nicolas; Gill, Phillipa (2020): "ICLab: A Global, Longitudinal Internet Censorship Measurement Platform", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]. Do not build a detector around it. And the EU's own geo-blocking law does not cover this case: Regulation (EU) 2018/302 addresses “unjustified geo-blocking” in the internal market — a trader refusing a customer in another member state — not a US publisher refusing the EU.3) Framing a GDPR-wall study as a 2018/302 compliance audit would be a category error.

Instruments and datasets, and the denominator each one hands you

You do not have to build a censorship measurement platform, and you should not. What you do have to know is that these are observatories with their own populations, and reusing one means inheriting its denominator — which is Existing datasets' subject, and that page counts a “censorship list (Citizen Lab, OONI)” family of 17 papers inside its own reuse population.

Checked by fetching each one on 2026-09-10; the commands and full output are on blocking_and_geodifference.

Instrument What it gives you State on 2026-09-10
OONI volunteer-run probe measurements, per-country, per-URL, since 2012; aggregation API and daily open-data dumps Live and current. api.ooni.io/api/v1/aggregation answers (a one-week Iran web_connectivity query returned 143,480 measurements, 42,483 of them confirmed); the ooni-data-eu-fra S3 bucket had all 24 hour-prefixes for the previous day. Probe CLI v3.30.0, 2026-07-27
Citizen Lab test lists the community-maintained URL lists nearly every study probes Live. github.com/citizenlab/test-lists last commit 2026-09-09, 150 per-country CSVs plus global.csv
Censored Planet remote-measurement observatory (Quack, Hyperquack, Satellite, Iris lineage), raw data, dashboard, BigQuery pipeline Live and current, but its URLs have moved: the censoredplanet.org/data/raw path cited by 2020–2023 papers is a 404, and the entry points are now dashboard.censoredplanet.org, data.censoredplanet.org (a GraphQL endpoint) and docs.censoredplanet.org. The dashboard reports 116,924,640,072 total measurements and 744,324,571 in the last 30 days
Censored Planet tooling geoinspector (geoblocking across DNS/TCP/TLS/HTTP), CenTrace (censorship traceroute), CenFuzz, dns_blockpage_fingerprint, geodiff-app Public, unevenly maintained. Last pushes: dns_blockpage_fingerprint and geodiff-app 2025-05-01, censoredplanet-analysis 2025-06-29, CenFuzz 2024-02-04, geoinspector 2023-05-19, CenTrace 2023-03-30. None archived
ICLab the 2020 IEEE S&P platform's data and dashboard The public site is unreachable; the project is not dead. iclab.org serves HTTP 200 only if you ignore certificate errors: its Let's Encrypt certificate expired 2025-07-18, so any client that verifies certificates fails. But the github.com/iclab organisation's centinel-prime (a containerised rewrite of the probing client) was last pushed 2026-05-05. Cite the paper, do not expect to download the data, and ask the authors rather than assuming the effort has stopped
GFWatch / GFWeb continuous DNS (GFWatch) and HTTP(S) (GFWeb) measurement of China's Great Firewall, with public dashboards Sites up, data stale. gfwatch.org renders, and its newest last_checked date is 2024/08/06; gfweb.ca's is 2024/11/19. Cite the papers [29Hoang, Nguyen Phong; Niaki, Arian Akhavan; Dalek, Jakub; Knockel, Jeffrey; Lin, Pellaeon; Marczak, Bill; Crete-Nishihata, Masashi; Gill, Phillipa; Polychronakis, Michalis (2021): "How Great is the Great Firewall? Measuring China's DNS Censorship", in: Proceedings of the USENIX Security Symposium. (Link)] [9Hoang, Nguyen Phong; Dalek, Jakub; Crete-Nishihata, Masashi; Christin, Nicolas; Yegneswaran, Vinod; Polychronakis, Michalis; Feamster, Nick (2024): "GFWeb: Measuring the Great Firewall's Web Censorship at Scale", in: Proceedings of the USENIX Security Symposium. (Link)]; do not cite the dashboards as current
IODA (Georgia Tech) outage and shutdown detection from BGP, telescope and active probing, with a public API Live and current: the dashboard showed data through 2026-09-10 and api.ioda.inetintel.cc.gatech.edu/v2/signals/raw/country/… answers without a key
M-Lab throughput and some interference tests from a large volunteer base Live; used in this corpus mostly for performance rather than blocking
Cloudflare Radar outage centre operator-side view of shutdowns and traffic anomalies Live, but bot-walled to scripted clients (HTTP 403 to curl). Cloudflare Radar covers the Radar API and says explicitly that it is not a catalogue of the outage products

The schema cannot see a dataset you only read. In the extraction, tools[] names OONI as used or produced in 7 papers of 5,859; 52 papers name it somewhere in their full text, 34 of them in this page's population. Censored Planet is 4 against 47, ICLab 1 against 42, and the Citizen Lab test lists are named in tools[] by zero papers and in 35 full texts. A structured query about instruments will therefore undercount reuse of an observatory by roughly an order of magnitude — and the full-text column is a mention count, not a usage count, since several of these names are homographs (Satellite, Geneva, Encore, Augur, and BERT inside the surname Deibert).

Use in publications

The figures below come from a structured extraction over 5,859 full-text papers from CCS, IMC, NDSS, PoPETs, USENIX Security, TheWebConf and IEEE S&P, 2010–2026. Every claim here is a claim about those seven venues: EuroS&P, ACSAC, RAID, AsiaCCS, CHI, SOUPS, PAM, FOCI and the Internet Measurement workshops are not in it, and FOCI in particular is where a large part of the censorship-measurement community publishes. The 2025 and 2026 venue-years are provisional — CCS 2026 and IMC 2026 have not been held, and IEEE S&P and WWW 2026 abstracts are not yet in the selection source — so treat them as under-represented by construction. Methodology and limitations are at the end of this section.

The population, and how it was drawn

There is no field in the extraction that says “this paper measures blocking”, so the population was built and then hand-audited. Nine probes — three over title and summary, one over detection[].phenomenon, one over tools[], three over the full text, and one hand-added paper with its reason recorded — produced 275 candidates, and every candidate carries a written verdict. The inclusion rule, as amended during review:

A paper is in if it measures whether some client could reach some content or service and attributes the failures to a deliberate blocking decision by someone other than the client — a state, an ISP, a resolver operator, a CDN, a platform or the content owner — where that decision is keyed on who or where the client is, or on the content's acceptability to an authority or platform.

…or on the traffic looking like an attempt to evade such a decision — Shadowsocks, a fully encrypted flow, an SNI. Measuring how a deployed censor detects and blocks such traffic is in; proposing a detector for it is not.

Out: blocking keyed on the content being malicious; blocking keyed on the client looking automated; blocking the client chose (its own ad blocker); a system proposed to evade blocking or a detector proposed for finding circumvention; and studies of what a takedown or post-moderation did to later behaviour.

That leaves 63 papers, 1.1% of the corpus:

Family Papers What it is
network interference 45 censorship, DNS manipulation, connection tampering, throttling, interception
server-side refusal or differentiation 11 geoblocking, geodifference, sanctions
platform-side moderation of access 5 IPFS, LLM services, a retailer's shipping restrictions
law-driven blocking 2 GDPR walls, CCPA geofencing
the probe list as the object of study 2 [45Weinberg, Zachary; Sharif, Mahmood; Szurdi, Janos; Christin, Nicolas (2017): "Topics of Controversy: An Empirical Analysis of Web Censorship Lists", in: Proceedings on Privacy Enhancing Technologies. (DOI)] [19Tang, Jenny; Alvarez, Léo; Brar, Arjun; Hoang, Nguyen Phong; Christin, Nicolas (2024): "Automatic Generation of Web Censorship Probe Lists", Proceedings on Privacy Enhancing Technologies 2024(4):44-60. (DOI)]
client-side filtering products 1 protective DNS, and its over-blocking [46Liu, Mingxuan; Zhang, Yiming; Li, Xiang; Lu, Chaoyi; Liu, Baojun; Duan, Haixin; Zheng, Xiaofeng (2024): "Understanding the Implementation and Security Implications of Protective DNS Services", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]
population (union) 63 three papers carry two families

Adjacent, counted and excluded: 29 circumvention or evasion systems and censor-side detectors (Geneva, domain fronting, MassBrowser, SpotProxy, QUICstep and so on — they measure blocking instrumentally, to defeat it), 12 outage-detection papers that measure unreachability without attributing it to anybody's decision (see below), 5 interview or survey studies of living with censorship, 3 papers that ran into an EU wall without measuring it, 2 instruments for measuring from a place, and 1 SoK. And 161 of the 275 candidates (58.5%) are off topic, each with a named reason on the provenance page: 115 incidental mentions, 8 VPN-ecosystem papers, 7 that name Tor Metrics, 5 takedown-effect studies, 4 proxy-detection papers, and homonyms including statistical censoring, concept censorship in a diffusion model, CPU throttling, and iRiS the iOS analyser colliding with Iris the DNS platform. That ratio is the honest cost of a keyword probe in this area, and it is why the queued estimate for this page (74 candidate papers from a title-and-summary probe) is neither an upper nor a lower bound on the population. Three of the nine probes, and the one hand-added paper, were added after a draft existed, each because a later check found a paper the candidate set could not see — which is the best available evidence that a tenth probe would find another one.

Where the work is

Venue Papers in population Corpus papers Share of venue
IMC 19 638 3.0%
USENIX Security 14 1,410 1.0%
NDSS 8 701 1.1%
PoPETs 7 510 1.4%
TheWebConf 6 843 0.7%
IEEE S&P 5 767 0.7%
CCS 4 990 0.4%

IMC carries the largest share of any venue — 3.0% of its papers, against 1.4% for the next one — and a third of the population by count. Note what the platform mix says: web is 32 of the 63 (50.8%) and other-online-service 47 — this is a networking literature as much as a web one, and only 12 of the 63 ran a browser crawl while 50 ran a network scan or probe and 34 reanalysed an existing dataset. If you arrive from web measurement, the methods here will feel like scanning, because they are: Internet scanning is the instrument page.

Window Population network server-side law-driven platform filtering probe list (circumvention)
2010–2014 5 5 0 0 0 0 0 1
2015–2019 14 11 2 0 0 0 1 8
2020–2024 32 22 5 2 2 1 1 15
2025–2026 (provisional) 12 7 4 0 3 0 0 4

Family columns overlap by three papers and the circumvention column is outside the population, so rows do not sum to the second column.

The composition moves even though the volume does not. Network-level censorship measurement is the constant; server-side and platform-side blocking is where the recent work is5 of those 14 papers (the two families overlap by two, so 11 + 5 is 14 distinct) are from 2025–2026, in a window that is under-populated by construction. All three of the platform papers in that window are about services rather than networks: Chinese LLM services [20Ablove, Anna; Chandrashekaran, Shreyas; Qiang, Xiao; Ensafi, Roya (2026): "Characterizing the Implementation of Censorship Policies in Chinese LLM Services", in: Proceedings of the Network and Distributed System Security Symposium. (Link)], moderation differences across 12 VPN locations in commercial models [21Lipphardt, Friedemann; Ali, Moonis; Banzer, Martin; Feldmann, Anja; Gosain, Devashish (2026): "There is No War in Ba Sing Se: A Global Analysis of Content Moderation in Large Language Models", in: Proceedings of the Network and Distributed System Security Symposium. (Link)], and region-restricted product availability at one retailer [22Knockel, Jeffrey; Dałek, Jakub; Aljizawi, Noura; Ahmed, Mohamed; Meletti, Levi; Lau, Justin (2026): "Banned Books: Analysis of Censorship on Amazon.com", Proceedings on Privacy Enhancing Technologies 2026(3):200-214. (DOI)]; the platform family's other two are the 2024 IPFS pair [47Sokoto, Saidu; Balduf, Leonhard; Trautwein, Dennis; Wei, Yiluo; Tyson, Gareth; Castro, Ignacio; Ascigil, Onur; Pavlou, George; Korczyński, Maciej; Scheuermann, Björn; Król, Michał (2024): "Guardians of the Galaxy: Content Moderation in the InterPlanetary File System", in: Proceedings of the USENIX Security Symposium. (Link)]. Read it as a direction, not as a trend: five papers in two thin venue-years cannot carry a slope, and the page does not test one.

Do these papers have a control?

vantage[] is the extraction's record of where a measurement was taken from; the location strings are free text and are folded through scripts/geo.mjs, with the 15 unmapped strings printed on the provenance page.

Of the 63 papers in the population Papers Share
carry at least one vantage tuple 62 98.4%
name at least one location, after folding 53 85.5% (of 62)
name two or more distinct locations, or an explicit multi-country vantage 40 64.5% (of 62)
name exactly one location and no multi-country claim 13 21.0% (of 62)
say nothing about vantage infrastructure on any tuple 9 14.5% (of 62)
say “control vantage” or an equivalent in the full text (tight probe) 32 50.8% (of 63)
same, loose probe — the tight patterns plus any “control” near a vantage word 47 74.6% (of 63)

Read the last two rows carefully: they are probes for the word, not for the design. A paper can compare against a control and never write “control vantage”, and the gap between 32 and 47 is exactly how much a probe's width decides the number. The loose probe is built from the tight one by construction, and the report throws if any tight hit escapes it — without that, “tight” and “loose” quietly become two different questions and their counts stop being comparable at all. The defensible claim is the one the numbers support: one in five states a single vantage location (13 of the 62 with a vantage tuple), and one in four never puts the word control near a vantage point even loosely (16 of 63) — which is a reporting finding, and a reason to make your own control explicit.

Infrastructure, where stated, is cloud (25 papers) and university (24) ahead of commercial VPN (12) and volunteer devices (8) — with the caveat crawling_location documents: a datacenter address is treated differently from a residential one, so a cloud-only blocking measurement can mistake bot management for censorship.

What the corpus does not cover

  • Shutdowns as a deliberate act. There is an outage-detection literature in these venues — 12 papers, from optical-layer failures in a backbone to detecting outages from Google Trends, and including a three-year active scan of all Ukrainian IPv4 space [48Holzbauer, Florian; Strobl, Sebastian; Ullrich, Johanna (2025): "Tracking Internet Disruptions in Ukraine: Insights from Three Years of Active Full Block Scans", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] — but it measures unreachability without attributing it to anybody's decision, which is why none of them is in the population. That Ukraine study [48Holzbauer, Florian; Strobl, Sebastian; Ullrich, Johanna (2025): "Tracking Internet Disruptions in Ukraine: Insights from Three Years of Active Full Block Scans", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], the closest thing in the corpus to shutdown measurement, correlates its outages with the national grid operator's power-outage data (r = 0.725 outside frontline regions) rather than with an order to disconnect. Exactly one paper in the 275-paper candidate set attributes a country-wide outage to censorship, and it is from 2011 [49Dainotti, Alberto; Squarcella, Claudio; Aben, Emile; Claffy, Kimberly C.; Chiesa, Marco; Russo, Michele; Pescapè, Antonio (2011): "Analysis of country-wide internet outages caused by censorship", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]. The instruments that carry the attributed version of this work today — IODA, and the NGO reporting around Access Now's #KeepItOn — publish outside these seven venues.
  • GDPR-driven blocking prevalence, as above: one incidental measurement on 20 hand-picked sites, plus one CCPA analogue.
  • Mobile network blocking. mobile is 2 of 63, and cellular-network measurement platforms appear as instruments rather than as blocking measurements.
  • Eleven venue-years are missing outright. The extraction has no papers at all from NDSS 2010, 2011, 2016 and 2018, or from PoPETs 2010–2014 (CCS and IMC 2026 have not been held). That is why Khattak et al.'s NDSS 2016 study of differential treatment of Tor users, cited above, is not in the population: it is not in the corpus. Any absence this page reports has to be read against that list, which the report prints.
  • The FOCI/PAM literature. A large part of censorship measurement publishes at FOCI, PAM and the Internet Society workshops, none of which are in this corpus. Nothing here should be read as “the field has not done X” — only as “these seven venues have not”.

Methodology and limitations of these figures

Every corpus figure on this page is produced by scripts/report_blocking_geodifference.mjs; the instrument table's dates and statuses come from scripts/bgd_external_checks.sh and scripts/bgd_render_checks.mjs; the quoted phrases are checked by scripts/bgd_quotecheck.py. All four scripts, their unedited output, the hand-verdict map for all 275 candidates, the probe patterns, the fold residue and the external-source rejections are published on blocking_and_geodifference. Corpus-level caveats — venue scope, the selection funnel, the provisional years, extraction stability — are on Corpus.

Four limitations belong on the page itself:

  1. The population is a judgement, not a query. 63 papers is the union of six hand-assigned families under the inclusion rule above. A reasonable person would draw a different line in at least three places, all named on the provenance page.
  2. Free-text fields are folded, and folds have residue. detection[].phenomenon agrees with itself run-to-run on about a fifth of exact strings, so it is used to find papers and never to count them.
  3. Signal counts are full-text probes, bounded in both directions as the table says.
  4. Prevalence figures are quoted from the papers with the papers' own denominators. Each was checked against the source text; the checks, and any that did not match, are on the provenance page.

What to report

For a blocking claim to be reusable, a methods section needs:

  1. The treatment vantage — country, network, infrastructure type, provider, and how you verified you were where you think you were (Design:Crawling location § Verify the vantage point).
  2. The control vantage, named as such, and what makes it a control.
  3. The definition of “blocked”, as a decision rule over observable signals, including what happens on a 200 with a block page.
  4. The layer at which you detected it: DNS, TCP, TLS, HTTP, application. Blocking migrates between layers, and a single-layer measurement understates it.
  5. Repeats: how many probes, over what interval, and the agreement rule that turns them into one verdict.
  6. Hand validation: the sample size, who labelled it, and the resulting false-positive rate.
  7. Source-address and port policy if you are probing a stateful filter [11Bhaskar, Abhishek; Pearce, Paul (2022): "Many Roads Lead To Rome: How Packet Headers Influence DNS Censorship Measurement", in: Proceedings of the USENIX Security Symposium. (Link)].
  8. The list, and its version — a Citizen Lab country list, global.csv, a top list, or your own — because “blocked domains” is meaningless without it (Website selection, Sampling).
  9. Ethics: whose machine ran inside the condition, what you exposed them to, and what your review said (Ethics).

Papers to read first

  1. [16McDonald, Allison; Bernhard, Matthew; Valenta, Luke; VanderSloot, Benjamin; Scott, Will; Sullivan, Nick; Halderman, J. Alex; Ensafi, Roya (2018): "403 Forbidden: A Global View of CDN Geoblocking", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] — 403 Forbidden: A Global View of CDN Geoblocking (IMC 2018). The geoblocking method, and the paper that publishes its own 27% false-positive rate.
  2. [1Ablove, Anna; Chandrashekaran, Shreyas; Le, Hieu; Raman, Ram Sundara; Ramesh, Reethika; Oppenheimer, Harry; Ensafi, Roya (2024): "Digital Discrimination of Users in Sanctioned States: The Case of the Cuba Embargo", in: Proceedings of the USENIX Security Symposium. (Link)] — Digital Discrimination of Users in Sanctioned States (USENIX Security 2024). Per-layer decomposition, control vantage, and the 200-OK block page.
  3. [14Tsai, Elisa; Kumar, Deepak; Sundara Raman, Ram; Li, Gavin; Eiger, Yael; Ensafi, Roya (2023): "CERTainty: Detecting DNS Manipulation at Scale using TLS Certificates", Proceedings on Privacy Enhancing Technologies 2023(3):122-137. (DOI)] — CERTainty (PoPETs 2023). The measurement that retired consistency-only DNS detection.
  4. [3Niaki, Arian Akhavan; Cho, Shinyoung; Weinberg, Zachary; Hoang, Nguyen Phong; Razaghpanah, Abbas; Christin, Nicolas; Gill, Phillipa (2020): "ICLab: A Global, Longitudinal Internet Censorship Measurement Platform", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] and [4Raman, Ram Sundara; Shenoy, Prerana; Kohls, Katharina; Ensafi, Roya (2020): "Censored Planet: An Internet-wide, Longitudinal Censorship Observatory", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] — the two observatory papers, and the clearest statements of the control design. Read the ICLab paper as history (its public site is certificate-dead, see the table above); Censored Planet is still running, so read that one as documentation too.
  5. [36Kumar, Renuka; Virkud, Apurva; Sundara Raman, Ram; Prakash, Atul; Ensafi, Roya (2022): "A Large-scale Investigation into Geodifferences in Mobile Apps", in: Proceedings of the USENIX Security Symposium. (Link)] — Geodifferences in Mobile Apps (USENIX Security 2022). What “the same service” means across 26 countries.
  6. [15Jones, Ben; Lee, Tzu-Wen; Feamster, Nick; Gill, Phillipa (2014): "Automated Detection and Fingerprinting of Censorship Block Pages", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] — Automated Detection and Fingerprinting of Censorship Block Pages (IMC 2014). Old, still the reference for what a block-page classifier can and cannot do.
  7. [9Hoang, Nguyen Phong; Dalek, Jakub; Crete-Nishihata, Masashi; Christin, Nicolas; Yegneswaran, Vinod; Polychronakis, Michalis; Feamster, Nick (2024): "GFWeb: Measuring the Great Firewall's Web Censorship at Scale", in: Proceedings of the USENIX Security Symposium. (Link)] — GFWeb (USENIX Security 2024). Scale, and the both-sides-of-the-border control.
  8. [40Take, Kejsi; Young, Jordyn; Bhalerao, Rasika; Gallagher, Kevin; Forte, Andrea; McCoy, Damon; Greenstadt, Rachel (2024): "What to Expect When You’re Accessing: An Exploration of User Privacy Rights in People Search Websites", in: Proceedings on Privacy Enhancing Technologies. (DOI)] — the EU-wall datapoint, in a paper about something else.
  9. [8Nourin, Sadia; Rye, Erik C.; Bock, Kevin; Hoang, Nguyen Phong; Levin, Dave (2025): "Is Nobody There? Good! Globally Measuring Connection Tampering Without Responsive Endhosts", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] — Is Nobody There? Good! (IEEE S&P 2025). The current frontier of remote measurement: no responsive endhost required.

Open questions

  • How common is GDPR-driven blocking of EU visitors in 2026? One incidental measurement on 20 sites in seven venues, no prevalence figure, and a straightforward method: a top list, an EEA vantage and a non-EEA control, a definition of “blocked” from this page, and a hand-validated sample.
  • How much of a geoblocking signal measured from the cloud is actually bot management? Telling the two apart per request is routine — re-request from a second vantage, which several of these papers do. What nobody in this population has done is quantify the confusion: what share of “geoblocked” verdicts collected from datacenter address space would disappear from a residential vantage in the same country.
  • What is the current state of GFWatch and GFWeb? Both dashboards render and neither has fresh data (newest dates 2024-08-06 and 2024-11-19). Whether the pipelines stopped or only the dashboards did is not answerable from outside — and ICLab is the cautionary case in the other direction: its website has been certificate-dead for over a year while its client codebase was pushed to in 2026. A dead website is not a dead project, and a live website is not live data. Check the repository and ask the authors before writing either off.
  • Nobody has published an LLM-based block-page classifier in these venues. The task — sparse positives, multilingual bodies, adversarial similarity to real error pages — is a plausible fit and an unoccupied slot; it also needs the false-positive discipline in the table above.
  • Reproducing a geoblocking result is unusually hard, because the phenomenon is jurisdictional and moves. No paper in this 63-paper population re-runs an earlier geoblocking study on the same population — so within these venues there is no evidence either way about how stable the numbers above are. That is a claim about this corpus, not about the field.
  • Crawling location — choosing and verifying the vantage point. This page tells you what to do with a vantage point; that one tells you how to get one.
  • DNS — the resolution layer, DNS manipulation detection, and why packet headers change your answer.
  • Crawler detection — blocked for looking like a crawler, which is a different cause with the same symptom.
  • Existing datasets — reusing an observatory, and inheriting its population.
  • Internet scanning — the instrument most of this population actually uses.
  • Consent — geo-targeted banners, the most common reason a privacy crawl cares about jurisdiction.
  • Website selection / Sampling — probe lists and top lists are not interchangeable.
  • Ethics — running a probe from inside a censoring network, and from someone else's device.
  • Corpus — venue scope, funnel, and the provisional years.

References

[1]
Ablove, Anna; Chandrashekaran, Shreyas; Le, Hieu; Raman, Ram Sundara; Ramesh, Reethika; Oppenheimer, Harry; Ensafi, Roya (2024): "Digital Discrimination of Users in Sanctioned States: The Case of the Cuba Embargo", in: Proceedings of the USENIX Security Symposium. (Link)
[2]
Roth, Sebastian; Calzavara, Stefano; Wilhelm, Moritz; Rabitti, Alvise; Stock, Ben (2022): "The Security Lottery: Measuring Client-Side Web Security Inconsistencies", in: Proceedings of the USENIX Security Symposium, pp. 2047-2064. (Link)
[3]
Niaki, Arian Akhavan; Cho, Shinyoung; Weinberg, Zachary; Hoang, Nguyen Phong; Razaghpanah, Abbas; Christin, Nicolas; Gill, Phillipa (2020): "ICLab: A Global, Longitudinal Internet Censorship Measurement Platform", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[4]
Raman, Ram Sundara; Shenoy, Prerana; Kohls, Katharina; Ensafi, Roya (2020): "Censored Planet: An Internet-wide, Longitudinal Censorship Observatory", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[5]
Pearce, Paul; Ensafi, Roya; Li, Frank; Feamster, Nick; Paxson, Vern (2017): "Augur: Internet-Wide Detection of Connectivity Disruptions", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[6]
VanderSloot, Benjamin; McDonald, Allison; Scott, Will; Halderman, J. Alex; Ensafi, Roya (2018): "Quack: Scalable Remote Measurement of Application-Layer Censorship", in: Proceedings of the USENIX Security Symposium. (Link)
[7]
Raman, Ram Sundara; Stoll, Adrian; Dalek, Jakub; Ramesh, Reethika; Scott, Will; Ensafi, Roya (2020): "Measuring the Deployment of Network Censorship Filters at Global Scale", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[8]
Nourin, Sadia; Rye, Erik C.; Bock, Kevin; Hoang, Nguyen Phong; Levin, Dave (2025): "Is Nobody There? Good! Globally Measuring Connection Tampering Without Responsive Endhosts", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[9]
Hoang, Nguyen Phong; Dalek, Jakub; Crete-Nishihata, Masashi; Christin, Nicolas; Yegneswaran, Vinod; Polychronakis, Michalis; Feamster, Nick (2024): "GFWeb: Measuring the Great Firewall's Web Censorship at Scale", in: Proceedings of the USENIX Security Symposium. (Link)
[10]
Wu, Mingshi; Zohaib, Ali; Durumeric, Zakir; Houmansadr, Amir; Wustrow, Eric (2025): "A Wall Behind A Wall: Emerging Regional Censorship in China", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[11]
Bhaskar, Abhishek; Pearce, Paul (2022): "Many Roads Lead To Rome: How Packet Headers Influence DNS Censorship Measurement", in: Proceedings of the USENIX Security Symposium. (Link)
[12]
Bhaskar, Abhishek; Pearce, Paul (2024): "Understanding Routing-Induced Censorship Changes Globally", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[13]
Pearce, Paul; Jones, Ben; Li, Frank; Ensafi, Roya; Feamster, Nick; Weaver, Nick; Paxson, Vern (2017): "Global Measurement of DNS Manipulation", in: Proceedings of the USENIX Security Symposium. (Link)
[14]
Tsai, Elisa; Kumar, Deepak; Sundara Raman, Ram; Li, Gavin; Eiger, Yael; Ensafi, Roya (2023): "CERTainty: Detecting DNS Manipulation at Scale using TLS Certificates", Proceedings on Privacy Enhancing Technologies 2023(3):122-137. (DOI)
[15]
Jones, Ben; Lee, Tzu-Wen; Feamster, Nick; Gill, Phillipa (2014): "Automated Detection and Fingerprinting of Censorship Block Pages", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[16]
McDonald, Allison; Bernhard, Matthew; Valenta, Luke; VanderSloot, Benjamin; Scott, Will; Sullivan, Nick; Halderman, J. Alex; Ensafi, Roya (2018): "403 Forbidden: A Global View of CDN Geoblocking", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[17]
Fan, Shencha; Sippe, Jackson; San, Sakamoto; Sheffey, Jade; Fifield, David; Houmansadr, Amir; Wedwards, Elson; Wustrow, Eric (2025): "Wallbleed: A Memory Disclosure Vulnerability in the Great Firewall of China", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[18]
Xue, Diwen; Huremagic, Armin; Wang, Wayne; Raman, Ram Sundara; Ensafi, Roya (2025): "Fingerprinting Deep Packet Inspection Devices by their Ambiguities", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[19]
Tang, Jenny; Alvarez, Léo; Brar, Arjun; Hoang, Nguyen Phong; Christin, Nicolas (2024): "Automatic Generation of Web Censorship Probe Lists", Proceedings on Privacy Enhancing Technologies 2024(4):44-60. (DOI)
[20]
Ablove, Anna; Chandrashekaran, Shreyas; Qiang, Xiao; Ensafi, Roya (2026): "Characterizing the Implementation of Censorship Policies in Chinese LLM Services", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[21]
Lipphardt, Friedemann; Ali, Moonis; Banzer, Martin; Feldmann, Anja; Gosain, Devashish (2026): "There is No War in Ba Sing Se: A Global Analysis of Content Moderation in Large Language Models", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[22]
Knockel, Jeffrey; Dałek, Jakub; Aljizawi, Noura; Ahmed, Mohamed; Meletti, Levi; Lau, Justin (2026): "Banned Books: Analysis of Censorship on Amazon.com", Proceedings on Privacy Enhancing Technologies 2026(3):200-214. (DOI)
[23]
Elmenhorst, Kathrin; Schütz, Bertram; Aschenbruck, Nils; Basso, Simone (2021): "Web censorship measurements of HTTP/3 over QUIC", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[24]
Wu, Mingshi; Sippe, Jackson; Sivakumar, Danesh; Burg, Jack; Anderson, Peter; Wang, Xiaokang; Bock, Kevin; Houmansadr, Amir; Levin, Dave; Wustrow, Eric (2023): "How the Great Firewall of China Detects and Blocks Fully Encrypted Traffic", in: Proceedings of the USENIX Security Symposium. (Link)
[25]
Lu, Chaoyi; Liu, Baojun; Li, Zhou; Hao, Shuang; Duan, Hai-Xin; Zhang, Mingming; Leng, Chunying; Liu, Ying; Zhang, Zaifeng; Wu, Jianping (2019): "An End-to-End, Large-Scale Measurement of DNS-over-Encryption: How Far Have We Come?", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[26]
Li, Ruixuan; Liu, Baojun; Lu, Chaoyi; Duan, Haixin; Shao, Jun (2024): "A Worldwide View on the Reachability of Encrypted DNS Services", in: Proceedings of the ACM Web Conference. (DOI)
[27]
Khattak, Sheharbano; Fifield, David; Afroz, Sadia; Javed, Mobin; Sundaresan, Srikanth; McCoy, Damon; Paxson, Vern; Murdoch, Steven J. (2016): "Do You See What I See? Differential Treatment of Anonymous Users", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[28]
Singh, Rachee; Nithyanand, Rishab; Afroz, Sadia; Pearce, Paul; Tschantz, Michael Carl; Gill, Phillipa; Paxson, Vern (2017): "Characterizing the Nature and Dynamics of Tor Exit Blocking", in: 26th USENIX Security Symposium (USENIX Security 17), pp. 325-341. USENIX Association. (Link)
[29]
Hoang, Nguyen Phong; Niaki, Arian Akhavan; Dalek, Jakub; Knockel, Jeffrey; Lin, Pellaeon; Marczak, Bill; Crete-Nishihata, Masashi; Gill, Phillipa; Polychronakis, Michalis (2021): "How Great is the Great Firewall? Measuring China's DNS Censorship", in: Proceedings of the USENIX Security Symposium. (Link)
[30]
Rambert, Raymond; Weinberg, Zachary; Barradas, Diogo; Christin, Nicolas (2021): "Chinese Wall or Swiss Cheese? Keyword filtering in the Great Firewall of China", in: Proceedings of the ACM Web Conference. (DOI)
[31]
Tai, Jonas; Sengottuvelavan, Karthik Nishanth; Whiting, Peter; Hoang, Nguyen Phong (2025): "IRBlock: A Large-Scale Measurement Study of the Great Firewall of Iran", in: Proceedings of the USENIX Security Symposium. (Link)
[32]
Yadav, Tarun Kumar; Sinha, Akshat; Gosain, Devashish; Sharma, Piyush Kumar; Chakravarty, Sambuddho (2018): "Where The Light Gets In: Analyzing Web Censorship Mechanisms in India", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[33]
Chaabane, Abdelberi; Chen, Terence; Cunche, Mathieu; De Cristofaro, Emiliano; Friedman, Arik; Kaafar, Mohamed Ali (2014): "Censorship in the Wild: Analyzing Internet Filtering in Syria", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[34]
Nourin, Sadia; Tran, Van; Jiang, Xi; Bock, Kevin; Feamster, Nick; Hoang, Nguyen Phong; Levin, Dave (2023): "Measuring and Evading Turkmenistan's Internet Censorship: A Case Study in Large-Scale Measurements of a Low-Penetration Country", in: Proceedings of the ACM Web Conference. (DOI)
[35]
Raman, Ram Sundara; Evdokimov, Leonid; Wustrow, Eric; Halderman, J. Alex; Ensafi, Roya (2020): "Investigating Large Scale HTTPS Interception in Kazakhstan", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[36]
Kumar, Renuka; Virkud, Apurva; Sundara Raman, Ram; Prakash, Atul; Ensafi, Roya (2022): "A Large-scale Investigation into Geodifferences in Mobile Apps", in: Proceedings of the USENIX Security Symposium. (Link)
[37]
Guo, Jiawei; Nong, Yu; Lin, Zhiqiang; Cai, Haipeng (2025): "Code Speaks Louder: Exploring Security and Privacy Relevant Regional Variations in Mobile Applications", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[38]
Singh, Sachin Kumar; Ricci, Robert; Gamero-Garrido, Alexander (2025): "Where in the World Are My Trackers? Mapping Web Tracking Flow Across Diverse Geographic Regions", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[39]
Ramesh, Reethika; Raman, Ram Sundara; Virkud, Apurva; Dirksen, Alexandra; Huremagic, Armin; Fifield, David; Rodenburg, Dirk; Hynes, Rod; Madory, Doug; Ensafi, Roya (2023): "Network Responses to Russia's Invasion of Ukraine in 2022: A Cautionary Tale for Internet Freedom", in: Proceedings of the USENIX Security Symposium. (Link)
[40]
Take, Kejsi; Young, Jordyn; Bhalerao, Rasika; Gallagher, Kevin; Forte, Andrea; McCoy, Damon; Greenstadt, Rachel (2024): "What to Expect When You’re Accessing: An Exploration of User Privacy Rights in People Search Websites", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[41]
Van Nortwick, Maggie; Wilson, Christo (2022): "Setting the Bar Low: Are Websites Complying With the Minimum Requirements of the CCPA?", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[42]
Dimova, Yana; Van Goethem, Tom; Joosen, Wouter (2023): "Everybody's Looking for SSOmething: A large-scale evaluation on the privacy of OAuth authentication on the web", Proceedings on Privacy Enhancing Technologies 2023(4). (DOI)
[43]
Shezan, Faysal Hossain; Su, Zihao; Kang, Mingqing; Phair, Nicholas; Thomas, Patrick William; van Dam, Michelangelo; Cao, Yinzhi; Tian, Yuan (2023): "CHKPLUG: Checking GDPR Compliance of WordPress Plugins via Cross-language Code Property Graph", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[44]
Haghighi, Sara; LaChance, Clark; Pourghasemi Fatideh, Ali; Breaux, Travis; Ghanavati, Sepideh (2026): "The Role of Online Forums in Developer Understanding of Privacy Law - A Reddit Case Study", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[45]
Weinberg, Zachary; Sharif, Mahmood; Szurdi, Janos; Christin, Nicolas (2017): "Topics of Controversy: An Empirical Analysis of Web Censorship Lists", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[46]
Liu, Mingxuan; Zhang, Yiming; Li, Xiang; Lu, Chaoyi; Liu, Baojun; Duan, Haixin; Zheng, Xiaofeng (2024): "Understanding the Implementation and Security Implications of Protective DNS Services", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[47]
Sokoto, Saidu; Balduf, Leonhard; Trautwein, Dennis; Wei, Yiluo; Tyson, Gareth; Castro, Ignacio; Ascigil, Onur; Pavlou, George; Korczyński, Maciej; Scheuermann, Björn; Król, Michał (2024): "Guardians of the Galaxy: Content Moderation in the InterPlanetary File System", in: Proceedings of the USENIX Security Symposium. (Link)
[48]
Holzbauer, Florian; Strobl, Sebastian; Ullrich, Johanna (2025): "Tracking Internet Disruptions in Ukraine: Insights from Three Years of Active Full Block Scans", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[49]
Dainotti, Alberto; Squarcella, Claudio; Aben, Emile; Claffy, Kimberly C.; Chiesa, Marco; Russo, Michele; Pescapè, Antonio (2011): "Analysis of country-wide internet outages caused by censorship", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
1)
The 2018 wave was widely reported at the time — the 403 Forbidden paper cites the BBC's “GDPR: US news sites unavailable to EU users under new rules”, May 2018, as its motivating example. That is a news report, not a measurement.
2)
https://www.rfc-editor.org/rfc/rfc7725.json, fetched 2026-09-10: status PROPOSED STANDARD, pub_date February 2016, no obsoleted_by.
3)
CELEX 32018R0302, fetched on 2026-09-10 from the Publications Office CELLAR service (http://publications.europa.eu/resource/celex/32018R0302), whose title reads “on addressing unjustified geo-blocking and other forms of discrimination based on customers' nationality, place of residence or place of establishment within the internal market”. CELLAR was used because this repository's notes record EUR-Lex answering automated clients with an AWS-WAF challenge — which could not be reproduced on the day: a plain curl with no User-Agent got HTTP 200 and the real page, with no challenge markers, so treat that wall as intermittent rather than as a fact. EUR-Lex's own Modified by metadata shows the Regulation has been touched exactly once in substance since 2018 — Article 10(1) was implicitly repealed by Regulation (EU) 2017/2394 with effect from 17/01/2020, an enforcement-cooperation cross-reference, not one of the non-discrimination articles — plus three language corrigenda. Neither the DSA (2022/2065) nor the Data Act appears in that metadata.
You could leave a comment if you were logged in.
design/blocking_and_geodifference.1789084047.txt.gz · Last modified: by karel.kubicek.claude