User Tools

Site Tools


provenance:design:blocking_and_geodifference

This is an old revision of the document!


A PCRE backtrack error occured. Try to increase the pcre.backtrack_limit in php.ini

Provenance: Blocking and Geodifference

Working log for blocking_and_geodifference. Every figure on that page, the query that produced it, the population it is a share of, the folds and their residue, the quotes that were checked against the source PDFs, the external sources that were fetched and the ones that were rejected, and the judgement calls — including the ones a reasonable person would have made differently.

Corpus-level caveats — venue scope, the selection funnel, the provisional 2025–2026 years, extraction stability — are on Corpus and are not restated here.

Run: 2026-09-10. Corpus at the time: data/extract/run1/extractions.jsonl, 5,859 extracted papers, 7 venues (CCS, IMC, NDSS, PoPETs, USENIX Security, TheWebConf, IEEE S&P), 2010–2026. Authored by Claude (Opus 5) with five review passes (see §12). No part of this page was carried over from an earlier run: this page did not exist before, and the queued estimate that scoped it (74 candidate papers) was re-derived rather than reused — see §3.

1. What the page had to establish, and what the schema gives you

The extraction has no field meaning “this paper measures blocking”. It has:

Field What it can do here What it cannot
detection[]phenomenon, technique, metric, prevalence find candidate papers, and supply the measured results with the paper's own denominator be counted: phenomenon is free text and agrees with itself run-to-run on roughly a fifth of exact strings
tools[] and otherToolsMentioned find papers that name an instrument they used or produced see a dataset a paper only read — OONI is in tools[] for 7 papers and in 52 full texts (§7)
vantage[]locations, infrastructure, serviceName count papers with one, two or many vantage points, after folding the free-text locations tell you which vantage point was the control
platforms, studyTypes split the population by web/scan/reuse define the population: no studyType means “measured blocking”
legal[] find the law-driven slice find a GDPR wall — 281 papers name GDPR, 4 of them are in this page's population

So the population is a hand-audited candidate set, and the audit is the artefact. §2 is the rule, §3 the probes, §4 the verdicts.

2. The inclusion rule

Written down before any count was published, and reproduced verbatim on the content page:

A paper is IN the population if it measures whether some client could reach some
content or service AND attributes the failures to a deliberate blocking decision
taken by someone other than the client — a state, an ISP, a resolver operator, a
CDN, a platform, or the content owner — where that decision is keyed on WHO OR
WHERE the client is, on the content's acceptability to an authority or platform,
or on the traffic looking like an attempt to evade such a decision.

Explicitly OUT, each with its own off-topic family:
  * blocking keyed on the content being malicious (malware, phishing, spam);
  * blocking keyed on the client looking automated (programming:crawler_detection);
  * blocking the client chose (its own ad blocker or filter list);
  * a system PROPOSED to evade blocking, and a DETECTOR proposed for finding
    circumvention traffic (both counted as the ADJACENT family "circumvention");
  * the effect of a takedown or of moderating user posts on later behaviour.

The third key was added after a review pass: this literature also measures
blocking keyed on THE TRAFFIC LOOKING LIKE AN ATTEMPT TO EVADE a blocking
decision — Shadowsocks, fully encrypted flows, an SNI. Measuring how a DEPLOYED
censor detects and blocks such traffic is in the population; proposing a
detector for it is not. The two sides are one sentence apart and the first
version of this rule did not separate them, so two papers sat in the population
under a rule that appeared to exclude them.

This is the version the script prints and the content page reproduces, and it is the second version. The first said only “who or where the client is, or the content's acceptability”, which did not cover blocking keyed on what the traffic looks like — and the population contained four such papers all along: the GFW's active probing of hidden circumvention servers (IMC 2015), its blocking of Shadowsocks (IMC 2020), its blocking of fully encrypted traffic [1Wu, Mingshi; Sippe, Jackson; Sivakumar, Danesh; Burg, Jack; Anderson, Peter; Wang, Xiaokang; Bock, Kevin; Houmansadr, Amir; Levin, Dave; Wustrow, Eric (2023): "How the Great Firewall of China Detects and Blocks Fully Encrypted Traffic", in: Proceedings of the USENIX Security Symposium. (Link)], and SNI-based QUIC censorship (USENIX Security 2025). The generic review pass found the mismatch. The choice made here was to amend the rule to the population rather than to cut the papers, because protocol-shaped blocking is a real and growing part of what censors do, and a page that excluded it would be describing a smaller phenomenon than the one it is named after. The alternative — moving the four to circumvention and publishing a population of 59 — is a defensible reading, and the effect of it can be read off §4.

One boundary case is deliberately in: liu2024_implementation, the protective-DNS study, whose motive is security rather than geography, because its claim has this page's shape — control resolver against treatment resolver, and a definition of “blocked” that has to survive an NXDOMAIN and a sinkhole. It is the whole filter family, so the effect of that decision on any figure can be checked by subtracting one.

3. The nine probes, and why there are nine

Each probe is a candidate generator, never a population. Counts are papers.

# Probe Over Papers Added by it
1 title title + summary: censorship, censored, OONI, Censored Planet, ICLab, GFW, Great Firewall, geoblock, blockpage, internet shutdown, geodifferen, geo-restrict, geofenc, network interference, connection tampering, DNS manipulation, throttl 83 the spine
2 detect detection[].phenomenon — censor, geoblock, blockpage, DNS manipulation/censorship/interference, url-filter, keyword filter, over-block, blocked site/domain/content, throttling, sanction 71 papers whose title says nothing
3 instr tools[] + otherToolsMentioned name-matched against OONI, Censored Planet, ICLab, Quack, Hyperquack, Satellite, Augur, GFWatch, Encore, Geneva, Iris, Tor Metrics, Citizen Lab 35 instrument users
4 ftinstr full text names an observatory or the Citizen Lab test lists 80 dataset readers, which tools[] cannot see
5 ftvocab full text: blockpage, geoblock, geodifferen, geo-restrict, HTTP 451, “unavailable for legal reasons” 103 papers that hit blocking in passing
6 wall four tight jurisdiction-wall patterns (§8) 9 7 papers no other probe caught
7 recall title + summary: blocking-resistant, unblock, reachab, inaccessib, “available across”, “not available in”, content moderation, takedown, deplatform, shadow ban, “across N countries”, country-level/regional variation 60 2 papers that belong in the population
8 outage title + summary: outage, shutdown, blackout, disruption, internet resilience, connectivity loss, depeering, route withdrawal 30 the whole outage adjacent family (12), and a corrected claim
9 manual papers added by hand, each with its reason recorded in the script 1 the OpenVPN-fingerprinting paper, which no regex reaches

Union: 275 candidates, 4.7% of the corpus. Probes 6, 7 and 8 exist because probes 1–5 are built around the words censor and geoblock, and each was added after a later check found a paper the candidate set could not see:

  • Probe 6 came from the jurisdiction section: four wall-shaped patterns caught 7 papers no censorship-shaped probe had.
  • Probe 7 came from asking what a paper would call this if it never used the word censorship. Its yield is the honest measure of the hole: 2 in-population papers ([2Lu, Chaoyi; Liu, Baojun; Li, Zhou; Hao, Shuang; Duan, Hai-Xin; Zhang, Mingming; Leng, Chunying; Liu, Ying; Zhang, Zaifeng; Wu, Jianping (2019): "An End-to-End, Large-Scale Measurement of DNS-over-Encryption: How Far Have We Come?", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], the 2019 DNS-over-encryption reachability measurement, and [3Sokoto, Saidu; Balduf, Leonhard; Trautwein, Dennis; Wei, Yiluo; Tyson, Gareth; Castro, Ignacio; Ascigil, Onur; Pavlou, George; Korczyński, Maciej; Scheuermann, Björn; Król, Michał (2024): "Guardians of the Galaxy: Content Moderation in the InterPlanetary File System", in: Proceedings of the USENIX Security Symposium. (Link)], IPFS moderation) against 44 off-topic hits.
  • Probe 9 is not a probe at all. The generic review pass named 2022/USENIX/openvpn-is-open-to-vpn-fingerprinting as having the same shape as two papers in the population. Its title and summary contain no blocking vocabulary, so no regex could ever reach it; the script now has a hand-added candidate list, with the reason recorded per paper, because a candidate set has to be able to admit one rather than pretend the regexes are complete. Its verdict is circumvention (it proposes the fingerprinting method rather than measuring a deployed censor).
  • Probe 8 came from reading the shared bibliography's newest entry. holzbauer2025_tracking, Tracking Internet Disruptions in Ukraine (IMC 2025), was in literature:bibliography already, is squarely about whether clients could reach the Internet, and no probe caught it — while a draft of the page asserted that shutdowns as an event class were absent from the corpus. They are not: there are 12 outage-detection papers, and what they lack is not measurement but attribution. The distinction became an adjacent family (§4) and the page's claim was rewritten around it. This is the finding to remember: a candidate set built from one vocabulary cannot see a literature that uses another, and the counter-example was already on the wiki.

The queued estimate was wrong in both directions and is not comparable. roadmap queued this page on a title-and-summary probe that matched 74 papers, 38 of them on the web platform, 30 in 2020–2023 and 22 in 2024–2026. That probe is not in scripts/gap_probe_roadmap.mjs — it came from the 2026-09-02 brainstorm pass and only its alternation is recorded in prose — so it cannot be re-run as such. Running the recorded alternation (censorship|censored|OONI|Censored Planet|GFW|geoblock|blockpage|internet shutdown) over the current corpus gives 68 / 38 / 27 / 21. The web column matches exactly and the other three do not, so either the corpus moved under it or the recorded alternation is not quite what was run; this run cannot distinguish those, and does not claim to. And 68 is not a floor for this page's population either. Of those 68 hits, 38 are in the population and 30 are not — circumvention systems, homonyms and incidental mentions. In the other direction, 16 of the population's 63 papers are caught by no title probe at all, and probe 1 (a wider title probe than the queued one) excludes 36 of its own 83 hits, 43.4%. A candidate count justifies looking, and nothing more.

4. Verdicts: all 275, and the off-topic families in full

Every candidate carries a hand verdict in the VERDICT map inside scripts/report_blocking_geodifference.mjs (§16). The script throws if a candidate has no verdict, if a verdict names a paper no probe caught, or if a family has no label — both directions, because a one-directional check rots silently the next time a probe moves.

Family Papers In population?
net — network-level interference 45 yes
server — server- or platform-side refusal by region 11 yes
platform — moderation of access inside one service 5 yes
jurisdiction — law-driven blocking 2 yes
probelist — the probe list as the object of study 2 yes
filter — client-side filtering products 1 yes (boundary case, §2)
union 63
circumvention 29 no
outage 12 no — measures unreachability without attributing it to anybody's decision
interview 5 no
wall-encounter 3 no
vantage-instrument 2 no
sok 1 no

Three papers carry two measurement families: [4Ramesh, Reethika; Raman, Ram Sundara; Virkud, Apurva; Dirksen, Alexandra; Huremagic, Armin; Fifield, David; Rodenburg, Dirk; Hynes, Rod; Madory, Doug; Ensafi, Roya (2023): "Network Responses to Russia's Invasion of Ukraine in 2022: A Cautionary Tale for Internet Freedom", in: Proceedings of the USENIX Security Symposium. (Link)] (net+server), [5Knockel, Jeffrey; Dałek, Jakub; Aljizawi, Noura; Ahmed, Mohamed; Meletti, Levi; Lau, Justin (2026): "Banned Books: Analysis of Censorship on Amazon.com", Proceedings on Privacy Enhancing Technologies 2026(3):200-214. (DOI)] (server+platform) and [6Lipphardt, Friedemann; Ali, Moonis; Banzer, Martin; Feldmann, Anja; Gosain, Devashish (2026): "There is No War in Ba Sing Se: A Global Analysis of Content Moderation in Large Language Models", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] (platform+server), which is why the family counts sum to 66 and the union is 63.

161 of 275 candidates (58.5%) are off topic, and every one has a named reason:

Off-topic family Papers Why
off:incidental 115 one incidental mention, or a related-work citation only
off:vpn-proxy-ecosystem 8 the VPN or proxy ecosystem itself → Crawling location
off:homonym-tor-metrics 7 names Tor Metrics, a Tor network dataset, not a blocking observatory
off:takedown-effects 5 what a takedown did to the target, not whether a client could reach content
off:proxy-detection 4 detecting proxy or VPN clients — the server side of vantage choice
off:blockchain-sanctions 3 sanctions and blocklists on a blockchain
off:homonym-statistical-censoring 2 “censored”/“censoring” in the statistics or ML sense (interval-censored Hawkes processes; adversarial censoring of demographic features)
off:adblocking 2 blocking by the user's own ad blocker → Filter lists
off:security-blocking 2 blocking malware or phishing
off:llm-guardrails 2 model refusals of what the user asked
off:moderation-effects 2 the effect of moderating user posts on later behaviour
off:homonym-iris 1 iRiS, an iOS private-API analyser, colliding with Iris the DNS-manipulation platform
off:homonym-augur 1 Augur named in a browser-fingerprinting paper, not Augur the disruption prober
off:homonym-throttling 1 bandwidth throttling as a Tor defence
off:homonym-ml-concept-censorship 1 concept censorship in a diffusion model
off:age-gate 1 age assurance → queued as privacy:age_assurance
off:bot-blocking 1 blocked for looking like a bot → Crawler detection
off:censor-as-attack-surface 1 censorship middleboxes abused as an amplifier
off:moderation-system 1 builds a moderation classifier rather than measuring blocking
off:resilience-modelling 1 models what would happen if countries disconnected, from routing graphs; no client reachability measured

That 58.5% is the page's own headline caution: in this subject area a keyword probe is roughly 40% precise, and the words that do the damage are ordinary English ones.

5. Folds and residue

Vantage locations

vantage[].locations is free text and is folded through scripts/geo.mjs (shared with Crawling location and IP classification, so the pages agree). Unmapped strings are still counted as distinct locations — an unmapped string is a place the paper named, and dropping it would make “two or more vantage points” an artefact of the alias list — and every one is printed:

  4x  "Guangzhou"
  2x  "Dominican Republic"
  2x  "Jamaica"
  2x  "Puerto Rico"
  2x  "Saint Barthelemy"
  2x  "Saint Kitts and Nevis"
  2x  "Guadeloupe"
  2x  "Trinidad and Tobago"
  2x  "Martinique"
  2x  "Grenada"
  2x  "around the world"
  2x  "Taichung"
  2x  "Hangzhou"
  2x  "Longmont?"
  2x  "San Jose"

Twelve of the fifteen are real places geo.mjs has no alias for — the Caribbean cluster comes from the Cuba-embargo study's control set, the Chinese and Taiwanese cities from the GFW work. geo.mjs was not extended: it is shared with two published pages whose figures would move. The consequence is disclosed rather than fixed, and it is small — every paper in the residue already has two or more mapped locations.

Detection phenomena: found with, never counted from

detection[].phenomenon was used only as probe 2. It is unusable as a count: probe 2's own pattern matches 157 distinct phenomenon strings across 71 papers (printed by the report), and a wider exploratory pattern over phenomenon + technique + metric during scoping matched several hundred more, almost all off topic. Reading either set shows CPU throttling, GPS interference, WLAN interference, ad blocking, image filtering, reviewer-assignment manipulation and “clip-on memory manipulator attack” alongside the real ones. The figures on the page that look like phenomenon counts are not: they are hand verdicts (§4) and full-text signal probes (§6).

6. The signal probes, and both of their bounds

The “what blocked is operationalised as” table is eleven full-text regexes evaluated over the 63 papers. Each is:

  • an upper bound on use — a sentence in related work matches; and
  • a lower bound on the idea — a paper can compare against a control and never write the word.

Both bounds are stated on the content page, in the table's own caption. Median signals matched per paper: 4; 45 of 63 match three or more; 6 papers match none and are named in the output (two geodifference studies, a CCPA study, an IPFS attack paper, a connectivity characterisation and a P2P bootstrapping study — all of them papers whose blocking claim is not about a page body).

Two probes were repaired after reading their hits, and both repairs changed a published number:

  • the LLM row matched an unanchored BERT, which matches the surname Deibert — Ron Deibert of the Citizen Lab, who is in the reference list of a large share of this literature. Before anchoring, the row read 3 papers with two of the three being reference-list noise. Every term in both the row and the widened probe is now \b-anchored.
  • the widened probe's first draft included prompt\w*, which matches “censorship prompted”. It came out.

The control-vantage probe is deliberately published at two widths, because the width decides the claim: tight gives 32 of 63 and loose gives 47 of 63. A review pass found that the first version of the loose probe did not contain the tight one — 6 of the 32 tight hits matched no loose pattern, because the tight probe's “uncensored control” and “vantage point outside country” branches contain no literal “control” near a vantage word. So “tight” and “loose” were two different questions whose counts could not be compared, and the guard could not see it: it compared sizes, and a size comparison passes on two disjoint sets. Both loose probes are now built from their tight counterparts by construction ([…TIGHT, …extra]), and both guards assert containment: they throw and name the escaping papers. The published loose figure moved from 41 to 47 as a result.

7. Instruments: the structural undercount, and the homograph warning

The instrument table crosses tools[] (used or produced) against a full-text mention sweep, corpus-wide:

Instrument tools[] full text of which in the 63
OONI 7 52 34
Censored Planet 4 47 28
ICLab 1 42 28
Citizen Lab test list 0 35 25
Quack 5 31 22
Augur 2 43 18
Hyperquack 3 6 5
GFWatch 1 7 7
Iris 3 157 9
Satellite 4 195 22
Geneva 6 191 10
M-Lab 9 44 2
OpenNet Initiative / ONI 0 32 23
Encore 0 30 14

The page uses this table for one claim — that a structured tool query undercounts observatory reuse by about an order of magnitude — and for no usage figure. The bottom half of the table is why: Satellite is satellite Internet, Geneva is a city and a convention, Iris is an iOS analyser and a given name, ONI is a substring, and Encore and Augur are ordinary words. Any per-instrument adoption figure would need the same hand audit as §4.

8. The jurisdiction-wall probe, at two widths

“Block” alone is useless in this corpus: it is overwhelmingly blockchain, blocklist and ad blocking. An early draft of this probe used (block|deny|refus)[^.]{0,80}(EU|European|GDPR) and returned 42 papers, of which most were about Ethereum. The published probe is four tight patterns requiring a wall-shaped construction; the loose variant drops the requirement that the two halves be close together.

Width Papers of 5,859 What they are
tight 9 2 measurements, 3 encounters, 4 incidental (an ad-blocker legal note, a filter-list paper's ISP statistics, an HbbTV denylist, an MV3 study's European vantage point)
loose 24 the 9 above plus 15 more, all incidental — a PRNG paper, Twitter event extraction, aircraft communications, anti-adblocker detection, LTE eavesdropping, an Alexa-skills study, two consent studies, an SMS-phishing study and others

The same containment bug applied here and was found by the same review pass: with the first loose probe, 6 of the 9 tight hits were not loose hits, the union was 24 and the page said “all 27 distinct hits” — 9 + 18, added as if the sets nested. The loose probe now contains the tight one by construction, the guard asserts containment, and the union it prints is 24. All 24 were read, and the disposition of each is in the VERDICT map or printed as “not a candidate” in the output. That is the evidence behind the page's claim that the prevalence of GDPR-driven blocking has not been measured in these seven venues: it is a claim about 5,859 papers under two probe widths and a hand read of every hit, and it is stated on the page with that scope.

9. Quotes and figures checked against the source PDFs

scripts/bgd_quotecheck.py (§16) looks for 94 strings — every phrase the page quotes, and every paper-sourced figure a reader could pull on — in the cited paper's own text, trying paper.cols.txt, paper.norm.txt and a live pypdf extraction of paper.pdf, each verbatim after normalisation and then folded to lowercase alphanumerics. It exits non-zero on any miss. Final run: 94 checked, 0 NOT FOUND.

Three strings failed on the first run, and all three were the same mistake — detection[].prevalence is a model summary, not a quote:

What the page said What the paper says Fix
“over 99% of global clients can normally access large DNS-over-Encryption servers” “Over 99% global users can normally access large DNS-over-Encryption servers” page corrected to the paper's words
“detecting 943K and 55K pay-level domains censored by…” the paper writes “pay-level domains (PLDs) censored by…” check string corrected; the page's own sentence paraphrases and does not quote
“we sent 10M DoEv4 queries … of which 592K (5.92%) queries were blocked” two sentences: “592K (5.92%) DoEv4 queries and 28K (4.91%) DoEv6 queries are blocked” (abstract) and “we perform 10M DoEv4 and 560K DoEv6 queries from 102 countries/regions” (method) both strings checked, and 5.92% of 10M = 592K, so the numerator, denominator and rate are one experiment

Two further quotes were replaced before publication because they were not in the papers at all — they were the extraction's summary field, which is model-written:

  • ICLab was quoted as comparing DNS, TCP, HTTP, TLS, traceroute, and packet observations to controls”. That is the summary. The paper's own words are “compares them with responses to matching DNS queries from our control node”, and its architecture figure labels a “Control vantage”. Both are now on the page and both are in the checker.
  • Censored Planet was described in the summary's terms; the page now quotes “four remote measurement techniques (Augur, Satellite/Iris, Quack, and Hyperquack)” and “synchronized measurements on 6 different Internet protocols (IP, DNS, HTTP, HTTPS, Echo and Discard)”.

One BibTeX record was wrong in the corpus index and was corrected by hand: IMC/2014/censorship-in-the-wild… carried author = {Abdelberi, Chaabane and …}, with given and family name swapped. The paper's own title block reads “Abdelberi Chaabane” and Crossref agrees, so the entry is chaabane2014_censorship with {Chaabane, Abdelberi}. bibgen.mjs derived the citekey from the swapped field and would have published abdelberi2014_censorship.

10. Number guard

node scripts/check_page_numbers.mjs pages/design_blocking_and_geodifference.txt work_bg/all_reports.txt — where all_reports.txt is the concatenation of the four committed outputs — ends “OK — every figure in the page traces”, with no ALLOW entries added for this page.

Two figures had to be made traceable rather than allow-listed, which is the point of the guard:

  • the Censored Planet dashboard prints its counters with dots as thousands separators (116.924.640.072), so a page quoting them with commas could not be matched. bgd_render_checks.mjs now prints both forms.
  • the page's “13 of 62 (21.0%)” single-vantage share was computed on the page and not by the script. The script now prints the percentage.

The guard is a presence check, not a binding check: it proves each number appears in an output, not that it is the right number for its sentence. The prose was therefore re-read by hand after every fold or population change, and by the review passes in §12.

11. External and industry sources

Two scripts, both committed with their unedited output. HTTP status is printed before any byte count or content, because a 302 to a redirect page and an empty 200 are indistinguishable in a size.

  • scripts/bgd_external_checks.shcurl checks on every instrument, dataset, standard and legal instrument the page names.
  • scripts/bgd_render_checks.mjs — Playwright renders for the four sources that are single-page apps, where “HTTP 200” says nothing about the data behind them. /usr/bin/chromium cannot be driven in this container, so Playwright's own build is used via PLAYWRIGHT_BROWSERS_PATH=/workspace/.playwright.

What the checks established, all on 2026-09-10:

Claim on the page How it was verified
OONI is live and current api.ooni.io/api/v1/aggregation returned 143,480 measurements for a one-week Iran web_connectivity query, 42,483 confirmed; ooni-data-eu-fra S3 listing for raw/20260909/ returned all 24 hour prefixes. The bucket's first listing page ends at 2023-07-16, so a paginated listing must not be used to answer “is it current” — ask for the date instead
OONI Probe CLI v3.30.0, 2026-07-27 GitHub releases API, and /releases?per_page=3 printed as well as /releases/latest, because a prerelease can lead the list
Citizen Lab test lists are maintained last commit 2026-09-09; 150 per-country CSVs plus global.csv
Censored Planet's URLs have moved
provenance/design/blocking_and_geodifference.1789084054.txt.gz · Last modified: by karel.kubicek.claude