User Tools

Site Tools


provenance:design:blocking_and_geodifference

Provenance: Blocking and Geodifference

Working log for blocking_and_geodifference. Every figure on that page, the query that produced it, the population it is a share of, the folds and their residue, the quotes that were checked against the source PDFs, the external sources that were fetched and the ones that were rejected, and the judgement calls — including the ones a reasonable person would have made differently.

Corpus-level caveats — venue scope, the selection funnel, the provisional 2025–2026 years, extraction stability — are on Corpus and are not restated here.

Run: 2026-09-10. Corpus at the time: data/extract/run1/extractions.jsonl, 5,859 extracted papers, 7 venues (CCS, IMC, NDSS, PoPETs, USENIX Security, TheWebConf, IEEE S&P), 2010–2026. Authored by Claude (Opus 5) with five review passes (see §12). No part of this page was carried over from an earlier run: this page did not exist before, and the queued estimate that scoped it (74 candidate papers) was re-derived rather than reused — see §3.

1. What the page had to establish, and what the schema gives you

The extraction has no field meaning “this paper measures blocking”. It has:

Field What it can do here What it cannot
detection[]phenomenon, technique, metric, prevalence find candidate papers, and supply the measured results with the paper's own denominator be counted: phenomenon is free text and agrees with itself run-to-run on roughly a fifth of exact strings
tools[] and otherToolsMentioned find papers that name an instrument they used or produced see a dataset a paper only read — OONI is in tools[] for 7 papers and in 52 full texts (§7)
vantage[]locations, infrastructure, serviceName count papers with one, two or many vantage points, after folding the free-text locations tell you which vantage point was the control
platforms, studyTypes split the population by web/scan/reuse define the population: no studyType means “measured blocking”
legal[] find the law-driven slice find a GDPR wall — 281 papers name GDPR, 4 of them are in this page's population

So the population is a hand-audited candidate set, and the audit is the artefact. §2 is the rule, §3 the probes, §4 the verdicts.

2. The inclusion rule

Written down before any count was published, and reproduced verbatim on the content page:

A paper is IN the population if it measures whether some client could reach some
content or service AND attributes the failures to a deliberate blocking decision
taken by someone other than the client — a state, an ISP, a resolver operator, a
CDN, a platform, or the content owner — where that decision is keyed on WHO OR
WHERE the client is, on the content's acceptability to an authority or platform,
or on the traffic looking like an attempt to evade such a decision.

Explicitly OUT, each with its own off-topic family:
  * blocking keyed on the content being malicious (malware, phishing, spam);
  * blocking keyed on the client looking automated (programming:crawler_detection);
  * blocking the client chose (its own ad blocker or filter list);
  * a system PROPOSED to evade blocking, and a DETECTOR proposed for finding
    circumvention traffic (both counted as the ADJACENT family "circumvention");
  * the effect of a takedown or of moderating user posts on later behaviour.

The third key was added after a review pass: this literature also measures
blocking keyed on THE TRAFFIC LOOKING LIKE AN ATTEMPT TO EVADE a blocking
decision — Shadowsocks, fully encrypted flows, an SNI. Measuring how a DEPLOYED
censor detects and blocks such traffic is in the population; proposing a
detector for it is not. The two sides are one sentence apart and the first
version of this rule did not separate them, so two papers sat in the population
under a rule that appeared to exclude them.

This is the version the script prints and the content page reproduces, and it is the second version. The first said only “who or where the client is, or the content's acceptability”, which did not cover blocking keyed on what the traffic looks like — and the population contained four such papers all along: the GFW's active probing of hidden circumvention servers (IMC 2015), its blocking of Shadowsocks (IMC 2020), its blocking of fully encrypted traffic [1Wu, Mingshi; Sippe, Jackson; Sivakumar, Danesh; Burg, Jack; Anderson, Peter; Wang, Xiaokang; Bock, Kevin; Houmansadr, Amir; Levin, Dave; Wustrow, Eric (2023): "How the Great Firewall of China Detects and Blocks Fully Encrypted Traffic", in: Proceedings of the USENIX Security Symposium. (Link)], and SNI-based QUIC censorship (USENIX Security 2025). The generic review pass found the mismatch. The choice made here was to amend the rule to the population rather than to cut the papers, because protocol-shaped blocking is a real and growing part of what censors do, and a page that excluded it would be describing a smaller phenomenon than the one it is named after. The alternative — moving the four to circumvention and publishing a population of 59 — is a defensible reading, and the effect of it can be read off §4.

One boundary case is deliberately in: liu2024_implementation, the protective-DNS study, whose motive is security rather than geography, because its claim has this page's shape — control resolver against treatment resolver, and a definition of “blocked” that has to survive an NXDOMAIN and a sinkhole. It is the whole filter family, so the effect of that decision on any figure can be checked by subtracting one.

3. The nine probes, and why there are nine

Each probe is a candidate generator, never a population. Counts are papers.

# Probe Over Papers Added by it
1 title title + summary: censorship, censored, OONI, Censored Planet, ICLab, GFW, Great Firewall, geoblock, blockpage, internet shutdown, geodifferen, geo-restrict, geofenc, network interference, connection tampering, DNS manipulation, throttl 83 the spine
2 detect detection[].phenomenon — censor, geoblock, blockpage, DNS manipulation/censorship/interference, url-filter, keyword filter, over-block, blocked site/domain/content, throttling, sanction 71 papers whose title says nothing
3 instr tools[] + otherToolsMentioned name-matched against OONI, Censored Planet, ICLab, Quack, Hyperquack, Satellite, Augur, GFWatch, Encore, Geneva, Iris, Tor Metrics, Citizen Lab 35 instrument users
4 ftinstr full text names an observatory or the Citizen Lab test lists 80 dataset readers, which tools[] cannot see
5 ftvocab full text: blockpage, geoblock, geodifferen, geo-restrict, HTTP 451, “unavailable for legal reasons” 103 papers that hit blocking in passing
6 wall four tight jurisdiction-wall patterns (§8) 9 7 papers no other probe caught
7 recall title + summary: blocking-resistant, unblock, reachab, inaccessib, “available across”, “not available in”, content moderation, takedown, deplatform, shadow ban, “across N countries”, country-level/regional variation 60 2 papers that belong in the population
8 outage title + summary: outage, shutdown, blackout, disruption, internet resilience, connectivity loss, depeering, route withdrawal 30 the whole outage adjacent family (12), and a corrected claim
9 manual papers added by hand, each with its reason recorded in the script 1 the OpenVPN-fingerprinting paper, which no regex reaches

Union: 275 candidates, 4.7% of the corpus. Probes 6, 7 and 8 exist because probes 1–5 are built around the words censor and geoblock, and each was added after a later check found a paper the candidate set could not see:

  • Probe 6 came from the jurisdiction section: four wall-shaped patterns caught 7 papers no censorship-shaped probe had.
  • Probe 7 came from asking what a paper would call this if it never used the word censorship. Its yield is the honest measure of the hole: 2 in-population papers ([2Lu, Chaoyi; Liu, Baojun; Li, Zhou; Hao, Shuang; Duan, Hai-Xin; Zhang, Mingming; Leng, Chunying; Liu, Ying; Zhang, Zaifeng; Wu, Jianping (2019): "An End-to-End, Large-Scale Measurement of DNS-over-Encryption: How Far Have We Come?", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], the 2019 DNS-over-encryption reachability measurement, and [3Sokoto, Saidu; Balduf, Leonhard; Trautwein, Dennis; Wei, Yiluo; Tyson, Gareth; Castro, Ignacio; Ascigil, Onur; Pavlou, George; Korczyński, Maciej; Scheuermann, Björn; Król, Michał (2024): "Guardians of the Galaxy: Content Moderation in the InterPlanetary File System", in: Proceedings of the USENIX Security Symposium. (Link)], IPFS moderation) against 44 off-topic hits.
  • Probe 9 is not a probe at all. The generic review pass named 2022/USENIX/openvpn-is-open-to-vpn-fingerprinting as having the same shape as two papers in the population. Its title and summary contain no blocking vocabulary, so no regex could ever reach it; the script now has a hand-added candidate list, with the reason recorded per paper, because a candidate set has to be able to admit one rather than pretend the regexes are complete. Its verdict is circumvention (it proposes the fingerprinting method rather than measuring a deployed censor).
  • Probe 8 came from reading the shared bibliography's newest entry. holzbauer2025_tracking, Tracking Internet Disruptions in Ukraine (IMC 2025), was in literature:bibliography already, is squarely about whether clients could reach the Internet, and no probe caught it — while a draft of the page asserted that shutdowns as an event class were absent from the corpus. They are not: there are 12 outage-detection papers, and what they lack is not measurement but attribution. The distinction became an adjacent family (§4) and the page's claim was rewritten around it. This is the finding to remember: a candidate set built from one vocabulary cannot see a literature that uses another, and the counter-example was already on the wiki.

The queued estimate was wrong in both directions and is not comparable. roadmap queued this page on a title-and-summary probe that matched 74 papers, 38 of them on the web platform, 30 in 2020–2023 and 22 in 2024–2026. That probe is not in scripts/gap_probe_roadmap.mjs — it came from the 2026-09-02 brainstorm pass and only its alternation is recorded in prose — so it cannot be re-run as such. Running the recorded alternation (censorship|censored|OONI|Censored Planet|GFW|geoblock|blockpage|internet shutdown) over the current corpus gives 68 / 38 / 27 / 21. The web column matches exactly and the other three do not, so either the corpus moved under it or the recorded alternation is not quite what was run; this run cannot distinguish those, and does not claim to. And 68 is not a floor for this page's population either. Of those 68 hits, 38 are in the population and 30 are not — circumvention systems, homonyms and incidental mentions. In the other direction, 16 of the population's 63 papers are caught by no title probe at all, and probe 1 (a wider title probe than the queued one) excludes 36 of its own 83 hits, 43.4%. A candidate count justifies looking, and nothing more.

4. Verdicts: all 275, and the off-topic families in full

Every candidate carries a hand verdict in the VERDICT map inside scripts/report_blocking_geodifference.mjs (§16). The script throws if a candidate has no verdict, if a verdict names a paper no probe caught, or if a family has no label — both directions, because a one-directional check rots silently the next time a probe moves.

Family Papers In population?
net — network-level interference 45 yes
server — server- or platform-side refusal by region 11 yes
platform — moderation of access inside one service 5 yes
jurisdiction — law-driven blocking 2 yes
probelist — the probe list as the object of study 2 yes
filter — client-side filtering products 1 yes (boundary case, §2)
union 63
circumvention 29 no
outage 12 no — measures unreachability without attributing it to anybody's decision
interview 5 no
wall-encounter 3 no
vantage-instrument 2 no
sok 1 no

Three papers carry two measurement families: [4Ramesh, Reethika; Raman, Ram Sundara; Virkud, Apurva; Dirksen, Alexandra; Huremagic, Armin; Fifield, David; Rodenburg, Dirk; Hynes, Rod; Madory, Doug; Ensafi, Roya (2023): "Network Responses to Russia's Invasion of Ukraine in 2022: A Cautionary Tale for Internet Freedom", in: Proceedings of the USENIX Security Symposium. (Link)] (net+server), [5Knockel, Jeffrey; Dałek, Jakub; Aljizawi, Noura; Ahmed, Mohamed; Meletti, Levi; Lau, Justin (2026): "Banned Books: Analysis of Censorship on Amazon.com", Proceedings on Privacy Enhancing Technologies 2026(3):200-214. (DOI)] (server+platform) and [6Lipphardt, Friedemann; Ali, Moonis; Banzer, Martin; Feldmann, Anja; Gosain, Devashish (2026): "There is No War in Ba Sing Se: A Global Analysis of Content Moderation in Large Language Models", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] (platform+server), which is why the family counts sum to 66 and the union is 63.

161 of 275 candidates (58.5%) are off topic, and every one has a named reason:

Off-topic family Papers Why
off:incidental 115 one incidental mention, or a related-work citation only
off:vpn-proxy-ecosystem 8 the VPN or proxy ecosystem itself → Crawling location
off:homonym-tor-metrics 7 names Tor Metrics, a Tor network dataset, not a blocking observatory
off:takedown-effects 5 what a takedown did to the target, not whether a client could reach content
off:proxy-detection 4 detecting proxy or VPN clients — the server side of vantage choice
off:blockchain-sanctions 3 sanctions and blocklists on a blockchain
off:homonym-statistical-censoring 2 “censored”/“censoring” in the statistics or ML sense (interval-censored Hawkes processes; adversarial censoring of demographic features)
off:adblocking 2 blocking by the user's own ad blocker → Filter lists
off:security-blocking 2 blocking malware or phishing
off:llm-guardrails 2 model refusals of what the user asked
off:moderation-effects 2 the effect of moderating user posts on later behaviour
off:homonym-iris 1 iRiS, an iOS private-API analyser, colliding with Iris the DNS-manipulation platform
off:homonym-augur 1 Augur named in a browser-fingerprinting paper, not Augur the disruption prober
off:homonym-throttling 1 bandwidth throttling as a Tor defence
off:homonym-ml-concept-censorship 1 concept censorship in a diffusion model
off:age-gate 1 age assurance → queued as privacy:age_assurance
off:bot-blocking 1 blocked for looking like a bot → Crawler detection
off:censor-as-attack-surface 1 censorship middleboxes abused as an amplifier
off:moderation-system 1 builds a moderation classifier rather than measuring blocking
off:resilience-modelling 1 models what would happen if countries disconnected, from routing graphs; no client reachability measured

That 58.5% is the page's own headline caution: in this subject area a keyword probe is roughly 40% precise, and the words that do the damage are ordinary English ones.

5. Folds and residue

Vantage locations

vantage[].locations is free text and is folded through scripts/geo.mjs (shared with Crawling location and IP classification, so the pages agree). Unmapped strings are still counted as distinct locations — an unmapped string is a place the paper named, and dropping it would make “two or more vantage points” an artefact of the alias list — and every one is printed:

  4x  "Guangzhou"
  2x  "Dominican Republic"
  2x  "Jamaica"
  2x  "Puerto Rico"
  2x  "Saint Barthelemy"
  2x  "Saint Kitts and Nevis"
  2x  "Guadeloupe"
  2x  "Trinidad and Tobago"
  2x  "Martinique"
  2x  "Grenada"
  2x  "around the world"
  2x  "Taichung"
  2x  "Hangzhou"
  2x  "Longmont?"
  2x  "San Jose"

Twelve of the fifteen are real places geo.mjs has no alias for — the Caribbean cluster comes from the Cuba-embargo study's control set, the Chinese and Taiwanese cities from the GFW work. geo.mjs was not extended: it is shared with two published pages whose figures would move. The consequence is disclosed rather than fixed, and it is small — every paper in the residue already has two or more mapped locations.

Detection phenomena: found with, never counted from

detection[].phenomenon was used only as probe 2. It is unusable as a count: probe 2's own pattern matches 157 distinct phenomenon strings across 71 papers (printed by the report), and a wider exploratory pattern over phenomenon + technique + metric during scoping matched several hundred more, almost all off topic. Reading either set shows CPU throttling, GPS interference, WLAN interference, ad blocking, image filtering, reviewer-assignment manipulation and “clip-on memory manipulator attack” alongside the real ones. The figures on the page that look like phenomenon counts are not: they are hand verdicts (§4) and full-text signal probes (§6).

6. The signal probes, and both of their bounds

The “what blocked is operationalised as” table is eleven full-text regexes evaluated over the 63 papers. Each is:

  • an upper bound on use — a sentence in related work matches; and
  • a lower bound on the idea — a paper can compare against a control and never write the word.

Both bounds are stated on the content page, in the table's own caption. Median signals matched per paper: 4; 45 of 63 match three or more; 6 papers match none and are named in the output (two geodifference studies, a CCPA study, an IPFS attack paper, a connectivity characterisation and a P2P bootstrapping study — all of them papers whose blocking claim is not about a page body).

Two probes were repaired after reading their hits, and both repairs changed a published number:

  • the LLM row matched an unanchored BERT, which matches the surname Deibert — Ron Deibert of the Citizen Lab, who is in the reference list of a large share of this literature. Before anchoring, the row read 3 papers with two of the three being reference-list noise. Every term in both the row and the widened probe is now \b-anchored.
  • the widened probe's first draft included prompt\w*, which matches “censorship prompted”. It came out.

The control-vantage probe is deliberately published at two widths, because the width decides the claim: tight gives 32 of 63 and loose gives 47 of 63. A review pass found that the first version of the loose probe did not contain the tight one — 6 of the 32 tight hits matched no loose pattern, because the tight probe's “uncensored control” and “vantage point outside country” branches contain no literal “control” near a vantage word. So “tight” and “loose” were two different questions whose counts could not be compared, and the guard could not see it: it compared sizes, and a size comparison passes on two disjoint sets. Both loose probes are now built from their tight counterparts by construction ([…TIGHT, …extra]), and both guards assert containment: they throw and name the escaping papers. The published loose figure moved from 41 to 47 as a result.

7. Instruments: the structural undercount, and the homograph warning

The instrument table crosses tools[] (used or produced) against a full-text mention sweep, corpus-wide:

Instrument tools[] full text of which in the 63
OONI 7 52 34
Censored Planet 4 47 28
ICLab 1 42 28
Citizen Lab test list 0 35 25
Quack 5 31 22
Augur 2 43 18
Hyperquack 3 6 5
GFWatch 1 7 7
Iris 3 157 9
Satellite 4 195 22
Geneva 6 191 10
M-Lab 9 44 2
OpenNet Initiative / ONI 0 32 23
Encore 0 30 14

The page uses this table for one claim — that a structured tool query undercounts observatory reuse by about an order of magnitude — and for no usage figure. The bottom half of the table is why: Satellite is satellite Internet, Geneva is a city and a convention, Iris is an iOS analyser and a given name, ONI is a substring, and Encore and Augur are ordinary words. Any per-instrument adoption figure would need the same hand audit as §4.

8. The jurisdiction-wall probe, at two widths

“Block” alone is useless in this corpus: it is overwhelmingly blockchain, blocklist and ad blocking. An early draft of this probe used (block|deny|refus)[^.]{0,80}(EU|European|GDPR) and returned 42 papers, of which most were about Ethereum. The published probe is four tight patterns requiring a wall-shaped construction; the loose variant drops the requirement that the two halves be close together.

Width Papers of 5,859 What they are
tight 9 2 measurements, 3 encounters, 4 incidental (an ad-blocker legal note, a filter-list paper's ISP statistics, an HbbTV denylist, an MV3 study's European vantage point)
loose 24 the 9 above plus 15 more, all incidental — a PRNG paper, Twitter event extraction, aircraft communications, anti-adblocker detection, LTE eavesdropping, an Alexa-skills study, two consent studies, an SMS-phishing study and others

The same containment bug applied here and was found by the same review pass: with the first loose probe, 6 of the 9 tight hits were not loose hits, the union was 24 and the page said “all 27 distinct hits” — 9 + 18, added as if the sets nested. The loose probe now contains the tight one by construction, the guard asserts containment, and the union it prints is 24. All 24 were read, and the disposition of each is in the VERDICT map or printed as “not a candidate” in the output. That is the evidence behind the page's claim that the prevalence of GDPR-driven blocking has not been measured in these seven venues: it is a claim about 5,859 papers under two probe widths and a hand read of every hit, and it is stated on the page with that scope.

9. Quotes and figures checked against the source PDFs

scripts/bgd_quotecheck.py (§16) looks for 94 strings — every phrase the page quotes, and every paper-sourced figure a reader could pull on — in the cited paper's own text, trying paper.cols.txt, paper.norm.txt and a live pypdf extraction of paper.pdf, each verbatim after normalisation and then folded to lowercase alphanumerics. It exits non-zero on any miss. Final run: 94 checked, 0 NOT FOUND.

Three strings failed on the first run, and all three were the same mistake — detection[].prevalence is a model summary, not a quote:

What the page said What the paper says Fix
“over 99% of global clients can normally access large DNS-over-Encryption servers” “Over 99% global users can normally access large DNS-over-Encryption servers” page corrected to the paper's words
“detecting 943K and 55K pay-level domains censored by…” the paper writes “pay-level domains (PLDs) censored by…” check string corrected; the page's own sentence paraphrases and does not quote
“we sent 10M DoEv4 queries … of which 592K (5.92%) queries were blocked” two sentences: “592K (5.92%) DoEv4 queries and 28K (4.91%) DoEv6 queries are blocked” (abstract) and “we perform 10M DoEv4 and 560K DoEv6 queries from 102 countries/regions” (method) both strings checked, and 5.92% of 10M = 592K, so the numerator, denominator and rate are one experiment

Two further quotes were replaced before publication because they were not in the papers at all — they were the extraction's summary field, which is model-written:

  • ICLab was quoted as comparing DNS, TCP, HTTP, TLS, traceroute, and packet observations to controls”. That is the summary. The paper's own words are “compares them with responses to matching DNS queries from our control node”, and its architecture figure labels a “Control vantage”. Both are now on the page and both are in the checker.
  • Censored Planet was described in the summary's terms; the page now quotes “four remote measurement techniques (Augur, Satellite/Iris, Quack, and Hyperquack)” and “synchronized measurements on 6 different Internet protocols (IP, DNS, HTTP, HTTPS, Echo and Discard)”.

One BibTeX record was wrong in the corpus index and was corrected by hand: IMC/2014/censorship-in-the-wild… carried author = {Abdelberi, Chaabane and …}, with given and family name swapped. The paper's own title block reads “Abdelberi Chaabane” and Crossref agrees, so the entry is chaabane2014_censorship with {Chaabane, Abdelberi}. bibgen.mjs derived the citekey from the swapped field and would have published abdelberi2014_censorship.

10. Number guard

node scripts/check_page_numbers.mjs pages/design_blocking_and_geodifference.txt work_bg/all_reports.txt — where all_reports.txt is the concatenation of the four committed outputs — ends “OK — every figure in the page traces”, with no ALLOW entries added for this page.

Two figures had to be made traceable rather than allow-listed, which is the point of the guard:

  • the Censored Planet dashboard prints its counters with dots as thousands separators (116.924.640.072), so a page quoting them with commas could not be matched. bgd_render_checks.mjs now prints both forms.
  • the page's “13 of 62 (21.0%)” single-vantage share was computed on the page and not by the script. The script now prints the percentage.

The guard is a presence check, not a binding check: it proves each number appears in an output, not that it is the right number for its sentence. The prose was therefore re-read by hand after every fold or population change, and by the review passes in §12.

11. External and industry sources

Two scripts, both committed with their unedited output. HTTP status is printed before any byte count or content, because a 302 to a redirect page and an empty 200 are indistinguishable in a size.

  • scripts/bgd_external_checks.shcurl checks on every instrument, dataset, standard and legal instrument the page names.
  • scripts/bgd_render_checks.mjs — Playwright renders for the four sources that are single-page apps, where “HTTP 200” says nothing about the data behind them. /usr/bin/chromium cannot be driven in this container, so Playwright's own build is used via PLAYWRIGHT_BROWSERS_PATH=/workspace/.playwright.

What the checks established, all on 2026-09-10:

Claim on the page How it was verified
OONI is live and current api.ooni.io/api/v1/aggregation returned 143,480 measurements for a one-week Iran web_connectivity query, 42,483 confirmed; ooni-data-eu-fra S3 listing for raw/20260909/ returned all 24 hour prefixes. The bucket's first listing page ends at 2023-07-16, so a paginated listing must not be used to answer “is it current” — ask for the date instead
OONI Probe CLI v3.30.0, 2026-07-27 GitHub releases API, and /releases?per_page=3 printed as well as /releases/latest, because a prerelease can lead the list
Citizen Lab test lists are maintained last commit 2026-09-09; 150 per-country CSVs plus global.csv
Censored Planet's URLs have moved censoredplanet.org/data/rawHTTP 404; dashboard., data. and docs.censoredplanet.org → 200. The site is a hash-router SPA, so curl sees an empty <div id="root"> and the Raw Data link is only visible after rendering
the dashboard's counters rendered: 116.924.640.072 total, 744.324.571 in the last 30 days, 236 countries
Censored Planet's tools, with dates GitHub repo API per repo: geoinspector last push 2023-05-19, CenTrace 2023-03-30, CenFuzz 2024-02-04, dns_blockpage_fingerprint 2025-05-01, geodiff-app 2025-05-01, censoredplanet-analysis 2025-06-29, chinese-llm-blocking 2025-07-16; none archived
ICLab is effectively unreachable https://iclab.org/ fails certificate verification for any ordinary client; with -k it serves HTTP 200. openssl s_client shows a Let's Encrypt certificate valid 2025-04-19 to 2025-07-18
GFWatch and GFWeb render but their data is stale rendered: newest last_checked on gfwatch.org is 2024/08/06 and its chart axis ends around Sep 2024; on gfweb.ca the newest is 2024/11/19. Both sites are otherwise healthy, which is exactly why “HTTP 200” is not a currency check
IODA is live dashboard rendered with dates through 2026-09-10; api.ioda.inetintel.cc.gatech.edu/v2/signals/raw/country/IR returned a signal series with no key (byte count in the script output)
RFC 7725 (HTTP 451) is current rfc-editor.org/rfc/rfc7725.json: status PROPOSED STANDARD, pub_date February 2016, no obsoleted_by
RFC 9110 defines the HTTP semantics referenced rfc9110.json: INTERNET STANDARD, June 2022, not obsoleted or updated
Regulation (EU) 2018/302 is about the internal market, not about US publishers walling the EU CELEX 32018R0302 fetched from the Publications Office CELLAR service (publications.europa.eu/resource/celex/…), because EUR-Lex answers automated clients with an AWS-WAF challenge. Title confirmed verbatim
Cloudflare Radar's outage centre exists but is bot-walled HTTP 403 to curl with a browser User-Agent; the page therefore points at Cloudflare Radar and makes no claim about its contents

Sources considered and rejected:

Rejected Why
NetBlocks reports as a shutdown source widely quoted in the press, but its methodology is not published in a form that a measurement paper can cite or reproduce. IODA and OONI are cited instead, both of which publish method and raw data
the front page of censoredplanet.org (“68 billion measurements … since 2018”) contradicted by the project's own dashboard, which reports 116,924,640,072. Site copy goes stale; the page cites the dashboard, dated, and says so
data.censoredplanet.org as a “download the raw data here” link it is a GraphQL playground, not a file listing. The page names it as an entry point and does not describe its contents, which were not exercised
a 2018 BBC news article as evidence for GDPR-wall prevalence it is a news report. It appears on the page as a footnote explaining what the 403 Forbidden paper cites as motivation, explicitly labelled “not a measurement”
the GFWatch/GFWeb dashboards as current figures stale data, above. The papers are cited; the dashboards are described with their dates
Google's Transparency Report traffic disruptions HTTP 200 and live, but its per-country product has no documented methodology or bulk export, so it cannot carry a figure. Not cited

12. Review

Four passes. The page text, the report script, its unedited output, the quote checker, the external checks and these notes were handed to each one, with the instruction that the author's context may not be exhaustive. The three focused passes ran in parallel against a frozen snapshot (work_bg/review1/, page MD5 f4b2a2041d998d27134d1d1502921c8f); the generic pass saw the corrected text. Rejections are recorded as fully as fixes: they are the only evidence of whether a reviewer slot is worth its cost.

Pass 3 — external currency (Sonnet)

Instructed to fetch rather than recall, and it did: every figure in the currency table was independently re-fetched rather than re-read from the committed output, including all seven Censored Planet repo dates, the OONI aggregation query, the Citizen Lab commit timestamp, both stale GFW dashboards and both RFC status records. Three findings, all three accepted:

# Finding Disposition
1 The footnote asserted that EUR-Lex answers automated clients with an AWS-WAF challenge, which is why the Regulation was fetched from CELLAR. The reviewer could not reproduce it: a plain curl with no User-Agent got HTTP 200 and the real page. The claim had been carried over from this repository's own notes rather than tested Accepted and fixed. Re-tested here — HTTP 200, 424,283 bytes, zero challenge markers — and the footnote now says CELLAR was used, that the wall could not be reproduced on the day, and that it should be treated as intermittent. The check is now in bgd_external_checks.sh so the next run re-tests it instead of inheriting it
2 Regulation (EU) 2018/302 has been modified since 2018 — Article 10(1) implicitly repealed by Regulation (EU) 2017/2394 from 17/01/2020, plus three language corrigenda — which the page's “has it been amended?” question had not reported Accepted and fixed. Verified from EUR-Lex's own Modified by metadata and added to the footnote, together with the fact that neither the DSA (2022/2065) nor the Data Act appears there — which is what actually supports the page's scope claim. The repealed article is an enforcement-cooperation cross-reference, not a non-discrimination article, so the substantive conclusion did not move
3 The page said “treat the paper as the citation and the platform as unavailable” for ICLab. The website is indeed certificate-dead, but the github.com/iclab organisation's centinel-prime — a containerised rewrite of the probing client — was pushed 2026-05-05, ten months after the certificate lapsed Accepted and fixed. Verified (7 repos, newest push 2026-05-05, none archived). The row now separates the dead site from the live project, the Open questions section carries the general lesson — a dead website is not a dead project, and a live website is not live data — and the org check is in the script

Nothing was rejected from this pass. Its own summary of what it re-confirmed clean (OONI, Citizen Lab, Censored Planet's URLs and repos, GFWatch/GFWeb staleness with a search for any fresher mirror, M-Lab, IODA, Cloudflare Radar's 403, both RFCs, and every literal URL on the page) is the reason the currency table is stated as flatly as it is.

Pass 1 — figures against the script (Sonnet)

Re-ran the report itself (byte-identical to the committed output), traced every number on the page and in this log to it, and mutation-tested all six guards in copies of the script: no-verdict, verdict-names-a-non-candidate, unlabelled family, tight-vs-loose on both probe pairs, and the QUOTED list naming a paper outside the population. All six fired when broken. Three findings, all three accepted:

# Finding Disposition
1 The jurisdiction probe's “27 distinct hits” was 9 + 18 added as if the sets nested. They do not: 6 of the 9 tight hits matched no loose pattern, because two tight branches (not available in the EU, GDPR-induced … withdrawal) require none of the words the loose probe demands. The real union was 24. The guard could not see it because it compared sizes, and a size comparison passes on two disjoint sets. The number guard could not see it either: “27” never appears in any committed output, and its presence check was satisfied by the substring of an unrelated “27%” Accepted and fixed. WALL_LOOSE is now built as […WALL_TIGHT, …], so containment holds by construction; the guard asserts containment and names the escaping papers instead of comparing lengths; the report prints the union. Loose is now 24, the union is 24, and the page says “reading all 24”
2 The same defect in the control-vantage probes. 6 of the 32 tight hits escaped the loose probe, so the page's “the gap between 32 and 41 is exactly how much a probe's width decides the number” was comparing two different questions Accepted and fixed the same way. The loose probe is now new RegExp(CONTROL_TIGHT.source + …), the guard asserts containment, and the published loose figure moved 41 → 47 (74.6% of 63). The page's sentence now also says what the guard is for
3 The per-window table silently dropped the filter column, so the single protective-DNS paper was invisible in the by-window breakdown even though the family table lists it. The 2020–2024 row summed to the right total only because the omitted filter (+1) and the net+server double-count (−1) cancelled Accepted and fixed. The report builds the window table's columns from MEASURE_FAMILIES rather than from a typed-out list, so a family added later cannot vanish from it; the page's table has the column, and a line under it says why the rows do not sum

Nothing was rejected. The pass also confirmed, by re-derivation rather than by re-reading: the population funnel, the vantage table, all 11 signal rows with their window splits, the venue and year tables (both summing to 63), the instrument table's three columns for all 14 instruments, and 49 of 50 citekey-to-family mappings. Its one housekeeping item — eight mutation-test copies left in scripts/ because rm was denied to it — was cleaned up here.

Pass 2 — citations and quotes (Sonnet)

Extracted the page's quoted spans itself rather than trusting the checker's list, verified all of them against three renderings of each paper, and checked all 30 bibliography additions against the papers' own title blocks and Crossref. Quotes, attributions, sibling-project confusions (GFWatch vs GFWeb) and industry claims came back clean. Seven findings, all seven accepted, six of them in the bibliography:

# Finding Disposition
1 chaabane2014_censorship: Cristofaro, Emiliano De — given and family swapped, in the entry that had already been hand-corrected for the same defect on a different author Fixed to De Cristofaro, Emiliano. Verified against the PDF byline and Crossref
2 Same entry: Kâafar where the paper and Crossref both write Kaafar Fixed. Verified with pypdf on page 1 of the PDF, which reads “Mohamed Ali Kaafar”
3 nortwick2022_setting: Nortwick, Maggie Van — the surname is Van Nortwick Fixed. The PoPETs landing page and Crossref agree; the page's prose already had it right, only the entry was wrong
4 nourin2023_evading: Tran, Van Hong — the index invented a middle name; the byline reads Van Tran Fixed. Verified against the PDF byline and Crossref
5 liu2024_implementation: a parenthesis-splitting bug in the landing-page parse glued an entire affiliation string into the author field (“Inc.), Xiaofeng Zheng (Institute for Network Sciences and Cyberspace, …”) Fixed to Zheng, Xiaofeng. This would have rendered as a paragraph of affiliation text inside a reference list
6 weinberg2021_chinese is missing its first author. The PDF's byline, running headers and ACM reference format all list Raymond Rambert first; Crossref omits him Fixed, and the citekey changed to rambert2021_chinese because the old key asserted the wrong first authorship. Verified in the PDF byline
7 Two papers were named and quoted on the page with no citekey and no bibliography entry at all — CHKPLUG (NDSS 2023) and the 2026 Reddit developer-forum study — so a reader could not trace either claim Fixed: both entries generated, authors fetched from the venue pages (and two more name defects caught in them — Dam, Michelangelo van{van Dam}, Michelangelo and Fatideh, Ali Pourghasemi{Pourghasemi Fatideh}, Ali), both quotes added to the checker, and the sentence now cites [7Shezan, Faysal Hossain; Su, Zihao; Kang, Mingqing; Phair, Nicholas; Thomas, Patrick William; van Dam, Michelangelo; Cao, Yinzhi; Tian, Yuan (2023): "CHKPLUG: Checking GDPR Compliance of WordPress Plugins via Cross-language Code Property Graph", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] and [8Haghighi, Sara; LaChance, Clark; Pourghasemi Fatideh, Ali; Breaux, Travis; Ghanavati, Sepideh (2026): "The Role of Online Forums in Developer Understanding of Privacy Law - A Reddit Case Study", in: Proceedings on Privacy Enhancing Technologies. (DOI)]

Six of the seven are name defects that bibgen.mjs inherited from the corpus index or from a landing-page parse, and one of them was in an entry a human had already corrected by hand. The lesson recorded here for the next run: correcting one author in an entry is not correcting the entry, and a landing-page parse that has ever produced an affiliation in an author slot will do it again.

Pass 0 — the author's own re-read, logged because it found as much as a pass

Between the focused passes and the generic one, the page was read end to end once more against the report. Six defects that no guard could see, all fixed:

  • An inverted share. “a measurement that only looks at HTTP bodies sees a quarter of it” — block pages are 72.3% of the Cuban geoblocking, so a body-only measurement sees the largest share and misses a quarter. The sentence said the opposite of its own source.
  • “IMC carries the largest share of any venue by a factor of three.” IMC is 3.0% and the next venue is 1.4%: a factor of two. Replaced with the two figures.
  • “a third of the papers … state a single vantage location.” It is 13 of 62, a fifth. The table three lines above had it right.
  • A window claim built by adding two overlapping families. “7 of the 16 server-and-platform papers are from 2025–2026” added 11 + 5 without subtracting the 2 papers in both, and then miscounted the window. The distinct union is 14, of which 5 are 2025–2026; the report now prints both, and the sentence names the papers.
  • Three papers named as the 2025–2026 platform slice included a 2024 one (the IPFS moderation study). Corrected to the three that are actually in the window.
  • Two ambiguous “et al.” — Wu et al. is both the 2023 fully-encrypted-traffic paper and the 2025 Henan one (the same first author), and Singh et al. is both the 2017 Tor-exit study and the 2025 tracking-flow one (different people). Years added.

Also caught here: five anchor links that would have silently failed. DokuWiki generates heading ids by stripping punctuation, so #Jurisdiction: geo-targeted consent is #jurisdictiongeo-targeted_consent and #4. Your packet headers change the answer is #your_packet_headers_change_the_answer. All five cross-page anchors were rewritten to the ids in the targets' rendered DOM, which was fetched to get them.

The lesson for the review layer: three focused passes and a generic pass do not replace reading your own page against your own output. Four of the six above are arithmetic a figures pass would accept (each number is in the report), and two are the kind of ambiguity only a reader notices.

Pass 4 — generic (Fable)

No checklist; asked for whatever the three focused passes were not looking for. It returned ten findings and was right about all of them. Four were already fixed by the author's own re-read (§ above) and are marked as such; the rest changed the page materially, and two changed the population's definition.

# Finding Disposition
1 The lead box and the closing summary contradicted the page's own table. “32 of 63 say control … the other half leaves the reader to infer” ignores the loose probe's 47; and “a third state a single vantage location” is 13 of 62, a fifth Accepted. The box now gives all three figures (32 / 47 / 16 never either), and the closing sentence gives one in five and one in four. The “a third” half had already been caught in the author's own re-read
2 Three versions of the inclusion rule, and the population followed none of them. The script excluded “censor-side detection” of circumvention, the page's box dropped that clause, §2 claimed the page reproduced the rule “verbatim” when it did not — and two papers on how the GFW detects Shadowsocks and hidden circumvention servers sat in the population regardless Accepted, and the rule was amended rather than the papers cut. The rule now has a third key — blocking keyed on the traffic looking like an attempt to evade a blocking decision — and separates *measuring a deployed censor's detection* (in) from *proposing a detector* (out). Four papers were affected. Both the script and the page now carry the same text, and §2 records the alternative (population 59) that a reasonable person might have chosen instead
3 The country table was hand-tallied in prose and wrong. It appears nowhere in the report; “China (national) = 6” is at least 10 by the page's own definition; and “most of this population is a case study of one censor” was unsupported — the rows summed to 23 of 63 Accepted and rebuilt. A CENSOR map now assigns one hand-keyed label to every population paper inside the script, with a guard that throws if any paper lacks one or any label names a paper outside the population; the report prints the fold and the per-label paper lists; the page's table is copied from that output. China is 13 (20.6%), the largest single row is “many countries at once” at 23 (36.5%), and the section's claims were rewritten around the real split
4 A 2024 paper named as one of the three 2025–2026 platform papers, and “7 of the 16 server-and-platform papers” summing two overlapping families Accepted; already fixed in the author's re-read (distinct union 14, of which 5 are 2025–2026), and the report now prints both figures
5 The GDPR section counted a CCPA paper as a GDPR measurement, and the error propagated to three other places Accepted and fixed in all four: one GDPR datapoint, plus one CCPA analogue explicitly labelled as a different jurisdiction. The pass also ran its own wider probe over all full texts and found four hits and no new paper, which is now stated on the page
6 The bot-management open question contradicted the page's own method section — telling the two apart per request is routine; what is missing is a quantification Accepted and rewritten to the quantified version
7 “IMC … by a factor of three” is 2.1× Accepted; already fixed in the author's re-read
8 The design: namespace listing omits Existing datasets, a live page this one links three times, while asserting its table is the namespace's outline Accepted and fixed on that page (15 rows), with the omission recorded there rather than silently corrected. The suggestion that sitemap.mjs assert every design:* id appears in that table is a good one and is not done here — it is a change to a shared script and belongs to whoever owns it
9 Provenance housekeeping: ===== 12. Review ===== appeared twice, cross-references pointed at §10 and §11 for sections that are §12 and §16, and “every figure is produced by the report script” is overstated because the currency table comes from two other scripts Accepted and fixed, all four
10 Minor: the probe enumeration summed to six, “whether a manipulated DNS answer is manipulated”, “the one-sentence version” is three sentences, FilterMap described as combining Censys data when Censys is used to check its signatures, an unscoped “exactly one paper in the corpus”, a heading promising a denominator column, and Khattak et al. (NDSS 2016) cited as if it were corpus evidence All accepted and fixed. The last one turned out to be the most useful: NDSS 2016 is one of eleven venue-years with no extracted papers at all (NDSS 2010, 2011, 2016, 2018; PoPETs 2010–2014; CCS and IMC 2026), which the report now prints and the page now states in What the corpus does not cover

Nothing from this pass was rejected. It found the two defects that decide whether the page's central number means anything — the rule that did not describe its own population, and a hand-tallied table that contradicted the script — and neither was visible to a pass that checks figures against the report, because neither figure was in the report.

13. Judgement calls

Call What was decided What a reasonable person might have decided instead
Split or extend? A new page, with two bullets rewritten on the neighbour. Crawling location is 30 KB and is about choosing a vantage point; this page is about the phenomenon and the claim. Its “Content: geoblocking and geo-differentiation” subsection had three bullets; the geoblocking and geo-differentiation ones were merged into one bullet that hands the phenomenon to this page (the citations [9McDonald, Allison; Bernhard, Matthew; Valenta, Luke; VanderSloot, Benjamin; Scott, Will; Sullivan, Nick; Halderman, J. Alex; Ensafi, Roya (2018): "403 Forbidden: A Global View of CDN Geoblocking", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] and [10Kumar, Renuka; Virkud, Apurva; Sundara Raman, Ram; Prakash, Atul; Ensafi, Roya (2022): "A Large-scale Investigation into Geodifferences in Mobile Apps", in: Proceedings of the USENIX Security Symposium. (Link)] stay there, because they are also the reason to care about vantage choice), and the personalisation bullet stayed and gained the distinction it was missing — personalisation changes the ordering of what you are served, not whether you are served. A link to this page was added to that page's Related pages list. Nothing numeric was moved, so no figure on that page went stale, and its citekey counts are unchanged (checked by diffing them before and after) Move the whole subsection out and leave a bare pointer. Rejected: the subsection is one of the three reasons that page gives for why the vantage point matters, and deleting it would leave that argument with a hole. Or leave the subsection untouched and duplicate it here — rejected as exactly the duplication the split was supposed to avoid
The page id design:blocking_and_geodifference, as fixed on roadmap on 2026-09-07 and enforced by scripts/sitemap.mjs. Writing it under any other id would have left the roadmap promising a page that does not exist design:blocking is shorter, and the geodifference half is arguably Crawling location' or Platforms'. Keeping the fixed id was the cheaper choice and the roadmap row was removed in the same sitting
Circumvention excluded The 28 circumvention papers are counted and named but not in the population: a paper that builds an evasion system measures blocking instrumentally, and including them would put Geneva and domain fronting in a table about how a blocking claim is made Include them — they contain some of the best blocking measurements in the corpus (Geneva's and GET /out's censor models especially). The compromise is that the count is published, so a reader who disagrees can add 28 and see what moves
Protective DNS included In, as the whole filter family (1 paper): its motive is security, but its claim structure is this page's Out, as security blocking. Subtracting one paper is the whole difference
Platform moderation partly in In where the availability of content depends on where or who the client is ([5Knockel, Jeffrey; Dałek, Jakub; Aljizawi, Noura; Ahmed, Mohamed; Meletti, Levi; Lau, Justin (2026): "Banned Books: Analysis of Censorship on Amazon.com", Proceedings on Privacy Enhancing Technologies 2026(3):200-214. (DOI)], [6Lipphardt, Friedemann; Ali, Moonis; Banzer, Martin; Feldmann, Anja; Gosain, Devashish (2026): "There is No War in Ba Sing Se: A Global Analysis of Content Moderation in Large Language Models", in: Proceedings of the Network and Distributed System Security Symposium. (Link)], [11Ablove, Anna; Chandrashekaran, Shreyas; Qiang, Xiao; Ensafi, Roya (2026): "Characterizing the Implementation of Censorship Policies in Chinese LLM Services", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]) or on a deliberate decision against the content ([3Sokoto, Saidu; Balduf, Leonhard; Trautwein, Dennis; Wei, Yiluo; Tyson, Gareth; Castro, Ignacio; Ascigil, Onur; Pavlou, George; Korczyński, Maciej; Scheuermann, Björn; Król, Michał (2024): "Guardians of the Galaxy: Content Moderation in the InterPlanetary File System", in: Proceedings of the USENIX Security Symposium. (Link)], and the IPFS Sybil-censorship paper); out where it is about refusals of what the user asked, or about what moderation did to later behaviour Draw the line at the network and exclude platforms entirely, which would take the population to 58 and remove the fastest-growing slice
The corpus is not the field The page says so twice, and names FOCI and PAM as the venues where much of this community publishes Say it once. Given that a “the literature has not done X” claim appears three times on the page, twice is the minimum
geo.mjs not extended The 15 unmapped location strings are disclosed, not fixed, because the fold is shared with two published pages Add the aliases and re-run the other pages. Correct, but out of scope for one page's sitting; the effect here is nil (every affected paper already has ≥2 mapped locations)
Which methods are “superseded” Two, and both on the strength of a later measurement rather than a hunch: consistency-only DNS detection (retired by [12Tsai, Elisa; Kumar, Deepak; Sundara Raman, Ram; Li, Gavin; Eiger, Yael; Ensafi, Roya (2023): "CERTainty: Detecting DNS Manipulation at Scale using TLS Certificates", Proceedings on Privacy Enhancing Technologies 2023(3):122-137. (DOI)]'s 72.45% false-positive measurement) and page-length thresholds alone (retired by the recall figure in [9McDonald, Allison; Bernhard, Matthew; Valenta, Luke; VanderSloot, Benjamin; Scott, Will; Sullivan, Nick; Halderman, J. Alex; Ensafi, Roya (2018): "403 Forbidden: A Global View of CDN Geoblocking", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] and by [13Jones, Ben; Lee, Tzu-Wen; Feamster, Nick; Gill, Phillipa (2014): "Automated Detection and Fingerprinting of Censorship Block Pages", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]'s own F1 comparison) Call block-page fingerprinting dated too. Rejected: the fingerprint corpora are actively maintained (dns_blockpage_fingerprint, last push 2025-05-01) and every 2024–2026 paper in the population still uses them

14. What could not be established

  • Whether GFWatch's and GFWeb's pipelines still run. Both dashboards render, both stop in 2024. Whether the measurement stopped or only the dashboard did is not answerable from outside, and neither project publishes a status page. The page says exactly that.
  • Whether ICLab's data is available anywhere. iclab.org cannot be reached by a client that verifies certificates, and no mirror was found. Not pursued further — the page treats the paper as the citation.
  • The prevalence of GDPR-driven blocking, which is the page's own headline gap: 2 incidental measurements in 5,859 papers.
  • How much bot management is misread as geoblocking. Both refuse a request on the basis of the client's address. No paper in the population separates them, and this run did not run the experiment that would.
  • Any per-instrument adoption figure. §7's full-text column is a mention count over homograph-heavy names; turning it into a usage figure needs the §4 treatment and was not done.
  • Whether the 2025–2026 growth in server-side and platform blocking is real. 7 of 16 such papers are in the two provisional venue-years. The page states the direction and the caveat, and does not test the trend — with 16 papers a trend test would be theatre.
  • The false-positive rate of this page's own hand audit. 275 verdicts, one auditor. A second auditor on a sample would have given an agreement figure; none exists, and the page does not claim one. This is the largest single weakness in the population figure.

15. The run

Date 2026-09-10
Corpus data/extract/run1/extractions.jsonl, 5,859 papers, 7 venues, 2010–2026
Full text 4 of the 5,859 papers have no paper.cols.txt; they are named in the report output and none is a candidate
Author Claude (Opus 5), one sitting
Reviewers 3 × Sonnet (figures-vs-script, citations-and-quotes, external currency) + 1 × Fable (generic), §12
Scripts committed report_blocking_geodifference.mjs, bgd_quotecheck.py, bgd_external_checks.sh, bgd_render_checks.mjs, each with its -output.txt
Bibliography 33 entries added, generated by scripts/bibgen.mjs from the corpus index with authors fetched from PETS/USENIX landing pages; one hand-corrected (§9). A DOI/title/surname duplicate scan over the combined file found no A- or B-class duplicates; the live page's entry count is printed by the check script

Mistakes made in this run, and what they cost

  1. A live GH_TOKEN was printed into the run log. Testing whether the GitHub API had credentials, this run used ${GH_TOKEN:+set}${GH_TOKEN:-unset} — a form whose second half expands to the value whenever the variable is set. The token appeared in the transcript. The only safe form is ${VAR:+set} alone. It is disclosed here because a leak that is not disclosed cannot be rotated, and the check script now carries the warning inline.
  2. otherToolsMentioned is an array of strings, not of objects. The first instrument probe read .name off each element, silently matched nothing, and reported 29 papers where the fixed version reports 35. Caught by reading a homonym hit, not by any guard.
  3. Two probes matched an English word rather than a method: prompt\w* matched “censorship prompted”, and an unanchored BERT matched the surname Deibert. Both had reached a draft figure.
  4. A first draft of the jurisdiction probe returned 42 papers, most of them about Ethereum, because “block” in this corpus is overwhelmingly blockchain. Replaced by four wall-shaped patterns and published at two widths.
  5. Two “quotes” came from the extraction's model-written summary field and were replaced with the papers' own words (§9). Nothing was published in that state, but the checker had to be extended to catch the class rather than the instance.
  6. scripts/sitemap.mjs reported a failure that was not one. It caches page source under out/pages/ and never expires it, so a roadmap row removed twenty minutes earlier still showed as “WRITTEN but still on the roadmap”. The cache entry was refreshed by hand; the script's cache policy is a known trap rather than a bug found here.
  7. A literal <div id="root"> inside monospace truncated this page from section 11 onward, and it was live in that state for about two minutes. DokuWiki's inline code does not escape a tag: everything after it was dropped, silently, with no error and no failing check — check_wrap.mjs and check_tables.mjs both passed on the source. It was caught by fetching the rendered DOM and comparing its heading, table and <file> counts against the source's. Nothing but reading the render catches this class.
  8. Every internal link on this page was namespace-relative. From provenance:design:, [[roadmap]] resolves to provenance:design:roadmap. Every link here is now root-anchored with a leading colon, and the rendered DOM was checked for wikilink2 (red-link) classes. Writing that sentence produced a third instance of the same bug: the example link, in monospace, was still a live link and rendered red, because inline code does not escape a wiki link any more than it escapes a tag. It needs — which is what the sentence you are reading now uses.
  9. report_blocking_geodifference.mjs exhausted the V8 heap on its first run, because it cached every paper's full text (0.5 GB on disk, over 4 GB as JS strings). It now does a single streaming pass for the whole-corpus regexes and caches only the 63 population texts.

16. The scripts, and their unedited output

Four scripts, each embedded below exactly as committed, followed by its output exactly as produced. The report script is long because the hand-verdict map for all 275 candidates lives inside it — that map is the audit trail, and putting it anywhere else would make the population unauditable without repo access.

report_blocking_geodifference.mjs

The report. Every corpus figure on the content page comes from this file; the hand-verdict map for every candidate is the VERDICT object below, and section 4 of this page tabulates it.

report_blocking_geodifference.mjs
#!/usr/bin/env node
// Report for design:blocking_and_geodifference — every figure the page publishes,
// with the population it is a share of.
//
// Run:  node scripts/report_blocking_geodifference.mjs
//       node scripts/report_blocking_geodifference.mjs --list <family>
//
// Three rules from data/extract/README.md are load-bearing here:
//   * count PAPERS, never tuples;
//   * a sentinel is never an answer;
//   * `detection.phenomenon` / `.technique` are free text (~20% run-to-run
//     agreement on exact strings), so they are folded into families and reported
//     as rankings, never as precise percentages of a phenomenon.
//
// The population is NOT a probe output. Five probes build a CANDIDATE set; every
// candidate carries a hand verdict in VERDICT below, and the script throws if a
// candidate has no verdict or a verdict names a paper no probe caught. The
// families that make up the page's population are MEASURE_FAMILIES.
 
import fs from 'node:fs';
import path from 'node:path';
import { loadExtractions, dataRoot, isSentinel } from './lib.mjs';
import { normalizeLocation } from './geo.mjs';
 
const rows = loadExtractions();
const KEY = (p) => `${p.year}/${p.venue}/${p.slug}`;
const byKey = new Map(rows.map((p) => [KEY(p), p]));
const T = (p) => `${p.title ?? ''} • ${p.summary ?? ''}`;
const pct = (n, d) => (d === 0 ? 'n/a' : `${((100 * n) / d).toFixed(1)}%`);
const norm = (s) => s.replace(/\s+/g, ' ');
 
// ---------------------------------------------------------------- full text
// 5,855 papers of .cols text is ~0.5 GB on disk and more than 4 GB as JS strings,
// so nothing is cached corpus-wide: the whole-corpus regexes are evaluated in ONE
// streaming pass (corpusScan below) and only the population's texts are kept.
function ftRead(p) {
  const f = path.join(dataRoot(), 'fulltext', String(p.year), p.venue, p.slug, 'paper.cols.txt');
  // Whitespace is collapsed before any regex runs: a PDF line break inside a
  // phrase otherwise makes a multi-word probe silently undercount.
  return fs.existsSync(f) ? norm(fs.readFileSync(f, 'latin1')) : null;
}
const popCache = new Map();
function ftPop(p) {
  const k = KEY(p);
  if (!popCache.has(k)) popCache.set(k, ftRead(p));
  return popCache.get(k);
}
 
// ---- one streaming pass for every whole-corpus regex --------------------
const CORPUS_RE = {
  // A negative claim ("nobody classifies block pages with an LLM") must not rest
  // on the narrow signal probe of section 5, so the widened version runs over the
  // whole corpus. Two alternatives had to be repaired after reading the hits:
  // "prompt\\w*" matches "censorship prompted", and an unanchored BERT matches the
  // surname **Deibert** — Ron Deibert of the Citizen Lab, who is in the reference
  // list of a large share of this literature. Every term is now \\b-anchored.
  llmWide: /\b(LLMs?|large language model|GPT-?[0-9]|ChatGPT|Claude|Gemini|Llama|DeBERTa|BERTopic|BERT)\b[^.]{0,200}(block ?page|censor\w*|geo-?block\w*|blocked (page|site|domain|content))|((block ?page|censor\w*|geo-?block\w*)[^.]{0,200}\b(LLMs?|large language model|GPT-?[0-9]|ChatGPT|DeBERTa|BERTopic)\b)/i,
  ftinstr: /\bOONI\b|Open Observatory of Network Interference|Censored ?Planet|\bICLab\b|GFWatch|Citizen ?Lab('s)? (global )?test list|citizenlab\/test-lists/i,
  ftvocab: /block ?page|blocking page|geo-?block|geo-?differen|geo-?restrict|HTTP 451|unavailable for legal reasons/i,
};
const INSTRUMENTS = {
  OONI: /\bOONI\b|Open Observatory of Network Interference/i,
  'Censored Planet': /Censored ?Planet/i,
  ICLab: /\bICLab\b/i,
  'Citizen Lab test list': /Citizen ?Lab('s)? (global )?(test|citizen) ?lists?|citizenlab\/test-lists/i,
  Quack: /\bQuack\b/i,
  Hyperquack: /\bHyperquack\b/i,
  Augur: /\bAugur\b/i,
  'Iris (DNS)': /\bIris\b/i,
  Satellite: /\bSatellite\b/i,
  GFWatch: /\bGFWatch\b/i,
  Geneva: /\bGeneva\b/i,
  'M-Lab': /\bM-?Lab\b|Measurement Lab/i,
  'OpenNet Initiative': /OpenNet Initiative|\bONI\b/i,
  Encore: /\bEncore\b/i,
};
// Tight, wall-specific probes. "block" alone is useless here: in this corpus it is
// overwhelmingly blockchain, blocklist and ad blocking.
const WALL_TIGHT = [
  /(?:geo-?block|block(?:ed|ing|s)?|deny|denied|denies|refus\w+|unavailable|inaccessible)[^.]{0,70}\b(?:EU|EEA|European(?: Union)?)\b[^.]{0,50}(?:visitor|user|reader|resident|traffic|IP address|audienc)/i,
  /\b(?:EU|EEA|European(?: Union)?)\b[^.]{0,50}(?:visitor|user|reader|resident|traffic)s?[^.]{0,70}(?:are |were |is |was )?(?:geo-?block\w*|block\w*|denied access|refused|shut out|unavailable)/i,
  /GDPR[- ](?:induced|driven|related|motivated)[^.]{0,40}(?:geo-?block|block|withdrawal|exit)/i,
  /not available[^.]{0,40}in (?:the )?(?:EU|Europe|European Union|EEA)/i,
];
// The loose probe must CONTAIN the tight one, or "tight" and "loose" are two
// different questions rather than two widths of one, and the union is not
// max(|tight|,|loose|). The first version did not: 6 of 9 tight hits matched no
// loose pattern, so a review pass found the page reporting 9+18=27 "distinct"
// hits where the real union was 24. It is now built from WALL_TIGHT by
// construction, and the guard below asserts containment rather than size.
const WALL_LOOSE = [
  ...WALL_TIGHT,
  /\b(?:EU|EEA|European)\b[^.]{0,120}\b(?:geo-?block\w*|geo-?fenc\w*|block(?:ed|ing)?|unavailable|inaccessible|refus\w+|denied)\b/i,
  /\b(?:geo-?block\w*|geo-?fenc\w*)\b[^.]{0,120}\b(?:GDPR|CCPA|ePrivacy|California|EU|EEA|European)\b/i,
  /\b(?:GDPR|CCPA)\b[^.]{0,120}\b(?:geo-?block\w*|geo-?fenc\w*)\b/i,
];
const scan = new Map(); // key -> { ftinstr, ftvocab, instr: Set, wallTight, wallLoose }
const ftMissing = [];
for (const p of rows) {
  const t = ftRead(p);
  if (t === null) { ftMissing.push(p); scan.set(KEY(p), null); continue; }
  const rec = { instr: new Set() };
  for (const [n, re] of Object.entries(CORPUS_RE)) rec[n] = re.test(t);
  for (const [n, re] of Object.entries(INSTRUMENTS)) if (re.test(t)) rec.instr.add(n);
  rec.wallTight = WALL_TIGHT.some((re) => re.test(t));
  rec.wallLoose = WALL_LOOSE.some((re) => re.test(t));
  scan.set(KEY(p), rec);
}
const S = (p) => scan.get(KEY(p));
 
// ------------------------------------------------------------------- probes
// NOTE: otherToolsMentioned is an array of STRINGS, not of objects. Reading
// `.name` off it matched nothing at all in the first version of this probe.
const toolNames = (p) => [...p.tools.map((t) => t.name), ...p.otherToolsMentioned].map((n) => String(n ?? '').trim());
 
const INSTR_RE = /^(OONI|OONI ?Probe|OONI ?Explorer|OONI data|Censored ?Planet|ICLab|Quack|Hyperquack|Satellite|Augur|GFWatch|Encore|Geneva|Iris|CensorScope|Tor Metrics|Citizen ?Lab.*|Block(page)? ?(corpus|list of blockpages))$/i;
 
// Hand-added candidates, with the reason each one was added.
const MANUAL_CANDIDATES = new Map([
  ['2022/USENIX/openvpn-is-open-to-vpn-fingerprinting',
   'named by the 2026-09-10 generic review as the same shape as two papers in the population; no probe reaches it because its title and summary never say censorship, blocking or VPN blocking'],
]);
 
const PROBES = {
  title: (p) => /censorship|censored|censoring|censors\b|OONI|Censored Planet|ICLab|GFW\b|Great Firewall|geo-?block|blockpage|block page|internet shutdown|geo-?differen|geo-?restrict|geofenc|network interference|connection tampering|DNS manipulation|throttl/i.test(T(p)),
  detect: (p) => p.detection.some((t) => /censor|geo-?block|geo-?differen|block ?page|internet (filter|shutdown)|dns (manipulat|censor|interference|redirection|tamper)|network interference|connection tamper|url[- ]filter|keyword[- ]?(based )?filter|content (filter|censor)|over-?block|site blocking|blocked (site|domain|web|content|url|categor)|throttling of|traffic throttl|sanction/i.test(String(t.phenomenon ?? ''))),
  instr: (p) => toolNames(p).some((n) => INSTR_RE.test(n)),
  ftinstr: (p) => S(p) !== null && S(p).ftinstr,
  ftvocab: (p) => S(p) !== null && S(p).ftvocab,
  // Sixth probe, added after the first run: the jurisdiction-wall patterns of
  // section 8 caught seven papers no other probe saw. A probe that finds a paper
  // the candidate set missed is a recall hole, so it becomes a candidate probe.
  wall: (p) => S(p) !== null && S(p).wallTight,
  // Seventh probe, added after the second run. The first six are built around the
  // words "censor" and "geoblock", so they miss a paper that says "reachability",
  // "content moderation" or "blocking-resistant" instead. It found two papers that
  // belong in the population (the 2019 DoT/DoH reachability measurement and the
  // IPFS moderation study) and 44 that do not, all verdicted below.
  // Eighth probe, added after review. The bibliography's own newest entry —
  // "Tracking Internet Disruptions in Ukraine" (IMC 2025) — was caught by none of
  // probes 1-7, and the page had claimed that shutdowns as an event class were
  // absent from the corpus. They are not: there is an outage-detection literature
  // here, it just does not attribute unreachability to anybody's decision. That
  // distinction is the inclusion rule, so these papers are an ADJACENT family
  // rather than off topic.
  outage: (p) => /\b(outage|outages|shutdown|shutdowns|blackout|disruption|disruptions|internet resilience|connectivity loss|loss of connectivity|depeering|route withdrawal)\b/i.test(T(p)),
  // Ninth "probe": papers added by hand. A review pass named a paper that no
  // regex could reach — OpenVPN fingerprinting, whose title and summary contain
  // no blocking vocabulary at all — and a candidate set has to be able to admit
  // one, with the reason recorded, rather than pretend the regexes are complete.
  manual: (p) => MANUAL_CANDIDATES.has(KEY(p)),
  recall: (p) => /blocking-?resist|unblock|un-?censor|reachab|inaccessib|availability across|available across|not available in|content moderation|takedown|deplatform|shadow ?ban|across (\d+ )?(countries|regions|jurisdictions)|country-?level (difference|variation)|regional (variation|difference)|per-country/i.test(T(p)),
};
 
const candidates = new Map(); // key -> string[] of probe names
for (const p of rows) {
  const hit = Object.entries(PROBES).filter(([, f]) => f(p)).map(([n]) => n);
  if (hit.length) candidates.set(KEY(p), hit);
}
 
// --------------------------------------------------------------- verdicts
// One hand verdict per candidate, from reading title + summary and, where that
// was not enough, the matched sentence in paper.cols.txt. A paper may carry two
// families (Network Responses measures censorship AND corporate geoblocking), so
// families overlap and their counts do not sum to the population.
//
// MEASURE families = the page's population: the paper measures who could not
// reach what. Adjacent families are counted and named but excluded, because a
// circumvention proposal or an interview study is not a blocking measurement.
const MEASURE_FAMILIES = ['net', 'server', 'jurisdiction', 'platform', 'filter', 'probelist'];
const ADJACENT_FAMILIES = ['circumvention', 'outage', 'wall-encounter', 'interview', 'sok', 'vantage-instrument'];
const FAMILY_LABEL = {
  net: 'network-level interference with access (censorship, DNS manipulation, connection tampering, throttling, interception)',
  server: 'server- or platform-side refusal or differentiation by region (geoblocking, geodifference, sanctions)',
  jurisdiction: 'blocking or differentiation driven by a privacy law (GDPR wall, CCPA geofencing)',
  platform: 'moderation or censorship enforced inside one platform or service',
  filter: 'filtering products deployed on the client side (protective DNS), and their over-blocking',
  probelist: 'the censorship probe list itself as the object of study',
  circumvention: 'ADJACENT: evasion or circumvention system, or censor-side detection of it',
  interview: 'ADJACENT: human-subjects study of living with censorship',
  sok: 'ADJACENT: SoK or survey of the censorship literature',
  'vantage-instrument': 'ADJACENT: instrument for measuring FROM a place (see design:crawling_location)',
  'wall-encounter': 'ADJACENT: ran into, worked around, or reported others proposing an EU wall, without measuring how common it is',
  'outage': 'ADJACENT: measures unreachability without attributing it to a deliberate blocking decision — power, fibre, weather, war damage, misconfiguration',
};
const OFF_LABEL = {
  'off:incidental': 'one incidental mention, or a related-work citation only',
  'off:homonym-tor-metrics': 'homonym: names Tor Metrics, a Tor network dataset, not a blocking observatory',
  'off:homonym-iris': 'homonym: iRiS, an iOS private-API analyser, not Iris the DNS-manipulation platform',
  'off:homonym-augur': 'homonym: names Augur in a browser-fingerprinting context, not Augur the disruption prober',
  'off:homonym-throttling': 'homonym: bandwidth throttling as a defence, not throttling as censorship',
  'off:homonym-statistical-censoring': 'homonym: "censored"/"censoring" in the statistics or ML sense',
  'off:homonym-ml-concept-censorship': 'homonym: concept censorship in a generative model',
  'off:blockchain-sanctions': 'sanctions and blocklists on a blockchain, not access to content',
  'off:adblocking': 'content blocking by the user’s own ad blocker (see programming:filter_lists)',
  'off:security-blocking': 'blocking malware or phishing, not blocking by who or where you are',
  'off:vpn-proxy-ecosystem': 'the VPN or proxy ecosystem itself (see design:crawling_location)',
  'off:proxy-detection': 'detecting proxy or VPN clients (the server side of vantage choice)',
  'off:age-gate': 'age assurance, a different access gate (queued as privacy:age_assurance)',
  'off:bot-blocking': 'blocked for looking like a bot (see programming:crawler_detection)',
  'off:censor-as-attack-surface': 'censorship middleboxes abused as an amplifier, not measured as a censor',
  'off:llm-guardrails': 'model guardrails and refusals of what the user asked, not access blocking',
  'off:takedown-effects': 'the effect of a takedown or deplatforming on the target, not whether a client could reach content',
  'off:moderation-effects': 'the effect of moderating user posts on later user behaviour',
  'off:moderation-system': 'proposes or evaluates a moderation classifier — builds the blocker rather than measuring blocking',
  'off:resilience-modelling': 'models what would happen if countries disconnected, from routing graphs; no client reachability measured',
};
 
const VERDICT = {
  '2010/WWW/measurement-and-analysis-of-an-online-content-voting-network-a-case-study-of-dig': ['off:incidental'],
  '2011/IMC/analysis-of-country-wide-internet-outages-caused-by-censorship': ['net'],
  '2012/CCS/context-aware-web-security-threat-prevention': ['off:security-blocking'],
  '2012/USENIX/throttling-tor-bandwidth-parasites': ['off:homonym-throttling'],
  '2013/CCS/polyglots-crossing-origins-by-crossing-formats': ['off:incidental'],
  '2013/USENIX/automatic-mediation-of-privacy-sensitive-resource-access-in-smartphone-applicati': ['off:incidental'],
  '2013/CCS/protocol-misidentification-made-easy-with-format-transforming-encryption': ['circumvention'],
  '2013/CCS/users-get-routed-traffic-correlation-on-tor-by-realistic-adversaries': ['off:homonym-tor-metrics'],
  '2013/IMC/a-method-for-identifying-and-confirming-the-use-of-url-filtering-products-for-ce': ['net'],
  '2014/IMC/a-look-at-the-consequences-of-internet-censorship-through-an-isp-lens': ['net'],
  '2014/IMC/automated-detection-and-fingerprinting-of-censorship-block-pages': ['net'],
  '2014/IMC/censorship-in-the-wild-analyzing-internet-filtering-in-syria': ['net'],
  '2014/USENIX/brahmastra-driving-apps-to-test-the-security-of-third-party-components': ['off:incidental'],
  '2015/CCS/cachebrowser-bypassing-chinese-censorship-without-proxies-using-cached-content': ['circumvention'],
  '2015/CCS/iris-vetting-private-api-abuse-in-ios-applications': ['off:homonym-iris'],
  '2015/IMC/examining-how-the-great-firewall-discovers-hidden-circumvention-servers': ['net'],
  '2015/IMC/going-wild-large-scale-classification-of-open-dns-resolvers': ['net'],
  '2015/IMC/in-and-out-of-cuba-characterizing-cubas-connectivity': ['net'],
  '2015/IMC/ting-measuring-and-exploiting-latencies-between-all-tor-nodes': ['off:homonym-tor-metrics'],
  '2015/PETS/analyzing-the-great-firewall-of-china-over-space-and-time': ['net'],
  '2016/CCS/practical-censorship-evasion-leveraging-content-delivery-networks': ['circumvention'],
  '2016/IMC/an-analysis-of-the-privacy-and-security-risks-of-android-vpn-permission-enabled': ['off:vpn-proxy-ecosystem'],
  '2016/IMC/tunneling-for-transparency-a-large-scale-analysis-of-end-to-end-violations-in-th': ['net'],
  '2016/PETS/denasa-destination-naive-as-awareness-in-anonymous-communications': ['off:homonym-tor-metrics'],
  '2016/IEEE-SP/sok-towards-grounding-censorship-circumvention-in-empiricism': ['sok'],
  '2017/IMC/initial-measurements-of-the-cuban-street-network': ['off:incidental'],
  '2017/IMC/your-state-is-not-mine-a-closer-look-at-evading-stateful-internet-censorship': ['circumvention'],
  '2017/NDSS/dissecting-tor-bridges-a-security-evaluation-of-their-private-and-public-infrast': ['off:homonym-tor-metrics'],
  '2017/PETS/a-usability-evaluation-of-tor-launcher': ['interview'],
  '2017/PETS/deltashaper-enabling-unobservable-censorship-resistant-tcp-tunneling-over-videoc': ['circumvention'],
  '2017/PETS/topics-of-controversy-an-empirical-analysis-of-web-censorship-lists': ['probelist'],
  '2017/USENIX/characterizing-the-nature-and-dynamics-of-tor-exit-blocking': ['server'],
  '2017/USENIX/global-measurement-of-dns-manipulation': ['net'],
  '2017/IEEE-SP/augur-internet-wide-detection-of-connectivity-disruptions': ['net'],
  '2018/IMC/403-forbidden-a-global-view-of-cdn-geoblocking': ['server'],
  '2018/PETS/nomoads-effective-and-efficient-cross-app-mobile-ad-blocking': ['off:incidental'],
  '2020/WWW/filter-list-generation-for-underserved-regions': ['off:incidental'],
  '2023/NDSS/i-still-know-what-you-watched-last-sunday-privacy-of-the-hbbtv-protocol-in-the-european-smart-tv-landscape': ['off:incidental'],
  '2023/NDSS/chkplug-checking-gdpr-compliance-of-wordpress-plugins-via-cross-language-code-property-graph': ['wall-encounter'],
  '2023/PETS/everybodys-looking-for-ssomething-a-large-scale-evaluation-on-the-privacy-of-oau': ['wall-encounter'],
  '2026/PETS/the-role-of-online-forums-in-developer-understanding-of-privacy-law-a-reddit-cas': ['wall-encounter'],
  '2026/PETS/privacy-vs-profit-the-impact-of-googles-manifest-version-3-mv3-update-on-ad-bloc': ['off:incidental'],
  '2018/IMC/an-empirical-study-of-the-i2p-anonymity-network-and-its-censorship-resistance': ['circumvention'],
  '2018/IMC/how-to-catch-when-proxies-lie-verifying-the-physical-locations-of-network-proxie': ['vantage-instrument'],
  '2018/IMC/where-the-light-gets-in-analyzing-web-censorship-mechanisms-in-india': ['net'],
  '2018/USENIX/fp-scanner-the-privacy-implications-of-browser-fingerprint-inconsistencies': ['off:homonym-augur'],
  '2018/USENIX/quack-scalable-remote-measurement-of-application-layer-censorship': ['net'],
  '2018/USENIX/the-dangers-of-key-reuse-practical-attacks-on-ipsec-ike': ['off:incidental'],
  '2018/USENIX/towards-predicting-efficient-and-anonymous-tor-circuits': ['off:homonym-tor-metrics'],
  '2018/WWW/proxytorrent-untangling-the-free-http-s-proxy-ecosystem': ['off:vpn-proxy-ecosystem'],
  '2019/CCS/geneva-evolving-censorship-evasion-strategies': ['circumvention'],
  '2019/CCS/you-shall-not-join-a-measurement-study-of-cryptocurrency-peer-to-peer-bootstrapp': ['net'],
  '2019/CCS/your-cache-has-fallen-cache-poisoned-denial-of-service-attack': ['off:incidental'],
  '2019/IEEE-SP/breaking-lte-on-layer-two': ['off:incidental'],
  '2019/IMC/tales-from-the-porn-a-comprehensive-privacy-analysis-of-the-web-porn-ecosystem': ['off:incidental'],
  '2019/NDSS/the-use-of-tls-in-censorship-circumvention': ['circumvention'],
  '2019/WWW/evaluating-anti-fingerprinting-privacy-enhancing-technologies': ['off:homonym-tor-metrics'],
  '2020/CCS/censored-planet-an-internet-wide-longitudinal-censorship-observatory': ['net'],
  '2020/CCS/continuous-and-multiregional-monitoring-of-malicious-hosts': ['off:incidental'],
  '2020/IEEE-SP/meddling-middlemen-empirical-analysis-of-the-risks-of-data-saving-mobile-browser': ['off:incidental'],
  '2020/IMC/how-china-detects-and-blocks-shadowsocks': ['net'],
  '2020/IMC/investigating-large-scale-https-interception-in-kazakhstan': ['net'],
  '2020/NDSS/a-practical-approach-for-taking-down-avalanche-botnets-under-real-world-constraints': ['off:incidental'],
  '2020/NDSS/decentralized-control-a-case-study-of-russia': ['net'],
  '2020/NDSS/encrypted-dns-privacy-a-traffic-analysis-perspective': ['circumvention'],
  '2020/NDSS/massbrowser-unblocking-the-censored-web-for-the-masses-by-the-masses': ['circumvention'],
  '2020/NDSS/measuring-the-deployment-of-network-censorship-filters-at-global-scale': ['net'],
  '2020/NDSS/on-using-application-layer-middlebox-protocols-for-peeking-behind-nat-gateways': ['off:incidental'],
  '2020/NDSS/symtcp-eluding-stateful-deep-packet-inspection-with-automated-discrepancy-discovery': ['circumvention'],
  '2020/PETS/no-boundaries-data-exfiltration-by-third-parties-embedded-on-web-pages': ['off:incidental'],
  '2020/PETS/running-refraction-networking-for-real': ['circumvention'],
  '2020/WWW/apophanies-or-epiphanies-how-crawlers-impact-our-understanding-of-the-web': ['off:incidental'],
  '2020/WWW/deconstructing-googles-web-light-service': ['server'],
  '2020/WWW/maddroid-characterizing-and-detecting-devious-ad-contents-for-android-apps': ['off:incidental'],
  '2020/IEEE-SP/iclab-a-global-longitudinal-internet-censorship-measurement-platform': ['net'],
  '2020/CCS/youve-changed-detecting-malicious-browser-extensions-through-their-update-deltas': ['off:incidental'],
  '2021/IMC/throttling-twitter-an-emerging-censorship-technique-in-russia': ['net'],
  '2021/IMC/web-censorship-measurements-of-http-3-over-quic': ['net'],
  '2021/NDSS/hey-alexa-is-this-skill-safe-taking-a-closer-look-at-the-alexa-skill-ecosystem': ['off:incidental'],
  '2021/PETS/oblivious-dns-over-https-odoh-a-practical-privacy-enhancement-to-dns': ['off:incidental'],
  '2021/PETS/too-close-for-comfort-morasses-of-anti-censorship-in-the-era-of-cdns': ['circumvention'],
  '2021/PETS/domain-name-encryption-is-not-enough-privacy-leakage-via-ip-based-website-finger': ['off:incidental'],
  '2021/USENIX/blind-in-on-path-attacks-and-applications-to-vpns': ['off:incidental'],
  '2021/PETS/holes-in-the-geofence-privacy-vulnerabilities-in-smart-dns-services': ['circumvention'],
  '2021/USENIX/domain-shadowing-leveraging-content-delivery-networks-for-robust-blocking-resist': ['circumvention'],
  '2021/USENIX/how-great-is-the-great-firewall-measuring-chinas-dns-censorship': ['net'],
  '2021/USENIX/its-stressful-having-all-these-phones-investigating-sex-workers-safety-goals-ris': ['off:incidental'],
  '2021/USENIX/weaponizing-middleboxes-for-tcp-reflected-amplification': ['off:censor-as-attack-surface'],
  '2021/WWW/chinese-wall-or-swiss-cheese-keyword-filtering-in-the-great-firewall-of-china': ['net'],
  '2021/WWW/security-of-alerting-authorities-in-the-www-measuring-namespaces-dnssec-and-web': ['off:incidental'],
  '2021/WWW/understanding-the-impact-of-encrypted-dns-on-internet-censorship': ['net'],
  '2022/CCS/ready-raider-one-exploring-the-misuse-of-cloud-gaming-services': ['off:incidental'],
  '2022/CCS/hidden-in-plain-sight-exploring-encrypted-channels-in-android-apps': ['off:incidental'],
  '2022/IMC/characterizing-permanently-dead-links-on-wikipedia': ['off:incidental'],
  '2022/IMC/respect-the-origin-a-best-case-evaluation-of-connection-coalescing-in-the-wild': ['off:incidental'],
  '2022/IMC/rusty-clusters-dusting-an-ipv6-research-foundation': ['net'],
  '2022/IMC/tspu-russias-decentralized-censorship-system': ['net'],
  '2022/NDSS/auto-draft-244': ['off:vpn-proxy-ecosystem'],
  '2022/PETS/from-onion-not-found-to-guard-discovery': ['off:homonym-tor-metrics'],
  '2022/PETS/setting-the-bar-low-are-websites-complying-with-the-minimum-requirements-of-the': ['jurisdiction'],
  '2022/PETS/toward-uncensorable-anonymous-and-private-access-over-satoshi-blockchains': ['circumvention'],
  '2022/PETS/understanding-utility-and-privacy-of-demographic-data-in-education-technology-by': ['off:homonym-statistical-censoring'],
  '2022/USENIX/a-large-scale-investigation-into-geodifferences-in-mobile-apps': ['server'],
  '2022/USENIX/get-out-automated-discovery-of-application-layer-censorship-evasion-strategies': ['circumvention'],
  '2022/USENIX/how-and-why-people-use-virtual-private-networks': ['off:vpn-proxy-ecosystem'],
  '2022/USENIX/many-roads-lead-to-rome-how-packet-headers-influence-dns-censorship-measurement': ['net'],
  '2022/USENIX/off-path-network-traffic-manipulation-via-revitalized-icmp-redirect-attacks': ['off:incidental'],
  '2022/USENIX/online-website-fingerprinting-evaluating-website-fingerprinting-attacks-on-tor-i': ['off:incidental'],
  '2022/USENIX/security-at-the-end-of-the-tunnel-the-anatomy-of-vpn-mental-models-among-experts': ['off:vpn-proxy-ecosystem'],
  '2022/USENIX/the-security-lottery-measuring-client-side-web-security-inconsistencies': ['server'],
  '2022/WWW/measuring-the-privacy-vs-compatibility-trade-off-in-preventing-third-party-state': ['off:incidental'],
  '2022/WWW/reproducibility-and-replicability-of-web-measurement-studies': ['off:incidental'],
  '2022/IEEE-SP/investigating-influencer-vpn-ads-on-youtube': ['off:vpn-proxy-ecosystem'],
  '2023/IMC/ptperf-on-the-performance-evaluation-of-tor-pluggable-transports': ['circumvention'],
  '2023/IMC/wolf-in-sheeps-clothing-evaluating-security-risks-of-the-undelegated-record-on-d': ['off:incidental'],
  '2023/NDSS/a-systematic-study-of-the-consistency-of-two-factor-authentication-user-journeys-on-top-ranked-websites': ['off:incidental'],
  '2023/CCS/poster-circumventing-the-gfw-with-tls-record-fragmentation': ['circumvention'],
  '2023/NDSS/your-router-is-my-prober-measuring-ipv6-networks-via-icmp-rate-limiting-side-channels': ['off:incidental'],
  '2023/PETS/certainty-detecting-dns-manipulation-at-scale-using-tls-certificates': ['net'],
  '2023/CCS/txphishscope-towards-detecting-and-understanding-transaction-based-phishing-on-e': ['off:incidental'],
  '2023/CCS/weve-disabled-mfa-for-you-an-evaluation-of-the-security-and-usability-of-multi-f': ['off:incidental'],
  '2023/IMC/ethereums-proposer-builder-separation-promises-and-realities': ['off:blockchain-sanctions'],
  '2023/USENIX/a-study-of-chinas-censorship-and-its-evasion-through-the-lens-of-online-gaming': ['interview'],
  '2023/USENIX/all-of-them-claim-to-be-the-best-multi-perspective-study-of-vpn-users-and-vpn-pr': ['off:vpn-proxy-ecosystem'],
  '2023/USENIX/bypassing-tunnels-leaking-vpn-client-traffic-by-abusing-routing-tables': ['off:incidental'],
  '2023/USENIX/deresistor-toward-detection-resistant-probing-for-evasion-of-internet-censorship': ['circumvention'],
  '2023/USENIX/how-the-great-firewall-of-china-detects-and-blocks-fully-encrypted-traffic': ['net'],
  '2023/WWW/on-how-zero-knowledge-proof-blockchain-mixers-improve-and-worsen-user-privacy': ['off:blockchain-sanctions'],
  '2023/USENIX/network-responses-to-russias-invasion-of-ukraine-in-2022-a-cautionary-tale-for-i': ['net', 'server'],
  '2023/USENIX/mobileatlas-geographically-decoupled-measurements-in-cellular-networks-for-secur': ['vantage-instrument'],
  '2023/WWW/interval-censored-transformer-hawkes-detecting-information-operations-using-the': ['off:homonym-statistical-censoring'],
  '2023/WWW/measuring-and-evading-turkmenistans-internet-censorship-a-case-study-in-large-sc': ['net'],
  '2024/IMC/poster-twinkle-twinkle-streaming-star-illuminating-cdn-performance-over-starlink': ['off:incidental'],
  '2024/NDSS/content-censorship-in-the-interplanetary-file-system': ['platform'],
  '2024/NDSS/masterkey-automated-jailbreaking-of-large-language-model-chatbots': ['off:llm-guardrails'],
  '2024/NDSS/on-precisely-detecting-censorship-circumvention-in-real-world-networks': ['circumvention'],
  '2024/CCS/the-privacy-utility-trade-off-in-the-topics-api': ['off:incidental'],
  '2024/NDSS/understanding-the-implementation-and-security-implications-of-protective-dns-services': ['filter'],
  '2024/PETS/generational-differences-in-understandings-of-privacy-terminology': ['off:incidental'],
  '2024/CCS/understanding-routing-induced-censorship-changes-globally': ['net'],
  '2024/IEEE-SP/sok-technical-implementation-and-human-impact-of-internet-privacy-regulations': ['off:incidental'],
  '2024/IMC/the-wisdom-of-the-measurement-crowd-building-the-internet-yellow-pages-a-knowled': ['off:incidental'],
  '2024/IMC/yesterday-once-more-global-measurement-of-internet-traffic-shadowing-behaviors': ['off:incidental'],
  '2024/USENIX/bridging-barriers-a-survey-of-challenges-and-priorities-in-the-censorship-circum': ['interview'],
  '2024/USENIX/calculatency-leveraging-cross-layer-network-latency-measurements-to-detect-proxy': ['off:proxy-detection'],
  '2024/NDSS/modeling-and-detecting-internet-censorship-events': ['net'],
  '2024/USENIX/diffie-hellman-picture-show-key-exchange-stories-from-commercial-vowifi-deployme': ['off:incidental'],
  '2024/USENIX/digital-discrimination-of-users-in-sanctioned-states-the-case-of-the-cuba-embarg': ['server'],
  '2024/PETS/automatic-generation-of-web-censorship-probe-lists': ['probelist'],
  '2024/PETS/what-to-expect-when-you-re-accessing-an-exploration-of-user-privacy-rights-in-pe': ['jurisdiction'],
  '2024/USENIX/abandon-all-hope-ye-who-enter-here-a-dynamic-longitudinal-investigation-of-andro': ['off:incidental'],
  '2024/PETS/cross-contextual-examination-of-older-adults-privacy-concerns-behaviors-and-vuln': ['off:incidental'],
  '2024/USENIX/fingerprinting-obfuscated-proxy-traffic-with-encapsulated-tls-handshakes': ['off:proxy-detection'],
  '2024/USENIX/gfweb-measuring-the-great-firewalls-web-censorship-at-scale': ['net'],
  '2024/USENIX/less-defined-knowledge-and-more-true-alarms-reference-based-phishing-detection-w': ['off:incidental'],
  '2024/USENIX/loopy-hell-ow-infinite-traffic-loops-at-the-application-layer': ['off:incidental'],
  '2024/USENIX/phishdecloaker-detecting-captcha-cloaked-phishing-websites-via-hybrid-vision-bas': ['off:incidental'],
  '2024/USENIX/spotproxy-rediscovering-the-cloud-for-censorship-circumvention': ['circumvention'],
  '2024/USENIX/is-it-a-trap-a-large-scale-empirical-study-and-comprehensive-assessment-of-onlin': ['off:incidental'],
  '2024/WWW/phishinwebview-analysis-of-anti-phishing-entities-in-mobile-apps-with-webview-ta': ['off:security-blocking'],
  '2024/WWW/a-worldwide-view-on-the-reachability-of-encrypted-dns-services': ['net'],
  '2024/IEEE-SP/to-auth-or-not-to-auth-a-comparative-analysis-of-the-pre-and-post-login-security': ['off:incidental'],
  '2024/IEEE-SP/pryde-a-modular-generalizable-workflow-for-uncovering-evasion-attacks-against-st': ['circumvention'],
  '2025/CCS/local-frames-exploiting-inherited-origins-to-bypass-content-blockers': ['off:adblocking'],
  '2025/CCS/security-and-privacy-measurements-in-cellular-networks-novel-approaches-in-a-glo': ['off:incidental'],
  '2025/IEEE-SP/data-to-infinity-and-beyond-examining-data-sharing-and-reuse-practices-in-the-co': ['off:incidental'],
  '2025/CCS/deep-dive-into-in-app-browsers-uncovering-hidden-pitfalls-in-certificate-validat': ['off:incidental'],
  '2025/IEEE-SP/learning-from-censored-experiences-social-media-discussions-around-censorship-ci': ['interview'],
  '2025/CCS/fingerprinting-deep-packet-inspection-devices-by-their-ambiguities': ['net'],
  '2025/IEEE-SP/born-with-a-silver-spoon-on-the-in-security-of-native-granted-app-privileges-in': ['off:incidental'],
  '2025/IMC/the-developer-the-rfc-and-the-middlebox-an-http-2-compliance-story': ['off:incidental'],
  '2025/NDSS/wallbleed-a-memory-disclosure-vulnerability-in-the-great-firewall-of-china': ['net'],
  '2025/IMC/where-in-the-world-are-my-trackers-mapping-web-tracking-flow-across-diverse-geog': ['server'],
  '2025/USENIX/double-edged-shield-on-the-fingerprintability-of-customized-ad-blockers': ['off:adblocking'],
  '2025/NDSS/themis-regulating-textual-inversion-for-personalized-concept-censorship': ['off:homonym-ml-concept-censorship'],
  '2025/USENIX/irblock-a-large-scale-measurement-study-of-the-great-firewall-of-iran': ['net'],
  '2025/USENIX/easy-as-childs-play-an-empirical-study-on-age-verification-of-adult-oriented-and': ['off:age-gate'],
  '2025/USENIX/exposing-and-circumventing-sni-based-quic-censorship-of-the-great-firewall-of-ch': ['net', 'circumvention'],
  '2025/WWW/harmful-terms-and-where-to-find-them-measuring-and-modeling-unfavorable-financia': ['off:incidental'],
  '2025/IEEE-SP/analyzing-the-ios-local-network-permission-from-a-technical-and-user-perspective': ['off:incidental'],
  '2025/IEEE-SP/is-nobody-there-good-globally-measuring-connection-tampering-without-responsive': ['net'],
  '2025/IEEE-SP/transport-layer-obscurity-circumventing-sni-censorship-on-the-tls-layer': ['circumvention'],
  '2025/IMC/somesite-i-used-to-crawl-awareness-agency-and-efficacy-in-protecting-content-cre': ['off:bot-blocking'],
  '2025/NDSS/the-discriminative-power-of-cross-layer-rtts-in-fingerprinting-proxy-traffic': ['off:proxy-detection'],
  '2025/IEEE-SP/a-wall-behind-a-wall-emerging-regional-censorship-in-china': ['net'],
  '2025/IMC/dive-into-the-cloud-unveiling-the-ab-usage-of-serverless-cloud-function-in-the-w': ['off:incidental'],
  '2025/IEEE-SP/code-speaks-louder-exploring-security-and-privacy-relevant-regional-variations-i': ['server'],
  '2026/NDSS/mirage-private-mobility-based-routing-for-censorship-evasion': ['circumvention'],
  '2026/NDSS/beyond-rtt-an-adversarially-robust-two-tiered-approach-for-residential-proxy-detection': ['off:proxy-detection'],
  '2026/NDSS/characterizing-the-implementation-of-censorship-policies-in-chinese-llm-services': ['platform'],
  '2026/PETS/precarious-but-active-a-look-at-privacy-behaviors-in-chinese-transformative-fand': ['interview'],
  '2026/PETS/quicstep-evaluating-connection-migration-based-quic-censorship-circumvention': ['circumvention'],
  '2026/USENIX/from-mirai-to-gorilla-deep-dive-into-a-long-lasting-ddos-for-hire-botnet': ['off:incidental'],
  '2026/WWW/evasion-under-blockchain-sanctions': ['off:blockchain-sanctions'],
  '2026/PETS/banned-books-analysis-of-censorship-on-amazon-com': ['server', 'platform'],
  '2026/PETS/exercising-the-ccpa-opt-out-right-on-android-legally-mandated-but-practically-ch': ['off:incidental'],
  '2026/NDSS/chameleoscan-demystifying-and-detecting-ios-chameleon-apps-via-llm-powered-ui-exploration': ['off:incidental'],
  '2026/WWW/tracking-the-stray-sheep-understanding-dns-response-manipulation-in-the-wild': ['net'],
  '2026/NDSS/there-is-no-war-in-ba-sing-se-a-global-analysis-of-content-moderation-in-large-language-models': ['platform', 'server'],
  '2026/PETS/disclosure-divergence-measuring-privacy-policy-and-data-safety-misalignment-at-s': ['off:incidental'],
  '2026/NDSS/mvpnalyzer-an-investigative-framework-for-auditing-the-security-privacy-of-mobile-vpns': ['off:vpn-proxy-ecosystem'],
  '2026/PETS/on-the-suitability-of-llm-driven-agents-for-dark-pattern-audits': ['off:incidental'],
  // --- added by hand after the generic review pass, 2026-09-10 ---
  // Censor-side detection of circumvention traffic: proposes the fingerprinting
  // method rather than measuring a deployed censor, so ADJACENT, not population.
  '2022/USENIX/openvpn-is-open-to-vpn-fingerprinting': ['circumvention'],
  // --- added by the eighth (outage) probe, 2026-09-10, after review ---
  '2010/USENIX/sepia-privacy-preserving-aggregation-of-multi-domain-network-events-and-statisti': ['off:incidental'],
  '2013/IEEE-SP/sok-p2pwned-modeling-and-evaluating-the-resilience-of-peer-to-peer-botnets': ['off:incidental'],
  '2015/CCS/from-system-services-freezing-to-system-server-shutdown-in-android-all-you-need': ['off:incidental'],
  '2015/NDSS/the-resilience-of-the-internet-to-colluding-country-induced-connectivity-disrupt': ['off:resilience-modelling'],
  '2015/IMC/leveraging-internet-background-radiation-for-opportunistic-network-analysis': ['off:incidental'],
  '2016/IMC/optical-layer-failures-in-a-large-backbone': ['outage'],
  '2016/IMC/reasons-dynamic-addresses-change': ['off:incidental'],
  '2017/IMC/pinpointing-delay-and-forwarding-anomalies-using-large-scale-traceroute-measurem': ['outage'],
  '2018/IMC/advancing-the-art-of-internet-edge-outage-detection': ['outage'],
  '2019/IMC/challenges-in-the-decentralised-web-the-mastodon-case': ['off:incidental'],
  '2020/IMC/five-alarms-assessing-the-vulnerability-of-us-cellular-communication-infrastruct': ['outage'],
  '2020/WWW/outage-detecting-power-and-communication-outages-from-social-networks': ['outage'],
  '2022/IMC/are-we-ready-for-metaverse-a-measurement-study-of-social-virtual-reality-platfor': ['off:incidental'],
  '2022/IMC/internet-outage-detection-using-passive-analysis': ['outage'],
  '2022/IMC/is-my-internet-down-sifting-through-user-affecting-outages-with-google-trends': ['outage'],
  '2023/NDSS/brokenwire-wireless-disruption-of-ccs-electric-vehicle-charging': ['off:incidental'],
  '2023/USENIX/access-denied-assessing-physical-risks-to-internet-access-networks': ['outage'],
  '2024/IMC/poster-enhancing-internet-disruption-investigation-via-path-monitoring': ['outage'],
  '2024/IMC/poster-investigating-network-security-post-outage-open-ports-vulnerabilities': ['outage'],
  '2024/CCS/poster-flashguard-real-time-disruption-of-non-price-flash-loan-attacks-in-defi': ['off:incidental'],
  '2024/IEEE-SP/surveilling-the-masses-with-wi-fi-based-positioning-systems': ['off:incidental'],
  '2024/USENIX/fledging-will-continue-until-privacy-improves-empirical-analysis-of-googles-priv': ['off:incidental'],
  '2025/IMC/poster-power-out-internet-down-measuring-internet-resilience-in-south-africa': ['outage'],
  '2025/IMC/tracking-internet-disruptions-in-ukraine-insights-from-three-years-of-active-ful': ['outage'],
  '2025/NDSS/rethink-reveal-the-threat-of-electromagnetic-interference-on-power-inverters': ['off:incidental'],
  '2025/IMC/hello-genai-dissecting-human-to-generative-ai-calling': ['off:incidental'],
  // --- added by the seventh (recall) probe, 2026-09-10 ---
  '2011/IMC/pingin-in-the-rain': ['off:incidental'],
  '2012/NDSS/you-can-run-but-you-can-t-hide-exposing-network-location-for-targeted-dos-attack': ['off:incidental'],
  '2013/CCS/beheading-hydras-performing-effective-botnet-takedowns': ['off:takedown-effects'],
  '2015/PETS/blocking-resistant-communication-through-domain-fronting': ['circumvention'],
  '2016/USENIX/hey-you-have-a-problem-on-the-feasibility-of-large-scale-web-vulnerability-notif': ['off:incidental'],
  '2017/CCS/faulds-a-non-parametric-iterative-classifier-for-internet-wide-os-fingerprinting': ['off:incidental'],
  '2017/IMC/the-record-route-option-is-an-option': ['off:incidental'],
  '2019/IMC/an-end-to-end-large-scale-measurement-of-dns-over-encryption-how-far-have-we-com': ['net'],
  '2019/IMC/ddos-hide-seek-on-the-effectiveness-of-a-booter-services-takedown': ['off:takedown-effects'],
  '2020/IMC/cloud-provider-connectivity-in-the-flat-internet': ['off:incidental'],
  '2020/NDSS/precisely-characterizing-security-impact-in-a-flood-of-patches-via-symbolic-rule-comparison': ['off:incidental'],
  '2020/NDSS/withdrawing-the-bgp-re-routing-curtain-understanding-the-security-impact-of-bgp-poisoning-through-real-world-measurements': ['off:incidental'],
  '2020/USENIX/fuzzguard-filtering-out-unreachable-inputs-in-directed-grey-box-fuzzing-through': ['off:incidental'],
  '2020/WWW/what-apps-did-you-use-understanding-the-long-term-evolution-of-mobile-app-usage': ['off:incidental'],
  '2021/CCS/ghost-in-the-binder-binder-transaction-redirection-attacks-in-android-system-ser': ['off:incidental'],
  '2021/CCS/revisiting-nakamoto-consensus-in-asynchronous-networks-a-comprehensive-analysis': ['off:incidental'],
  '2021/IMC/measuring-dns-over-https-performance-around-the-world': ['off:incidental'],
  '2021/USENIX/a-large-scale-study-of-user-behavior-expectations-and-engagement-with-android-pe': ['off:incidental'],
  '2021/WWW/surrounded-by-the-clouds-a-comprehensive-cloud-reachability-study': ['off:incidental'],
  '2022/IMC/performance-characterization-of-videoconferencing-in-the-wild': ['off:incidental'],
  '2023/IMC/on-the-importance-of-being-an-as-an-approach-to-country-level-as-rankings': ['off:incidental'],
  '2023/USENIX/collide-power-leaking-inaccessible-data-with-software-based-power-side-channels': ['off:incidental'],
  '2023/USENIX/dscope-a-cloud-native-internet-telescope': ['off:incidental'],
  '2023/WWW/automated-content-moderation-increases-adherence-to-community-guidelines': ['off:moderation-effects'],
  '2024/IEEE-SP/a-representative-study-on-human-detection-of-artificially-generated-media-across': ['off:incidental'],
  '2024/IEEE-SP/no-easy-way-out-the-effectiveness-of-deplatforming-an-extremist-forum-to-suppres': ['off:takedown-effects'],
  '2024/IMC/destination-reachable-what-icmpv6-error-messages-reveal-about-their-sources': ['off:incidental'],
  '2024/IMC/whats-in-the-dataset-unboxing-the-apnic-per-as-user-population-dataset': ['off:incidental'],
  '2024/USENIX/guardians-of-the-galaxy-content-moderation-in-the-interplanetary-file-system': ['platform'],
  '2024/USENIX/dvsorder-ballot-randomization-flaws-threaten-voter-privacy': ['off:incidental'],
  '2024/USENIX/sok-or-solk-on-the-quantitative-study-of-sociodemographic-factors-and-computer-s': ['off:incidental'],
  '2025/CCS/poster-eris-evaluating-rov-via-icmpv6-rate-limiting-side-channels': ['off:incidental'],
  '2025/USENIX/i-cannot-write-this-because-it-violates-our-content-policy-understanding-content': ['off:llm-guardrails'],
  '2025/USENIX/assessing-the-aftermath-the-effects-of-a-global-takedown-against-ddos-for-hire-s': ['off:takedown-effects'],
  '2025/WWW/causal-insights-into-parlers-content-moderation-shift-effects-on-toxicity-and-fa': ['off:moderation-effects'],
  '2025/WWW/unveiling-network-performance-in-the-wild-an-ad-driven-analysis-of-mobile-downlo': ['off:incidental'],
  '2025/NDSS/a-multifaceted-study-on-the-use-of-tls-and-auto-detect-in-email-ecosystems': ['off:incidental'],
  '2025/USENIX/darkgram-a-large-scale-analysis-of-cybercriminal-activity-channels-on-telegram': ['off:takedown-effects'],
  '2026/NDSS/actively-understanding-the-dynamics-and-risks-of-the-threat-intelligence-ecosystem': ['off:incidental'],
  '2026/NDSS/vulsca-a-community-level-sca-approach-for-accurate-c-c-supply-chain-vulnerability-analysis': ['off:incidental'],
  '2026/NDSS/from-noise-to-signal-precisely-identify-affected-packages-of-known-vulnerabilities-in-npm-ecosystem': ['off:incidental'],
  '2026/NDSS/crack-in-the-armor-underlying-infrastructure-threats-to-rpki-publication-point-reachability': ['off:incidental'],
  '2026/NDSS/enhancing-legal-document-security-and-accessibility-with-taf': ['off:incidental'],
  '2026/PETS/understanding-how-university-guidelines-address-privacy-and-security-issues-of-g': ['off:incidental'],
  '2026/PETS/the-empire-strikes-back-at-your-privacy-an-archaeology-of-tracking-on-government': ['off:incidental'],
  '2026/WWW/visual-content-moderation-in-messaging-systems': ['off:moderation-system'],
};
 
// Both directions, or the map rots silently the next time the probe moves.
const noVerdict = [...candidates.keys()].filter((k) => !(k in VERDICT));
const notCandidate = Object.keys(VERDICT).filter((k) => !candidates.has(k));
if (noVerdict.length || notCandidate.length) {
  throw new Error(
    `hand-verdict map is out of step with the probes:\n` +
      `  ${noVerdict.length} candidate(s) with no verdict:\n${noVerdict.map((k) => '    ' + k).join('\n')}\n` +
      `  ${notCandidate.length} verdict(s) for a paper no probe caught:\n${notCandidate.map((k) => '    ' + k).join('\n')}`,
  );
}
for (const [k, v] of Object.entries(VERDICT)) {
  for (const fam of v) {
    if (!MEASURE_FAMILIES.includes(fam) && !ADJACENT_FAMILIES.includes(fam) && !(fam in OFF_LABEL)) {
      throw new Error(`${k}: family "${fam}" has no label`);
    }
  }
}
 
const inFamily = (fam) => Object.entries(VERDICT).filter(([, v]) => v.includes(fam)).map(([k]) => byKey.get(k));
const BLOCK = [...new Set(MEASURE_FAMILIES.flatMap(inFamily))].sort((a, b) => a.year - b.year || a.venue.localeCompare(b.venue));
 
 
// ---------------------------------------------------------------- censor fold
// Who is doing the blocking, one label per paper in the population, assigned by
// hand from the paper. A review pass caught the page hand-tallying this table in
// prose and getting China wrong by a factor of two, so the split is produced
// here and the page copies it. Every population paper must have a label.
const CENSOR = {
  '2011/IMC/analysis-of-country-wide-internet-outages-caused-by-censorship': 'Egypt and Libya',
  '2013/IMC/a-method-for-identifying-and-confirming-the-use-of-url-filtering-products-for-ce': 'many countries at once',
  '2014/IMC/a-look-at-the-consequences-of-internet-censorship-through-an-isp-lens': 'Pakistan',
  '2014/IMC/automated-detection-and-fingerprinting-of-censorship-block-pages': 'many countries at once',
  '2014/IMC/censorship-in-the-wild-analyzing-internet-filtering-in-syria': 'Syria',
  '2015/IMC/examining-how-the-great-firewall-discovers-hidden-circumvention-servers': 'China',
  '2015/IMC/going-wild-large-scale-classification-of-open-dns-resolvers': 'many countries at once',
  '2015/IMC/in-and-out-of-cuba-characterizing-cubas-connectivity': 'Cuba',
  '2015/PETS/analyzing-the-great-firewall-of-china-over-space-and-time': 'China',
  '2016/IMC/tunneling-for-transparency-a-large-scale-analysis-of-end-to-end-violations-in-th': 'many countries at once',
  '2017/IEEE-SP/augur-internet-wide-detection-of-connectivity-disruptions': 'many countries at once',
  '2017/PETS/topics-of-controversy-an-empirical-analysis-of-web-censorship-lists': 'many countries at once',
  '2017/USENIX/global-measurement-of-dns-manipulation': 'many countries at once',
  '2017/USENIX/characterizing-the-nature-and-dynamics-of-tor-exit-blocking': 'website operators (Tor exits)',
  '2018/IMC/where-the-light-gets-in-analyzing-web-censorship-mechanisms-in-india': 'India',
  '2018/IMC/403-forbidden-a-global-view-of-cdn-geoblocking': 'website operators and CDNs',
  '2018/USENIX/quack-scalable-remote-measurement-of-application-layer-censorship': 'many countries at once',
  '2019/CCS/you-shall-not-join-a-measurement-study-of-cryptocurrency-peer-to-peer-bootstrapp': 'many countries at once',
  '2019/IMC/an-end-to-end-large-scale-measurement-of-dns-over-encryption-how-far-have-we-com': 'many countries at once',
  '2020/CCS/censored-planet-an-internet-wide-longitudinal-censorship-observatory': 'many countries at once',
  '2020/IEEE-SP/iclab-a-global-longitudinal-internet-censorship-measurement-platform': 'many countries at once',
  '2020/IMC/how-china-detects-and-blocks-shadowsocks': 'China',
  '2020/IMC/investigating-large-scale-https-interception-in-kazakhstan': 'Kazakhstan',
  '2020/NDSS/decentralized-control-a-case-study-of-russia': 'Russia',
  '2020/NDSS/measuring-the-deployment-of-network-censorship-filters-at-global-scale': 'many countries at once',
  '2020/WWW/deconstructing-googles-web-light-service': 'website operators and CDNs',
  '2021/IMC/throttling-twitter-an-emerging-censorship-technique-in-russia': 'Russia',
  '2021/IMC/web-censorship-measurements-of-http-3-over-quic': 'many countries at once',
  '2021/USENIX/how-great-is-the-great-firewall-measuring-chinas-dns-censorship': 'China',
  '2021/WWW/chinese-wall-or-swiss-cheese-keyword-filtering-in-the-great-firewall-of-china': 'China',
  '2021/WWW/understanding-the-impact-of-encrypted-dns-on-internet-censorship': 'many countries at once',
  '2022/IMC/rusty-clusters-dusting-an-ipv6-research-foundation': 'China',
  '2022/IMC/tspu-russias-decentralized-censorship-system': 'Russia',
  '2022/PETS/setting-the-bar-low-are-websites-complying-with-the-minimum-requirements-of-the': 'website operators (law-driven)',
  '2022/USENIX/many-roads-lead-to-rome-how-packet-headers-influence-dns-censorship-measurement': 'many countries at once',
  '2022/USENIX/a-large-scale-investigation-into-geodifferences-in-mobile-apps': 'app stores and developers',
  '2022/USENIX/the-security-lottery-measuring-client-side-web-security-inconsistencies': 'website operators and CDNs',
  '2023/PETS/certainty-detecting-dns-manipulation-at-scale-using-tls-certificates': 'many countries at once',
  '2023/USENIX/how-the-great-firewall-of-china-detects-and-blocks-fully-encrypted-traffic': 'China',
  '2023/USENIX/network-responses-to-russias-invasion-of-ukraine-in-2022-a-cautionary-tale-for-i': 'Russia',
  '2023/WWW/measuring-and-evading-turkmenistans-internet-censorship-a-case-study-in-large-sc': 'Turkmenistan',
  '2024/CCS/understanding-routing-induced-censorship-changes-globally': 'many countries at once',
  '2024/NDSS/modeling-and-detecting-internet-censorship-events': 'many countries at once',
  '2024/NDSS/content-censorship-in-the-interplanetary-file-system': 'a platform (IPFS, LLM services, a retailer)',
  '2024/NDSS/understanding-the-implementation-and-security-implications-of-protective-dns-services': 'resolver operators',
  '2024/PETS/what-to-expect-when-you-re-accessing-an-exploration-of-user-privacy-rights-in-pe': 'website operators (law-driven)',
  '2024/PETS/automatic-generation-of-web-censorship-probe-lists': 'China',
  '2024/USENIX/gfweb-measuring-the-great-firewalls-web-censorship-at-scale': 'China',
  '2024/USENIX/digital-discrimination-of-users-in-sanctioned-states-the-case-of-the-cuba-embarg': 'Cuba',
  '2024/USENIX/guardians-of-the-galaxy-content-moderation-in-the-interplanetary-file-system': 'a platform (IPFS, LLM services, a retailer)',
  '2024/WWW/a-worldwide-view-on-the-reachability-of-encrypted-dns-services': 'many countries at once',
  '2025/CCS/fingerprinting-deep-packet-inspection-devices-by-their-ambiguities': 'many countries at once',
  '2025/IEEE-SP/is-nobody-there-good-globally-measuring-connection-tampering-without-responsive': 'many countries at once',
  '2025/IEEE-SP/a-wall-behind-a-wall-emerging-regional-censorship-in-china': 'China',
  '2025/IEEE-SP/code-speaks-louder-exploring-security-and-privacy-relevant-regional-variations-i': 'app stores and developers',
  '2025/IMC/where-in-the-world-are-my-trackers-mapping-web-tracking-flow-across-diverse-geog': 'website operators and CDNs',
  '2025/NDSS/wallbleed-a-memory-disclosure-vulnerability-in-the-great-firewall-of-china': 'China',
  '2025/USENIX/irblock-a-large-scale-measurement-study-of-the-great-firewall-of-iran': 'Iran',
  '2025/USENIX/exposing-and-circumventing-sni-based-quic-censorship-of-the-great-firewall-of-ch': 'China',
  '2026/NDSS/there-is-no-war-in-ba-sing-se-a-global-analysis-of-content-moderation-in-large-language-models': 'a platform (IPFS, LLM services, a retailer)',
  '2026/NDSS/characterizing-the-implementation-of-censorship-policies-in-chinese-llm-services': 'China',
  '2026/PETS/banned-books-analysis-of-censorship-on-amazon-com': 'a platform (IPFS, LLM services, a retailer)',
  '2026/WWW/tracking-the-stray-sheep-understanding-dns-response-manipulation-in-the-wild': 'many countries at once',
};
 
// ------------------------------------------------------------------ printing
const H = (s) => console.log(`\n\n===== ${s} =====`);
const table = (head, rowsIn) => {
  console.log(head.join('\t'));
  for (const r of rowsIn) console.log(r.join('\t'));
};
 
const listFam = process.argv.includes('--list') ? process.argv[process.argv.indexOf('--list') + 1] : null;
if (listFam) {
  const ps = listFam === 'BLOCK' ? BLOCK : inFamily(listFam);
  for (const p of ps.sort((a, b) => a.year - b.year)) console.log(`${KEY(p)}\t${p.title}`);
  process.exit(0);
}
 
console.log(`report_blocking_geodifference.mjs — run ${new Date().toISOString().slice(0, 10)}`);
console.log(`corpus: ${rows.length} extracted papers, ${new Set(rows.map((p) => p.venue)).size} venues, ${Math.min(...rows.map((p) => p.year))}-${Math.max(...rows.map((p) => p.year))}`);
console.log(`paper.cols.txt missing for ${ftMissing.length} paper(s): ${ftMissing.map(KEY).join(', ') || 'none'}`);
console.log(`corpus reference denominators: crawled ${rows.filter((p) => p.crawlConfig !== null || p.studyTypes.includes('automated-web-crawl')).length}, web platform ${rows.filter((p) => p.platforms.includes('web')).length}, measuredFrom ${rows.filter((p) => p.vantage.length > 0).length}, scanned ${rows.filter((p) => p.studyTypes.includes('network-scan-or-probe')).length}`);
 
H('0. The inclusion rule, written down before the counts');
console.log(`A paper is IN the population if it measures whether some client could reach
some content or service AND attributes the failures to a deliberate blocking
decision taken by someone other than the client — a state, an ISP, a resolver
operator, a CDN, a platform, or the content owner — where that decision is keyed
on WHO OR WHERE the client is, on the content's acceptability to an authority or
platform, or on the traffic looking like an attempt to evade such a decision.
 
Explicitly OUT, each with its own off-topic family below:
  * blocking keyed on the content being malicious (malware, phishing, spam);
  * blocking keyed on the client looking automated (programming:crawler_detection);
  * blocking the client chose (its own ad blocker or filter list);
  * a system PROPOSED to evade blocking, and a DETECTOR proposed for finding
    circumvention traffic (both counted as the ADJACENT family "circumvention");
  * the effect of a takedown or of moderating user posts on later behaviour.
 
The third key was added after a review pass: this literature also measures
blocking keyed on THE TRAFFIC LOOKING LIKE AN ATTEMPT TO EVADE a blocking
decision — Shadowsocks, fully encrypted flows, an SNI. Measuring how a DEPLOYED
censor detects and blocks such traffic is in the population; proposing a
detector for it is not. The two sides are one sentence apart and the first
version of this rule did not separate them, so two papers sat in the population
under a rule that appeared to exclude them.
 
One deliberate boundary case is IN: the protective-DNS study, whose motive is
security rather than geography, because its claim has exactly this page's shape
(control resolver vs treatment resolver, and a definition of "blocked" that has
to survive an NXDOMAIN and a sinkhole). It is the whole "filter" family, so its
effect on any figure can be checked by subtracting one.`);
 
H('1. Candidate probes (floors, not populations)');
table(['probe', 'papers', 'of corpus'], Object.keys(PROBES).map((n) => {
  const c = [...candidates.values()].filter((v) => v.includes(n)).length;
  return [n, c, pct(c, rows.length)];
}));
console.log(`union of all ${Object.keys(PROBES).length} probes: ${candidates.size} candidates (${pct(candidates.size, rows.length)} of corpus)`);
const bySize = new Map();
for (const v of candidates.values()) bySize.set(v.length, (bySize.get(v.length) ?? 0) + 1);
console.log('candidates by how many probes caught them: ' + [...bySize.entries()].sort().map(([k, n]) => `${k} probe(s): ${n}`).join(', '));
 
H(`2. Hand verdicts over all ${candidates.size} candidates`);
const famCounts = [...MEASURE_FAMILIES, ...ADJACENT_FAMILIES].map((f) => [f, inFamily(f).length, FAMILY_LABEL[f]]);
table(['family', 'papers', 'what it is'], famCounts);
console.log(`\nPOPULATION (union of the ${MEASURE_FAMILIES.length} measurement families): ${BLOCK.length} papers`);
const doubleFamily = Object.entries(VERDICT).filter(([, v]) => v.filter((f) => MEASURE_FAMILIES.includes(f)).length > 1);
console.log(`  sum of the measurement family counts: ${MEASURE_FAMILIES.reduce((a, f) => a + inFamily(f).length, 0)} — higher than ${BLOCK.length} because ${doubleFamily.length} paper(s) carry two measurement families: ${doubleFamily.map(([k, v]) => `${k} [${v.join('+')}]`).join('; ')}`);
const offCounts = Object.keys(OFF_LABEL).map((f) => [f, inFamily(f).length, OFF_LABEL[f]]).filter((r) => r[1] > 0);
console.log('');
table(['off-topic family', 'papers', 'why it is off topic'], offCounts);
const accounted = new Set(Object.keys(VERDICT));
console.log(`\nverdicts: ${accounted.size}; candidates: ${candidates.size}; unaccounted: ${candidates.size - accounted.size}`);
console.log(`off topic in total: ${Object.keys(OFF_LABEL).flatMap(inFamily).length} of ${candidates.size} candidates (${pct(Object.keys(OFF_LABEL).flatMap(inFamily).length, candidates.size)}) — the price of a wide probe`);
 
H('3. The population over time and across venues');
const years = [...new Set(rows.map((p) => p.year))].sort();
table(['year', 'papers in population', 'corpus papers', 'share of corpus year'],
  years.map((y) => [y + (y >= 2025 ? ' (provisional)' : ''), BLOCK.filter((p) => p.year === y).length, rows.filter((p) => p.year === y).length, pct(BLOCK.filter((p) => p.year === y).length, rows.filter((p) => p.year === y).length)]));
const windows = [[2010, 2014], [2015, 2019], [2020, 2024], [2025, 2026]];
// Every MEASURE family gets a column, computed from the list rather than typed
// out, so a family added later cannot silently vanish from this table — which is
// what happened to `filter` on the first draft of the page.
const WINDOW_COLS = [...MEASURE_FAMILIES, 'circumvention'];
table(['window', 'population', ...WINDOW_COLS],
  windows.map(([a, b]) => [`${a}-${b}${a === 2025 ? ' (provisional)' : ''}`,
    BLOCK.filter((p) => p.year >= a && p.year <= b).length,
    ...WINDOW_COLS.map((f) => inFamily(f).filter((p) => p.year >= a && p.year <= b).length)]));
table(['venue', 'population', 'corpus', 'share of venue'],
  [...new Set(rows.map((p) => p.venue))].sort().map((v) => [v, BLOCK.filter((p) => p.venue === v).length, rows.filter((p) => p.venue === v).length, pct(BLOCK.filter((p) => p.venue === v).length, rows.filter((p) => p.venue === v).length)]));
table(['platform (multi-valued)', 'papers in population', 'share of ' + BLOCK.length],
  ['web', 'mobile', 'other-online-service', 'iot', 'offline', 'not-applicable'].map((pl) => [pl, BLOCK.filter((p) => p.platforms.includes(pl)).length, pct(BLOCK.filter((p) => p.platforms.includes(pl)).length, BLOCK.length)]));
// server and platform overlap by two papers, so the page's "server-and-platform"
// claim needs the DISTINCT union, not the sum of two family counts.
const srvPlat = [...new Set([...inFamily('server'), ...inFamily('platform')])];
console.log(`\nserver + platform, distinct papers: ${srvPlat.length} (server ${inFamily('server').length} + platform ${inFamily('platform').length} - ${inFamily('server').filter((p) => inFamily('platform').includes(p)).length} in both)`);
console.log(`  of those, 2025-2026: ${srvPlat.filter((p) => p.year >= 2025).length} — ${srvPlat.filter((p) => p.year >= 2025).map(KEY).join(', ')}`);
console.log(`  platform family in 2025-2026: ${inFamily('platform').filter((p) => p.year >= 2025).map(KEY).join(', ')}`);
console.log(`\nin the crawled population (crawlConfig or automated-web-crawl): ${BLOCK.filter((p) => p.crawlConfig !== null || p.studyTypes.includes('automated-web-crawl')).length} of ${BLOCK.length}`);
console.log(`ran a network scan or probe: ${BLOCK.filter((p) => p.studyTypes.includes('network-scan-or-probe')).length} of ${BLOCK.length}`);
console.log(`reanalyses an existing dataset: ${BLOCK.filter((p) => p.studyTypes.includes('existing-dataset-analysis')).length} of ${BLOCK.length}`);
 
H('3b. Who is doing the blocking (hand-keyed, one label per paper)');
const censorMissing = BLOCK.filter((p) => !(KEY(p) in CENSOR));
const censorExtra = Object.keys(CENSOR).filter((k) => !BLOCK.some((p) => KEY(p) === k));
if (censorMissing.length || censorExtra.length) {
  throw new Error(
    `CENSOR map out of step with the population: ${censorMissing.length} paper(s) with no label (${censorMissing.map(KEY).join(', ')}); ` +
      `${censorExtra.length} label(s) for a paper outside it (${censorExtra.join(', ')})`,
  );
}
const censorCounts = new Map();
for (const p of BLOCK) censorCounts.set(CENSOR[KEY(p)], [...(censorCounts.get(CENSOR[KEY(p)]) ?? []), p]);
table(['who is blocking', 'papers', `share of ${BLOCK.length}`, 'years'],
  [...censorCounts.entries()].sort((a, b) => b[1].length - a[1].length).map(([k, ps]) => [
    k, ps.length, pct(ps.length, BLOCK.length), `${Math.min(...ps.map((p) => p.year))}-${Math.max(...ps.map((p) => p.year))}`]));
console.log(`labels sum to ${[...censorCounts.values()].reduce((a, ps) => a + ps.length, 0)}, which must equal ${BLOCK.length}`);
for (const [k, ps] of [...censorCounts.entries()].sort()) {
  console.log(`\n-- ${k} (${ps.length})`);
  for (const p of ps.sort((a, b) => a.year - b.year)) console.log(`   ${KEY(p)}`);
}
 
H('4. Vantage points: does a blocking claim have a control?');
// Two vantage points is the structural minimum for a blocking claim: one inside
// the condition, one outside. `vantage.locations` is free text, so it is folded
// through geo.mjs and the residue is printed.
const residue = new Map();
function locs(p) {
  const out = new Set();
  let multi = false;
  for (const v of p.vantage) {
    for (const raw of v.locations ?? []) {
      if (isSentinel(raw)) continue;
      const n = normalizeLocation(raw);
      if (n.kind === 'country' || n.kind === 'region') out.add(n.value);
      else if (n.kind === 'multi') multi = true;
      else {
        // An unmapped string is still a location the paper named ("Guangzhou",
        // "Trinidad and Tobago" — geo.mjs has no alias for either). Counting it
        // as distinct keeps the "two or more vantage points" figure from being
        // an artefact of the alias list; the string is printed as residue below.
        residue.set(raw, (residue.get(raw) ?? 0) + 1);
        out.add(`unmapped:${String(raw).toLowerCase().trim()}`);
      }
    }
  }
  return { set: out, multi };
}
const withVantage = BLOCK.filter((p) => p.vantage.length > 0);
const noVantageTuple = BLOCK.filter((p) => p.vantage.length === 0);
const statedLoc = withVantage.filter((p) => { const l = locs(p); return l.set.size > 0 || l.multi; });
const twoPlus = withVantage.filter((p) => { const l = locs(p); return l.multi || l.set.size >= 2; });
console.log(`population: ${BLOCK.length}`);
console.log(`  has at least one vantage tuple: ${withVantage.length} (${pct(withVantage.length, BLOCK.length)})`);
console.log(`  no vantage tuple at all (pure reanalysis, or never said): ${noVantageTuple.length}`);
console.log(`  names at least one location, folded: ${statedLoc.length} of ${withVantage.length} (${pct(statedLoc.length, withVantage.length)})`);
console.log(`  two or more distinct locations, or an explicit multi-country vantage: ${twoPlus.length} of ${withVantage.length} (${pct(twoPlus.length, withVantage.length)})`);
console.log(`  exactly one location and no multi-country claim: ${statedLoc.length - twoPlus.length} of ${withVantage.length} (${pct(statedLoc.length - twoPlus.length, withVantage.length)})`);
const infra = new Map();
for (const p of BLOCK) for (const v of p.vantage) if (!isSentinel(v.infrastructure)) infra.set(v.infrastructure, (infra.get(v.infrastructure) ?? new Set()).add(KEY(p)));
table(['vantage infrastructure', 'papers', `share of ${withVantage.length}`], [...infra.entries()].sort((a, b) => b[1].size - a[1].size).map(([k, s]) => [k, s.size, pct(s.size, withVantage.length)]));
const infraSentinel = BLOCK.filter((p) => p.vantage.length > 0 && p.vantage.every((v) => isSentinel(v.infrastructure)));
console.log(`says nothing about infrastructure on any tuple: ${infraSentinel.length} of ${withVantage.length} (${pct(infraSentinel.length, withVantage.length)})`);
console.log(`\nlocation strings the fold could not map (residue, printed in full):`);
if (residue.size === 0) console.log('  (none)');
for (const [k, n] of [...residue.entries()].sort((a, b) => b[1] - a[1])) console.log(`  ${n}x  ${JSON.stringify(k)}`);
 
// A control vantage in the paper's own words. Tight requires "control" next to a
// vantage word; loose accepts either. Tight must be <= loose or the probe is broken.
const CONTROL_TIGHT = /control (vantage|VP|vantage point|machine|node|probe|server|host|site|resolver|network|location|measurement|group|country|client)s?|(vantage point|VP|machine|node|probe|server|host|client)s? (outside|located outside)[^.]{0,40}(censor|country|China|Iran|Russia|blocked)|uncensored (control|vantage|network|location)|non-?censor(ed|ing) (vantage|control|network|country|location)/i;
// Built from CONTROL_TIGHT for the same reason as WALL_LOOSE: the first version
// missed 6 of the 32 tight hits (the "uncensored control" and "vantage point
// outside <country>" branches contain no literal "control" near a vantage word),
// which made the page's "the gap between 32 and 41" sentence describe a
// difference between two DIFFERENT questions rather than two widths of one.
const CONTROL_LOOSE = new RegExp(
  CONTROL_TIGHT.source +
    '|\\bcontrol\\b[^.]{0,60}(vantage|VP|measurement|probe|machine|node|host|country|network)' +
    '|(vantage|VP|probe|machine|node|host)[^.]{0,60}\\bcontrol\\b',
  'i',
);
const ctrlTight = BLOCK.filter((p) => { const t = ftPop(p); return t !== null && CONTROL_TIGHT.test(t); });
const ctrlLoose = BLOCK.filter((p) => { const t = ftPop(p); return t !== null && CONTROL_LOOSE.test(t); });
const ctrlEscapes = ctrlTight.filter((p) => !ctrlLoose.includes(p));
if (ctrlEscapes.length) {
  throw new Error(
    `${ctrlEscapes.length} tight control-probe hit(s) escape the loose probe, so "tight" and "loose" are not two widths of one question: ` +
      ctrlEscapes.map(KEY).join(', '),
  );
}
console.log(`\nsays "control vantage" (or equivalent) in the full text, tight probe: ${ctrlTight.length} of ${BLOCK.length} (${pct(ctrlTight.length, BLOCK.length)})`);
console.log(`same, loose probe (any "control" near a vantage word): ${ctrlLoose.length} of ${BLOCK.length} (${pct(ctrlLoose.length, BLOCK.length)})`);
console.log('papers matching the tight control probe:');
for (const p of ctrlTight) console.log(`  ${KEY(p)}`);
 
H('5. What "blocked" is operationalised as');
// Signal families, folded from the paper's own text. These are what the paper
// checks to call something blocked. Multi-valued: most papers use several.
const SIGNAL = {
  'status code / HTTP error': /\b(HTTP )?(403|451|404|503)\b|status code|forbidden response|error (code|page)/i,
  'block page fingerprint or keyword': /block ?page|blocking page|blockpage (signature|fingerprint)|block-page/i,
  'page similarity / length outlier': /(page|response|body) (length|size)[^.]{0,40}(outlier|threshold|differ|distribution)|(cosine|jaccard|TF-?IDF|term frequency)[^.]{0,60}(similar|page|response)|similarity (score|threshold)[^.]{0,40}(page|response)/i,
  'DNS answer consistency': /(consistent|consistency|inconsisten\w+)[^.]{0,60}(DNS|resolution|response|answer)|(DNS|resolution)[^.]{0,60}(consistency|inconsistency) (metric|check|heuristic)|forged (DNS )?respons/i,
  'TLS certificate check': /(certificate|TLS)[^.]{0,60}(verif|validat|mismatch|chain)[^.]{0,60}(manipulat|block|censor|inject)|(manipulat|block|censor)[^.]{0,60}(certificate|TLS)[^.]{0,40}(verif|validat|mismatch)/i,
  'injected RST / connection reset': /\b(TCP )?RST\b|reset (packet|injection)|injected reset/i,
  'control-vantage comparison': CONTROL_TIGHT,
  'repeated probe / retry': /(repeat|retr(y|ies|ied)|re-?probe|multiple (probes|trials|measurements))[^.]{0,60}(confirm|verif|block|censor|persist)/i,
  'manual or browser validation': /manual(ly)? (verif|validat|check|inspect|review)[^.]{0,80}(block|censor|403|page)|(verif|validat)\w+[^.]{0,40}in a (real )?(web )?browser/i,
  'supervised classifier or clustering': /\b(classifier|random forest|logistic regression|SVM|neural network|clustering)\b[^.]{0,70}(block ?page|censor|blocked)/i,
  // \b-anchored: an unanchored BERT matches the surname Deibert, and the row read
  // 3 papers before that was fixed. See the widened probe below for the claim.
  'LLM or transformer model': /\b(LLMs?|large language model|GPT-?[0-9]|ChatGPT|Claude|Gemini|Llama|DeBERTa|BERTopic|BERT)\b[^.]{0,80}(block ?page|censor|blocked|classif)/i,
};
const sigPapers = {};
for (const [name, re] of Object.entries(SIGNAL)) {
  sigPapers[name] = BLOCK.filter((p) => { const t = ftPop(p); return t !== null && re.test(t); });
}
table(['signal the paper checks', `papers of ${BLOCK.length}`, 'share', 'first', 'last', '2010-19', '2020-24', '2025-26'],
  Object.entries(sigPapers).sort((a, b) => b[1].length - a[1].length).map(([n, ps]) => [
    n, ps.length, pct(ps.length, BLOCK.length),
    ps.length ? Math.min(...ps.map((p) => p.year)) : '-', ps.length ? Math.max(...ps.map((p) => p.year)) : '-',
    ps.filter((p) => p.year <= 2019).length, ps.filter((p) => p.year >= 2020 && p.year <= 2024).length, ps.filter((p) => p.year >= 2025).length]));
console.log('NOTE: these are full-text probes over the population, so each is an UPPER bound on');
console.log('papers that use the signal (a sentence in related work counts) and a LOWER bound on');
console.log('the idea being present (a paper can compare against a control without the word).');
const noSignal = BLOCK.filter((p) => !Object.values(sigPapers).some((ps) => ps.includes(p)));
console.log(`papers in the population matching none of the ${Object.keys(SIGNAL).length} signal probes: ${noSignal.length} — ${noSignal.map(KEY).join(', ') || 'none'}`);
const perPaper = BLOCK.map((p) => Object.values(sigPapers).filter((ps) => ps.includes(p)).length).sort((a, b) => a - b);
console.log(`signals matched per paper: min ${perPaper[0]}, median ${perPaper[Math.floor(perPaper.length / 2)]}, max ${perPaper.at(-1)}; ${perPaper.filter((n) => n >= 3).length} of ${BLOCK.length} match three or more`);
 
// The LLM row above is 3 papers, and reading them shows all three are about LLM
// SERVICES rather than an LLM used as the classifier. The widened probe (defined
// with the corpus regexes) tests that claim against the whole corpus.
const llmWide = rows.filter((p) => S(p) !== null && S(p).llmWide);
console.log(`\nwidened LLM/transformer-near-blocking probe, corpus-wide: ${llmWide.length} of ${rows.length}; in the population: ${llmWide.filter((p) => BLOCK.includes(p)).length} of ${BLOCK.length}`);
for (const p of llmWide.sort((a, b) => a.year - b.year)) console.log(`  ${p.year} ${p.venue} ${BLOCK.includes(p) ? '[in population]' : '[outside]'} ${p.title.slice(0, 85)}`);
 
H('6. detection[] tuples: the measured results, with their own denominators');
const detTuples = [];
for (const p of BLOCK) {
  for (const t of p.detection) {
    if (isSentinel(t.prevalence) || !t.prevalence) continue;
    detTuples.push({ p, t });
  }
}
console.log(`${detTuples.length} detection tuples with a stated prevalence, across ${new Set(detTuples.map((d) => KEY(d.p))).size} of ${BLOCK.length} papers in the population`);
const NUM = /\d/;
console.log(`of those, ${detTuples.filter((d) => NUM.test(d.t.prevalence)).length} carry a digit (i.e. could be a publishable figure)`);
console.log('\nevery prevalence tuple for the papers the page quotes:');
const QUOTED = [
  '2018/IMC/403-forbidden-a-global-view-of-cdn-geoblocking',
  '2014/IMC/automated-detection-and-fingerprinting-of-censorship-block-pages',
  '2024/USENIX/digital-discrimination-of-users-in-sanctioned-states-the-case-of-the-cuba-embarg',
  '2023/PETS/certainty-detecting-dns-manipulation-at-scale-using-tls-certificates',
  '2024/USENIX/gfweb-measuring-the-great-firewalls-web-censorship-at-scale',
  '2020/IEEE-SP/iclab-a-global-longitudinal-internet-censorship-measurement-platform',
  '2022/USENIX/a-large-scale-investigation-into-geodifferences-in-mobile-apps',
  '2024/PETS/automatic-generation-of-web-censorship-probe-lists',
  '2026/PETS/banned-books-analysis-of-censorship-on-amazon-com',
  '2024/WWW/a-worldwide-view-on-the-reachability-of-encrypted-dns-services',
  '2024/PETS/what-to-expect-when-you-re-accessing-an-exploration-of-user-privacy-rights-in-pe',
  '2022/USENIX/the-security-lottery-measuring-client-side-web-security-inconsistencies',
];
for (const k of QUOTED) {
  const p = byKey.get(k);
  if (!p) throw new Error(`QUOTED names a paper that is not in the corpus: ${k}`);
  if (!BLOCK.includes(p)) throw new Error(`QUOTED names a paper outside the population: ${k}`);
  console.log(`\n--- ${k}`);
  for (const t of p.detection) {
    if (!t.prevalence || isSentinel(t.prevalence)) continue;
    console.log(`    phenomenon: ${t.phenomenon}`);
    console.log(`    technique : ${t.technique}`);
    console.log(`    metric    : ${t.metric}`);
    console.log(`    prevalence: ${t.prevalence}`);
    console.log(`    quote [${t.evidence.section}]: ${t.evidence.quote}`);
  }
}
 
H('7. Instruments and observatories: who is used, and who is only cited');
// tools[] under-reports a dataset: the schema sees an instrument, not a corpus.
// So each name is counted twice, structurally and over the full text (one pass, above).
table(['instrument or dataset', 'named in tools[] as used/produced', 'papers whose full text names it (corpus-wide)', 'of which in the population'],
  Object.entries(INSTRUMENTS).map(([n, re]) => {
    const structural = rows.filter((p) => p.tools.some((t) => (t.usedOrMentioned === 'used' || t.usedOrMentioned === 'produced') && re.test(String(t.name ?? '')))).length;
    const fulltext = rows.filter((p) => S(p) !== null && S(p).instr.has(n));
    return [n, structural, fulltext.length, fulltext.filter((p) => BLOCK.includes(p)).length];
  }));
console.log('CAUTION: the full-text column is a mention count and several of these names are');
console.log('homographs — "Satellite" is satellite Internet, "Geneva" a city, "ONI" a substring,');
console.log('"Iris" an iOS analyser, "Encore" and "Augur" ordinary words. Use the column to show');
console.log('that tools[] under-counts a dataset, never as a usage figure.');
 
H('8. The jurisdiction slice: GDPR and CCPA walls');
// The two wall probes are defined with the other whole-corpus regexes above and
// evaluated in the same single pass.
const wt = rows.filter((p) => S(p) !== null && S(p).wallTight);
const wl = rows.filter((p) => S(p) !== null && S(p).wallLoose);
const wallEscapes = wt.filter((p) => !wl.includes(p));
if (wallEscapes.length) {
  throw new Error(
    `${wallEscapes.length} tight wall-probe hit(s) are not matched by the loose probe, so the two are different questions rather than two widths of one: ` +
      wallEscapes.map(KEY).join(', '),
  );
}
console.log(`corpus-wide, tight probe: ${wt.length} papers; loose probe: ${wl.length} papers (denominator ${rows.length})`);
console.log(`the loose probe contains the tight one by construction, so the distinct union is ${new Set([...wt, ...wl]).size} papers and every one of them was read`);
console.log('tight-probe hits, all of them, with the hand verdict:');
for (const p of wt.sort((a, b) => a.year - b.year)) console.log(`  ${KEY(p)}  [${(VERDICT[KEY(p)] ?? ['NOT A CANDIDATE']).join('+')}]`);
console.log('loose-probe hits that the tight probe missed:');
for (const p of wl.filter((p) => !wt.includes(p)).sort((a, b) => a.year - b.year)) console.log(`  ${KEY(p)}  [${(VERDICT[KEY(p)] ?? ['not a candidate']).join('+')}]  ${p.title.slice(0, 80)}`);
console.log(`\nlegal population for reference: ${rows.filter((p) => p.legal.length > 0).length} papers assess a law; of the population, ${BLOCK.filter((p) => p.legal.length > 0).length} do`);
const gdprLegal = rows.filter((p) => p.legal.some((l) => /GDPR/i.test(String(l.law ?? ''))));
console.log(`papers whose legal[] names GDPR: ${gdprLegal.length}; of those, in the population: ${gdprLegal.filter((p) => BLOCK.includes(p)).length}`);
 
H('9. Cross-checks against neighbouring pages');
console.log(`design:existing_datasets reports a "Censorship list (Citizen Lab, OONI)" family of 17 papers inside ITS OWN reuse population. That population is defined on that page and is not re-derived here — the two definitions below both differ from it, which is exactly why the page cites the family and not the denominator.`);
// That page's population is the UNION of the studyTypes flag and the temporal
// mode, not the studyTypes flag alone; using the wrong one prints 2,615 where the
// page says 3,389 and makes a shared figure look stale.
const reuse = rows.filter((p) => p.studyTypes.includes('existing-dataset-analysis') || p.temporal.some((t) => t.mode === 'existing-dataset'));
console.log(`  reuse population, that page's definition (studyTypes OR temporal.mode): ${reuse.length}`);
console.log(`  studyTypes flag alone: ${rows.filter((p) => p.studyTypes.includes('existing-dataset-analysis')).length}`);
console.log(`  of the population, in the reuse population: ${BLOCK.filter((p) => reuse.includes(p)).length}`);
console.log(`design:crawling_location's population is the ${rows.filter((p) => p.vantage.length > 0).length} papers with a vantage tuple; ${BLOCK.filter((p) => p.vantage.length > 0).length} of the population are in it.`);
console.log(`programming:crawler_detection owns "blocked because you look like a bot": ${inFamily('off:bot-blocking').length} candidate(s) were handed to it.`);
 
H('10. Probe recall, and the queued estimate');
// Two claims on the page and its provenance need these numbers: how much of the
// population a title probe alone would find, and what the roadmap's queued
// estimate re-derives to on the current corpus. The queued regex is reproduced
// here EXACTLY as the 2026-09-02 brainstorm pass ran it.
const titleHits = rows.filter(PROBES.title);
console.log(`probe 1 (title+summary) hits: ${titleHits.length}; of those in the population: ${titleHits.filter((p) => BLOCK.includes(p)).length}; excluded: ${titleHits.filter((p) => !BLOCK.includes(p)).length} (${pct(titleHits.filter((p) => !BLOCK.includes(p)).length, titleHits.length)} of the probe's hits)`);
console.log(`population papers NO title probe caught: ${BLOCK.filter((p) => !PROBES.title(p)).length} of ${BLOCK.length}`);
const QUEUED_RE = /censorship|censored|OONI|Censored Planet|GFW|geoblock|blockpage|internet shutdown/i;
const q = rows.filter((p) => QUEUED_RE.test(T(p)));
console.log(`the roadmap's queued probe, re-run on this corpus: ${q.length} papers (it reported 74 on 2026-09-02), web ${q.filter((p) => p.platforms.includes('web')).length}, 2020-2023 ${q.filter((p) => p.year >= 2020 && p.year <= 2023).length}, 2024-2026 ${q.filter((p) => p.year >= 2024).length}`);
console.log(`of those ${q.length}, in the population: ${q.filter((p) => BLOCK.includes(p)).length}; excluded: ${q.filter((p) => !BLOCK.includes(p)).length}`);
 
// A diagnostic the provenance page quotes: how unusable detection[].phenomenon is
// as a COUNT in this area, as opposed to as a way to find candidates.
const PHEN = /censor|geo-?block|geo-?differen|block ?page|internet (filter|shutdown)|dns (manipulat|censor|interference|redirection|tamper)|network interference|connection tamper|url[- ]filter|keyword[- ]?(based )?filter|content (filter|censor)|over-?block|site blocking|blocked (site|domain|web|content|url|categor)|throttling of|traffic throttl|sanction/i;
const phenStrings = new Set();
const phenPapers = new Set();
for (const p of rows) {
  for (const t of p.detection) {
    if (PHEN.test(String(t.phenomenon ?? ''))) { phenStrings.add(t.phenomenon); phenPapers.add(KEY(p)); }
  }
}
console.log(`\ndetection[].phenomenon, raw: ${phenStrings.size} distinct blocking-ish strings across ${phenPapers.size} papers — used to FIND candidates (probe 2) and never to count`);
 
H('11. Venue-years with no extracted papers at all');
// A review pass observed that Khattak et al. (NDSS 2016) is cited by this page
// and cannot be in the population. It is not a selection decision: NDSS 2016 is
// simply absent from the extraction, and so are several other venue-years. Any
// "the corpus does not contain X" claim has to be read against this list.
{
  const have = new Set(rows.map((p) => `${p.venue}/${p.year}`));
  const venues = [...new Set(rows.map((p) => p.venue))].sort();
  const gaps = [];
  for (const v of venues) for (let y = 2010; y <= 2026; y++) if (!have.has(`${v}/${y}`)) gaps.push(`${v} ${y}`);
  console.log(`venue-years with zero extracted papers: ${gaps.length} — ${gaps.join(', ')}`);
  const popYears = new Set(BLOCK.map((p) => `${p.venue}/${p.year}`));
  console.log(`(the population itself spans ${popYears.size} distinct venue-years)`);
}

report_blocking_geodifference-output.txt

Its unedited output.

report_blocking_geodifference-output.txt
report_blocking_geodifference.mjs — run 2026-09-10
corpus: 5859 extracted papers, 7 venues, 2010-2026
paper.cols.txt missing for 4 paper(s): 2010/USENIX/idle-port-scanning-and-non-interference-analysis-of-network-protocol-stacks-usin, 2014/CCS/beware-your-hands-reveal-your-secrets, 2020/IMC/bgp-beacons-network-tomography-and-bayesian-computation-to-locate-route-flap-dam, 2020/IEEE-SP/burglars-iot-paradise-understanding-and-mitigating-security-risks-of-general-mes
corpus reference denominators: crawled 1120, web platform 1622, measuredFrom 3908, scanned 930


===== 0. The inclusion rule, written down before the counts =====
A paper is IN the population if it measures whether some client could reach
some content or service AND attributes the failures to a deliberate blocking
decision taken by someone other than the client — a state, an ISP, a resolver
operator, a CDN, a platform, or the content owner — where that decision is keyed
on WHO OR WHERE the client is, on the content's acceptability to an authority or
platform, or on the traffic looking like an attempt to evade such a decision.

Explicitly OUT, each with its own off-topic family below:
  * blocking keyed on the content being malicious (malware, phishing, spam);
  * blocking keyed on the client looking automated (programming:crawler_detection);
  * blocking the client chose (its own ad blocker or filter list);
  * a system PROPOSED to evade blocking, and a DETECTOR proposed for finding
    circumvention traffic (both counted as the ADJACENT family "circumvention");
  * the effect of a takedown or of moderating user posts on later behaviour.

The third key was added after a review pass: this literature also measures
blocking keyed on THE TRAFFIC LOOKING LIKE AN ATTEMPT TO EVADE a blocking
decision — Shadowsocks, fully encrypted flows, an SNI. Measuring how a DEPLOYED
censor detects and blocks such traffic is in the population; proposing a
detector for it is not. The two sides are one sentence apart and the first
version of this rule did not separate them, so two papers sat in the population
under a rule that appeared to exclude them.

One deliberate boundary case is IN: the protective-DNS study, whose motive is
security rather than geography, because its claim has exactly this page's shape
(control resolver vs treatment resolver, and a definition of "blocked" that has
to survive an NXDOMAIN and a sinkhole). It is the whole "filter" family, so its
effect on any figure can be checked by subtracting one.


===== 1. Candidate probes (floors, not populations) =====
probe	papers	of corpus
title	83	1.4%
detect	71	1.2%
instr	35	0.6%
ftinstr	80	1.4%
ftvocab	103	1.8%
wall	9	0.2%
outage	30	0.5%
manual	1	0.0%
recall	60	1.0%
union of all 9 probes: 275 candidates (4.7% of corpus)
candidates by how many probes caught them: 1 probe(s): 191, 2 probe(s): 31, 3 probe(s): 17, 4 probe(s): 15, 5 probe(s): 18, 6 probe(s): 3


===== 2. Hand verdicts over all 275 candidates =====
family	papers	what it is
net	45	network-level interference with access (censorship, DNS manipulation, connection tampering, throttling, interception)
server	11	server- or platform-side refusal or differentiation by region (geoblocking, geodifference, sanctions)
jurisdiction	2	blocking or differentiation driven by a privacy law (GDPR wall, CCPA geofencing)
platform	5	moderation or censorship enforced inside one platform or service
filter	1	filtering products deployed on the client side (protective DNS), and their over-blocking
probelist	2	the censorship probe list itself as the object of study
circumvention	29	ADJACENT: evasion or circumvention system, or censor-side detection of it
outage	12	ADJACENT: measures unreachability without attributing it to a deliberate blocking decision — power, fibre, weather, war damage, misconfiguration
wall-encounter	3	ADJACENT: ran into, worked around, or reported others proposing an EU wall, without measuring how common it is
interview	5	ADJACENT: human-subjects study of living with censorship
sok	1	ADJACENT: SoK or survey of the censorship literature
vantage-instrument	2	ADJACENT: instrument for measuring FROM a place (see design:crawling_location)

POPULATION (union of the 6 measurement families): 63 papers
  sum of the measurement family counts: 66 — higher than 63 because 3 paper(s) carry two measurement families: 2023/USENIX/network-responses-to-russias-invasion-of-ukraine-in-2022-a-cautionary-tale-for-i [net+server]; 2026/PETS/banned-books-analysis-of-censorship-on-amazon-com [server+platform]; 2026/NDSS/there-is-no-war-in-ba-sing-se-a-global-analysis-of-content-moderation-in-large-language-models [platform+server]

off-topic family	papers	why it is off topic
off:incidental	115	one incidental mention, or a related-work citation only
off:homonym-tor-metrics	7	homonym: names Tor Metrics, a Tor network dataset, not a blocking observatory
off:homonym-iris	1	homonym: iRiS, an iOS private-API analyser, not Iris the DNS-manipulation platform
off:homonym-augur	1	homonym: names Augur in a browser-fingerprinting context, not Augur the disruption prober
off:homonym-throttling	1	homonym: bandwidth throttling as a defence, not throttling as censorship
off:homonym-statistical-censoring	2	homonym: "censored"/"censoring" in the statistics or ML sense
off:homonym-ml-concept-censorship	1	homonym: concept censorship in a generative model
off:blockchain-sanctions	3	sanctions and blocklists on a blockchain, not access to content
off:adblocking	2	content blocking by the user’s own ad blocker (see programming:filter_lists)
off:security-blocking	2	blocking malware or phishing, not blocking by who or where you are
off:vpn-proxy-ecosystem	8	the VPN or proxy ecosystem itself (see design:crawling_location)
off:proxy-detection	4	detecting proxy or VPN clients (the server side of vantage choice)
off:age-gate	1	age assurance, a different access gate (queued as privacy:age_assurance)
off:bot-blocking	1	blocked for looking like a bot (see programming:crawler_detection)
off:censor-as-attack-surface	1	censorship middleboxes abused as an amplifier, not measured as a censor
off:llm-guardrails	2	model guardrails and refusals of what the user asked, not access blocking
off:takedown-effects	5	the effect of a takedown or deplatforming on the target, not whether a client could reach content
off:moderation-effects	2	the effect of moderating user posts on later user behaviour
off:moderation-system	1	proposes or evaluates a moderation classifier — builds the blocker rather than measuring blocking
off:resilience-modelling	1	models what would happen if countries disconnected, from routing graphs; no client reachability measured

verdicts: 275; candidates: 275; unaccounted: 0
off topic in total: 161 of 275 candidates (58.5%) — the price of a wide probe


===== 3. The population over time and across venues =====
year	papers in population	corpus papers	share of corpus year
2010	0	119	0.0%
2011	1	116	0.9%
2012	0	151	0.0%
2013	1	125	0.8%
2014	3	166	1.8%
2015	4	190	2.1%
2016	1	182	0.5%
2017	4	231	1.7%
2018	3	254	1.2%
2019	2	402	0.5%
2020	7	404	1.7%
2021	5	379	1.3%
2022	6	546	1.1%
2023	4	719	0.6%
2024	10	690	1.4%
2025 (provisional)	8	770	1.0%
2026 (provisional)	4	415	1.0%
window	population	net	server	jurisdiction	platform	filter	probelist	circumvention
2010-2014	5	5	0	0	0	0	0	1
2015-2019	14	11	2	0	0	0	1	8
2020-2024	32	22	5	2	2	1	1	16
2025-2026 (provisional)	12	7	4	0	3	0	0	4
venue	population	corpus	share of venue
CCS	4	990	0.4%
IEEE-SP	5	767	0.7%
IMC	19	638	3.0%
NDSS	8	701	1.1%
PETS	7	510	1.4%
USENIX	14	1410	1.0%
WWW	6	843	0.7%
platform (multi-valued)	papers in population	share of 63
web	32	50.8%
mobile	2	3.2%
other-online-service	47	74.6%
iot	0	0.0%
offline	2	3.2%
not-applicable	0	0.0%

server + platform, distinct papers: 14 (server 11 + platform 5 - 2 in both)
  of those, 2025-2026: 5 — 2025/IMC/where-in-the-world-are-my-trackers-mapping-web-tracking-flow-across-diverse-geog, 2025/IEEE-SP/code-speaks-louder-exploring-security-and-privacy-relevant-regional-variations-i, 2026/PETS/banned-books-analysis-of-censorship-on-amazon-com, 2026/NDSS/there-is-no-war-in-ba-sing-se-a-global-analysis-of-content-moderation-in-large-language-models, 2026/NDSS/characterizing-the-implementation-of-censorship-policies-in-chinese-llm-services
  platform family in 2025-2026: 2026/NDSS/characterizing-the-implementation-of-censorship-policies-in-chinese-llm-services, 2026/PETS/banned-books-analysis-of-censorship-on-amazon-com, 2026/NDSS/there-is-no-war-in-ba-sing-se-a-global-analysis-of-content-moderation-in-large-language-models

in the crawled population (crawlConfig or automated-web-crawl): 12 of 63
ran a network scan or probe: 50 of 63
reanalyses an existing dataset: 34 of 63


===== 3b. Who is doing the blocking (hand-keyed, one label per paper) =====
who is blocking	papers	share of 63	years
many countries at once	23	36.5%	2013-2026
China	13	20.6%	2015-2026
website operators and CDNs	4	6.3%	2018-2025
Russia	4	6.3%	2020-2023
a platform (IPFS, LLM services, a retailer)	4	6.3%	2024-2026
Cuba	2	3.2%	2015-2024
website operators (law-driven)	2	3.2%	2022-2024
app stores and developers	2	3.2%	2022-2025
Egypt and Libya	1	1.6%	2011-2011
Pakistan	1	1.6%	2014-2014
Syria	1	1.6%	2014-2014
website operators (Tor exits)	1	1.6%	2017-2017
India	1	1.6%	2018-2018
Kazakhstan	1	1.6%	2020-2020
Turkmenistan	1	1.6%	2023-2023
resolver operators	1	1.6%	2024-2024
Iran	1	1.6%	2025-2025
labels sum to 63, which must equal 63

-- China (13)
   2015/IMC/examining-how-the-great-firewall-discovers-hidden-circumvention-servers
   2015/PETS/analyzing-the-great-firewall-of-china-over-space-and-time
   2020/IMC/how-china-detects-and-blocks-shadowsocks
   2021/USENIX/how-great-is-the-great-firewall-measuring-chinas-dns-censorship
   2021/WWW/chinese-wall-or-swiss-cheese-keyword-filtering-in-the-great-firewall-of-china
   2022/IMC/rusty-clusters-dusting-an-ipv6-research-foundation
   2023/USENIX/how-the-great-firewall-of-china-detects-and-blocks-fully-encrypted-traffic
   2024/PETS/automatic-generation-of-web-censorship-probe-lists
   2024/USENIX/gfweb-measuring-the-great-firewalls-web-censorship-at-scale
   2025/IEEE-SP/a-wall-behind-a-wall-emerging-regional-censorship-in-china
   2025/NDSS/wallbleed-a-memory-disclosure-vulnerability-in-the-great-firewall-of-china
   2025/USENIX/exposing-and-circumventing-sni-based-quic-censorship-of-the-great-firewall-of-ch
   2026/NDSS/characterizing-the-implementation-of-censorship-policies-in-chinese-llm-services

-- Cuba (2)
   2015/IMC/in-and-out-of-cuba-characterizing-cubas-connectivity
   2024/USENIX/digital-discrimination-of-users-in-sanctioned-states-the-case-of-the-cuba-embarg

-- Egypt and Libya (1)
   2011/IMC/analysis-of-country-wide-internet-outages-caused-by-censorship

-- India (1)
   2018/IMC/where-the-light-gets-in-analyzing-web-censorship-mechanisms-in-india

-- Iran (1)
   2025/USENIX/irblock-a-large-scale-measurement-study-of-the-great-firewall-of-iran

-- Kazakhstan (1)
   2020/IMC/investigating-large-scale-https-interception-in-kazakhstan

-- Pakistan (1)
   2014/IMC/a-look-at-the-consequences-of-internet-censorship-through-an-isp-lens

-- Russia (4)
   2020/NDSS/decentralized-control-a-case-study-of-russia
   2021/IMC/throttling-twitter-an-emerging-censorship-technique-in-russia
   2022/IMC/tspu-russias-decentralized-censorship-system
   2023/USENIX/network-responses-to-russias-invasion-of-ukraine-in-2022-a-cautionary-tale-for-i

-- Syria (1)
   2014/IMC/censorship-in-the-wild-analyzing-internet-filtering-in-syria

-- Turkmenistan (1)
   2023/WWW/measuring-and-evading-turkmenistans-internet-censorship-a-case-study-in-large-sc

-- a platform (IPFS, LLM services, a retailer) (4)
   2024/NDSS/content-censorship-in-the-interplanetary-file-system
   2024/USENIX/guardians-of-the-galaxy-content-moderation-in-the-interplanetary-file-system
   2026/NDSS/there-is-no-war-in-ba-sing-se-a-global-analysis-of-content-moderation-in-large-language-models
   2026/PETS/banned-books-analysis-of-censorship-on-amazon-com

-- app stores and developers (2)
   2022/USENIX/a-large-scale-investigation-into-geodifferences-in-mobile-apps
   2025/IEEE-SP/code-speaks-louder-exploring-security-and-privacy-relevant-regional-variations-i

-- many countries at once (23)
   2013/IMC/a-method-for-identifying-and-confirming-the-use-of-url-filtering-products-for-ce
   2014/IMC/automated-detection-and-fingerprinting-of-censorship-block-pages
   2015/IMC/going-wild-large-scale-classification-of-open-dns-resolvers
   2016/IMC/tunneling-for-transparency-a-large-scale-analysis-of-end-to-end-violations-in-th
   2017/IEEE-SP/augur-internet-wide-detection-of-connectivity-disruptions
   2017/PETS/topics-of-controversy-an-empirical-analysis-of-web-censorship-lists
   2017/USENIX/global-measurement-of-dns-manipulation
   2018/USENIX/quack-scalable-remote-measurement-of-application-layer-censorship
   2019/CCS/you-shall-not-join-a-measurement-study-of-cryptocurrency-peer-to-peer-bootstrapp
   2019/IMC/an-end-to-end-large-scale-measurement-of-dns-over-encryption-how-far-have-we-com
   2020/CCS/censored-planet-an-internet-wide-longitudinal-censorship-observatory
   2020/IEEE-SP/iclab-a-global-longitudinal-internet-censorship-measurement-platform
   2020/NDSS/measuring-the-deployment-of-network-censorship-filters-at-global-scale
   2021/IMC/web-censorship-measurements-of-http-3-over-quic
   2021/WWW/understanding-the-impact-of-encrypted-dns-on-internet-censorship
   2022/USENIX/many-roads-lead-to-rome-how-packet-headers-influence-dns-censorship-measurement
   2023/PETS/certainty-detecting-dns-manipulation-at-scale-using-tls-certificates
   2024/CCS/understanding-routing-induced-censorship-changes-globally
   2024/NDSS/modeling-and-detecting-internet-censorship-events
   2024/WWW/a-worldwide-view-on-the-reachability-of-encrypted-dns-services
   2025/CCS/fingerprinting-deep-packet-inspection-devices-by-their-ambiguities
   2025/IEEE-SP/is-nobody-there-good-globally-measuring-connection-tampering-without-responsive
   2026/WWW/tracking-the-stray-sheep-understanding-dns-response-manipulation-in-the-wild

-- resolver operators (1)
   2024/NDSS/understanding-the-implementation-and-security-implications-of-protective-dns-services

-- website operators (Tor exits) (1)
   2017/USENIX/characterizing-the-nature-and-dynamics-of-tor-exit-blocking

-- website operators (law-driven) (2)
   2022/PETS/setting-the-bar-low-are-websites-complying-with-the-minimum-requirements-of-the
   2024/PETS/what-to-expect-when-you-re-accessing-an-exploration-of-user-privacy-rights-in-pe

-- website operators and CDNs (4)
   2018/IMC/403-forbidden-a-global-view-of-cdn-geoblocking
   2020/WWW/deconstructing-googles-web-light-service
   2022/USENIX/the-security-lottery-measuring-client-side-web-security-inconsistencies
   2025/IMC/where-in-the-world-are-my-trackers-mapping-web-tracking-flow-across-diverse-geog


===== 4. Vantage points: does a blocking claim have a control? =====
population: 63
  has at least one vantage tuple: 62 (98.4%)
  no vantage tuple at all (pure reanalysis, or never said): 1
  names at least one location, folded: 53 of 62 (85.5%)
  two or more distinct locations, or an explicit multi-country vantage: 40 of 62 (64.5%)
  exactly one location and no multi-country claim: 13 of 62 (21.0%)
vantage infrastructure	papers	share of 62
cloud-provider	25	40.3%
university-network	24	38.7%
commercial-vpn	12	19.4%
volunteer-devices	8	12.9%
research-testbed	6	9.7%
residential	5	8.1%
tor	2	3.2%
proxy-service	1	1.6%
mobile-network	1	1.6%
says nothing about infrastructure on any tuple: 9 of 62 (14.5%)

location strings the fold could not map (residue, printed in full):
  4x  "Guangzhou"
  2x  "Dominican Republic"
  2x  "Jamaica"
  2x  "Puerto Rico"
  2x  "Saint Barthelemy"
  2x  "Saint Kitts and Nevis"
  2x  "Guadeloupe"
  2x  "Trinidad and Tobago"
  2x  "Martinique"
  2x  "Grenada"
  2x  "around the world"
  2x  "Taichung"
  2x  "Hangzhou"
  2x  "Longmont?"
  2x  "San Jose"

says "control vantage" (or equivalent) in the full text, tight probe: 32 of 63 (50.8%)
same, loose probe (any "control" near a vantage word): 47 of 63 (74.6%)
papers matching the tight control probe:
  2015/IMC/examining-how-the-great-firewall-discovers-hidden-circumvention-servers
  2015/PETS/analyzing-the-great-firewall-of-china-over-space-and-time
  2017/IEEE-SP/augur-internet-wide-detection-of-connectivity-disruptions
  2017/PETS/topics-of-controversy-an-empirical-analysis-of-web-censorship-lists
  2017/USENIX/global-measurement-of-dns-manipulation
  2017/USENIX/characterizing-the-nature-and-dynamics-of-tor-exit-blocking
  2018/IMC/where-the-light-gets-in-analyzing-web-censorship-mechanisms-in-india
  2018/IMC/403-forbidden-a-global-view-of-cdn-geoblocking
  2020/CCS/censored-planet-an-internet-wide-longitudinal-censorship-observatory
  2020/IEEE-SP/iclab-a-global-longitudinal-internet-censorship-measurement-platform
  2020/IMC/how-china-detects-and-blocks-shadowsocks
  2020/NDSS/decentralized-control-a-case-study-of-russia
  2020/NDSS/measuring-the-deployment-of-network-censorship-filters-at-global-scale
  2021/IMC/throttling-twitter-an-emerging-censorship-technique-in-russia
  2021/IMC/web-censorship-measurements-of-http-3-over-quic
  2021/USENIX/how-great-is-the-great-firewall-measuring-chinas-dns-censorship
  2021/WWW/chinese-wall-or-swiss-cheese-keyword-filtering-in-the-great-firewall-of-china
  2021/WWW/understanding-the-impact-of-encrypted-dns-on-internet-censorship
  2022/IMC/tspu-russias-decentralized-censorship-system
  2023/PETS/certainty-detecting-dns-manipulation-at-scale-using-tls-certificates
  2023/USENIX/how-the-great-firewall-of-china-detects-and-blocks-fully-encrypted-traffic
  2023/USENIX/network-responses-to-russias-invasion-of-ukraine-in-2022-a-cautionary-tale-for-i
  2024/CCS/understanding-routing-induced-censorship-changes-globally
  2024/NDSS/understanding-the-implementation-and-security-implications-of-protective-dns-services
  2024/USENIX/gfweb-measuring-the-great-firewalls-web-censorship-at-scale
  2024/USENIX/digital-discrimination-of-users-in-sanctioned-states-the-case-of-the-cuba-embarg
  2024/WWW/a-worldwide-view-on-the-reachability-of-encrypted-dns-services
  2025/CCS/fingerprinting-deep-packet-inspection-devices-by-their-ambiguities
  2025/IEEE-SP/a-wall-behind-a-wall-emerging-regional-censorship-in-china
  2025/NDSS/wallbleed-a-memory-disclosure-vulnerability-in-the-great-firewall-of-china
  2025/USENIX/irblock-a-large-scale-measurement-study-of-the-great-firewall-of-iran
  2025/USENIX/exposing-and-circumventing-sni-based-quic-censorship-of-the-great-firewall-of-ch


===== 5. What "blocked" is operationalised as =====
signal the paper checks	papers of 63	share	first	last	2010-19	2020-24	2025-26
status code / HTTP error	41	65.1%	2014	2026	13	22	6
injected RST / connection reset	35	55.6%	2011	2026	9	20	6
block page fingerprint or keyword	33	52.4%	2013	2026	10	18	5
control-vantage comparison	32	50.8%	2015	2025	8	19	5
DNS answer consistency	22	34.9%	2014	2026	4	14	4
page similarity / length outlier	14	22.2%	2014	2026	5	7	2
manual or browser validation	13	20.6%	2015	2024	7	6	0
repeated probe / retry	11	17.5%	2013	2025	4	5	2
supervised classifier or clustering	11	17.5%	2014	2026	2	7	2
TLS certificate check	3	4.8%	2023	2024	0	3	0
LLM or transformer model	3	4.8%	2024	2026	0	1	2
NOTE: these are full-text probes over the population, so each is an UPPER bound on
papers that use the signal (a sentence in related work counts) and a LOWER bound on
the idea being present (a paper can compare against a control without the word).
papers in the population matching none of the 11 signal probes: 6 — 2015/IMC/in-and-out-of-cuba-characterizing-cubas-connectivity, 2019/CCS/you-shall-not-join-a-measurement-study-of-cryptocurrency-peer-to-peer-bootstrapp, 2022/PETS/setting-the-bar-low-are-websites-complying-with-the-minimum-requirements-of-the, 2024/NDSS/content-censorship-in-the-interplanetary-file-system, 2025/IEEE-SP/code-speaks-louder-exploring-security-and-privacy-relevant-regional-variations-i, 2025/IMC/where-in-the-world-are-my-trackers-mapping-web-tracking-flow-across-diverse-geog
signals matched per paper: min 0, median 4, max 8; 45 of 63 match three or more

widened LLM/transformer-near-blocking probe, corpus-wide: 11 of 5859; in the population: 3 of 63
  2024 PETS [in population] Automatic Generation of Web Censorship Probe Lists
  2024 USENIX [outside] Malla: Demystifying Real-world Large Language Model Integrated Malicious Services
  2025 NDSS [outside] THEMIS: Regulating Textual Inversion for Personalized Concept Censorship
  2025 USENIX [outside] "I Cannot Write This Because It Violates Our Content Policy": Understanding Content M
  2025 USENIX [outside] Exposing the Guardrails: Reverse-Engineering and Jailbreaking Safety Filters in DALL·
  2025 IMC [outside] Do Spammers Dream of Electric Sheep? Characterizing the Prevalence of LLM-Generated M
  2025 USENIX [outside] Malicious LLM-Based Conversational AI Makes Users Reveal Personal Information
  2026 NDSS [in population] Characterizing the Implementation of Censorship Policies in Chinese LLM Services
  2026 NDSS [outside] Beyond Jailbreak: Unveiling Risks in LLM Applications Arising from Blurred Capability
  2026 USENIX [outside] When Memory Becomes a Vulnerability: Towards Multi-turn Jailbreak Attacks against Tex
  2026 NDSS [in population] There is No War in Ba Sing Se: A Global Analysis of Content Moderation in Large Langu


===== 6. detection[] tuples: the measured results, with their own denominators =====
403 detection tuples with a stated prevalence, across 63 of 63 papers in the population
of those, 344 carry a digit (i.e. could be a publishable figure)

every prevalence tuple for the papers the page quotes:

--- 2018/IMC/403-forbidden-a-global-view-of-cdn-geoblocking
    phenomenon: CDN geoblocking
    technique : Block-page signatures, page-length outliers, clustering, and repeated probes.
    metric    : share of domains or domain-country pairs
    prevalence: Of the 8,000 Alexa Top 10K domains tested globally, the median was 3 inaccessible domains per country, with a maximum of 71 in Syria.
    quote [results]: Of the 8,000 Alexa Top 10K domains we tested globally, we observed a median of 3 domains inaccessible due to geoblocking per country, with a maximum of 71 domains blocked in Syria.
    phenomenon: CDN geoblocking
    technique : Repeated global HTTP probing through Luminati residential exits.
    metric    : percentage of sampled domains
    prevalence: 4.4% of Alexa Top Million domains used their CDN's geoblocking feature in at least one country.
    quote [results]: Of domains in the Alexa Top Million, we observed an overall rate of 4.4% of domains utilizing their CDN's geoblocking feature in at least one country.
    phenomenon: Geoblocking by country
    technique : Country-domain probes with 23 total samples and an 80% agreement threshold.
    metric    : number of blocked domains per country
    prevalence: Syria 71, Iran 67, Sudan 66, and Cuba 66 blocked domains in the Alexa Top 10K.
    quote [results]: The top four countries are Syria, Iran, Sudan, and Cuba, by a wide margin.
    phenomenon: Geoblocking in Alexa Top 1M
    technique : Three baseline probes and 20 follow-up probes for block-page observations.
    metric    : percentage of sampled domains
    prevalence: 238 unique domains, or 4.4% of 5,462 tested categorized domains, showed explicit geoblocking.
    quote [results]: Total 5,462 238 (4.4%)
    phenomenon: Geoblocking in OONI data
    technique : Matched known explicit geoblocking signals in stored responses.
    metric    : share of global test list
    prevalence: 8,313 cases in 139 countries; 97 domains, or 9% of the global test list.
    quote [discussion]: We find 8,313 cases in 139 countries where OONI responses match the explicit signals of geoblocking we describe in Section 4.
    phenomenon: Block-page classifier false positives
    technique : Manual browser validation of automated 403 and block-page detections.
    metric    : false-positive rate
    prevalence: 27% of initially reported Alexa Top 1M block-page instances were false positives.
    quote [results]: Of the 1,068 instances of likely geoblocking across all domain and country pairs initially reported by our automated data classifier, 286 (27%) proved to be false positives upon manual inspection
    phenomenon: Page-length heuristic performance
    technique : Compared 30% page-length outliers against manually identified block pages.
    metric    : recall
    prevalence: Overall recall was 58.3%.
    quote [evaluation]: We found that the overall recall was only 58.3%.

--- 2014/IMC/automated-detection-and-fingerprinting-of-censorship-block-pages
    phenomenon: automated block-page detection
    technique : Compares test pages with known unblocked versions using similarity metrics
    metric    : true-positive and false-positive rates
    prevalence: 95% true positive rate and 1.371% false positive rate
    quote [conclusion]: Using these techniques, we built a block page detection method with a 95.03% true positive rate and a 1.371% false positive rate
    phenomenon: accessible-page classification
    technique : Uses page-length difference thresholding
    metric    : true-positive and false-positive rates
    prevalence: 95% true positive rate; 1.37% false positive rate at a 30% threshold
    quote [results]: A threshold that marks any difference in size over 30% as blocked achieves a true positive rate of 95% and a false positive rate of 1.37%.
    phenomenon: filtering-tool fingerprinting
    technique : Clusters page length and term-frequency features, then matches signatures
    metric    : cluster F-1 measure
    prevalence: Five filtering tools identified from 7 of 36 clusters
    quote [results]: Using this method, we identified five filtering tools that generated 7 out of 36 clusters from the dataset.
    phenomenon: template clustering
    technique : Partitions block pages by unique HTML term-frequency vectors
    metric    : F-1 measure
    prevalence: Term-frequency clustering F-1 = 0.98; page-length clustering F-1 = 0.64
    quote [results]: Term frequency clustering performs well, with an F-1 measure of 0.98; clustering based on page length is much worse, with an F-1 measure of 0.64.

--- 2024/USENIX/digital-discrimination-of-users-in-sanctioned-states-the-case-of-the-cuba-embarg
    phenomenon: geoblocking across network layers
    technique : DNS, TCP, TLS, HTTP measurements with control comparisons
    metric    : number of domains
    prevalence: 546 domains
    quote [abstract]: We identify 546 domains subject to geoblocking across all layers of the network stack
    phenomenon: DNS geoblocking
    technique : Recursive DNS traces and public-resolver comparison
    metric    : share of geoblocked domains
    prevalence: 37/546 domains (6.8%)
    quote [results]: We find 37 domains geoblocking in the DNS layer.
    phenomenon: TCP geoblocking
    technique : TCP handshakes and TCP traceroutes
    metric    : share of geoblocked domains
    prevalence: 97/546 domains (17.7%)
    quote [results]: We find 97 domains implementing geoblocking via TCP handshake failures
    phenomenon: TLS geoblocking
    technique : TLS handshakes and TLS traceroutes
    metric    : share of geoblocked domains
    prevalence: 23/546 domains (4.2%)
    quote [results]: 23 domains via TLS connection failures.
    phenomenon: HTTP geoblocking errors
    technique : HTTP GET requests and response-status analysis
    metric    : share of geoblocked domains
    prevalence: 24/546 domains (4.4%)
    quote [results]: 24 domains were geoblocked at the HTTP GET stage.
    phenomenon: HTTP(S) blockpages
    technique : Response-length comparison, clustering, and language inspection
    metric    : share of geoblocked domains
    prevalence: 395/546 domains (72.3%)
    quote [results]: We identify 395 domains implementing geoblocking via blockpage HTTP(S) responses.
    phenomenon: uninformative geoblocking notices
    technique : Manual categorization of response fingerprints
    metric    : share of geoblocked domains
    prevalence: 88% (480) of 546 domains
    quote [results]: Still, we emphasize that 88% (480) of geoblocked domains of 546 do not serve informative notice of why they are blocked.
    phenomenon: 200 OK blockpages
    technique : Fingerprint classification and manual inspection
    metric    : number of domains
    prevalence: 32 domains
    quote [results]: Of these, 32 domains serve blockpages with 200 OK status codes.

--- 2023/PETS/certainty-detecting-dns-manipulation-at-scale-using-tls-certificates
    phenomenon: DNS manipulation
    technique : Validated TLS certificates and matched returned pages against blockpage fingerprints.
    metric    : share of DNS resolutions
    prevalence: 17 TLS proxy vendors in 52 countries and ISP-level manipulation in 26 countries
    quote [results]: Globally, CERTainty identifies 17 TLS proxy vendors in 52 countries ... CERTainty also detects 55 ASes in 26 countries with ISP-level DNS manipulation
    phenomenon: Invalid TLS certificates
    technique : Checked root trust and requested-domain hostname matching.
    metric    : share of manipulated responses
    prevalence: 82.39% of invalid certificates came without a blockpage
    quote [results]: Among all the invalid certificates CERTainty detected, 82.39% come without a blockpage.
    phenomenon: HTTP blockpages
    technique : Clustered HTTP pages by page length and HTML structure, then manually fingerprinted clusters.
    metric    : number of fingerprints
    prevalence: 226 new blockpage fingerprints
    quote [methodology]: we observe that clustering the pages in the HTTP response based on page length and HTML structure works the best.
    phenomenon: False positives in consistency heuristics
    technique : Compared control-matching heuristics against certificate validity and blockpage matching.
    metric    : false-positive rate
    prevalence: 72.45% of responses labeled manipulation were false positives
    quote [results]: a staggering number of 72.45% DNS resolutions that are tagged as "manipulation" by consistency-based heuristics are false positives.
    phenomenon: False negatives in consistency heuristics
    technique : Compared heuristic matches with invalid certificates or blockpage fingerprints.
    metric    : false-negative rate
    prevalence: 9.70% of true manipulated responses were missed
    quote [evaluation]: 9.70% of true manipulated responses-having an invalid certificate or matching a blockpage fingerprint-are erroneously tagged as correct resolution
    phenomenon: TLS proxy vendors
    technique : Attributed manipulation using certificate issuer fields and blockpage fingerprints.
    metric    : number of vendors and countries
    prevalence: 17 vendors deployed in 52 countries
    quote [results]: CERTainty identifies 17 DNS manipulation filtering product vendors deployed in 52 countries
    phenomenon: ISP-level DNS manipulation
    technique : Used certificate validation and blockpage matching across resolver results.
    metric    : countries with detected manipulation
    prevalence: 26 countries
    quote [results]: In total, CERTainty discovers ISP-level DNS manipulation in 26 countries.
    phenomenon: NXDOMAIN manipulation
    technique : Filtered erroneous resolvers and inspected nonzero RCODE responses.
    metric    : resolver/domain patterns
    prevalence: Gamban returned NXDOMAIN for exactly 47 domains
    quote [appendix]: 4 resolvers (0.*.dns.gamban.com) all return RCODE:3 for exactly 47 domains.

--- 2024/USENIX/gfweb-measuring-the-great-firewalls-web-censorship-at-scale
    phenomenon: HTTP censorship
    technique : SYN and PSH/ACK probes eliciting injected RST/ACK packets
    metric    : censored FQDNs and PLDs
    prevalence: 943K pay-level domains censored by the GFW's HTTP filter
    quote [introduction]: Over a period of 20 months, from February 2022 to September 2023, GFWeb tested over one billion fully qualified domains (FQDNs), detecting 943K and 55K pay-level domains censored by the GFW's HTTP and HTTPS filters, respectively.
    phenomenon: HTTPS censorship
    technique : TLS Client Hello SNI probes eliciting injected RST/ACK packets
    metric    : censored FQDNs and PLDs
    prevalence: 55K pay-level domains censored by the GFW's HTTPS filter
    quote [introduction]: Over a period of 20 months, from February 2022 to September 2023, GFWeb tested over one billion fully qualified domains (FQDNs), detecting 943K and 55K pay-level domains censored by the GFW's HTTP and HTTPS filters, respectively.
    phenomenon: Asymmetric interference
    technique : Compared domain-triggered injections from inside and outside China
    metric    : domains triggering only from inside China
    prevalence: about 1K domains triggered HTTPS interference only when probed from inside China
    quote [results]: Comparing the sets of domains that trigger the GFW from both sides, we found about 1K domains that only trigger the HTTPS filter to inject RST packets when probed from inside the country.
    phenomenon: Residual traffic dropping
    technique : Repeated probes sharing TCP three-tuples over time
    metric    : maximum traffic-dropping duration
    prevalence: up to 350 seconds
    quote [results]: We discover that this traffic dropping behavior will continue happening for up to 350 seconds for TCP packets that share the same three-tuple.
    phenomenon: Domain overblocking
    technique : Compared observed blocking patterns with reverse-engineered regular expressions
    metric    : share of previously overblocked rules corrected
    prevalence: Nine out of ten most overblocked rules were correctly implemented with an additional dot
    quote [results]: Nine out of the ten most overblocked rules reported in [47] are now correctly implemented with an additional dot (\.) character at the beginning of the regular expressions.
    phenomenon: Domain category distribution
    technique : VirusTotal classification of discovered base domains
    metric    : classified base censored domains
    prevalence: 79.5K domains categorized; five leading categories comprised 60%
    quote [results]: Of more than 1M based censored domains discovered, we could only categorize 79.5K domains because many domains no longer exist or do not currently host any content.
    phenomenon: Cloud-provider redirection
    technique : Observed injected HTTP redirections from Chinese hosting providers
    metric    : interfered FQDNs
    prevalence: Aliyun interfered with 36.5M and QCloud with 39.1M FQDNs
    quote [results]: Over the course of our study, Aliyun and QCloud middleboxes have interfered with 36.5M and 39.1M FQDNs, respectively.
    phenomenon: ISP anti-fraud redirection
    technique : Limited TTL probing from a China Telecom-connected vantage point
    metric    : unique triggering FQDNs
    prevalence: 478K unique FQDNs triggered these injections
    quote [results]: GFWeb observed 478K unique FQDNs that trigger these injections.

--- 2020/IEEE-SP/iclab-a-global-longitudinal-internet-censorship-measurement-platform
    phenomenon: DNS manipulation
    technique : Compared vantage DNS responses with control-node responses using AS, routability, and temporal heuristics.
    metric    : unique URLs with observed manipulations
    prevalence: 15,007 DNS manipulations in 56 countries, applied to 489 unique URLs
    quote [results]: We observe 15,007 DNS manipulations in 56 countries, applied to 489 unique URLs.
    phenomenon: TCP packet injection
    technique : Analyzed packet traces for sequence collisions, conflicting payloads, RST/FIN flags, and control-node comparisons.
    metric    : definite censorship-causing injections
    prevalence: 143,225 injections in 54 countries, applied to 1,205 unique URLs
    quote [results]: only 0.7% of these are definitely due to censorship: 143,225 injections, in 54 countries, applied to 1,205 unique URLs.
    phenomenon: Censorship block pages
    technique : Applied verified regular expressions and clustered anomalous HTTP responses using HTML structure, textual similarity, and URL-to-country ratios.
    metric    : block pages observed
    prevalence: 232,183 block pages across 50 countries, applied to 2,782 unique URLs
    quote [results]: We observe 232,183 block pages across 50 countries, applied to 2,782 unique URLs.
    phenomenon: Previously unknown block pages
    technique : Manually inspected HTML-structure and LSH clusters, then added signatures to the curated regular-expression set.
    metric    : new block-page signatures
    prevalence: 48 previously unknown block-page signatures from 13 countries
    quote [introduction]: Using these classifiers we discovered 48 previously undetected block page signatures from 13 countries.
    phenomenon: User-tracking script injection
    technique : Manually inspected block-page-detector clusters and identified injected client-fingerprinting scripts.
    metric    : share of test page loads
    prevalence: 5–30% of test page loads from three major Korean ISPs over five months
    quote [results]: We observed injections of this script over a five-month period from Oct. 2016 through Feb. 2017, from vantage points within three major Korean ISPs, into 5-30% of all our test page loads
    phenomenon: Cryptocurrency-mining malware injection
    technique : Inspected suspicious HTTP responses and identified injected malware associated with infected MikroTik routers.
    metric    : earliest observation date
    prevalence: Observed in Brazil as early as July 21, 2018
    quote [results]: The malware appears in ICLab's records as early as July 21st, 2018-ten days before the earliest public report on the MikroTik botnet
    phenomenon: HTTP 451 geoblocking
    technique : Recorded HTTP status codes and manually examined accompanying block-page HTML.
    metric    : unique websites returning status 451
    prevalence: 23 unique websites from vantage points in 21 countries
    quote [results]: We observe 23 unique websites that return status 451, from vantages in 21 countries.

--- 2022/USENIX/a-large-scale-investigation-into-geodifferences-in-mobile-apps
    phenomenon: app geoblocking
    technique : Compared Google Play metadata and APK download success across countries.
    metric    : number and percentage of apps blocked per country
    prevalence: 3,672 of 5,385 apps were geoblocked in at least one country.
    quote [results]: Of the 5,385 apps, 3,672 apps are geoblocked in at least one country.
    phenomenon: developer-blocking
    technique : Mapped controlled Google Play download errors to country targeting.
    metric    : number and percentage of apps developer-blocked
    prevalence: 2,419 apps (44.9%) were developer-blocked in at least one country.
    quote [results]: 2,419 (44.9%) unique apps are developer-blocked in at least one country
    phenomenon: government-requested takedown
    technique : Compared error patterns against known takedown control apps.
    metric    : number of affected apps
    prevalence: 61 unique apps were subject to government-requested takedowns.
    quote [results]: We observe that 61 unique apps are subject to government-requested takedowns
    phenomenon: APK geodifferences
    technique : Binary-diffed APKs collected from different countries.
    metric    : number of apps with geodifferences
    prevalence: 596 apps exhibited geodifferences.
    quote [results]: we observe geodifferences in 596 apps as seen by a binary diff of our apks across countries.
    phenomenon: permission geodifferences
    technique : Compared country-specific permission sets against their intersection.
    metric    : number of apps with differing permissions
    prevalence: 127 apps exhibited permission geodifferences.
    quote [results]: We found 127 apps that exhibit geodifferences in permissions requested.
    phenomenon: ad-tracker geodifferences
    technique : Compared country-specific third-party tracker sets.
    metric    : number of apps with additional ad trackers
    prevalence: 118 apps had additional ad trackers.
    quote [results]: We found 118 apps with additional ad trackers
    phenomenon: encrypted-communication geodifferences
    technique : Compared network-security configuration settings across APKs.
    metric    : number of apps with differing settings
    prevalence: 23 apps selectively used unencrypted communication settings.
    quote [results]: We find 23 apps that selectively use unencrypted communication settings for some countries.
    phenomenon: privacy-policy geodifferences
    technique : Compared downloaded and cleaned policy texts across countries.
    metric    : number of apps with differing policies
    prevalence: 103 apps had geodifferences in privacy policies.
    quote [results]: We find 103 apps with geodifferences in privacy policies.

--- 2024/PETS/automatic-generation-of-web-censorship-probe-lists
    phenomenon: potential web censorship
    technique : Compared repeated curl results across regional vantage points against a freedom-based baseline.
    metric    : number of potentially blocked domains
    prevalence: 1,490 unique domains potentially faced blocking
    quote [results]: We identified 1,490 unique domains that potentially faced blocking, as they remained inaccessible for over four months of curl measurements and triggered anomalies in the OONI tests.
    phenomenon: China web blocking
    technique : Combined curl failures, OONI anomalies, and direct GFW DNS/HTTP/HTTPS tests.
    metric    : confirmed blocked domains
    prevalence: 527 unique domains detected blocked by the GFW
    quote [results]: In total, 527 unique domains between Beijing and Shanghai were detected to be blocked by the GFW.
    phenomenon: new blocked domains
    technique : Filtered generated URLs to domains absent from the original source list.
    metric    : new domains with suspected blocking
    prevalence: over 1,200 new domains in Beijing and Shanghai triggered OONI anomalies and consistently failed curl
    quote [results]: In Beijing and Shanghai, over 1,200 domains not present in our original source list returned anomalies detected by OONI and consistently failed to connect via curl.
    phenomenon: regional accessibility differences
    technique : Repeated URL requests and classified HTTP status and curl exit codes.
    metric    : URL accessibility proportion
    prevalence: Beijing 66.88% accessible; Shanghai 64.40% accessible
    quote [results]: Beijing 64,518 (66.88%) ... Shanghai 55,969 (64.40%).

--- 2026/PETS/banned-books-analysis-of-censorship-on-amazon-com
    phenomenon: Amazon shipment restrictions
    technique : Automated location switching, offer inspection, and cart-button side-channel testing
    metric    : number of restricted products
    prevalence: 17,842 products restricted from shipment to at least one region
    quote [abstract]: We found 17,842 products that Amazon restricted from being shipped to at least one world region.
    phenomenon: Book shipment censorship
    technique : Amazon category analysis of tested and restricted products
    metric    : share of tested books restricted
    prevalence: 8,965 of 796,081 books (1.1%) restricted to at least one Middle Eastern country
    quote [results]: Amazon restricted shipment of 8,965 out of the 796,081 (1.1%) books in that sample to at least one of Saudi Arabia, the UAE, Qatar, or Yemen.
    phenomenon: Misleading availability messages
    technique : Recorded Amazon availability and error messages for restricted products
    metric    : message share
    prevalence: 74% currently unavailable, 23% temporarily out of stock, 3% cannot ship to location
    quote [results]: Currently unavailable 74% Temporarily out of stock 23% Cannot ship to location 3%
    phenomenon: Censorship regimes
    technique : Hierarchical clustering using Hamming-distance availability comparisons
    metric    : variance explained by nine-cluster model
    prevalence: Clusters explained 83.9% of availability-matrix variance (R² = 0.839)
    quote [results]: The clusters model explains 83.9% of the variance in entries of the availability matrix (𝑅 2 = 0.839)
    phenomenon: False-positive restrictions
    technique : Manual comparison of perceived triggers with actual product content
    metric    : share of manually reviewed products classified as false positives
    prevalence: 47% of temporarily-out-of-stock messages and 11% of cannot-ship messages
    quote [results]: 47% of the "Temporarily out of stock messages" products and 11% of the "This item cannot be shipped" products in our sample were categorized as false positives.

--- 2024/WWW/a-worldwide-view-on-the-reachability-of-encrypted-dns-services
    phenomenon: DoEv4 service blocking
    technique : Seven-stage connection and response testing from global vantage points.
    metric    : blocked-query ratio
    prevalence: 592K (5.92%) of 10M DoEv4 queries were blocked.
    quote [results]: During our measurement period, we sent 10M DoEv4 queries to same 1302 DoEv4 domains from 5K VPs, of which 592K (5.92%) queries were blocked.
    phenomenon: DoEv6 service blocking
    technique : IPv6 DNS resolution, connectivity, handshake, and response testing.
    metric    : blocked-query ratio
    prevalence: 28K (4.91%) of 560K DoEv6 queries were blocked.
    quote [results]: During our measurement period, we sent 560K DoEv6 queries to 448 DoEv6 domains from 473 VPs, of which 28K (4.91%) queries were blocked.
    phenomenon: DoE blocking types
    technique : Classified failures at resolution, ping, transport, handshake, QUIC, and response stages.
    metric    : share of blocked queries by blocking type
    prevalence: 62.83% of DoEv4 services were inaccessible due to Ping blocking.
    quote [results]: Surprisingly, 62.83% of DoEv4 services are inaccessible due to Ping blocking.
    phenomenon: Censorship-indicative blocking
    technique : Checked forged IPs, injected packets, certificates, blockpages, and HTTP 403 responses.
    metric    : share of blocked queries meeting indicators
    prevalence: 27.18% of blocked DoEv4 and 19.73% of blocked DoEv6 queries met at least one condition.
    quote [results]: Our results indicate that 27.18% of blocked DoEv4 queries and 19.73% of blocked DoEv6 queries meet at least one of the aforementioned conditions.
    phenomenon: Incomplete DoE-domain blocking
    technique : Compared access through alternate IP addresses and DoE protocols.
    metric    : share of blocked cases bypassable
    prevalence: In 59.31% of cases, VPs could use other IP addresses or DoE protocols.
    quote [results]: The results show that in 59.31% of cases, VPs can use other IP addresses or DoE protocols to access blocked DoE domains.
    phenomenon: Stable operational DoE domains
    technique : Intersected monthly scan results across three consecutive months.
    metric    : count of stable domains
    prevalence: About 1K stable DoE domains provided services for three consecutive months.
    quote [introduction]: Throughout 15 monthly scans, we find about 1K stable DoE domains, which provide DoE services for three consecutive months.

--- 2024/PETS/what-to-expect-when-you-re-accessing-an-exploration-of-user-privacy-rights-in-pe
    phenomenon: data access request outcomes
    technique : Researchers submitted forms and emails, recorded responses, and obtained paid reports.
    metric    : share of sites providing access reports
    prevalence: Only one group of connected sites provided access to the same report given to paying customers.
    quote [introduction]: Only one of these groups studied (BeenVerified) responded to access requests with the same report given to their paying customers.
    phenomenon: self-search visibility
    technique : Researchers searched their own names and recorded displayed information.
    metric    : number of researchers listed per site
    prevalence: All 20 sites included at least one researcher; nine included all four.
    quote [results]: During the initial self-search we found that all sites included at least one of the researchers and nine sites included all four.
    phenomenon: EU IP blocking
    technique : EU-based researcher attempted access, including through a VPN.
    metric    : share of sites attempting EU blocking
    prevalence: Nine sites attempted to block EU IPs.
    quote [results]: We found that nine of the sites attempt to block EU IPs.
    phenomenon: paid-report contents
    technique : Researchers purchased and inspected 12 reports for information types and accuracy.
    metric    : number of purchased reports
    prevalence: 12 reports purchased from fee and hybrid sites.
    quote [results]: We purchased 12 reports, from the fee and hybrid sites in Table 1.
    phenomenon: cross-site removal effects
    technique : Removal from four sites was followed by verification on remaining sites after two weeks.
    metric    : number of indirectly affected sites
    prevalence: Information was also removed from at least five sites.
    quote [results]: After removing from BeenVerified, PeopleFinders and Whitepages, information was also removed from at least five sites.
    phenomenon: information reappearance
    technique : Researchers repeated self-searches two months after the first removal phase.
    metric    : reappearance after two months
    prevalence: No reappearance was observed after two months.
    quote [results]: No reappearance was observed after two months.

--- 2022/USENIX/the-security-lottery-measuring-client-side-web-security-inconsistencies
    phenomenon: client-side security inconsistencies
    technique : Compared semantic security policies across compatible repeated responses and tests.
    metric    : number of affected sites
    prevalence: 321 sites in total
    quote [results]: In total we detected some inconsistency in 321 sites.
    phenomenon: intra-test inconsistencies
    technique : Compared five responses under identical test conditions.
    metric    : number of affected sites
    prevalence: 127 sites in the first crawl; 210 across all crawls
    quote [results]: Overall, our crawl detected 127 sites which have some type of intra-test inconsistency
    phenomenon: inter-test inconsistencies
    technique : Compared responses across user-agent, language, VPN, and Onion tests.
    metric    : number of affected sites
    prevalence: 194 to 267 sites
    quote [results]: from 194 to 267 sites with inter-test inconsistencies.
    phenomenon: Content Security Policy inconsistencies
    technique : Evaluated CSP semantics for XSS, framing, and TLS enforcement.
    metric    : number of sites
    prevalence: 36 intra-test; 47 inter-test inconsistencies
    quote [results]: Content Security Policy 1,998 12 11 31 23 36
    phenomenon: X-Frame-Options inconsistencies
    technique : Compared normalized framing-protection classes across responses.
    metric    : number of sites
    prevalence: 50 intra-test; 37 inter-test inconsistencies
    quote [results]: X-Frame-Options 5,692 20 18 43 22 50
    phenomenon: HSTS inconsistencies
    technique : Compared four HSTS protection classes and includeSubDomains/preload directives.
    metric    : number of sites
    prevalence: 38 intra-test; 35 inter-test inconsistencies
    quote [results]: Strict-Transport-Security 4,562 15 13 28 23 38
    phenomenon: cookie security inconsistencies
    technique : Compared HttpOnly, Secure, and SameSite attributes for identified cookies.
    metric    : number of sites
    prevalence: 16 intra-test; 167 inter-test inconsistencies
    quote [results]: Cookie Security 3,876 10 9 11 12 16
    phenomenon: page-similarity validation errors
    technique : Combined script hosts, script counts, title overlap, and response size.
    metric    : false-negative and false-positive rates
    prevalence: 1/1,939 (0.05%) false negatives; 1/93 (1.08%) false positives
    quote [results]: 1/1,939 (0.05%) 1/93 (1.08%)


===== 7. Instruments and observatories: who is used, and who is only cited =====
instrument or dataset	named in tools[] as used/produced	papers whose full text names it (corpus-wide)	of which in the population
OONI	7	52	34
Censored Planet	4	47	28
ICLab	1	42	28
Citizen Lab test list	0	35	25
Quack	5	31	22
Hyperquack	3	6	5
Augur	2	43	18
Iris (DNS)	3	157	9
Satellite	4	195	22
GFWatch	1	7	7
Geneva	6	191	10
M-Lab	9	44	2
OpenNet Initiative	0	32	23
Encore	0	30	14
CAUTION: the full-text column is a mention count and several of these names are
homographs — "Satellite" is satellite Internet, "Geneva" a city, "ONI" a substring,
"Iris" an iOS analyser, "Encore" and "Augur" ordinary words. Use the column to show
that tools[] under-counts a dataset, never as a usage figure.


===== 8. The jurisdiction slice: GDPR and CCPA walls =====
corpus-wide, tight probe: 9 papers; loose probe: 24 papers (denominator 5859)
the loose probe contains the tight one by construction, so the distinct union is 24 papers and every one of them was read
tight-probe hits, all of them, with the hand verdict:
  2018/IMC/403-forbidden-a-global-view-of-cdn-geoblocking  [server]
  2018/PETS/nomoads-effective-and-efficient-cross-app-mobile-ad-blocking  [off:incidental]
  2020/WWW/filter-list-generation-for-underserved-regions  [off:incidental]
  2023/NDSS/i-still-know-what-you-watched-last-sunday-privacy-of-the-hbbtv-protocol-in-the-european-smart-tv-landscape  [off:incidental]
  2023/NDSS/chkplug-checking-gdpr-compliance-of-wordpress-plugins-via-cross-language-code-property-graph  [wall-encounter]
  2023/PETS/everybodys-looking-for-ssomething-a-large-scale-evaluation-on-the-privacy-of-oau  [wall-encounter]
  2024/PETS/what-to-expect-when-you-re-accessing-an-exploration-of-user-privacy-rights-in-pe  [jurisdiction]
  2026/PETS/the-role-of-online-forums-in-developer-understanding-of-privacy-law-a-reddit-cas  [wall-encounter]
  2026/PETS/privacy-vs-profit-the-impact-of-googles-manifest-version-3-mv3-update-on-ad-bloc  [off:incidental]
loose-probe hits that the tight probe missed:
  2013/CCS/security-analysis-of-pseudo-random-number-generators-with-input-dev-random-is-no  [not a candidate]  Security analysis of pseudo-random number generators with input: /dev/random is 
  2015/WWW/weakly-supervised-extraction-of-computer-security-events-from-twitter  [not a candidate]  Weakly Supervised Extraction of Computer Security Events from Twitter.
  2016/PETS/access-denied-contrasting-data-access-in-the-united-states-and-ireland  [not a candidate]  Access Denied! Contrasting Data Access in the United States and Ireland
  2017/PETS/detecting-anti-ad-blockers-in-the-wild  [not a candidate]  Detecting Anti Ad-blockers in the Wild
  2018/PETS/undermining-privacy-in-the-aircraft-communications-addressing-and-reporting-syst  [not a candidate]  Undermining Privacy in the Aircraft Communications Addressing and Reporting Syst
  2020/USENIX/call-me-maybe-eavesdropping-encrypted-lte-calls-with-revolte  [not a candidate]  Call Me Maybe: Eavesdropping Encrypted LTE Calls With ReVoLTE
  2021/NDSS/hey-alexa-is-this-skill-safe-taking-a-closer-look-at-the-alexa-skill-ecosystem  [off:incidental]  Hey Alexa, is this Skill Safe?: Taking a Closer Look at the Alexa Skill Ecosyste
  2021/WWW/understanding-the-impact-of-encrypted-dns-on-internet-censorship  [net]  Understanding the Impact of Encrypted DNS on Internet Censorship.
  2022/PETS/setting-the-bar-low-are-websites-complying-with-the-minimum-requirements-of-the  [jurisdiction]  Setting the Bar Low: Are Websites Complying With the Minimum Requirements of the
  2022/USENIX/a-large-scale-investigation-into-geodifferences-in-mobile-apps  [server]  A Large-scale Investigation into Geodifferences in Mobile Apps
  2024/PETS/two-steps-forward-and-one-step-back-the-right-to-opt-out-of-sale-under-cpra  [not a candidate]  Two Steps Forward and One Step Back: The Right to Opt-out of Sale under CPRA
  2024/USENIX/the-effect-of-design-patterns-on-present-and-future-cookie-consent-decisions  [not a candidate]  The Effect of Design Patterns on (Present and Future) Cookie Consent Decisions
  2025/IEEE-SP/not-the-right-question-a-study-on-attitudes-toward-client-side-scanning-with-sec  [not a candidate]  "Not the Right Question?" A Study on Attitudes Toward Client-Side Scanning with 
  2025/USENIX/exposing-and-circumventing-sni-based-quic-censorship-of-the-great-firewall-of-ch  [net+circumvention]  Exposing and Circumventing SNI-based QUIC Censorship of the Great Firewall of Ch
  2025/IMC/fishing-for-smishing-understanding-sms-phishing-infrastructure-and-strategies-by  [not a candidate]  Fishing for Smishing: Understanding SMS Phishing Infrastructure and Strategies b

legal population for reference: 402 papers assess a law; of the population, 8 do
papers whose legal[] names GDPR: 281; of those, in the population: 4


===== 9. Cross-checks against neighbouring pages =====
design:existing_datasets reports a "Censorship list (Citizen Lab, OONI)" family of 17 papers inside ITS OWN reuse population. That population is defined on that page and is not re-derived here — the two definitions below both differ from it, which is exactly why the page cites the family and not the denominator.
  reuse population, that page's definition (studyTypes OR temporal.mode): 3186
  studyTypes flag alone: 2615
  of the population, in the reuse population: 34
design:crawling_location's population is the 3908 papers with a vantage tuple; 62 of the population are in it.
programming:crawler_detection owns "blocked because you look like a bot": 1 candidate(s) were handed to it.


===== 10. Probe recall, and the queued estimate =====
probe 1 (title+summary) hits: 83; of those in the population: 47; excluded: 36 (43.4% of the probe's hits)
population papers NO title probe caught: 16 of 63
the roadmap's queued probe, re-run on this corpus: 68 papers (it reported 74 on 2026-09-02), web 38, 2020-2023 27, 2024-2026 21
of those 68, in the population: 38; excluded: 30

detection[].phenomenon, raw: 157 distinct blocking-ish strings across 71 papers — used to FIND candidates (probe 2) and never to count


===== 11. Venue-years with no extracted papers at all =====
venue-years with zero extracted papers: 11 — CCS 2026, IMC 2026, NDSS 2010, NDSS 2011, NDSS 2016, NDSS 2018, PETS 2010, PETS 2011, PETS 2012, PETS 2013, PETS 2014
(the population itself spans 40 distinct venue-years)

bgd_quotecheck.py

Quote and figure spot-check against three renderings of each cited paper; the string count is in its output.

bgd_quotecheck.py
#!/usr/bin/env python3
"""Quote and figure spot-check for design:blocking_and_geodifference.
 
    python3 scripts/bgd_quotecheck.py
 
Every phrase quoted on the page, and every paper-sourced figure that a reader
could check, is looked for in the cited paper's own text. Coverage is the
figures a reviewer would actually pull on: it does not extend to every numeral
in every table, and the provenance page says which are covered.  Three renderings are tried in order, because no
single one is reliable for two-column ACM/IEEE PDFs:
 
  1. ``paper.cols.txt``  -- the de-columned rendering the corpus ships
  2. ``paper.norm.txt``  -- the raw single-stream rendering
  3. ``pypdf``           -- extracted here, live, from ``paper.pdf``
 
and each rendering is tried twice: verbatim after NFKC/whitespace/quote/dash
normalisation, then folded to lowercase alphanumerics with ``fi``/``fl``
collapsed to ``f``.  The fold is not cosmetic: several of these PDFs drop the
``fi`` ligature outright, so the paper literally reads "we fnd 44 949 ads", and
a verbatim check on the published sentence FAILS on a correct quote.
 
Exits non-zero on any NOT-FOUND.
"""
import re
import sys
import pathlib
import unicodedata
 
ROOTS = [pathlib.Path("/workspace/publications_dataset/data"),
         pathlib.Path("/workspace/publications_dataset")]
ROOT = next(r for r in ROOTS if (r / "extract/run1/extractions.jsonl").exists())
FT = ROOT / "fulltext"
 
 
def norm(s: str) -> str:
    # NFKC folds the mathematical-italic letters that LaTeX emits for inline
    # maths: benzaamia2026 writes its own sample size as U+1D441 ("\U0001d441 =
    # 48,511"), not ASCII "N", so a verbatim check on "N = 48,511" fails.
    s = unicodedata.normalize("NFKC", s)
    s = s.replace("fi", "fi").replace("fl", "fl")
    s = s.replace("‘", "'").replace("’", "'").replace("ʼ", "'")
    s = s.replace("“", '"').replace("”", '"')
    s = re.sub(r"[‐-―−]", "-", s)
    return re.sub(r"\s+", " ", s).strip()
 
 
def fold(s: str) -> str:
    s = norm(s).lower().replace("fi", "f").replace("fl", "f")
    return re.sub(r"[^a-z0-9]", "", s)
 
 
_cache: dict[str, dict[str, str]] = {}
 
 
def renderings(paper: str) -> dict[str, str]:
    if paper in _cache:
        return _cache[paper]
    venue, year, slug = paper.split("/")
    d = FT / year / venue / slug
    out = {}
    for name, f in (("cols", "paper.cols.txt"), ("norm", "paper.norm.txt")):
        if (d / f).exists():
            out[name] = norm((d / f).read_text(encoding="utf8", errors="replace"))
    if (d / "paper.pdf").exists():
        from pypdf import PdfReader
        out["pypdf"] = norm("\n".join(p.extract_text() or "" for p in PdfReader(d / "paper.pdf").pages))
    _cache[paper] = out
    return out
 
 
CHECKS = {
    # ---- the geoblocking method paper, and the error rates it publishes
    "IMC/2018/403-forbidden-a-global-view-of-cdn-geoblocking": [
        "Of domains in the Alexa Top Million, we observed an overall rate of 4.4% of domains utilizing their CDN's geoblocking feature in at least one country",
        "we observed a median of 3 domains inaccessible due to geoblocking per country, with a maximum of 71 domains blocked in Syria",
        "The top four countries are Syria, Iran, Sudan, and Cuba",
        "Of the 1,068 instances of likely geoblocking across all domain and country pairs initially reported by our automated data classifier, 286 (27%) proved to be false positives upon manual inspection",
        "We found that the overall recall was only 58.3%",
        # the repeat-and-agree rule the page attributes to this paper
        "23",
        "80%",
    ],
    # ---- the block-page classifier the field still cites
    "IMC/2014/automated-detection-and-fingerprinting-of-censorship-block-pages": [
        "A threshold that marks any difference in size over 30% as blocked achieves a true positive rate of 95% and a false positive rate of 1.37%",
        "Term frequency clustering performs well, with an F-1 measure of 0.98; clustering based on page length is much worse, with an F-1 measure of 0.64",
    ],
    # ---- per-layer decomposition, and the 200-OK block page
    "USENIX/2024/digital-discrimination-of-users-in-sanctioned-states-the-case-of-the-cuba-embarg": [
        "We identify 546 domains subject to geoblocking across all layers of the network stack",
        "We find 37 domains geoblocking in the DNS layer",
        "We find 97 domains implementing geoblocking via TCP handshake failures",
        "23 domains via TLS connection failures",
        "24 domains were geoblocked at the HTTP GET stage",
        "We identify 395 domains implementing geoblocking via blockpage HTTP(S) responses",
        "88% (480) of geoblocked domains of 546 do not serve informative notice of why they are blocked",
        "Of these, 32 domains serve blockpages with 200 OK status codes",
        "10,093",
    ],
    # ---- the measurement that retired consistency-only DNS detection
    "PETS/2023/certainty-detecting-dns-manipulation-at-scale-using-tls-certificates": [
        "a staggering number of 72.45% DNS resolutions that are tagged as \"manipulation\" by consistency-based heuristics are false positives",
        "9.70% of true manipulated responses",
        "Among all the invalid certificates CERTainty detected, 82.39% come without a blockpage",
        "226",
    ],
    # ---- the design CERTainty replaced
    "USENIX/2017/global-measurement-of-dns-manipulation": [
        "41,778 responses (0.31%) as manipulated",
        "58 countries",
        "13,594,683",
    ],
    # ---- the two platform papers: the control-vantage construction
    "IEEE-SP/2020/iclab-a-global-longitudinal-internet-censorship-measurement-platform": [
        "responses to matching DNS queries from our control node",
        "Control vantage",
        "we discovered 48 previously undetected block page signatures from 13 countries",
        "We observe 23 unique websites that return status 451, from vantages in 21 countries",
        "We observe 232,183 block pages across 50 countries, applied to 2,782 unique URLs",
    ],
    "CCS/2020/censored-planet-an-internet-wide-longitudinal-censorship-observatory": [
        "four remote measurement techniques (Augur, Satellite/Iris, Quack, and Hyperquack)",
        "synchronized measurements on 6 different Internet protocols (IP, DNS, HTTP, HTTPS, Echo and Discard)",
    ],
    # ---- remote measurement: echo servers, then unresponsive hosts
    "USENIX/2018/quack-scalable-remote-measurement-of-application-layer-censorship": [
        "echo servers were present in 184 countries with 4458 unique ASes",
        "the number of blocked domains in Iran increases from 25 to 374",
    ],
    "NDSS/2020/measuring-the-deployment-of-network-censorship-filters-at-global-scale": [
        "FilterMap's data analysis phase generated 90 blockpage clusters",
        "103 countries",
        "each unique blockpage is manually verified to avoid false positives",
    ],
    "IEEE-SP/2017/augur-internet-wide-detection-of-connectivity-disruptions": [
        "side channels to measure reachability between two Internet locations without directly controlling a measurement",
        "shared IP ID counter",
    ],
    "IEEE-SP/2025/is-nobody-there-good-globally-measuring-connection-tampering-without-responsive": [
        "We were able to successfully trigger HTTP interference to 1,303,570 distinct /24s (8.6% of the 15,158,447 that we were able to probe)",
        "For IPv6, we were able to trigger HTTP interference to 862,292 /48s (36.9% of the /48s we could probe)",
    ],
    # ---- both sides of the border, and residual censorship
    "USENIX/2024/gfweb-measuring-the-great-firewalls-web-censorship-at-scale": [
        "detecting 943K and 55K pay-level domains (PLDs) censored by the GFW's HTTP and HTTPS filters, respectively",
        "we found about 1K domains that only trigger the HTTPS filter to inject RST packets when probed from inside the country",
        "this traffic dropping behavior will continue happening for up to 350 seconds for TCP packets that share the same three-tuple",
    ],
    "IEEE-SP/2025/a-wall-behind-a-wall-emerging-regional-censorship-in-china": [
        "We found that only traffic going out of Henan was blocked by the regional firewall",
        "the Henan Firewall blocked 4,196,532 domains-more than five times the 741,542 domains ever blocked by the GFW",
        "The most distinctive fingerprint of the Henan Firewall's RST packets is their 10-byte TCP payload",
        "We found that the Henan Firewall does not perform any residual censorship",
    ],
    # ---- geodifference
    "USENIX/2022/a-large-scale-investigation-into-geodifferences-in-mobile-apps": [
        "Of the 5,385 apps, 3,672 apps are geoblocked in at least one country",
        "2,419 (44.9%) unique apps are developer-blocked in at least one country",
        "We observe that 61 unique apps are subject to government-requested takedowns",
        "we observe geodifferences in 596 apps as seen by a binary diff of our apks across countries",
        "We found 127 apps that exhibit geodifferences in permissions requested",
        "We found 118 apps with additional ad trackers",
        "We find 103 apps with geodifferences in privacy policies",
        "26",
    ],
    "USENIX/2022/the-security-lottery-measuring-client-side-web-security-inconsistencies": [
        "not every URL returns the same content on each load, in particular in the presence of errors or block pages",
        "In total we detected some inconsistency in 321 sites",
        "the 10,000 highest-ranking sites available through HTTPS",
    ],
    "USENIX/2023/network-responses-to-russias-invasion-of-ukraine-in-2022-a-cautionary-tale-for-i": [
        "136 Russian government domains (25.09%) block access to all tested countries outside Russia, and a further 112 government domains (20.66%) cannot be accessed from tests outside Russia and Kazakhstan",
        # the print above is truncated at 112 characters, so the second figure would
        # not appear in this script's own output and the number guard could not see it
        "112 government domains (20.66%)",
    ],
    "PETS/2026/banned-books-analysis-of-censorship-on-amazon-com": [
        "We found 17,842 products that Amazon restricted from being shipped to at least one world region",
        "Amazon restricted shipment of 8,965 out of the 796,081 (1.1%) books in that sample to at least one of Saudi Arabia, the UAE, Qatar, or Yemen",
        "47% of the \"Temporarily out of stock messages\" products and 11% of the \"This item cannot be shipped\" products in our sample were categorized as false positives",
    ],
    "IMC/2025/where-in-the-world-are-my-trackers-mapping-web-tracking-flow-across-diverse-geog": [
        "23",
    ],
    # ---- the jurisdiction wall
    "PETS/2024/what-to-expect-when-you-re-accessing-an-exploration-of-user-privacy-rights-in-pe": [
        "We found that nine of the sites attempt to block EU IPs",
        "20",
    ],
    "PETS/2023/everybodys-looking-for-ssomething-a-large-scale-evaluation-on-the-privacy-of-oau": [
        "did not explicitly block EU-users",
    ],
    "PETS/2022/setting-the-bar-low-are-websites-complying-with-the-minimum-requirements-of-the": [
        "geofencing",
    ],
    "NDSS/2023/chkplug-checking-gdpr-compliance-of-wordpress-plugins-via-cross-language-code-property-graph": [
        "unless they specifically block EU traffic",
    ],
    "PETS/2026/the-role-of-online-forums-in-developer-understanding-of-privacy-law-a-reddit-cas": [
        "proposing to block EU users",
    ],
    # ---- LLMs, where they actually appear
    "PETS/2024/automatic-generation-of-web-censorship-probe-lists": [
        "we further conduct topic expansion using large language models and Google Trends",
        "BERTopic",
    ],
    "NDSS/2026/characterizing-the-implementation-of-censorship-policies-in-chinese-llm-services": [
        "only 29 out of 349 output-blocked queries subject to output blocking across all 5 samples",
        "All services block at the input and output stages",
    ],
    "NDSS/2026/there-is-no-war-in-ba-sing-se-a-global-analysis-of-content-moderation-in-large-language-models": [
        "DeBERTa",
        "two human annotators manually examined over 12k classifications",
        "700k",
    ],
    # ---- protocol migration
    "USENIX/2023/how-the-great-firewall-of-china-detects-and-blocks-fully-encrypted-traffic": [
        "fully encrypted",
    ],
    "WWW/2024/a-worldwide-view-on-the-reachability-of-encrypted-dns-services": [
        # the page's "592K of 10M DoEv4 queries (5.92%)" pairs the abstract's rate with
        # the method's total. Both strings are checked, and 5.92% of 10M = 592K, so
        # the numerator, denominator and percentage are from one experiment.
        "592K (5.92%) DoEv4 queries and 28K (4.91%) DoEv6 queries are blocked",
        "we perform 10M DoEv4 and 560K DoEv6 queries from 102 countries/regions over two months",
        "5031",
        "62.83% of DoEv4 services are inaccessible due to Ping blocking",
        "we observe seven blocking types, including Pre-resolve, Ping, TCP, TLS, QUIC version negotiation, QUIC, and Response blocking",
        "comparison of measurement results from VPs and control nodes",
        "102",
    ],
    "IMC/2019/an-end-to-end-large-scale-measurement-of-dns-over-encryption-how-far-have-we-com": [
        "Over 99% global users can normally access large DNS-over-Encryption servers",
    ],
    # ---- the ground-truth case
    "IMC/2014/censorship-in-the-wild-analyzing-internet-filtering-in-syria": [
        "600GB",
        "Blue Coat",
    ],
    # ---- the memory-disclosure bug, quoted only as a fact about the injector
    "NDSS/2025/wallbleed-a-memory-disclosure-vulnerability-in-the-great-firewall-of-china": [
        "up to 125 bytes",
    ],
    # ---- the outage literature the population deliberately excludes: the page's
    # claim is that it measures unreachability without attributing it to a decision
    "IMC/2025/tracking-internet-disruptions-in-ukraine-insights-from-three-years-of-active-ful": [
        "there is indeed a strong positive correlation between the hours of Internet and power outages in non-frontline regions",
        "0.725",
        "we show that outages in nonfrontline regions strongly correlate with power outages",
    ],
    # ---- per-ISP mechanism differences
    "IMC/2018/where-the-light-gets-in-analyzing-web-censorship-mechanisms-in-india": [
        "nine",
    ],
}
 
 
def _guard_no_duplicate_keys() -> None:
    """A repeated dict literal key silently discards the earlier value, so an
    edit that re-adds a paper drops its existing checks with no error. Count the
    keys in this file's own source instead of trusting the dict."""
    src = pathlib.Path(__file__).read_text(encoding="utf8")
    keys = re.findall(r'^    "([A-Za-z0-9/_.-]+)":', src, re.M)
    dupes = {k for k in keys if keys.count(k) > 1}
    if dupes:
        sys.exit(f"ABORT: duplicate CHECKS keys, earlier quotes would be silently dropped: {sorted(dupes)}")
    if len(keys) != len(CHECKS):
        sys.exit(f"ABORT: {len(keys)} keys in source but {len(CHECKS)} in the dict.")
 
 
def main() -> int:
    _guard_no_duplicate_keys()
    fails = 0
    total = 0
    for paper, quotes in CHECKS.items():
        rend = renderings(paper)
        print(f"\n### {paper}")
        for q in quotes:
            total += 1
            how = None
            for name in ("cols", "norm", "pypdf"):
                if name not in rend:
                    continue
                if norm(q) in rend[name]:
                    how = f"{name}-exact"
                    break
                if fold(q) in fold(rend[name]):
                    how = f"{name}-folded"
                    break
            if how is None:
                how = "NOT FOUND"
                fails += 1
            print(f"  [{how:12}] {q[:112]}")
    print(f"\n{total} strings checked, {fails} NOT FOUND.")
    return 1 if fails else 0
 
 
if __name__ == "__main__":
    sys.exit(main())

bgd_quotecheck-output.txt

Its unedited output. Exits non-zero on any NOT FOUND.

bgd_quotecheck-output.txt
could not convert string to float: b'0.000000000000-5684342' : FloatObject (b'0.000000000000-5684342') invalid; use 0.0 instead
could not convert string to float: b'0.00-34172053' : FloatObject (b'0.00-34172053') invalid; use 0.0 instead
could not convert string to float: b'0.00-30084234' : FloatObject (b'0.00-30084234') invalid; use 0.0 instead
Exceeded 5000 form XObject invocations while extracting text; further form content is skipped.

### IMC/2018/403-forbidden-a-global-view-of-cdn-geoblocking
  [pypdf-exact ] Of domains in the Alexa Top Million, we observed an overall rate of 4.4% of domains utilizing their CDN's geoblo
  [pypdf-exact ] we observed a median of 3 domains inaccessible due to geoblocking per country, with a maximum of 71 domains bloc
  [cols-exact  ] The top four countries are Syria, Iran, Sudan, and Cuba
  [cols-exact  ] Of the 1,068 instances of likely geoblocking across all domain and country pairs initially reported by our autom
  [pypdf-exact ] We found that the overall recall was only 58.3%
  [cols-exact  ] 23
  [cols-exact  ] 80%

### IMC/2014/automated-detection-and-fingerprinting-of-censorship-block-pages
  [cols-exact  ] A threshold that marks any difference in size over 30% as blocked achieves a true positive rate of 95% and a fal
  [cols-exact  ] Term frequency clustering performs well, with an F-1 measure of 0.98; clustering based on page length is much wo

### USENIX/2024/digital-discrimination-of-users-in-sanctioned-states-the-case-of-the-cuba-embarg
  [cols-exact  ] We identify 546 domains subject to geoblocking across all layers of the network stack
  [cols-exact  ] We find 37 domains geoblocking in the DNS layer
  [cols-exact  ] We find 97 domains implementing geoblocking via TCP handshake failures
  [cols-exact  ] 23 domains via TLS connection failures
  [cols-exact  ] 24 domains were geoblocked at the HTTP GET stage
  [pypdf-folded] We identify 395 domains implementing geoblocking via blockpage HTTP(S) responses
  [cols-exact  ] 88% (480) of geoblocked domains of 546 do not serve informative notice of why they are blocked
  [pypdf-exact ] Of these, 32 domains serve blockpages with 200 OK status codes
  [cols-exact  ] 10,093

### PETS/2023/certainty-detecting-dns-manipulation-at-scale-using-tls-certificates
  [cols-exact  ] a staggering number of 72.45% DNS resolutions that are tagged as "manipulation" by consistency-based heuristics 
  [cols-exact  ] 9.70% of true manipulated responses
  [cols-exact  ] Among all the invalid certificates CERTainty detected, 82.39% come without a blockpage
  [cols-exact  ] 226

### USENIX/2017/global-measurement-of-dns-manipulation
  [cols-exact  ] 41,778 responses (0.31%) as manipulated
  [cols-exact  ] 58 countries
  [cols-exact  ] 13,594,683

### IEEE-SP/2020/iclab-a-global-longitudinal-internet-censorship-measurement-platform
  [cols-exact  ] responses to matching DNS queries from our control node
  [cols-exact  ] Control vantage
  [pypdf-exact ] we discovered 48 previously undetected block page signatures from 13 countries
  [pypdf-exact ] We observe 23 unique websites that return status 451, from vantages in 21 countries
  [pypdf-exact ] We observe 232,183 block pages across 50 countries, applied to 2,782 unique URLs

### CCS/2020/censored-planet-an-internet-wide-longitudinal-censorship-observatory
  [cols-exact  ] four remote measurement techniques (Augur, Satellite/Iris, Quack, and Hyperquack)
  [cols-exact  ] synchronized measurements on 6 different Internet protocols (IP, DNS, HTTP, HTTPS, Echo and Discard)

### USENIX/2018/quack-scalable-remote-measurement-of-application-layer-censorship
  [pypdf-folded] echo servers were present in 184 countries with 4458 unique ASes
  [pypdf-exact ] the number of blocked domains in Iran increases from 25 to 374

### NDSS/2020/measuring-the-deployment-of-network-censorship-filters-at-global-scale
  [cols-exact  ] FilterMap's data analysis phase generated 90 blockpage clusters
  [cols-exact  ] 103 countries
  [cols-exact  ] each unique blockpage is manually verified to avoid false positives

### IEEE-SP/2017/augur-internet-wide-detection-of-connectivity-disruptions
  [pypdf-exact ] side channels to measure reachability between two Internet locations without directly controlling a measurement
  [cols-exact  ] shared IP ID counter

### IEEE-SP/2025/is-nobody-there-good-globally-measuring-connection-tampering-without-responsive
  [cols-exact  ] We were able to successfully trigger HTTP interference to 1,303,570 distinct /24s (8.6% of the 15,158,447 that w
  [cols-exact  ] For IPv6, we were able to trigger HTTP interference to 862,292 /48s (36.9% of the /48s we could probe)

### USENIX/2024/gfweb-measuring-the-great-firewalls-web-censorship-at-scale
  [cols-exact  ] detecting 943K and 55K pay-level domains (PLDs) censored by the GFW's HTTP and HTTPS filters, respectively
  [pypdf-exact ] we found about 1K domains that only trigger the HTTPS filter to inject RST packets when probed from inside the c
  [pypdf-folded] this traffic dropping behavior will continue happening for up to 350 seconds for TCP packets that share the same

### IEEE-SP/2025/a-wall-behind-a-wall-emerging-regional-censorship-in-china
  [cols-exact  ] We found that only traffic going out of Henan was blocked by the regional firewall
  [cols-exact  ] the Henan Firewall blocked 4,196,532 domains-more than five times the 741,542 domains ever blocked by the GFW
  [cols-exact  ] The most distinctive fingerprint of the Henan Firewall's RST packets is their 10-byte TCP payload
  [cols-exact  ] We found that the Henan Firewall does not perform any residual censorship

### USENIX/2022/a-large-scale-investigation-into-geodifferences-in-mobile-apps
  [pypdf-exact ] Of the 5,385 apps, 3,672 apps are geoblocked in at least one country
  [pypdf-exact ] 2,419 (44.9%) unique apps are developer-blocked in at least one country
  [pypdf-exact ] We observe that 61 unique apps are subject to government-requested takedowns
  [pypdf-folded] we observe geodifferences in 596 apps as seen by a binary diff of our apks across countries
  [pypdf-folded] We found 127 apps that exhibit geodifferences in permissions requested
  [pypdf-exact ] We found 118 apps with additional ad trackers
  [pypdf-exact ] We find 103 apps with geodifferences in privacy policies
  [cols-exact  ] 26

### USENIX/2022/the-security-lottery-measuring-client-side-web-security-inconsistencies
  [cols-exact  ] not every URL returns the same content on each load, in particular in the presence of errors or block pages
  [cols-exact  ] In total we detected some inconsistency in 321 sites
  [cols-exact  ] the 10,000 highest-ranking sites available through HTTPS

### USENIX/2023/network-responses-to-russias-invasion-of-ukraine-in-2022-a-cautionary-tale-for-i
  [cols-exact  ] 136 Russian government domains (25.09%) block access to all tested countries outside Russia, and a further 112 g
  [cols-exact  ] 112 government domains (20.66%)

### PETS/2026/banned-books-analysis-of-censorship-on-amazon-com
  [cols-exact  ] We found 17,842 products that Amazon restricted from being shipped to at least one world region
  [cols-exact  ] Amazon restricted shipment of 8,965 out of the 796,081 (1.1%) books in that sample to at least one of Saudi Arab
  [cols-exact  ] 47% of the "Temporarily out of stock messages" products and 11% of the "This item cannot be shipped" products in

### IMC/2025/where-in-the-world-are-my-trackers-mapping-web-tracking-flow-across-diverse-geog
  [cols-exact  ] 23

### PETS/2024/what-to-expect-when-you-re-accessing-an-exploration-of-user-privacy-rights-in-pe
  [cols-exact  ] We found that nine of the sites attempt to block EU IPs
  [cols-exact  ] 20

### PETS/2023/everybodys-looking-for-ssomething-a-large-scale-evaluation-on-the-privacy-of-oau
  [cols-exact  ] did not explicitly block EU-users

### PETS/2022/setting-the-bar-low-are-websites-complying-with-the-minimum-requirements-of-the
  [cols-exact  ] geofencing

### NDSS/2023/chkplug-checking-gdpr-compliance-of-wordpress-plugins-via-cross-language-code-property-graph
  [cols-exact  ] unless they specifically block EU traffic

### PETS/2026/the-role-of-online-forums-in-developer-understanding-of-privacy-law-a-reddit-cas
  [cols-exact  ] proposing to block EU users

### PETS/2024/automatic-generation-of-web-censorship-probe-lists
  [cols-exact  ] we further conduct topic expansion using large language models and Google Trends
  [cols-exact  ] BERTopic

### NDSS/2026/characterizing-the-implementation-of-censorship-policies-in-chinese-llm-services
Multiple definitions in dictionary at byte 0x3d5c for key /ToUnicode
Multiple definitions in dictionary at byte 0x3e06 for key /ToUnicode
Multiple definitions in dictionary at byte 0x3eb0 for key /ToUnicode
Multiple definitions in dictionary at byte 0x3f5a for key /ToUnicode
Multiple definitions in dictionary at byte 0x4004 for key /ToUnicode
Multiple definitions in dictionary at byte 0x40ad for key /ToUnicode
Multiple definitions in dictionary at byte 0x4156 for key /ToUnicode
Multiple definitions in dictionary at byte 0x4200 for key /ToUnicode
Multiple definitions in dictionary at byte 0x42a6 for key /ToUnicode
Multiple definitions in dictionary at byte 0x4350 for key /ToUnicode
Multiple definitions in dictionary at byte 0x43f8 for key /ToUnicode
Multiple definitions in dictionary at byte 0x44a0 for key /ToUnicode
Multiple definitions in dictionary at byte 0x454a for key /ToUnicode
Multiple definitions in dictionary at byte 0x45f2 for key /ToUnicode
  [cols-exact  ] only 29 out of 349 output-blocked queries subject to output blocking across all 5 samples
  [pypdf-exact ] All services block at the input and output stages

### NDSS/2026/there-is-no-war-in-ba-sing-se-a-global-analysis-of-content-moderation-in-large-language-models
  [cols-exact  ] DeBERTa
  [cols-exact  ] two human annotators manually examined over 12k classifications
  [cols-exact  ] 700k

### USENIX/2023/how-the-great-firewall-of-china-detects-and-blocks-fully-encrypted-traffic
  [cols-exact  ] fully encrypted

### WWW/2024/a-worldwide-view-on-the-reachability-of-encrypted-dns-services
  [cols-exact  ] 592K (5.92%) DoEv4 queries and 28K (4.91%) DoEv6 queries are blocked
  [cols-exact  ] we perform 10M DoEv4 and 560K DoEv6 queries from 102 countries/regions over two months
  [cols-exact  ] 5031
  [pypdf-folded] 62.83% of DoEv4 services are inaccessible due to Ping blocking
  [cols-exact  ] we observe seven blocking types, including Pre-resolve, Ping, TCP, TLS, QUIC version negotiation, QUIC, and Resp
  [cols-exact  ] comparison of measurement results from VPs and control nodes
  [cols-exact  ] 102

### IMC/2019/an-end-to-end-large-scale-measurement-of-dns-over-encryption-how-far-have-we-com
  [pypdf-exact ] Over 99% global users can normally access large DNS-over-Encryption servers

### IMC/2014/censorship-in-the-wild-analyzing-internet-filtering-in-syria
  [cols-exact  ] 600GB
  [cols-exact  ] Blue Coat

### NDSS/2025/wallbleed-a-memory-disclosure-vulnerability-in-the-great-firewall-of-china
  [cols-exact  ] up to 125 bytes

### IMC/2025/tracking-internet-disruptions-in-ukraine-insights-from-three-years-of-active-ful
  [cols-exact  ] there is indeed a strong positive correlation between the hours of Internet and power outages in non-frontline r
  [cols-exact  ] 0.725
  [cols-exact  ] we show that outages in nonfrontline regions strongly correlate with power outages

### IMC/2018/where-the-light-gets-in-analyzing-web-censorship-mechanisms-in-india
  [cols-exact  ] nine

94 strings checked, 0 NOT FOUND.

bgd_external_checks.sh

Every external source the page names, fetched. Status before content, always.

bgd_external_checks.sh
#!/bin/sh
# External-source checks for design:blocking_and_geodifference.
# Every claim the page makes about a live instrument, dataset or standard is
# checked here, with the HTTP status printed BEFORE any content, because a 302 to
# a redirect page and an empty 200 look identical in a byte count.
UA='Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/130.0 Safari/537.36'
say() { printf '\n=== %s\n' "$1"; }
code() { curl -sS -o /tmp/bgd_body -w 'HTTP %{http_code}  final=%{url_effective}  bytes=%{size_download}\n' -L -A "$UA" "$1"; }
 
say 'OONI: project site'
code 'https://ooni.org/'
grep -oiE '<title>[^<]*' /tmp/bgd_body | head -1
 
say 'OONI: measurement aggregation API (is the public API alive?)'
code 'https://api.ooni.io/api/v1/aggregation?since=2026-09-01&until=2026-09-08&test_name=web_connectivity&probe_cc=IR'
head -c 400 /tmp/bgd_body; echo
 
say 'OONI: raw measurement listing API'
code 'https://api.ooni.io/api/v1/measurements?probe_cc=CN&limit=1'
head -c 300 /tmp/bgd_body; echo
 
say 'OONI: S3 open data bucket index page'
code 'https://ooni.org/post/mining-ooni-data/'
grep -oiE 'ooni-data-eu-fra[^"< ]*' /tmp/bgd_body | head -3
 
# The unauthenticated GitHub API is 60 requests/hour per IP and a rate-limited
# call that prints nothing reads like a pass, so every gh() call prints its own
# status and says FAILED out loud. GH_TOKEN is used if present and never echoed:
# `${GH_TOKEN:+set}${GH_TOKEN:-unset}` prints the token itself, and did, in this
# very run — see provenance:design:blocking_and_geodifference.
# The body goes to a file and the status to stderr, so a caller can pipe the body
# into a parser without the status line ending up inside the JSON — the first
# version of this function printed "HTTP 200" into the pipe and every parse said
# FAILED, which looked exactly like a rate limit.
gh() {
  st=$(curl -sS -o /tmp/bgd_gh -w '%{http_code}' -H 'Accept: application/vnd.github+json' \
       ${GH_TOKEN:+-H "Authorization: Bearer $GH_TOKEN"} "$1")
  printf 'HTTP %s  %s\n' "$st" "$1" >&2
  [ "$st" = "200" ] || { echo "FAILED — HTTP $st from $1, no figure from this call" >&2; return 1; }
  cat /tmp/bgd_gh
}
 
say 'OONI Probe CLI: latest release'
gh 'https://api.github.com/repos/ooni/probe-cli/releases/latest' | python3 -c 'import sys,json;d=json.load(sys.stdin);print(d["tag_name"], d["published_at"], "prerelease="+str(d["prerelease"]))' 2>/dev/null || echo 'FAILED to parse'
 
say 'OONI Probe CLI: newest 3 releases (a /releases/latest can hide a prerelease)'
gh 'https://api.github.com/repos/ooni/probe-cli/releases?per_page=3' | python3 -c 'import sys,json
for r in json.load(sys.stdin): print(r["tag_name"], r["published_at"], "prerelease="+str(r["prerelease"]))' 2>/dev/null || echo 'FAILED to parse'
 
say 'Citizen Lab test lists: latest commit, and how many country lists exist'
gh 'https://api.github.com/repos/citizenlab/test-lists/commits?per_page=1' | python3 -c 'import sys,json;d=json.load(sys.stdin);print(d[0]["commit"]["committer"]["date"], "|", d[0]["commit"]["message"].splitlines()[0][:70])' 2>/dev/null || echo 'FAILED to parse'
gh 'https://api.github.com/repos/citizenlab/test-lists/contents/lists' | python3 -c 'import sys,json;d=json.load(sys.stdin);print(len([x for x in d if x["name"].endswith(".csv")]), "csv files;", "global.csv present:", any(x["name"]=="global.csv" for x in d))' 2>/dev/null || echo 'FAILED to parse'
 
say 'Censored Planet: is the observatory still publishing? (repo activity)'
gh 'https://api.github.com/repos/censoredplanet/censoredplanet-analysis/commits?per_page=1' | python3 -c 'import sys,json;d=json.load(sys.stdin);print(d[0]["commit"]["committer"]["date"], "|", d[0]["commit"]["message"].splitlines()[0][:70])' 2>/dev/null || echo 'FAILED to parse'
 
say 'Censored Planet: site'
code 'https://censoredplanet.org/'
grep -oiE '<title>[^<]*' /tmp/bgd_body | head -1
 
say 'Censored Planet: data page (what is downloadable, and how recent)'
code 'https://censoredplanet.org/data'
grep -oE 'CP_[A-Za-z]+-[0-9]{4}-[0-9]{2}-[0-9]{2}[^"< ]*' /tmp/bgd_body | tail -5
grep -oiE 'href="[^"]*data[^"]*"' /tmp/bgd_body | head -8
 
say 'ICLab: site (2020 IEEE S&P platform — still reachable?)'
echo '-- with certificate verification (what an ordinary client does):'
code 'https://iclab.org/'
echo '-- ignoring the certificate:'
curl -sSk -o /tmp/bgd_body -w 'HTTP %{http_code}  bytes=%{size_download}\n' -L -A "$UA" 'https://iclab.org/'
grep -oiE '<title>[^<]*' /tmp/bgd_body | head -1
echo '-- certificate validity:'
echo | openssl s_client -servername iclab.org -connect iclab.org:443 2>/dev/null | openssl x509 -noout -dates -issuer -subject
 
say 'GFWatch'
code 'https://gfwatch.org/'
grep -oiE '<title>[^<]*' /tmp/bgd_body | head -1
 
say 'M-Lab'
code 'https://www.measurementlab.net/data/'
grep -oiE '<title>[^<]*' /tmp/bgd_body | head -1
 
say 'IODA (Georgia Tech / CAIDA outage detection)'
code 'https://ioda.inetintel.cc.gatech.edu/'
grep -oiE '<title>[^<]*' /tmp/bgd_body | head -1
 
say 'IODA API: is there a public entities/signals endpoint?'
code 'https://api.ioda.inetintel.cc.gatech.edu/v2/signals/raw/country/IR?from=1757000000&until=1757086400'
head -c 300 /tmp/bgd_body; echo
 
say 'OONI open data: is the S3 bucket still receiving daily dumps?'
code 'https://ooni-data-eu-fra.s3.eu-central-1.amazonaws.com/?list-type=2&delimiter=/&prefix=raw/'
echo "newest date prefixes returned by the first page of the listing (the listing is paginated, so this is the first 1000 keys, NOT the newest date):"
grep -oE '<Prefix>raw/[0-9]+/</Prefix>' /tmp/bgd_body | tail -2
echo "direct test for yesterday's prefix (a listing that is paginated cannot answer 'is it current'; asking for the date can):"
Y=$(date -u -d yesterday +%Y%m%d)
code "https://ooni-data-eu-fra.s3.eu-central-1.amazonaws.com/?list-type=2&delimiter=/&prefix=raw/${Y}/"
echo "prefix asked for: raw/${Y}/ ; hour prefixes found:"
grep -oE '<Prefix>raw/[0-9]+/[0-9]+/</Prefix>' /tmp/bgd_body | wc -l
 
say 'Censored Planet current entry points (the paths a 2020-2023 paper cites are gone)'
for u in https://censoredplanet.org/data/raw https://dashboard.censoredplanet.org/ https://data.censoredplanet.org/ https://docs.censoredplanet.org/; do
  printf '%s -> ' "$u"; curl -sS -o /dev/null -w 'HTTP %{http_code}\n' -L -A "$UA" "$u"
done
 
say 'Censored Planet tool repos named on its Tools page'
for r in geoinspector CenTrace CenFuzz dns_blockpage_fingerprint chinese-llm-blocking censoredplanet-analysis geodiff-app; do
  printf '%-28s ' "$r"
  gh "https://api.github.com/repos/CensoredPlanet/$r" > /tmp/bgd_repo 2>/dev/null \
    && python3 -c "
import json;d=json.load(open('/tmp/bgd_repo'))
print('last push', d['pushed_at'][:10], '| created', d['created_at'][:10], '| archived', d['archived'], '|', (d['description'] or '(no description)')[:58])" \
    || echo 'FAILED'
done
 
say 'Cloudflare Radar outage centre (industry instrument for shutdowns)'
code 'https://radar.cloudflare.com/outage-center'
 
say 'Google Transparency Report traffic disruptions'
code 'https://transparencyreport.google.com/traffic/overview'
 
say 'Access Now KeepItOn shutdown dataset'
code 'https://www.accessnow.org/campaign/keepiton/'
grep -oiE '<title>[^<]*' /tmp/bgd_body | head -1
 
say 'RFC 7725 (HTTP 451) status from the RFC Editor index'
code 'https://www.rfc-editor.org/rfc/rfc7725.json'
head -c 600 /tmp/bgd_body; echo
 
say 'RFC 9110 (HTTP semantics, defines 403) status'
code 'https://www.rfc-editor.org/rfc/rfc9110.json'
python3 -c 'import json;d=json.load(open("/tmp/bgd_body"));print({k:d.get(k) for k in ("doc_id","status","pub_date","obsoleted_by","updated_by")})' 2>/dev/null || echo FAILED
 
say 'The wiki bibliography this page appends to: how many entries does it hold now?'
curl -sS -o /tmp/bgd_bib -w 'HTTP %{http_code}  bytes=%{size_download}\n' -L 'https://measuretheweb.org/literature/bibliography?do=export_raw'
printf 'BibTeX entries live on literature:bibliography: '
grep -c '^@' /tmp/bgd_bib
 
say 'ICLab: is the PROJECT alive even though the site is not? (added after review)'
# A dead website is not a dead project. The public dashboard is unreachable, but the
# codebase may still be moving, and a reader told "the platform is unavailable" would
# not think to look.
gh 'https://api.github.com/orgs/iclab/repos?per_page=100&sort=pushed' > /tmp/bgd_iclab 2>/dev/null \
  && python3 -c "
import json
d=json.load(open('/tmp/bgd_iclab'))
print(len(d), 'repos in the iclab org, newest first:')
for r in sorted(d, key=lambda x: x['pushed_at'], reverse=True)[:6]:
    print('  ', r['name'].ljust(24), 'pushed', r['pushed_at'][:10], 'archived', r['archived'], '|', (r['description'] or '(no description)')[:58])" \
  || echo 'FAILED'
 
say 'EUR-Lex reachability, re-tested (the claim that it bot-walls automated clients)'
# The published note in this repo says EUR-Lex answers automated fetchers with an
# AWS-WAF challenge. A review pass could not reproduce that, and neither can this
# script today, so the check is here rather than the assertion being carried forward.
curl -sS -o /tmp/bgd_eurlex -w 'HTTP %{http_code}  bytes=%{size_download}\n' -L \
  'https://eur-lex.europa.eu/legal-content/EN/ALL/?uri=CELEX:32018R0302'
grep -oiE '<title>[^<]*' /tmp/bgd_eurlex | head -1
printf 'challenge markers (aws-waf / just a moment / captcha) found: '
grep -ciE 'aws-waf|just a moment|captcha' /tmp/bgd_eurlex
 
say 'Regulation 2018/302: has it been modified since 2018? (from EUR-Lex metadata)'
python3 - <<'PYEOF'
import re, html
t = open('/tmp/bgd_eurlex', encoding='utf8', errors='replace').read()
t = re.sub(r'\s+', ' ', html.unescape(re.sub(r'<[^>]+>', ' ', t)))
m = re.search(r'Modified by:(.{0,400})', t)
print('Modified by:', m.group(1).strip()[:380] if m else 'NOT FOUND')
for k in ('32065', '2022/2065', '32023R2854'):
    print(f'  mentions {k}:', k in t)
PYEOF
 
say 'EU Geo-blocking Regulation 2018/302 via the Publications Office CELLAR'
curl -sS -o /tmp/bgd_celex -w 'HTTP %{http_code}  bytes=%{size_download}\n' -L -H 'Accept: application/xhtml+xml' -H 'Accept-Language: eng' 'http://publications.europa.eu/resource/celex/32018R0302'
grep -oiE 'Regulation \(EU\) 2018/302[^<]{0,120}' /tmp/bgd_celex | head -2

bgd_external_checks-output.txt

Its unedited output.

bgd_external_checks-output.txt
=== OONI: project site
HTTP 200  final=https://ooni.org/  bytes=23969
<title>OONI: Open Observatory of Network Interference | OONI

=== OONI: measurement aggregation API (is the public API alive?)
HTTP 200  final=https://api.ooni.io/api/v1/aggregation?since=2026-09-01&until=2026-09-08&test_name=web_connectivity&probe_cc=IR  bytes=265
{"v":0,"dimension_count":0,"db_stats":{"row_count":13030747,"bytes":464938932,"total_row_count":13030747,"elapsed_seconds":0.1298508644104004},"result":{"anomaly_count":19156,"confirmed_count":42483,"failure_count":4532,"ok_count":77309,"measurement_count":143480}}

=== OONI: raw measurement listing API
HTTP 200  final=https://api.ooni.io/api/v1/measurements?probe_cc=CN&limit=1  bytes=810
{"metadata":{"count":-1,"current_page":1,"limit":1,"next_url":"https://api.ooni.io/api/v1/measurements?probe_cc=CN&limit=1&offset=1","offset":0,"pages":-1,"query_time":0.06128096580505371},"results":[{"anomaly":false,"confirmed":false,"failure":true,"input":null,"probe_asn":"AS4837","probe_cc":"CN",

=== OONI: S3 open data bucket index page
HTTP 200  final=https://ooni.org/post/mining-ooni-data/  bytes=18548

=== OONI Probe CLI: latest release
HTTP 200  https://api.github.com/repos/ooni/probe-cli/releases/latest
v3.30.0 2026-07-27T15:34:35Z prerelease=False

=== OONI Probe CLI: newest 3 releases (a /releases/latest can hide a prerelease)
HTTP 200  https://api.github.com/repos/ooni/probe-cli/releases?per_page=3
v3.30.0 2026-07-27T15:34:35Z prerelease=False
v3.30.0-beta 2026-07-26T20:11:47Z prerelease=True
v3.30.0-alpha 2026-07-24T15:14:35Z prerelease=True

=== Citizen Lab test lists: latest commit, and how many country lists exist
HTTP 200  https://api.github.com/repos/citizenlab/test-lists/commits?per_page=1
2026-09-09T12:26:27Z | Merge pull request #2266 from ooni/user-contribution/2865a2f52728bdf80
HTTP 200  https://api.github.com/repos/citizenlab/test-lists/contents/lists
150 csv files; global.csv present: True

=== Censored Planet: is the observatory still publishing? (repo activity)
HTTP 200  https://api.github.com/repos/censoredplanet/censoredplanet-analysis/commits?per_page=1
2025-06-29T20:30:37Z | Skip rows with empty start_time or end_time to prevent IndexError in f

=== Censored Planet: site
HTTP 200  final=https://censoredplanet.org/  bytes=1644
<title>Censored Planet

=== Censored Planet: data page (what is downloadable, and how recent)
HTTP 404  final=https://censoredplanet.org/data  bytes=2010

=== ICLab: site (2020 IEEE S&P platform — still reachable?)
-- with certificate verification (what an ordinary client does):
curl: (60) SSL certificate problem: certificate has expired
More details here: https://curl.se/docs/sslcerts.html

curl failed to verify the legitimacy of the server and therefore could not
establish a secure connection to it. To learn more about this situation and
how to fix it, please visit the web page mentioned above.
HTTP 000  final=https://iclab.org/  bytes=0
-- ignoring the certificate:
HTTP 200  bytes=48777
<title>ICLab
-- certificate validity:
notBefore=Apr 19 01:52:39 2025 GMT
notAfter=Jul 18 01:52:38 2025 GMT
issuer=C = US, O = Let's Encrypt, CN = R11
subject=CN = iclab.org

=== GFWatch
HTTP 200  final=https://gfwatch.org/  bytes=4078
<title>GFWatch Dashboard

=== M-Lab
HTTP 200  final=https://www.measurementlab.net/data/  bytes=26222
<title>M-Lab Data - M-Lab

=== IODA (Georgia Tech / CAIDA outage detection)
HTTP 200  final=https://ioda.inetintel.cc.gatech.edu/  bytes=3120
<title>IODA

=== IODA API: is there a public entities/signals endpoint?
HTTP 200  final=https://api.ioda.inetintel.cc.gatech.edu/v2/signals/raw/country/IR?from=1757000000&until=1757086400  bytes=61922
{"type":"signals","metadata":{"requestTime":"2026-09-10T23:39:50+00:00","responseTime":"2026-09-10T23:39:51+00:00"},"requestParameters":{"from":1757000000,"until":1757086400,"datasource":null,"sourceParams":null,"maxPoints":null},"error":null,"perf":[{"datasource":"gtr","backend":"influxv2","timeUse

=== OONI open data: is the S3 bucket still receiving daily dumps?
HTTP 200  final=https://ooni-data-eu-fra.s3.eu-central-1.amazonaws.com/?list-type=2&delimiter=/&prefix=raw/  bytes=63381
newest date prefixes returned by the first page of the listing (the listing is paginated, so this is the first 1000 keys, NOT the newest date):
<Prefix>raw/20230715/</Prefix>
<Prefix>raw/20230716/</Prefix>
direct test for yesterday's prefix (a listing that is paginated cannot answer 'is it current'; asking for the date can):
HTTP 200  final=https://ooni-data-eu-fra.s3.eu-central-1.amazonaws.com/?list-type=2&delimiter=/&prefix=raw/20260909/  bytes=1869
prefix asked for: raw/20260909/ ; hour prefixes found:
24

=== Censored Planet current entry points (the paths a 2020-2023 paper cites are gone)
https://censoredplanet.org/data/raw -> HTTP 404
https://dashboard.censoredplanet.org/ -> HTTP 200
https://data.censoredplanet.org/ -> HTTP 200
https://docs.censoredplanet.org/ -> HTTP 200

=== Censored Planet tool repos named on its Tools page
geoinspector                 last push 2023-05-19 | created 2023-02-24 | archived False | Geoblocking Measurement Toolkit
CenTrace                     last push 2023-03-30 | created 2022-10-07 | archived False | Run HTTP and HTTPS traceroutes to detect the network locat
CenFuzz                      last push 2024-02-04 | created 2022-10-07 | archived False | Tool for fuzzing HTTP and HTTPS requests to endpoints, and
dns_blockpage_fingerprint    last push 2025-05-01 | created 2022-05-04 | archived False | Code and notebooks for generating DNS blockpage fingerprin
chinese-llm-blocking         last push 2025-07-16 | created 2025-07-16 | archived False | (no description)
censoredplanet-analysis      last push 2025-06-29 | created 2020-11-03 | archived False | Analysis of the CensoredPlanet data.
geodiff-app                  last push 2025-05-01 | created 2021-09-30 | archived False | A Large-scale Investigation into Geodifferences in Mobile 

=== Cloudflare Radar outage centre (industry instrument for shutdowns)
HTTP 403  final=https://radar.cloudflare.com/outage-center  bytes=5671

=== Google Transparency Report traffic disruptions
HTTP 200  final=https://transparencyreport.google.com/traffic/overview  bytes=7814

=== Access Now KeepItOn shutdown dataset
HTTP 200  final=https://www.accessnow.org/campaign/keepiton/  bytes=485572
<title>#KeepItOn: fighting internet shutdowns around the world - Access Now

=== RFC 7725 (HTTP 451) status from the RFC Editor index
HTTP 200  final=https://www.rfc-editor.org/rfc/rfc7725.json  bytes=770
{
  "draft": "draft-ietf-httpbis-legally-restricted-status-04",
  "doc_id": "RFC7725",
  "title": "An HTTP Status Code to Report Legal Obstacles",
  "authors": [
    "T. Bray"
  ],
  "format": [
    "TEXT",
    "HTML"
  ],
  "page_count": "5",
  "pub_status": "PROPOSED STANDARD",
  "status": "PROPOSED STANDARD",
  "source": "HTTP",
  "abstract": "This document specifies a Hypertext Transfer Protocol (HTTP) status code for use when resource access is denied as a consequence of legal demands.",
  "pub_date": "February 2016",
  "keywords": [
    "Hypertext Transfer Protocol"
  ],
  "obsoletes": [

=== RFC 9110 (HTTP semantics, defines 403) status
HTTP 200  final=https://www.rfc-editor.org/rfc/rfc9110.json  bytes=1479
{'doc_id': 'RFC9110', 'status': 'INTERNET STANDARD', 'pub_date': 'June 2022', 'obsoleted_by': [], 'updated_by': []}

=== The wiki bibliography this page appends to: how many entries does it hold now?
HTTP 200  bytes=437218
BibTeX entries live on literature:bibliography: 952

=== ICLab: is the PROJECT alive even though the site is not? (added after review)
7 repos in the iclab org, newest first:
   centinel-prime           pushed 2026-05-05 archived False | (no description)
   historical-data-tools    pushed 2023-03-14 archived False | Tools for working with the historic data collected by ICLa
   centinel                 pushed 2022-03-02 archived False | (no description)
   centinel-server          pushed 2017-03-03 archived False | (no description)
   DifferentiationDetector  pushed 2016-12-26 archived False | (no description)
   test-lists               pushed 2016-05-03 archived False | URL testing lists intended for discovering website censors

=== EUR-Lex reachability, re-tested (the claim that it bot-walls automated clients)
HTTP 200  bytes=424283
<title>Regulation - 2018/302 - EN - EUR-Lex
challenge markers (aws-waf / just a moment / captcha) found: 0

=== Regulation 2018/302: has it been modified since 2018? (from EUR-Lex metadata)
Modified by: Relation Act Comment Subdivision concerned From To Implicitly repealed by 32017R2394 Partial repeal article 10 point 1 17/01/2020 Corrected by 32018R0302R(01) Corrected by 32018R0302R(02) (HU) Corrected by 32018R0302R(03) (NL) Instruments cited: 12016E/TXT 12016E057 12016E101 12016E102 12016M005 12016P/TXT 12016P011 12016P016 12016P017 12016P038 31999L0044 32001L0029 32006L0112
  mentions 32065: False
  mentions 2022/2065: False
  mentions 32023R2854: False

=== EU Geo-blocking Regulation 2018/302 via the Publications Office CELLAR
HTTP 200  bytes=123529
REGULATION (EU) 2018/302 OF THE EUROPEAN PARLIAMENT AND OF THE COUNCIL
Regulation (EU) 2018/302 of the European Parliament and of the Council of 28 February 2018 on addressing unjustified geo-blocking and other form

bgd_render_checks.mjs

The four sources that are single-page apps, rendered, because HTTP 200 is not a currency check.

bgd_render_checks.mjs
#!/usr/bin/env node
// The three public dashboards this page points a reader at are single-page apps:
// curl gets an empty <div id="root"> from censoredplanet.org and a Dash bootstrap
// from gfwatch.org, so "HTTP 200" says nothing about whether the data behind them
// is current. This renders each one and prints the newest date it shows.
//
//   PLAYWRIGHT_BROWSERS_PATH=/workspace/.playwright node scripts/bgd_render_checks.mjs
//
// Playwright's own Chromium build is required: /usr/bin/chromium cannot be driven
// here (Debian wrapper vs crashpad flags).
import { chromium } from 'playwright';
 
const TARGETS = [
  ['https://censoredplanet.org/', 2000, 1200],
  ['https://censoredplanet.org/#/tools', 2500, 2000],
  ['https://dashboard.censoredplanet.org/', 7000, 900],
  ['https://gfwatch.org/', 7000, 1800],
  ['https://gfweb.ca/', 7000, 1800],
  ['https://ioda.inetintel.cc.gatech.edu/dashboard', 8000, 900],
];
const DATE = /\b(19|20)\d{2}[-/](0?[1-9]|1[0-2])[-/](0?[1-9]|[12]\d|3[01])\b|\b(Jan|Feb|Mar|Apr|May|Jun|Jul|Aug|Sep|Oct|Nov|Dec)[a-z]* \d{1,2},? ?(19|20)?\d{2}\b/g;
 
const b = await chromium.launch();
const page = await b.newPage();
console.log(`bgd_render_checks.mjs — run ${new Date().toISOString().slice(0, 10)}`);
for (const [url, wait, chars] of TARGETS) {
  try {
    const r = await page.goto(url, { waitUntil: 'networkidle', timeout: 90000 });
    await page.waitForTimeout(wait);
    const text = (await page.innerText('body')).replace(/\n{2,}/g, '\n');
    const dates = [...new Set((text.match(DATE) ?? []))];
    // The Censored Planet dashboard prints its counters with DOTS as thousands
    // separators (116.924.640.072). A page that quotes them with commas cannot be
    // matched against this output by check_page_numbers.mjs, so print both forms.
    const counters = [...new Set((text.match(/\d{1,3}(?:[.,]\d{3}){2,}/g) ?? []))];
    if (counters.length) {
      console.log('big counters as displayed: ' + counters.join(' | '));
      console.log('same, with ASCII commas   : ' + counters.map((c) => c.replace(/[.,]/g, ',')).join(' | '));
    }
    console.log(`\n#### ${url}   HTTP ${r === null ? '(hash navigation, no response)' : r.status()}`);
    console.log(`dates visible on the rendered page: ${dates.join(' | ') || '(none)'}`);
    console.log('--- rendered text, first ' + chars + ' characters ---');
    console.log(text.slice(0, chars));
  } catch (e) {
    console.log(`\n#### ${url}   RENDER FAILED: ${e.message.split('\n')[0].slice(0, 160)}`);
  }
}
await b.close();

bgd_render_checks-output.txt

Its unedited output, including the dashboard dates the page's currency table rests on.

bgd_render_checks-output.txt
bgd_render_checks.mjs — run 2026-09-10

#### https://censoredplanet.org/   HTTP 200
dates visible on the rendered page: (none)
--- rendered text, first 1200 characters ---
About
Research
Blogs
Events
Publications
Tools
News
Talks
Join Us
Censored Planet
We are a research team at the University of Michigan building scalable systems and novel techniques to protect users from online censorship, surveillance, and digital divide.
Publications
About
Latest News
18-Aug-2026– Read London Daily News article featuring our work on VPN apps leaking user data.
What We Do
Our research lies at the intersection of Networking, Security & Privacy, and Internet Measurements. We take a data-driven approach to detect and defend against powerful network intermediaries and government threat actors.
We have won numerous awards including the Internet Defense Prize, IRTF Applied Networking Research Prizes, and Distinguished and Practical Paper awards. Additionaly, our collaboration with Consumer Reports has been cited by members of Congress in urging the Federal Trade Commission (FTC) to regulate the VPN ecosystem.
Censored Planet has a track-record of perseverance, pragmatism, and broad collaboration—attributes that have helped us achieve positive impacts within the complex landscape of Internet Freedom research.
Conducted
68
billion
measurements across over 220 countries si

#### https://censoredplanet.org/#/tools   HTTP (hash navigation, no response)
dates visible on the rendered page: (none)
--- rendered text, first 2000 characters ---
About
Research
Blogs
Events
Publications
Tools
News
Talks
Join Us
Tools
Censored Planet develops and maintains a suite of open source tools for investigating internet censorship and network interference.
Censored Planet Observatory
A global-scale measurement platform that monitors internet censorship and network interference.
VPNalyzer
A tool that enables systematic, semi-automated investigation of the security and effectiveness of VPNs, revealing vulnerabilities and guiding better design practices.
GeoInspector
Measurement tool for detecting geoblocking and server-side blocking across DNS, TCP, TLS, and HTTP protocols.
Chinese LLM Censorship
Data from analysis of overt blocking embedded in Chinese LLM services, including query lists and blocking fingerprints.
CenFuzz
Tool for fuzzing HTTP/HTTPS requests to uncover censorship rules and trigger conditions.
CenTrace
Application-layer censorship traceroute tool using TTL-limited HTTP and TLS packets to locate censorship infrastructure.
CenProbe
Set of scripts for actively probing and collecting data about particular network devices, including those that perform censorship.
ZeroTrace
Go package implementing the 0trace technique to estimate network-layer RTT using TTL-limited segments in existing TCP connections.
CenAnalysis
Pipeline for transforming raw Censored Planet Observatory data into structured BigQuery tables.
AppMap
Scripts for collecting Google Play metadata, APKs, and associated privacy policies.
DNSBlockpage
Tools for generating DNS blockpage fingerprints and analyzing Satellite v2 censorship data.
Do you have a question?
Get in touch
4908 Bob and Betty Beyster Building
2260 Hayward St.,
Ann Arbor, MI 48109
ensafi@umich.edu
About
Blog
Publications
Tools
Press
Observatory
Dashboard
Raw Data
Documentation
Social media
© 2025 Censored Planet
Terms of Service
big counters as displayed: 116.924.640.072 | 744.324.571
same, with ASCII commas   : 116,924,640,072 | 744,324,571

#### https://dashboard.censoredplanet.org/   HTTP 200
dates visible on the rendered page: (none)
--- rendered text, first 900 characters ---
Censored Planet Dashboard
📏 Total Measurements
116.924.640.072
🌐 Countries
236
Measurements last 30 days
744.324.571
Explore Censored Planet Data ↗︎
Interference Rate Last 30 Days i
Iran
17.98
China
4.23
Russia
2.97
Cambodia
2.45
Cyprus
2.25
Afghanistan
2.09
Tanzania
2.03
Myanmar
1.99
Iraq
1.82
Pakistan
1.78
Turkmenistan
1.69
United Arab Emirates
1.68
Guyana
1.67
Philippines
1.62
Nepal
1.58
Jersey
1.56
Kazakhstan
1.56
Kuwait
1.53
New Caledonia
1.52
Ivory Coast
1.51
Cayman Islands
1.51
Morocco
1.49
Oman
1.46
Uzbekistan
1.46
Honduras
1.45
Mali
1.44
Laos
1.44
Slovenia
1.44
Latvia
1.42
Namibia
1.41
Djibouti
1.41
Aruba
1.40
Niue
1.40
Equatorial Guinea
1.40
U.S. Virgin Islands
1.40
South Sudan
1.40
São Tomé and Príncipe
1.40
Solomon Islands
1.40
Cuba
1.40
Papua New Guinea
1.40
Syria
1.40
Guatemala
1.40
Liberia
1.40
Chad
1.40
Somalia
1.40
Madagascar
1.40
Antigua and Barbuda
1.40
Guernsey
1.40

#### https://gfwatch.org/   HTTP 200
dates visible on the rendered page: March 2020 | Aug 112024 | 2020/03/20 | 2024/08/06 | 2024/05/23
--- rendered text, first 1800 characters ---
OverviewCensored domainsFake IP addressesPublications
GFWatch Dashboard
GFWatch is a measurement platform capable of testing hundreds of millions of domains daily, enabling the continuous monitoring of the Great Firewall's DNS filtering behavior. The newly censored domains discovered by GFWatch provide a useful insight into China's information control policies. This project is a result of an academic collaboration between researchers from Stony Brook University, University of Massachusetts - Amherst, University of California - Berkeley, and the Citizen Lab at the University of Toronto.
This project is a sibling project of the GFWeb project, which monitors the Great Firewall's HTTP(S) censorship. Data from GFWatch can be found at https://GFWeb.ca.
Since its launch in March 2020, GFWatch has discovered:
DNS base censored domains	669,508
Newly blocked DNS base domains (last 30 days)	2,430
Fake IPv4 Addresses	2,057
Fake IPv6 Addresses	4,472
Number of censored domains observed daily within the last 30 days
Aug 112024
Aug 18
Aug 25
Sep 1
164k
166k
168k
170k
172k
Censored domains
Base censored domains
Top classified categories of base censored domains within the last 30 days
file sharing/storage
adult content
malicious
info.tech
business
proxy/anonymizers
news/media
entertainment
parked domains
gambling
0
5k
10k
15k
Categories classified via Virus Total
# base censored domains
Top notable newly blocked domains
Tranco_rank
	
base_censored_domain
	
first_checked
	
last_checked
1	google.com	2020/03/20	2024/08/06
5	facebook.com	2020/03/20	2024/08/06
8	youtube.com	2020/03/20	2024/08/06
12	twitter.com	2020/03/20	2024/08/06
15	instagram.com	2020/03/20	2024/08/06
23	tiktokcdn.com	2024/05/23	2024/08/06
24	googlevideo.com	2020/03/20	2024/08/06
34	wikipedia.org	2020/03/20	2024/08/06
54	p

#### https://gfweb.ca/   HTTP 200
dates visible on the rendered page: Feb 2022 | Jul 2024 | Sep 2024 | Nov 2024 | 2024/09/24 | 2024/11/19 | 2024/10/10 | 2024/10/31 | 2024/10/15 | 2024/09/09 | 2024/09/10 | 2024/11/16 | 2024/11/18
--- rendered text, first 1800 characters ---
HTTP blocking overviewHTTPS blocking overviewHTTP censored domainsHTTPS censored domainsPublications
GFWeb Dashboard
GFWeb is a measurement platform capable of testing hundreds of millions of domains monthly, enabling the continuous monitoring of the Great Firewall's HTTP(S) filtering behavior. The newly censored domains discovered by GFWeb provide a useful insight into China's information control policies. This project is a result of an academic collaboration between researchers from the University of British Columbia, University of Chicago, the Citizen Lab at University of Toronto, Carnegie Mellon University, SRI International, and Stony Brook University.
This project is a sibling project of the GFWatch project, which monitors the Great Firewall's DNS censorship. Data from GFWatch can be found at https://GFWatch.org.
Since its launch in Feb 2022, GFWeb has discovered:
HTTP base censored domains	966,842
Newly blocked HTTP base domains (last 90 days)	2,028
Number of censored domains observed monthly within the last 6 months
Jul 2024
Sep 2024
Nov 2024
300k
320k
340k
Censored domains
Base censored domains
Top classified categories of base censored domains within the last 3 months
adult content
file sharing/storage
parked domains
business
malicious
proxy/anonymizers
gambling
info.tech
news/media
entertainment
0
5k
10k
15k
Categories classified via Virus Total
# base censored domains
Top notable newly blocked domains
Tranco_rank
	
base_censored_domain
	
first_checked
	
last_checked
3046	docker.io	2024/09/24	2024/11/19
9454	audacy.com	2024/10/10	2024/11/19
10294	spicychat.ai	2024/10/31	2024/11/19
21323	start.me	2024/10/15	2024/11/19
26224	opensocietyfoundations.org	2024/10/15	2024/11/19
31558	youjizz.sex	2024/09/09	2024/11/19
35049	desiringgod.org	2024/09/10	2024/11/19
35553

#### https://ioda.inetintel.cc.gatech.edu/dashboard   HTTP 200
dates visible on the rendered page: Sep 9, 2026 | Sep 10, 2026 | August 13, 2026 | June 18, 2026 | June 10, 2026 | April 27, 2026 | June 2025 | April 6, 2026
--- rendered text, first 900 characters ---
Dashboard
API
About
Reports
Resources
English
Simple
Country
All Countries
Region
All Regions
Networks
All Networks
Select a Time Range
Country View
Region View
ASN/ISP View
Country Outages
+
-
Leaflet
Outage Severity Score:
1K
7M
15M
29M
Max
Sep 9, 2026 11:12pm UTC - Sep 10, 2026 11:12pm UTC
News
IODA
@eldomador.bsky.social
August 13, 2026
Kamchatka outage captured in @ioda.live: ioda.inetintel.cc.gatech.edu/region/3492?...
❤️ 5💬 0
View on Bluesky →
IODA
@ioda.live
June 18, 2026
The ongoing Internet shutdown in Pakistan's Azad Jammu and Kashmir is nearing 2-week duration. The shutdown was put into place after protests related to electoral representation for Pakistan-administered Kashmir. Follow connectivity in near realtime: ioda.inetintel.cc.gatech.edu/region/3083?...
❤️ 4💬 0
View on Bluesky →
IODA
@ioda.live
June 10, 2026
Hey, David, can you email us at ioda-info@cc.gatech.edu
❤️ 0�
  • corpus — venue scope, the selection funnel, and the provisional 2025–2026 years
  • provenance:design:crawling_location does not exist (deliberately not created here — that page's figures were not touched by this run); the namespace's own log is design
  • roadmap / roadmap — where this page's id and queued estimate came from
[1]
Wu, Mingshi; Sippe, Jackson; Sivakumar, Danesh; Burg, Jack; Anderson, Peter; Wang, Xiaokang; Bock, Kevin; Houmansadr, Amir; Levin, Dave; Wustrow, Eric (2023): "How the Great Firewall of China Detects and Blocks Fully Encrypted Traffic", in: Proceedings of the USENIX Security Symposium. (Link)
[2]
Lu, Chaoyi; Liu, Baojun; Li, Zhou; Hao, Shuang; Duan, Hai-Xin; Zhang, Mingming; Leng, Chunying; Liu, Ying; Zhang, Zaifeng; Wu, Jianping (2019): "An End-to-End, Large-Scale Measurement of DNS-over-Encryption: How Far Have We Come?", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[3]
Sokoto, Saidu; Balduf, Leonhard; Trautwein, Dennis; Wei, Yiluo; Tyson, Gareth; Castro, Ignacio; Ascigil, Onur; Pavlou, George; Korczyński, Maciej; Scheuermann, Björn; Król, Michał (2024): "Guardians of the Galaxy: Content Moderation in the InterPlanetary File System", in: Proceedings of the USENIX Security Symposium. (Link)
[4]
Ramesh, Reethika; Raman, Ram Sundara; Virkud, Apurva; Dirksen, Alexandra; Huremagic, Armin; Fifield, David; Rodenburg, Dirk; Hynes, Rod; Madory, Doug; Ensafi, Roya (2023): "Network Responses to Russia's Invasion of Ukraine in 2022: A Cautionary Tale for Internet Freedom", in: Proceedings of the USENIX Security Symposium. (Link)
[5]
Knockel, Jeffrey; Dałek, Jakub; Aljizawi, Noura; Ahmed, Mohamed; Meletti, Levi; Lau, Justin (2026): "Banned Books: Analysis of Censorship on Amazon.com", Proceedings on Privacy Enhancing Technologies 2026(3):200-214. (DOI)
[6]
Lipphardt, Friedemann; Ali, Moonis; Banzer, Martin; Feldmann, Anja; Gosain, Devashish (2026): "There is No War in Ba Sing Se: A Global Analysis of Content Moderation in Large Language Models", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[7]
Shezan, Faysal Hossain; Su, Zihao; Kang, Mingqing; Phair, Nicholas; Thomas, Patrick William; van Dam, Michelangelo; Cao, Yinzhi; Tian, Yuan (2023): "CHKPLUG: Checking GDPR Compliance of WordPress Plugins via Cross-language Code Property Graph", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[8]
Haghighi, Sara; LaChance, Clark; Pourghasemi Fatideh, Ali; Breaux, Travis; Ghanavati, Sepideh (2026): "The Role of Online Forums in Developer Understanding of Privacy Law - A Reddit Case Study", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[9]
McDonald, Allison; Bernhard, Matthew; Valenta, Luke; VanderSloot, Benjamin; Scott, Will; Sullivan, Nick; Halderman, J. Alex; Ensafi, Roya (2018): "403 Forbidden: A Global View of CDN Geoblocking", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[10]
Kumar, Renuka; Virkud, Apurva; Sundara Raman, Ram; Prakash, Atul; Ensafi, Roya (2022): "A Large-scale Investigation into Geodifferences in Mobile Apps", in: Proceedings of the USENIX Security Symposium. (Link)
[11]
Ablove, Anna; Chandrashekaran, Shreyas; Qiang, Xiao; Ensafi, Roya (2026): "Characterizing the Implementation of Censorship Policies in Chinese LLM Services", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[12]
Tsai, Elisa; Kumar, Deepak; Sundara Raman, Ram; Li, Gavin; Eiger, Yael; Ensafi, Roya (2023): "CERTainty: Detecting DNS Manipulation at Scale using TLS Certificates", Proceedings on Privacy Enhancing Technologies 2023(3):122-137. (DOI)
[13]
Jones, Ben; Lee, Tzu-Wen; Feamster, Nick; Gill, Phillipa (2014): "Automated Detection and Fingerprinting of Censorship Block Pages", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
provenance/design/blocking_and_geodifference.txt · Last modified: by karel.kubicek.claude