User Tools

Site Tools


design:crawling_location

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
design:crawling_location [2026/08/12 09:06] – Refresh all corpus figures for the extended 2010-2026 extraction (4,322 -> 5,859 papers; crawled population 859 -> 1,120). New provider-family fold (scripts/vantage_fold.mjs) replaces the hand-fold; 2025-2026 column added and marked provisional; IEEE S&P karel.kubicek.claudedesign:crawling_location [2026/09/17 10:45] (current) – ConsentAction audit propagation and EU/EEA figures; Authored by Claude karel.kubicek.claude
Line 1: Line 1:
 ====== Crawling Location ====== ====== Crawling Location ======
  
-The //vantage point// of a measurement is the network position your traffic appears to originate from: its IP address, and everything a website can infer from it — country, city, network type, and whether it looks like a person or a datacenter. It is chosen at least implicitly by every study that touches the live web, and it is the design decision least often reported: of the crawling papers in our corpus of seven major security and privacy venues, **70.9% never state where they measured from** (see [[#Use in Publications]]).+The //vantage point// of a measurement is the network position your traffic appears to originate from: its IP address, and everything a website can infer from it — country, city, network type, and whether it looks like a person or a datacenter. It is chosen at least implicitly by every study that touches the live web, and it is the design decision least often reported: of the crawling papers in our corpus of seven major security and privacy venues, **70.8record a vantage point and never say where it was** (see [[#Use in Publications]]).
  
 That silence matters because the vantage point changes three separable things, and only the first is obvious: That silence matters because the vantage point changes three separable things, and only the first is obvious:
Line 9: Line 9:
   - **How your IP is treated.** Datacenter, university, Tor and residential addresses receive measurably different treatment from bot management and from trackers.   - **How your IP is treated.** Datacenter, university, Tor and residential addresses receive measurably different treatment from bot management and from trackers.
  
-This page covers all three, how researchers have actually chosen vantage points, the practical options, and how to verify that the vantage point you think you have is the one you got. It pairs with [[Design:IP classification]] (classifying //other people's// addresses), [[Design:Website selection]] (which sites), and [[Programming:Crawler]] (which tool).+This page covers all three, how researchers have actually chosen vantage points, the practical options, and how to verify that the vantage point you think you have is the one you got. It pairs with [[Design:IP classification]] (classifying //other people's// addresses), [[Design:Website selection]] (which sites), and [[Programming:Crawler]] (which tool). [[Privacy:Age assurance]] is the sharpest current case of the first bullet: since 2025 several jurisdictions //require// a site to behave differently for visitors they believe are theirs, so an age-gate prevalence figure without a stated vantage point is not a quantity.
  
 ===== Why the Vantage Point Changes Your Results ===== ===== Why the Vantage Point Changes Your Results =====
Line 27: Line 27:
 ==== Content: geoblocking and geo-differentiation ==== ==== Content: geoblocking and geo-differentiation ====
  
-Even setting privacy law aside, the web is not the same everywhere:+Even setting privacy law aside, the web is not the same everywhere, and a crawl from the wrong place records an absence that is a property of **your vantage point** rather than of the site:
  
-  * **Outright geoblocking.** McDonald et al. {[mcdonald2018_forbidden]} measured CDN-level geoblocking from 177 countries, finding server-side blocking of entire regions. A crawl from a blocked country records an absence that is a property of your vantage point, not of the site's tracking behaviour. +  * **Outright geoblocking and geo-differentiated content.** McDonald et al. {[mcdonald2018_forbidden]} measured CDN-level geoblocking from 177 countries, and Kumar et al. {[kumar2022_investigation]} found systematic differences in availability and behaviour across 26 countries chosen "to have reliable direct vantage points". Both are the subject of **[[Design:Blocking and geodifference]]**, which owns the phenomenon, the instruments, and what a blocking claim has to contain — including the error rates these detectors have been measured to have. Read it if your vantage point may not be allowed to see what you are counting; this page stops at what that means for //choosing// the vantage point
-  * **Geo-differentiated content.** Kumar et al. {[kumar2022_investigation]} compared mobile apps across 26 countries chosen "to have reliable direct vantage points", finding systematic differences in availability and behaviour+  * **Personalisation.** Kliman-Silver et al. {[klimansilver2015_location]} showed that geolocation drives measurable web-search personalisation, so location is a confound in any study of ranked or targeted output — a difference in the //ordering// of what you are served rather than in whether you are served at all.
-  * **Personalisation.** Kliman-Silver et al. {[klimansilver2015_location]} showed that geolocation drives measurable web-search personalisation, so location is a confound in any study of ranked or targeted output.+
  
 ==== Treatment: what your IP says about you ==== ==== Treatment: what your IP says about you ====
Line 127: Line 126:
 | EU/EEA vantage //(of those stating)// | 19.0% | 25.0% | 38.6% | 46.3% | 59.6% | | EU/EEA vantage //(of those stating)// | 19.0% | 25.0% | 38.6% | 46.3% | 59.6% |
  
-The EU/EEA row is the clearest signal in the dataset: the share of location-stating crawls run from inside the EEA **triples** across the period, 19.0% to 59.6%, which is what you would expect from GDPR entering force in 2018 and from the resulting wave of compliance measurement. Location reporting itself improves far less — from 20.6% to 28.8% over sixteen years. Read the last column with care: CCS and IMC 2026 have not been held and two more 2026 venue-years are incompletely selected, so it rests on fewer papers than the 198 count suggests it should.+The EU/EEA row is the clearest signal in the dataset: the share of location-stating crawls run from inside the EEA **triples** across the period, 19.0% to 59.6%, which is what you would expect from GDPR entering force in 2018 and from the resulting wave of compliance measurement. Location reporting itself improves far less — from 20.6% to 28.8% across the corpus 2010-2026. Read the last column with care: CCS and IMC 2026 have not been held and two more 2026 venue-years are incompletely selected, so it rests on fewer papers than the 198 count suggests it should.
  
 ==== Legal framing does predict the vantage point ==== ==== Legal framing does predict the vantage point ====
Line 141: Line 140:
  
 <WRAP important> <WRAP important>
-The gap is worse where it matters most. Of the **349 papers that state what their crawler did with the consent banner** — accept, reject, a CMP-specific choice, or explicitly no interaction — only **68 (19.5%) state an EU/EEA vantage, and 210 give no location at all.** Consent behaviour is the single most geo-dependent thing on the web, and the majority of papers interacting with it do not report the jurisdiction they observed it from.+The gap is worse where it matters most. Of the **55 papers whose audited full-text record states what their crawler did with the consent banner** — accept, reject, a CMP-specific choice, or explicitly no interaction — **37 (67.3%) state an EU/EEA vantage, and 11 state no location at all.** Consent behaviour is the single most geo-dependent thing on the web, and the jurisdiction is still missing for a substantial minority of those audited paper claims.
  
-Both figures come from the same per-paper record, so the 349 are a subset of the 1,120 crawling papers rather than a separately sampled group: the consent action and the vantage point were extracted in one pass from the same full text, and a paper counts here only if it stated its consent action explicitly.+Both figures use the same per-paper record, but the consent denominator is the audited one, not the raw schema count: the 2026-09-05 audit read all 349 non-sentinel ''crawlConfig.consentAction'' labels and rejected 279 of the 313 ''no-interaction'' defaults. The resulting 55 are a subset of the 1,120 crawling papers, and a paper counts here only if the audit found a consent action claim in its full text.
 </WRAP> </WRAP>
  
Line 149: Line 148:
  
   * "This crawl was made from France on September 20th and 21st 2019." {[matte2020_cookie]}   * "This crawl was made from France on September 20th and 21st 2019." {[matte2020_cookie]}
-  * "We crawl websites using 30 German datacenter IP addresses provided by The Bright Initiative from Bright Data." {[bouhoula2024automated]}+  * "We crawl websites using 30 German datacenter IP addresses provided by The Bright Initiative from Bright Data." {[bouhoula2024_automated]}
   * "all but one use an IP address associated with a European server (Frankfurt am Main; DEU) … the other is run from a US IP address (Council Bluffs, IA; USA)." {[demir2024_bannertools]}   * "all but one use an IP address associated with a European server (Frankfurt am Main; DEU) … the other is run from a US IP address (Council Bluffs, IA; USA)." {[demir2024_bannertools]}
   * "We choose three geolocations for our measurement: (1) Germany (EU), (2) Japan (AS), and (3) the United States (NA)." — via ProtonVPN {[demir2022_reproducibility]}   * "We choose three geolocations for our measurement: (1) Germany (EU), (2) Japan (AS), and (3) the United States (NA)." — via ProtonVPN {[demir2022_reproducibility]}
Line 159: Line 158:
   * **How they were produced.** One structured record per paper was extracted from full text, each tuple carrying a verbatim evidence quote and its section, so any figure here can be traced to the sentence that supports it. Locations are free text ("Frankfurt, Germany", "US-East", "61 countries across all world regions") and were normalised to countries or regions before counting by ''scripts/geo.mjs''; 15 distinct strings across 13 papers could not be mapped and are excluded, and 5 of those papers are left with no place at all. The full residue list is on the [[provenance:design:crawling_location|provenance page]]. Papers, never tuples, are counted.   * **How they were produced.** One structured record per paper was extracted from full text, each tuple carrying a verbatim evidence quote and its section, so any figure here can be traced to the sentence that supports it. Locations are free text ("Frankfurt, Germany", "US-East", "61 countries across all world regions") and were normalised to countries or regions before counting by ''scripts/geo.mjs''; 15 distinct strings across 13 papers could not be mapped and are excluded, and 5 of those papers are left with no place at all. The full residue list is on the [[provenance:design:crawling_location|provenance page]]. Papers, never tuples, are counted.
   * **Silence is not absence.** "Does not state a location" means the paper did not say, not that the authors did not know. These are reporting figures.   * **Silence is not absence.** "Does not state a location" means the paper did not say, not that the authors did not know. These are reporting figures.
-  * **Venue coverage.** EuroS&PACSAC, RAID, AsiaCCS, CHI and SOUPS are absent entirely; any claim here is a claim about seven venues. The 2025 and 2026 venue-years are incomplete for reasons of calendar and indexing rather than relevance, so per-year trends should be read as ending in 2024.+  * **Venue coverage.** Seven venues only, and 20252026 are incomplete for reasons of calendar and indexing rather than relevance, so per-year trends should be read as ending in 2024. Which venues, which years and what each stage of the selection funnel costs are on [[literature:corpus]].
   * **Field stability.** The stated/not-stated distinction and the infrastructure enum are reproducible to within a few points on a repeat extraction. Free-text service names are much less stable and are reported folded into families and as rankings, not as precise figures. The ''Amazon AWS / EC2'' row, for instance, merges at least five spellings.   * **Field stability.** The stated/not-stated distinction and the infrastructure enum are reproducible to within a few points on a repeat extraction. Free-text service names are much less stable and are reported folded into families and as rankings, not as precise figures. The ''Amazon AWS / EC2'' row, for instance, merges at least five spellings.
   * **''crawled'' is defined** as a paper whose crawl configuration was recorded or whose study types include an automated web crawl (1,120 papers, 19.1% of the corpus). This corpus is seven broad security venues, not a web-measurement corpus, so shares of all 5,859 papers would be meaningless here.   * **''crawled'' is defined** as a paper whose crawl configuration was recorded or whose study types include an automated web crawl (1,120 papers, 19.1% of the corpus). This corpus is seven broad security venues, not a web-measurement corpus, so shares of all 5,859 papers would be meaningless here.
Line 189: Line 188:
 The risk is current, not historical. On **2 July 2026** the FBI, with the IRS Criminal Investigation division, Google's Threat Intelligence Group and Lumen, [[https://krebsonsecurity.com/2026/07/fbi-seizes-netnut-proxy-platform-popa-botnet/|seized hundreds of domains belonging to NetNut]], a major residential proxy provider operated by Alarum Technologies, over an alleged overlap between its exit-node pool and the "Popa" botnet of at least two million compromised devices. Alarum [[https://alarum.io/alarum-technologies-responds-to-inquiry-into-residential-proxy-networks/|disputes the allegations]], stated it had not been formally contacted by the FBI as of 3 July 2026, and paused parts of the network. Treat the case as unresolved — but note that a researcher who had bought bandwidth from that pool would now be explaining it to their ethics board. The risk is current, not historical. On **2 July 2026** the FBI, with the IRS Criminal Investigation division, Google's Threat Intelligence Group and Lumen, [[https://krebsonsecurity.com/2026/07/fbi-seizes-netnut-proxy-platform-popa-botnet/|seized hundreds of domains belonging to NetNut]], a major residential proxy provider operated by Alarum Technologies, over an alleged overlap between its exit-node pool and the "Popa" botnet of at least two million compromised devices. Alarum [[https://alarum.io/alarum-technologies-responds-to-inquiry-into-residential-proxy-networks/|disputes the allegations]], stated it had not been formally contacted by the FBI as of 3 July 2026, and paused parts of the network. Treat the case as unresolved — but note that a researcher who had bought bandwidth from that pool would now be explaining it to their ethics board.
  
-If you use one, we suggest: name the provider in the paper, state what the provider claims about consent, say whether your IRB or ethics board reviewed that claim specifically, and prefer providers used by prior peer-reviewed work. Bright Data's research arm ("The Bright Initiative") is the route taken by {[bouhoula2024automated]} — and note that they used **datacenter** IPs from it, not residential ones. See [[Practices:Ethics]].+If you use one, we suggest: name the provider in the paper, state what the provider claims about consent, say whether your IRB or ethics board reviewed that claim specifically, and prefer providers used by prior peer-reviewed work. Bright Data's research arm ("The Bright Initiative") is the route taken by {[bouhoula2024_automated]} — and note that they used **datacenter** IPs from it, not residential ones. See [[Practices:Ethics]].
 </WRAP> </WRAP>
  
 ==== Research measurement platforms ==== ==== Research measurement platforms ====
  
-[[https://atlas.ripe.net/|RIPE Atlas]] (probes and anchors worldwide, credit-based), [[https://www.measurementlab.net/|M-Lab]] and [[https://www.caida.org/projects/ark/|CAIDA Ark]] give wide, citable, reproducible geographic coverage. The catch is layer: they are built for network measurement, not for driving a browser, so they suit DNS, reachability and latency questions rather than tracking or consent. **PlanetLab is discontinued** — it appears in papers in the corpus and is not an option for new work; the [[https://www.edge-net.org/|EdgeNet]] project is the nearest successor.+[[https://atlas.ripe.net/|RIPE Atlas]] (probes and anchors worldwide, credit-based), [[https://www.measurementlab.net/|M-Lab]] and [[https://www.caida.org/projects/ark/|CAIDA Ark]] give wide, citable, reproducible geographic coverage. The catch is layer: they are built for network measurement, not for driving a browser, so they suit DNS, reachability and latency questions rather than tracking or consent. **PlanetLab is discontinued** — it appears in papers in the corpus and is not an option for new work; the [[https://www.edge-net.org/|EdgeNet]] project is the nearest successor.
  
 ===== Verify the Vantage Point ===== ===== Verify the Vantage Point =====
Line 203: Line 202:
   * **Geolocation databases disagree with each other.** They agree well at country level and poorly below it, which is the subject of [[Design:IP classification]].   * **Geolocation databases disagree with each other.** They agree well at country level and poorly below it, which is the subject of [[Design:IP classification]].
  
-The following script checks both. It queries several free services for your egress IP, reports whether they agree, flags datacenter/VPN/proxy addresses, and exits non-zero if the country is not the one you expected — so it can gate a crawl rather than merely inform you. Free tiers are rate-limited, so call it once when a vantage point comes up and again when it goes down, not per request. Adding a ''--json'' branch that dumps the same report as a dict is a two-line change if you want to log it alongside the crawl.+The following script checks both. It queries several free services for your egress IP, reports whether they agree, flags datacenter/VPN/proxy addresses, and exits non-zero if the country is not the one you expected — so it can gate a crawl rather than merely inform you. Free tiers are rate-limited, so call it once when a vantage point comes up and again when it goes down, not per request. Adding a ''%%--json%%'' branch that dumps the same report as a dict is a two-line change if you want to log it alongside the crawl.
  
 <file python verify_vantage.py> <file python verify_vantage.py>
Line 332: Line 331:
 ===== Open Questions ===== ===== Open Questions =====
  
-  * <wrap todo>No peer-reviewed cross-vendor measurement of how much datacenter IP address space is penalised by bot-management vendors. Jueckstock et al. {[jueckstock2021_realistic]} measure the effect on privacy metrics but not the mechanism per vendor.</wrap> +<WRAP todo> 
-  * <wrap todo>How stable is CMP geo-targeting configuration over time? All the industry documentation describes the feature; nobody appears to have measured how often operators change the region-to-template binding.</wrap> +  * No peer-reviewed cross-vendor measurement of how much datacenter IP address space is penalised by bot-management vendors. Jueckstock et al. {[jueckstock2021_realistic]} measure the effect on privacy metrics but not the mechanism per vendor. 
-  * <wrap todo>Whether commercial VPN mislabelling has been re-measured academically since {[weinberg2018_catch]} (2018). The only recent figures we found are a vendor study.</wrap> +  * How stable is CMP geo-targeting configuration over time? All the industry documentation describes the feature; nobody appears to have measured how often operators change the region-to-template binding. 
-  * <wrap todo>Residential proxy pool overlap with known botnets, post-NetNut. Spur, Synthient and Nokia Deepfield published attributions in June 2026; a systematic academic treatment would be valuable for ethics review.</wrap>+  * Whether commercial VPN mislabelling has been re-measured academically since {[weinberg2018_catch]} (2018). The only recent figures we found are a vendor study. 
 +  * Residential proxy pool overlap with known botnets, post-NetNut. Spur, Synthient and Nokia Deepfield published attributions in June 2026; a systematic academic treatment would be valuable for ethics review. 
 +</WRAP>
  
 ===== Related Pages ===== ===== Related Pages =====
  
 +  * [[Design:Blocking and geodifference]] — what your vantage point was **not allowed to see**: geoblocking, GDPR walls, censorship and network interference, and how a blocking claim is made.
   * [[Design:IP classification]] — classifying the addresses you observe, and the geolocation databases this page's script exercises.   * [[Design:IP classification]] — classifying the addresses you observe, and the geolocation databases this page's script exercises.
   * [[Design:Website selection]] — country-specific top lists (CrUX has country breakdowns; SecRank is Chinese-DNS-based) interact with vantage choice.   * [[Design:Website selection]] — country-specific top lists (CrUX has country breakdowns; SecRank is Chinese-DNS-based) interact with vantage choice.
   * [[Design:Archives]] — web archives sidestep the vantage point and introduce their own biases.   * [[Design:Archives]] — web archives sidestep the vantage point and introduce their own biases.
   * [[Programming:Crawler]] — proxy and per-context network configuration per crawling library.   * [[Programming:Crawler]] — proxy and per-context network configuration per crawling library.
-  * [[Programming:Stateful stateless]] — the other axis Jueckstock et al. {[jueckstock2021_realistic]} vary.+  * [[Programming:Stateful stateless]] — the design choice that decides what a crawl can observe at all. Note that Jueckstock et al. {[jueckstock2021_realistic]} do //not// vary it: every crawl in that paper launches "with a clean user profile (i.e., no cookies or cached content)". The second axis they vary is the browser configuration, naive against stealth.
   * [[Privacy:Consent]] — what to do with the banner once you are in the right jurisdiction.   * [[Privacy:Consent]] — what to do with the banner once you are in the right jurisdiction.
   * [[Practices:Ethics]] — residential proxies, volunteer devices, and acceptable-use policies.   * [[Practices:Ethics]] — residential proxies, volunteer devices, and acceptable-use policies.
design/crawling_location.1786525614.txt.gz · Last modified: by karel.kubicek.claude