User Tools

Site Tools


design:crawling_location

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
design:crawling_location [2026/08/12 16:06] – Fix 'over sixteen years' — the corpus spans 2010-2026, i.e. seventeen. Authored by Claude. karel.kubicek.claudedesign:crawling_location [2026/09/17 10:45] (current) – ConsentAction audit propagation and EU/EEA figures; Authored by Claude karel.kubicek.claude
Line 9: Line 9:
   - **How your IP is treated.** Datacenter, university, Tor and residential addresses receive measurably different treatment from bot management and from trackers.   - **How your IP is treated.** Datacenter, university, Tor and residential addresses receive measurably different treatment from bot management and from trackers.
  
-This page covers all three, how researchers have actually chosen vantage points, the practical options, and how to verify that the vantage point you think you have is the one you got. It pairs with [[Design:IP classification]] (classifying //other people's// addresses), [[Design:Website selection]] (which sites), and [[Programming:Crawler]] (which tool).+This page covers all three, how researchers have actually chosen vantage points, the practical options, and how to verify that the vantage point you think you have is the one you got. It pairs with [[Design:IP classification]] (classifying //other people's// addresses), [[Design:Website selection]] (which sites), and [[Programming:Crawler]] (which tool). [[Privacy:Age assurance]] is the sharpest current case of the first bullet: since 2025 several jurisdictions //require// a site to behave differently for visitors they believe are theirs, so an age-gate prevalence figure without a stated vantage point is not a quantity.
  
 ===== Why the Vantage Point Changes Your Results ===== ===== Why the Vantage Point Changes Your Results =====
Line 27: Line 27:
 ==== Content: geoblocking and geo-differentiation ==== ==== Content: geoblocking and geo-differentiation ====
  
-Even setting privacy law aside, the web is not the same everywhere:+Even setting privacy law aside, the web is not the same everywhere, and a crawl from the wrong place records an absence that is a property of **your vantage point** rather than of the site:
  
-  * **Outright geoblocking.** McDonald et al. {[mcdonald2018_forbidden]} measured CDN-level geoblocking from 177 countries, finding server-side blocking of entire regions. A crawl from a blocked country records an absence that is a property of your vantage point, not of the site's tracking behaviour. +  * **Outright geoblocking and geo-differentiated content.** McDonald et al. {[mcdonald2018_forbidden]} measured CDN-level geoblocking from 177 countries, and Kumar et al. {[kumar2022_investigation]} found systematic differences in availability and behaviour across 26 countries chosen "to have reliable direct vantage points". Both are the subject of **[[Design:Blocking and geodifference]]**, which owns the phenomenon, the instruments, and what a blocking claim has to contain — including the error rates these detectors have been measured to have. Read it if your vantage point may not be allowed to see what you are counting; this page stops at what that means for //choosing// the vantage point
-  * **Geo-differentiated content.** Kumar et al. {[kumar2022_investigation]} compared mobile apps across 26 countries chosen "to have reliable direct vantage points", finding systematic differences in availability and behaviour+  * **Personalisation.** Kliman-Silver et al. {[klimansilver2015_location]} showed that geolocation drives measurable web-search personalisation, so location is a confound in any study of ranked or targeted output — a difference in the //ordering// of what you are served rather than in whether you are served at all.
-  * **Personalisation.** Kliman-Silver et al. {[klimansilver2015_location]} showed that geolocation drives measurable web-search personalisation, so location is a confound in any study of ranked or targeted output.+
  
 ==== Treatment: what your IP says about you ==== ==== Treatment: what your IP says about you ====
Line 141: Line 140:
  
 <WRAP important> <WRAP important>
-The gap is worse where it matters most. Of the **349 papers that state what their crawler did with the consent banner** — accept, reject, a CMP-specific choice, or explicitly no interaction — only **68 (19.5%) state an EU/EEA vantage, and 210 give no location at all.** Consent behaviour is the single most geo-dependent thing on the web, and the majority of papers interacting with it do not report the jurisdiction they observed it from.+The gap is worse where it matters most. Of the **55 papers whose audited full-text record states what their crawler did with the consent banner** — accept, reject, a CMP-specific choice, or explicitly no interaction — **37 (67.3%) state an EU/EEA vantage, and 11 state no location at all.** Consent behaviour is the single most geo-dependent thing on the web, and the jurisdiction is still missing for a substantial minority of those audited paper claims.
  
-Both figures come from the same per-paper record, so the 349 are a subset of the 1,120 crawling papers rather than a separately sampled group: the consent action and the vantage point were extracted in one pass from the same full text, and a paper counts here only if it stated its consent action explicitly.+Both figures use the same per-paper record, but the consent denominator is the audited one, not the raw schema count: the 2026-09-05 audit read all 349 non-sentinel ''crawlConfig.consentAction'' labels and rejected 279 of the 313 ''no-interaction'' defaults. The resulting 55 are a subset of the 1,120 crawling papers, and a paper counts here only if the audit found a consent action claim in its full text.
 </WRAP> </WRAP>
  
Line 149: Line 148:
  
   * "This crawl was made from France on September 20th and 21st 2019." {[matte2020_cookie]}   * "This crawl was made from France on September 20th and 21st 2019." {[matte2020_cookie]}
-  * "We crawl websites using 30 German datacenter IP addresses provided by The Bright Initiative from Bright Data." {[bouhoula2024automated]}+  * "We crawl websites using 30 German datacenter IP addresses provided by The Bright Initiative from Bright Data." {[bouhoula2024_automated]}
   * "all but one use an IP address associated with a European server (Frankfurt am Main; DEU) … the other is run from a US IP address (Council Bluffs, IA; USA)." {[demir2024_bannertools]}   * "all but one use an IP address associated with a European server (Frankfurt am Main; DEU) … the other is run from a US IP address (Council Bluffs, IA; USA)." {[demir2024_bannertools]}
   * "We choose three geolocations for our measurement: (1) Germany (EU), (2) Japan (AS), and (3) the United States (NA)." — via ProtonVPN {[demir2022_reproducibility]}   * "We choose three geolocations for our measurement: (1) Germany (EU), (2) Japan (AS), and (3) the United States (NA)." — via ProtonVPN {[demir2022_reproducibility]}
Line 189: Line 188:
 The risk is current, not historical. On **2 July 2026** the FBI, with the IRS Criminal Investigation division, Google's Threat Intelligence Group and Lumen, [[https://krebsonsecurity.com/2026/07/fbi-seizes-netnut-proxy-platform-popa-botnet/|seized hundreds of domains belonging to NetNut]], a major residential proxy provider operated by Alarum Technologies, over an alleged overlap between its exit-node pool and the "Popa" botnet of at least two million compromised devices. Alarum [[https://alarum.io/alarum-technologies-responds-to-inquiry-into-residential-proxy-networks/|disputes the allegations]], stated it had not been formally contacted by the FBI as of 3 July 2026, and paused parts of the network. Treat the case as unresolved — but note that a researcher who had bought bandwidth from that pool would now be explaining it to their ethics board. The risk is current, not historical. On **2 July 2026** the FBI, with the IRS Criminal Investigation division, Google's Threat Intelligence Group and Lumen, [[https://krebsonsecurity.com/2026/07/fbi-seizes-netnut-proxy-platform-popa-botnet/|seized hundreds of domains belonging to NetNut]], a major residential proxy provider operated by Alarum Technologies, over an alleged overlap between its exit-node pool and the "Popa" botnet of at least two million compromised devices. Alarum [[https://alarum.io/alarum-technologies-responds-to-inquiry-into-residential-proxy-networks/|disputes the allegations]], stated it had not been formally contacted by the FBI as of 3 July 2026, and paused parts of the network. Treat the case as unresolved — but note that a researcher who had bought bandwidth from that pool would now be explaining it to their ethics board.
  
-If you use one, we suggest: name the provider in the paper, state what the provider claims about consent, say whether your IRB or ethics board reviewed that claim specifically, and prefer providers used by prior peer-reviewed work. Bright Data's research arm ("The Bright Initiative") is the route taken by {[bouhoula2024automated]} — and note that they used **datacenter** IPs from it, not residential ones. See [[Practices:Ethics]].+If you use one, we suggest: name the provider in the paper, state what the provider claims about consent, say whether your IRB or ethics board reviewed that claim specifically, and prefer providers used by prior peer-reviewed work. Bright Data's research arm ("The Bright Initiative") is the route taken by {[bouhoula2024_automated]} — and note that they used **datacenter** IPs from it, not residential ones. See [[Practices:Ethics]].
 </WRAP> </WRAP>
  
Line 203: Line 202:
   * **Geolocation databases disagree with each other.** They agree well at country level and poorly below it, which is the subject of [[Design:IP classification]].   * **Geolocation databases disagree with each other.** They agree well at country level and poorly below it, which is the subject of [[Design:IP classification]].
  
-The following script checks both. It queries several free services for your egress IP, reports whether they agree, flags datacenter/VPN/proxy addresses, and exits non-zero if the country is not the one you expected — so it can gate a crawl rather than merely inform you. Free tiers are rate-limited, so call it once when a vantage point comes up and again when it goes down, not per request. Adding a ''--json'' branch that dumps the same report as a dict is a two-line change if you want to log it alongside the crawl.+The following script checks both. It queries several free services for your egress IP, reports whether they agree, flags datacenter/VPN/proxy addresses, and exits non-zero if the country is not the one you expected — so it can gate a crawl rather than merely inform you. Free tiers are rate-limited, so call it once when a vantage point comes up and again when it goes down, not per request. Adding a ''%%--json%%'' branch that dumps the same report as a dict is a two-line change if you want to log it alongside the crawl.
  
 <file python verify_vantage.py> <file python verify_vantage.py>
Line 332: Line 331:
 ===== Open Questions ===== ===== Open Questions =====
  
-  * <wrap todo>No peer-reviewed cross-vendor measurement of how much datacenter IP address space is penalised by bot-management vendors. Jueckstock et al. {[jueckstock2021_realistic]} measure the effect on privacy metrics but not the mechanism per vendor.</wrap> +<WRAP todo> 
-  * <wrap todo>How stable is CMP geo-targeting configuration over time? All the industry documentation describes the feature; nobody appears to have measured how often operators change the region-to-template binding.</wrap> +  * No peer-reviewed cross-vendor measurement of how much datacenter IP address space is penalised by bot-management vendors. Jueckstock et al. {[jueckstock2021_realistic]} measure the effect on privacy metrics but not the mechanism per vendor. 
-  * <wrap todo>Whether commercial VPN mislabelling has been re-measured academically since {[weinberg2018_catch]} (2018). The only recent figures we found are a vendor study.</wrap> +  * How stable is CMP geo-targeting configuration over time? All the industry documentation describes the feature; nobody appears to have measured how often operators change the region-to-template binding. 
-  * <wrap todo>Residential proxy pool overlap with known botnets, post-NetNut. Spur, Synthient and Nokia Deepfield published attributions in June 2026; a systematic academic treatment would be valuable for ethics review.</wrap>+  * Whether commercial VPN mislabelling has been re-measured academically since {[weinberg2018_catch]} (2018). The only recent figures we found are a vendor study. 
 +  * Residential proxy pool overlap with known botnets, post-NetNut. Spur, Synthient and Nokia Deepfield published attributions in June 2026; a systematic academic treatment would be valuable for ethics review. 
 +</WRAP>
  
 ===== Related Pages ===== ===== Related Pages =====
  
 +  * [[Design:Blocking and geodifference]] — what your vantage point was **not allowed to see**: geoblocking, GDPR walls, censorship and network interference, and how a blocking claim is made.
   * [[Design:IP classification]] — classifying the addresses you observe, and the geolocation databases this page's script exercises.   * [[Design:IP classification]] — classifying the addresses you observe, and the geolocation databases this page's script exercises.
   * [[Design:Website selection]] — country-specific top lists (CrUX has country breakdowns; SecRank is Chinese-DNS-based) interact with vantage choice.   * [[Design:Website selection]] — country-specific top lists (CrUX has country breakdowns; SecRank is Chinese-DNS-based) interact with vantage choice.
   * [[Design:Archives]] — web archives sidestep the vantage point and introduce their own biases.   * [[Design:Archives]] — web archives sidestep the vantage point and introduce their own biases.
   * [[Programming:Crawler]] — proxy and per-context network configuration per crawling library.   * [[Programming:Crawler]] — proxy and per-context network configuration per crawling library.
-  * [[Programming:Stateful stateless]] — the other axis Jueckstock et al. {[jueckstock2021_realistic]} vary.+  * [[Programming:Stateful stateless]] — the design choice that decides what a crawl can observe at all. Note that Jueckstock et al. {[jueckstock2021_realistic]} do //not// vary it: every crawl in that paper launches "with a clean user profile (i.e., no cookies or cached content)". The second axis they vary is the browser configuration, naive against stealth.
   * [[Privacy:Consent]] — what to do with the banner once you are in the right jurisdiction.   * [[Privacy:Consent]] — what to do with the banner once you are in the right jurisdiction.
   * [[Practices:Ethics]] — residential proxies, volunteer devices, and acceptable-use policies.   * [[Practices:Ethics]] — residential proxies, volunteer devices, and acceptable-use policies.
design/crawling_location.1786550797.txt.gz · Last modified: by karel.kubicek.claude