User Tools

Site Tools


design:crawling_location

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Next revision
Previous revision
design:crawling_location [2026/08/05 15:51] – Create Design:Crawling location. Covers why the vantage point changes results (geo-targeted consent, geoblocking, datacenter-IP treatment), a survey of 859 crawling papers from 7 venues 2010-2024 showing 70.9% never state where they measured from, the pra karel.kubicek.claudedesign:crawling_location [2026/09/17 10:45] (current) – ConsentAction audit propagation and EU/EEA figures; Authored by Claude karel.kubicek.claude
Line 1: Line 1:
 ====== Crawling Location ====== ====== Crawling Location ======
  
-The //vantage point// of a measurement is the network position your traffic appears to originate from: its IP address, and everything a website can infer from it — country, city, network type, and whether it looks like a person or a datacenter. It is chosen at least implicitly by every study that touches the live web, and it is the design decision least often reported: of the crawling papers in our corpus of seven major security and privacy venues, **70.9% never state where they measured from** (see [[#Use in Publications]]).+The //vantage point// of a measurement is the network position your traffic appears to originate from: its IP address, and everything a website can infer from it — country, city, network type, and whether it looks like a person or a datacenter. It is chosen at least implicitly by every study that touches the live web, and it is the design decision least often reported: of the crawling papers in our corpus of seven major security and privacy venues, **70.8record a vantage point and never say where it was** (see [[#Use in Publications]]).
  
 That silence matters because the vantage point changes three separable things, and only the first is obvious: That silence matters because the vantage point changes three separable things, and only the first is obvious:
Line 9: Line 9:
   - **How your IP is treated.** Datacenter, university, Tor and residential addresses receive measurably different treatment from bot management and from trackers.   - **How your IP is treated.** Datacenter, university, Tor and residential addresses receive measurably different treatment from bot management and from trackers.
  
-This page covers all three, how researchers have actually chosen vantage points, the practical options, and how to verify that the vantage point you think you have is the one you got. It pairs with [[Design:IP classification]] (classifying //other people's// addresses), [[Design:Website selection]] (which sites), and [[Programming:Crawler]] (which tool).+This page covers all three, how researchers have actually chosen vantage points, the practical options, and how to verify that the vantage point you think you have is the one you got. It pairs with [[Design:IP classification]] (classifying //other people's// addresses), [[Design:Website selection]] (which sites), and [[Programming:Crawler]] (which tool). [[Privacy:Age assurance]] is the sharpest current case of the first bullet: since 2025 several jurisdictions //require// a site to behave differently for visitors they believe are theirs, so an age-gate prevalence figure without a stated vantage point is not a quantity.
  
 ===== Why the Vantage Point Changes Your Results ===== ===== Why the Vantage Point Changes Your Results =====
Line 27: Line 27:
 ==== Content: geoblocking and geo-differentiation ==== ==== Content: geoblocking and geo-differentiation ====
  
-Even setting privacy law aside, the web is not the same everywhere:+Even setting privacy law aside, the web is not the same everywhere, and a crawl from the wrong place records an absence that is a property of **your vantage point** rather than of the site:
  
-  * **Outright geoblocking.** McDonald et al. {[mcdonald2018_forbidden]} measured CDN-level geoblocking from 177 countries, finding server-side blocking of entire regions. A crawl from a blocked country records an absence that is a property of your vantage point, not of the site's tracking behaviour. +  * **Outright geoblocking and geo-differentiated content.** McDonald et al. {[mcdonald2018_forbidden]} measured CDN-level geoblocking from 177 countries, and Kumar et al. {[kumar2022_investigation]} found systematic differences in availability and behaviour across 26 countries chosen "to have reliable direct vantage points". Both are the subject of **[[Design:Blocking and geodifference]]**, which owns the phenomenon, the instruments, and what a blocking claim has to contain — including the error rates these detectors have been measured to have. Read it if your vantage point may not be allowed to see what you are counting; this page stops at what that means for //choosing// the vantage point
-  * **Geo-differentiated content.** Kumar et al. {[kumar2022_investigation]} compared mobile apps across 26 countries chosen "to have reliable direct vantage points", finding systematic differences in availability and behaviour+  * **Personalisation.** Kliman-Silver et al. {[klimansilver2015_location]} showed that geolocation drives measurable web-search personalisation, so location is a confound in any study of ranked or targeted output — a difference in the //ordering// of what you are served rather than in whether you are served at all.
-  * **Personalisation.** Kliman-Silver et al. {[klimansilver2015_location]} showed that geolocation drives measurable web-search personalisation, so location is a confound in any study of ranked or targeted output.+
  
 ==== Treatment: what your IP says about you ==== ==== Treatment: what your IP says about you ====
Line 47: Line 46:
 ===== Use in Publications ===== ===== Use in Publications =====
  
-The figures below come from a structured extraction over **4,322 full-text papers** from CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf and IEEE S&P, 2010–2024. Unless stated otherwise the population is the **859 papers that ran a crawl**, and sentinel values (''not-stated'') are counted as what they are rather than as answers. Methodology and limitations are at the end of this section.+The figures below come from a structured extraction over **5,859 full-text papers** from CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf and IEEE S&P, 2010–2026. Unless stated otherwise the population is the **1,120 papers that ran a crawl**, and sentinel values (''not-stated'') are counted as what they are rather than as answers. The 2025 and 2026 venue-years are provisional — CCS and IMC 2026 have not been held, and IEEE S&P and WWW 2026 abstracts are not yet in the selection source — so they are under-represented by construction. Methodology and limitations are at the end of this section.
  
 ==== Most papers do not say where they measured from ==== ==== Most papers do not say where they measured from ====
  
-^ Outcome ^ Papers ^ Share of 859 +^ Outcome ^ Papers ^ Share of 1,120 
-| No vantage point recorded at all | 21 | 2.4% | +| No vantage point recorded at all | 31 | 2.8% | 
-| Vantage point recorded, location **not stated** | 609 | 70.9% | +| Vantage point recorded, location **not stated** | 793 | 70.8% | 
-| **States at least one place** | 229 | 26.7% |+| **States at least one place** | 296 | 26.4% |
  
 ==== Of those that do, most use a single vantage point ==== ==== Of those that do, most use a single vantage point ====
  
-^ Distinct places ^ Papers ^ Share of 229 +^ Distinct places ^ Papers ^ Share of 296 
-| 1 | 137 59.8% | +| 1 | 174 58.8% | 
-| 2 | 36 | 15.7% | +| 2 | 45 | 15.2% | 
-| 3 | 25 | 10.9% | +| 3 | 30 | 10.1% | 
-| 4 or more | 31 13.5% |+| 4 or more | 47 15.9% |
  
-Multi-vantage measurement is therefore **40.2% of papers that state a location, but only 10.7% of all crawling papers**. Reporting it the second way is the honest framing: the denominator is every paper that could have said something.+Multi-vantage measurement is therefore **41.2% of papers that state a location, but only 10.9% of all crawling papers**. Reporting it the second way is the honest framing: the denominator is every paper that could have said something.
  
 ==== Where from ==== ==== Where from ====
  
-^ Place ^ Papers ^ Share of 229 stating ^ +^ Place ^ Papers ^ Share of 296 stating ^ 
-| United States | 139 60.7% | +| United States | 182 61.5% | 
-| Germany | 41 17.9% | +| Germany | 60 20.3% | 
-| //multi-country// (e.g. "61 countries") | 33 | 14.4% | +| //Europe// (no country given) | 42 | 14.2% | 
-| //Europe// (no country given) | 31 13.5% | +| //multi-country// (e.g. "61 countries") | 37 12.5% | 
-India 17 7.4% | +United Kingdom 26 8.8% | 
-| China | 16 | 7.0% | +| China | 23 | 7.8% | 
-United Kingdom 15 6.6% | +India 23 7.8% | 
-Canada 14 | 6.1% | +Singapore 19 | 6.4% | 
-Singapore 13 5.7% | +Canada 18 6.1% | 
-| Australia | 12 5.2% |+| Australia | 18 6.1% |
  
-Aggregated: **United States or North America 148 papers (64.6%)**, EU/EEA 88 (38.4%), both 54. Shares exceed 100% because a paper can name several places. Sixty-four distinct places appear in total, so the long tail is thin — the field measures the web overwhelmingly from the United States and from Germany.+Aggregated: **United States or North America 193 papers (65.2%)**, EU/EEA 124 (41.9%), both 74. Shares exceed 100% because a paper can name several places. Ninety-one distinct places appear in total, so the long tail is thin — the field measures the web overwhelmingly from the United States and from Germany.
  
 ==== From what kind of infrastructure ==== ==== From what kind of infrastructure ====
  
-Only **344 of 859 (40.0%)** state an infrastructure type at all.+Only **445 of 1,120 (39.7%)** state an infrastructure type at all.
  
-^ Infrastructure ^ Papers ^ Share of 344 stating ^ +^ Infrastructure ^ Papers ^ Share of 445 stating ^ 
-| University network | 107 31.1% | +| University network | 134 30.1% | 
-| Cloud provider | 95 | 27.6% | +| Cloud provider | 123 | 27.6% | 
-| Research testbed((A dedicated measurement platform or lab deployment rather than a general-purpose network: RIPE Atlas, M-Lab, CAIDA Ark, PlanetLab, and one-off testbeds built for the paper. The boundary against //university network// is not always crisp, so read these two rows together.)) | 89 25.9% | +| Research testbed((A dedicated measurement platform or lab deployment rather than a general-purpose network: RIPE Atlas, M-Lab, CAIDA Ark, PlanetLab, and one-off testbeds built for the paper. The boundary against //university network// is not always crisp, so read these two rows together.)) | 122 27.4% | 
-| Commercial VPN | 38 11.0% | +| Commercial VPN | 48 10.8% | 
-| Volunteer devices | 24 7.0% | +| Volunteer devices | 29 6.5% | 
-Tor 14 | 4.1% | +Residential 19 | 4.3% | 
-Residential 13 3.8% | +Proxy service 18 4.0% | 
-Proxy service 12 | 3.5% | +Tor 15 | 3.4% | 
-| Mobile network | | 2.3% |+| Mobile network | | 2.0% |
  
-Note how rare **residential** is (3.8% of the papers that say anything) against Jueckstock et al.'s finding that it is the realism best-case. The named providers tell the same story — of the 149 papers naming a platform, Amazon dominates:+Note how rare **residential** is (4.3% of the papers that say anything) against Jueckstock et al.'s finding that it is the realism best-case. The named providers tell the same story — of the 165 papers naming a platform that folds to a provider family, Amazon dominates:
  
-^ Provider family ^ Papers ^ Share of 149 naming ^ +^ Provider family ^ Papers ^ Share of 165 naming ^ Spellings folded 
-| Amazon AWS / EC2 | 52 34.9% | +| Amazon AWS / EC2 | 68 41.2| 16 
-| Commercial VPN (named) | 11 | 7.4% | +| Commercial VPN (named) | 20 | 12.1% | 17 | 
-| Tor | | 6.0% | +| Google Cloud | 11 | 6.7| 6 
-Google Cloud 4.7% | +| Tor | 11 | 6.7| 2 
-| DigitalOcean | 4.7% | +University network 10 6.1| 10 
-| PlanetLab //(discontinued)//| 4.7% | +| DigitalOcean | 5.5| 3 | 
-| Residential/datacenter proxy vendor | 3.4% | +| Other research testbed | 9 | 5.5% | 9 
-Microsoft Azure 2.7% | +| PlanetLab //(discontinued)//| 4.8| 1 
-| Linode / Vultr / OVH / Hetzner | 4 | 2.7% | +| Residential/datacenter proxy vendor | 4.8% | 
-| RIPE Atlas | 3 | 2.0% | +| Alibaba / Aliyun / Baidu | 6 | 3.6| 6 
-| M-Lab / CAIDA Ark | 2 | 1.3% |+ISP / mobile carrier 3.6| 6 
 +| Linode / Vultr / OVH / Hetzner | 6 | 3.6% | 4 | 
 +| Microsoft Azure | 5 | 3.0| 1 
 +| RIPE Atlas | 5 | 3.0% | 1 | 
 +| CDN (Cloudflare / Akamai) | 4 | 2.4| 3 
 +| M-Lab / CAIDA Ark | 2 | 1.2% | 2 | 
 + 
 +A service string can name more than one platform ("Amazon EC2 and Aliyun"), so it counts in each family and the shares exceed 100%. 186 papers name a platform in total; the 21 not in the table named something that is not a vantage point at all — ''vantage.serviceName'' also catches services a paper //queried// (VirusTotal, Google Translate, Safe Browsing). The full residue is on the [[provenance:design:crawling_location|provenance page]].
  
 ==== Reporting is improving, slowly ==== ==== Reporting is improving, slowly ====
  
-^ Indicator ^ 2010–2013 ^ 2014–2017 ^ 2018–2021 ^ 2022–2024 ^ +^ Indicator ^ 2010–2013 ^ 2014–2017 ^ 2018–2021 ^ 2022–2024 ^ 2025–2026 //(provisional)// 
-| Papers that crawled | 89 158 298 314 +| Papers that crawled | 102 | 167 308 345 198 
-| States a place | 22.5% | 24.1% | 27.5% | 28.3% | +| States a place | 20.6% | 24.0% | 26.9% | 27.5% | 28.8% | 
-| More than one place //(of those stating)// | 55.0% | 31.6% | 34.1% | 46.1% | +| More than one place //(of those stating)// | 52.4% | 30.0% | 34.9% | 45.3% | 47.4% | 
-| EU/EEA vantage //(of those stating)// | 20.0% | 26.3% | 37.8% | 48.3% |+| EU/EEA vantage //(of those stating)// | 19.0% | 25.0% | 38.6% | 46.3% | 59.6% |
  
-The EU/EEA row is the clearest signal in the dataset: the share of location-stating crawls run from inside the EEA **more than doubles** across the period, 20.0% to 48.3%, which is what you would expect from GDPR entering force in 2018 and from the resulting wave of compliance measurement. Location reporting itself improves far less — from 22.5% to 28.3over fifteen years.+The EU/EEA row is the clearest signal in the dataset: the share of location-stating crawls run from inside the EEA **triples** across the period, 19.0% to 59.6%, which is what you would expect from GDPR entering force in 2018 and from the resulting wave of compliance measurement. Location reporting itself improves far less — from 20.6% to 28.8across the corpus 2010-2026. Read the last column with care: CCS and IMC 2026 have not been held and two more 2026 venue-years are incompletely selected, so it rests on fewer papers than the 198 count suggests it should.
  
 ==== Legal framing does predict the vantage point ==== ==== Legal framing does predict the vantage point ====
  
-Grouping the 859 crawling papers by which law, if any, they assessed:+Grouping the 1,120 crawling papers by which law, if any, they assessed:
  
 ^ Crawling papers that… ^ N ^ State a place ^ EU/EEA vantage ^ US vantage ^ ^ Crawling papers that… ^ N ^ State a place ^ EU/EEA vantage ^ US vantage ^
-| assessed an EU law (GDPR / ePrivacy / DSA) | 62 54.8% | 46.8% | 24.2% | +| assessed an EU law (GDPR / ePrivacy / DSA) | 93 55.9% | 48.4% | 26.9% | 
-| assessed a US law (CCPA / COPPA / …) | 29 48.3% | 20.7% | 44.8% | +| assessed a US law (CCPA / COPPA / …) | 43 55.8% | 25.6% | 51.2% | 
-| assessed no law | 773 23.9% | 7.4% | 16.0% |+| assessed no law | 989 22.9% | 7.7% | 15.4% |
  
-Papers with a legal question are **more than twice as likely to say where they measured from**, and they line their vantage point up with the jurisdiction they are assessing. Of the 62 crawling papers assessing an EU law, 34 state a place and **29 of those (85.3%) measured from the EU/EEA** — good practice, clearly visible in the data. The residue is the interesting part: **papers assessed an EU law from outside the EU/EEA**, and 28 assessing an EU law never said where they were.+Papers with a legal question are **more than twice as likely to say where they measured from**, and they line their vantage point up with the jurisdiction they are assessing. Of the 93 crawling papers assessing an EU law, 52 state a place and **45 of those (86.5%) measured from the EU/EEA** — good practice, clearly visible in the data. The residue is the interesting part: **papers assessed an EU law from outside the EU/EEA**, and 41 assessing an EU law never said where they were.
  
 <WRAP important> <WRAP important>
-The gap is worse where it matters most. Of the **270 papers that state what their crawler did with the consent banner** — accept, reject, a CMP-specific choice, or explicitly no interaction — only **48 (17.8%) state an EU/EEA vantage, and 166 give no location at all.** Consent behaviour is the single most geo-dependent thing on the web, and the majority of papers interacting with it do not report the jurisdiction they observed it from.+The gap is worse where it matters most. Of the **55 papers whose audited full-text record states what their crawler did with the consent banner** — accept, reject, a CMP-specific choice, or explicitly no interaction — **37 (67.3%) state an EU/EEA vantage, and 11 state no location at all.** Consent behaviour is the single most geo-dependent thing on the web, and the jurisdiction is still missing for a substantial minority of those audited paper claims.
  
-Both figures come from the same per-paper record, so the 270 are a subset of the 859 crawling papers rather than a separately sampled group: the consent action and the vantage point were extracted in one pass from the same full text, and a paper counts here only if it stated its consent action explicitly.+Both figures use the same per-paper record, but the consent denominator is the audited one, not the raw schema count: the 2026-09-05 audit read all 349 non-sentinel ''crawlConfig.consentAction'' labels and rejected 279 of the 313 ''no-interaction'' defaults. The resulting 55 are a subset of the 1,120 crawling papers, and a paper counts here only if the audit found a consent action claim in its full text.
 </WRAP> </WRAP>
  
Line 142: Line 148:
  
   * "This crawl was made from France on September 20th and 21st 2019." {[matte2020_cookie]}   * "This crawl was made from France on September 20th and 21st 2019." {[matte2020_cookie]}
-  * "We crawl websites using 30 German datacenter IP addresses provided by The Bright Initiative from Bright Data." {[bouhoula2024automated]}+  * "We crawl websites using 30 German datacenter IP addresses provided by The Bright Initiative from Bright Data." {[bouhoula2024_automated]}
   * "all but one use an IP address associated with a European server (Frankfurt am Main; DEU) … the other is run from a US IP address (Council Bluffs, IA; USA)." {[demir2024_bannertools]}   * "all but one use an IP address associated with a European server (Frankfurt am Main; DEU) … the other is run from a US IP address (Council Bluffs, IA; USA)." {[demir2024_bannertools]}
   * "We choose three geolocations for our measurement: (1) Germany (EU), (2) Japan (AS), and (3) the United States (NA)." — via ProtonVPN {[demir2022_reproducibility]}   * "We choose three geolocations for our measurement: (1) Germany (EU), (2) Japan (AS), and (3) the United States (NA)." — via ProtonVPN {[demir2022_reproducibility]}
Line 150: Line 156:
 ==== Methodology and limitations of these figures ==== ==== Methodology and limitations of these figures ====
  
-  * **How they were produced.** One structured record per paper was extracted from full text, each tuple carrying a verbatim evidence quote and its section, so any figure here can be traced to the sentence that supports it. Locations are free text ("Frankfurt, Germany", "US-East", "61 countries across all world regions") and were normalised to countries or regions before counting; strings across papers could not be mapped and are excluded. Papers, never tuples, are counted.+  * **How they were produced.** One structured record per paper was extracted from full text, each tuple carrying a verbatim evidence quote and its section, so any figure here can be traced to the sentence that supports it. Locations are free text ("Frankfurt, Germany", "US-East", "61 countries across all world regions") and were normalised to countries or regions before counting by ''scripts/geo.mjs''15 distinct strings across 13 papers could not be mapped and are excluded, and 5 of those papers are left with no place at all. The full residue list is on the [[provenance:design:crawling_location|provenance page]]. Papers, never tuples, are counted.
   * **Silence is not absence.** "Does not state a location" means the paper did not say, not that the authors did not know. These are reporting figures.   * **Silence is not absence.** "Does not state a location" means the paper did not say, not that the authors did not know. These are reporting figures.
-  * **Venue coverage.** IEEE S&P is only 43% retrieved (paywall), so it is under-representedEuroS&PACSAC, RAID, AsiaCCS, CHI and SOUPS are absent entirely; any claim here is a claim about seven venues.+  * **Venue coverage.** Seven venues only, and 2025–2026 are incomplete for reasons of calendar and indexing rather than relevance, so per-year trends should be read as ending in 2024Which venueswhich years and what each stage of the selection funnel costs are on [[literature:corpus]].
   * **Field stability.** The stated/not-stated distinction and the infrastructure enum are reproducible to within a few points on a repeat extraction. Free-text service names are much less stable and are reported folded into families and as rankings, not as precise figures. The ''Amazon AWS / EC2'' row, for instance, merges at least five spellings.   * **Field stability.** The stated/not-stated distinction and the infrastructure enum are reproducible to within a few points on a repeat extraction. Free-text service names are much less stable and are reported folded into families and as rankings, not as precise figures. The ''Amazon AWS / EC2'' row, for instance, merges at least five spellings.
-  * **''crawled'' is defined** as a paper whose crawl configuration was recorded or whose study types include an automated web crawl (859 papers, 19.9% of the corpus). This corpus is seven broad security venues, not a web-measurement corpus, so shares of all 4,322 papers would be meaningless here.+  * **''crawled'' is defined** as a paper whose crawl configuration was recorded or whose study types include an automated web crawl (1,120 papers, 19.1% of the corpus). This corpus is seven broad security venues, not a web-measurement corpus, so shares of all 5,859 papers would be meaningless here
 +  * **Every query behind this section, its report script and its unedited output** are on [[provenance:design:crawling_location]]; corpus-level caveats are on [[literature:corpus]].
  
 ===== Choosing a Vantage Point ===== ===== Choosing a Vantage Point =====
Line 181: Line 188:
 The risk is current, not historical. On **2 July 2026** the FBI, with the IRS Criminal Investigation division, Google's Threat Intelligence Group and Lumen, [[https://krebsonsecurity.com/2026/07/fbi-seizes-netnut-proxy-platform-popa-botnet/|seized hundreds of domains belonging to NetNut]], a major residential proxy provider operated by Alarum Technologies, over an alleged overlap between its exit-node pool and the "Popa" botnet of at least two million compromised devices. Alarum [[https://alarum.io/alarum-technologies-responds-to-inquiry-into-residential-proxy-networks/|disputes the allegations]], stated it had not been formally contacted by the FBI as of 3 July 2026, and paused parts of the network. Treat the case as unresolved — but note that a researcher who had bought bandwidth from that pool would now be explaining it to their ethics board. The risk is current, not historical. On **2 July 2026** the FBI, with the IRS Criminal Investigation division, Google's Threat Intelligence Group and Lumen, [[https://krebsonsecurity.com/2026/07/fbi-seizes-netnut-proxy-platform-popa-botnet/|seized hundreds of domains belonging to NetNut]], a major residential proxy provider operated by Alarum Technologies, over an alleged overlap between its exit-node pool and the "Popa" botnet of at least two million compromised devices. Alarum [[https://alarum.io/alarum-technologies-responds-to-inquiry-into-residential-proxy-networks/|disputes the allegations]], stated it had not been formally contacted by the FBI as of 3 July 2026, and paused parts of the network. Treat the case as unresolved — but note that a researcher who had bought bandwidth from that pool would now be explaining it to their ethics board.
  
-If you use one, we suggest: name the provider in the paper, state what the provider claims about consent, say whether your IRB or ethics board reviewed that claim specifically, and prefer providers used by prior peer-reviewed work. Bright Data's research arm ("The Bright Initiative") is the route taken by {[bouhoula2024automated]} — and note that they used **datacenter** IPs from it, not residential ones. See [[Practices:Ethics]].+If you use one, we suggest: name the provider in the paper, state what the provider claims about consent, say whether your IRB or ethics board reviewed that claim specifically, and prefer providers used by prior peer-reviewed work. Bright Data's research arm ("The Bright Initiative") is the route taken by {[bouhoula2024_automated]} — and note that they used **datacenter** IPs from it, not residential ones. See [[Practices:Ethics]].
 </WRAP> </WRAP>
  
 ==== Research measurement platforms ==== ==== Research measurement platforms ====
  
-[[https://atlas.ripe.net/|RIPE Atlas]] (probes and anchors worldwide, credit-based), [[https://www.measurementlab.net/|M-Lab]] and [[https://www.caida.org/projects/ark/|CAIDA Ark]] give wide, citable, reproducible geographic coverage. The catch is layer: they are built for network measurement, not for driving a browser, so they suit DNS, reachability and latency questions rather than tracking or consent. **PlanetLab is discontinued** — it appears in papers in the corpus and is not an option for new work; the [[https://www.edge-net.org/|EdgeNet]] project is the nearest successor.+[[https://atlas.ripe.net/|RIPE Atlas]] (probes and anchors worldwide, credit-based), [[https://www.measurementlab.net/|M-Lab]] and [[https://www.caida.org/projects/ark/|CAIDA Ark]] give wide, citable, reproducible geographic coverage. The catch is layer: they are built for network measurement, not for driving a browser, so they suit DNS, reachability and latency questions rather than tracking or consent. **PlanetLab is discontinued** — it appears in papers in the corpus and is not an option for new work; the [[https://www.edge-net.org/|EdgeNet]] project is the nearest successor.
  
 ===== Verify the Vantage Point ===== ===== Verify the Vantage Point =====
Line 195: Line 202:
   * **Geolocation databases disagree with each other.** They agree well at country level and poorly below it, which is the subject of [[Design:IP classification]].   * **Geolocation databases disagree with each other.** They agree well at country level and poorly below it, which is the subject of [[Design:IP classification]].
  
-The following script checks both. It queries several free services for your egress IP, reports whether they agree, flags datacenter/VPN/proxy addresses, and exits non-zero if the country is not the one you expected — so it can gate a crawl rather than merely inform you. Free tiers are rate-limited, so call it once when a vantage point comes up and again when it goes down, not per request. Adding a ''--json'' branch that dumps the same report as a dict is a two-line change if you want to log it alongside the crawl.+The following script checks both. It queries several free services for your egress IP, reports whether they agree, flags datacenter/VPN/proxy addresses, and exits non-zero if the country is not the one you expected — so it can gate a crawl rather than merely inform you. Free tiers are rate-limited, so call it once when a vantage point comes up and again when it goes down, not per request. Adding a ''%%--json%%'' branch that dumps the same report as a dict is a two-line change if you want to log it alongside the crawl.
  
 <file python verify_vantage.py> <file python verify_vantage.py>
Line 324: Line 331:
 ===== Open Questions ===== ===== Open Questions =====
  
-  * <wrap todo>No peer-reviewed cross-vendor measurement of how much datacenter IP address space is penalised by bot-management vendors. Jueckstock et al. {[jueckstock2021_realistic]} measure the effect on privacy metrics but not the mechanism per vendor.</wrap> +<WRAP todo> 
-  * <wrap todo>How stable is CMP geo-targeting configuration over time? All the industry documentation describes the feature; nobody appears to have measured how often operators change the region-to-template binding.</wrap> +  * No peer-reviewed cross-vendor measurement of how much datacenter IP address space is penalised by bot-management vendors. Jueckstock et al. {[jueckstock2021_realistic]} measure the effect on privacy metrics but not the mechanism per vendor. 
-  * <wrap todo>Whether commercial VPN mislabelling has been re-measured academically since {[weinberg2018_catch]} (2018). The only recent figures we found are a vendor study.</wrap> +  * How stable is CMP geo-targeting configuration over time? All the industry documentation describes the feature; nobody appears to have measured how often operators change the region-to-template binding. 
-  * <wrap todo>Residential proxy pool overlap with known botnets, post-NetNut. Spur, Synthient and Nokia Deepfield published attributions in June 2026; a systematic academic treatment would be valuable for ethics review.</wrap>+  * Whether commercial VPN mislabelling has been re-measured academically since {[weinberg2018_catch]} (2018). The only recent figures we found are a vendor study. 
 +  * Residential proxy pool overlap with known botnets, post-NetNut. Spur, Synthient and Nokia Deepfield published attributions in June 2026; a systematic academic treatment would be valuable for ethics review. 
 +</WRAP>
  
 ===== Related Pages ===== ===== Related Pages =====
  
 +  * [[Design:Blocking and geodifference]] — what your vantage point was **not allowed to see**: geoblocking, GDPR walls, censorship and network interference, and how a blocking claim is made.
   * [[Design:IP classification]] — classifying the addresses you observe, and the geolocation databases this page's script exercises.   * [[Design:IP classification]] — classifying the addresses you observe, and the geolocation databases this page's script exercises.
   * [[Design:Website selection]] — country-specific top lists (CrUX has country breakdowns; SecRank is Chinese-DNS-based) interact with vantage choice.   * [[Design:Website selection]] — country-specific top lists (CrUX has country breakdowns; SecRank is Chinese-DNS-based) interact with vantage choice.
   * [[Design:Archives]] — web archives sidestep the vantage point and introduce their own biases.   * [[Design:Archives]] — web archives sidestep the vantage point and introduce their own biases.
   * [[Programming:Crawler]] — proxy and per-context network configuration per crawling library.   * [[Programming:Crawler]] — proxy and per-context network configuration per crawling library.
-  * [[Programming:Stateful stateless]] — the other axis Jueckstock et al. {[jueckstock2021_realistic]} vary.+  * [[Programming:Stateful stateless]] — the design choice that decides what a crawl can observe at all. Note that Jueckstock et al. {[jueckstock2021_realistic]} do //not// vary it: every crawl in that paper launches "with a clean user profile (i.e., no cookies or cached content)". The second axis they vary is the browser configuration, naive against stealth.
   * [[Privacy:Consent]] — what to do with the banner once you are in the right jurisdiction.   * [[Privacy:Consent]] — what to do with the banner once you are in the right jurisdiction.
   * [[Practices:Ethics]] — residential proxies, volunteer devices, and acceptable-use policies.   * [[Practices:Ethics]] — residential proxies, volunteer devices, and acceptable-use policies.
design/crawling_location.1785945117.txt.gz · Last modified: by karel.kubicek.claude