provenance:design:ip_classification
Differences
This shows you the differences between two versions of the page.
| Both sides previous revisionPrevious revision | |||
| provenance:design:ip_classification [2026/09/21 13:30] – Quote-check refresh 2026-09-21: re-ran quote_check.mjs --classification ip-address with the pypdf fallback; 94 below threshold -> 67 rescued + 27 below in both; new provenance section. Authored by Claude karel.kubicek.claude | provenance:design:ip_classification [2026/09/22 22:27] (current) – Add §14: the 2026-09-22 Fable review (findings on disk), Sonnet currency pass, settlement of the 2026-09-03 self-served pass, revision ledger, classify_ips.py change with the preserved 2026-08-06 output, vendor re-check, embedded report script/fold/output karel.kubicek.claude | ||
|---|---|---|---|
| Line 17: | Line 17: | ||
| | Data | '' | | Data | '' | ||
| | Refreshed | 2026-08-12 | | | Refreshed | 2026-08-12 | | ||
| + | | Reviewed | 2026-09-22, Fable plus a Sonnet currency pass — §14 | | ||
| ===== 2. Populations and denominators ===== | ===== 2. Populations and denominators ===== | ||
| Line 35: | Line 36: | ||
| <code bash> | <code bash> | ||
| cd / | cd / | ||
| - | node scripts/ | ||
| node scripts/ | node scripts/ | ||
| node scripts/ | node scripts/ | ||
| Line 45: | Line 45: | ||
| </ | </ | ||
| - | '' | + | The fold has no self-test of its own; its residue is printed at the foot of the report script' |
| + | |||
| + | Re-run 2026-09-22 on the corrected page, '' | ||
| + | |||
| + | ^ Figures ^ Where they come from ^ | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | |||
| + | Before 2026-09-22 this paragraph said 11, then was three edits behind at 16; the list above replaces it. '' | ||
| ===== 4. What the refresh changed ===== | ===== 4. What the refresh changed ===== | ||
| Line 151: | Line 167: | ||
| ===== 8. What could not be established ===== | ===== 8. What could not be established ===== | ||
| - | * **Whether the 88 unread below-threshold quotes check out.** See §6. | + | * **Whether the 24 unread below-threshold quotes check out.** See §6 and // |
| * **Free vs paid MaxMind.** The fold does not separate GeoLite2 from GeoIP2, and most papers do not say which they used. The page says so. Nothing in the extraction can close this; only reading the 134 papers can. | * **Free vs paid MaxMind.** The fold does not separate GeoLite2 from GeoIP2, and most papers do not say which they used. The page says so. Nothing in the extraction can close this; only reading the 134 papers can. | ||
| * **Whether "no validation" | * **Whether "no validation" | ||
| Line 159: | Line 175: | ||
| ===== 10. Review pass, 2026-08-12 ===== | ===== 10. Review pass, 2026-08-12 ===== | ||
| - | // | + | // |
| * '' | * '' | ||
| Line 177: | Line 193: | ||
| | Caveats deleted | "IEEE S&P is only 43% retrieved (paywall)" | | Caveats deleted | "IEEE S&P is only 43% retrieved (paywall)" | ||
| | Mistake caught in review | The hardcoded '' | | Mistake caught in review | The hardcoded '' | ||
| - | | Review | Reviewed by Claude Fable 5 on 2026-08-12 with the instruction that the summary might not be exhaustive. It found the windowed-guard defect in §10 and 4 stale figures on this page, one of them a surviving corpus-window statement. All fixes were applied and re-saved the same day. | | + | | Review | Reviewed by Claude Fable 5 on 2026-08-12 with the instruction that the summary might not be exhaustive. It found the windowed-guard defect in §10 and 4 stale figures on this page, one of them a surviving corpus-window statement. All fixes were applied and re-saved the same day. Later content-page revisions are in §14.3. | |
| Line 356: | Line 372: | ||
| Mechanical rendering repair only: a fresh live raw/XHTML export of 188 pages was checked with '' | Mechanical rendering repair only: a fresh live raw/XHTML export of 188 pages was checked with '' | ||
| + | |||
| + | ===== 14. Fable review and fixes, 2026-09-22 ===== | ||
| + | |||
| + | // | ||
| + | |||
| + | ^ Item ^ Value ^ | ||
| + | | Date | 2026-09-22, unsupervised | | ||
| + | | Revisions reviewed | content page rev '' | ||
| + | | Corpus | '' | ||
| + | | Reviewers | '' | ||
| + | | Author of the fixes | Claude Opus 5.5, which also settled the 2026-09-03 self-served pass (F1–F6) against the live revisions | | ||
| + | | Artefacts | '' | ||
| + | | Script changes | '' | ||
| + | | Pages saved | [[design: | ||
| + | |||
| + | §9 was never used; the numbering is kept so that revision summaries citing §10–§13 still resolve. | ||
| + | |||
| + | ==== 14.1 Corpus figures ==== | ||
| + | |||
| + | **No stale or mis-denominated corpus figure on the content page.** '' | ||
| + | |||
| + | One figure was **true but compared against the wrong base** (found by the fix author, not by a reviewer): "196 of 295 (66.4%) report no validation … against 29.9% across all 4,439 papers" | ||
| + | |||
| + | Guard caveat for the next run: the page's new '' | ||
| + | |||
| + | ==== 14.2 Findings and what was done with each ==== | ||
| + | |||
| + | Severity is the reviewer' | ||
| + | |||
| + | ^ ID ^ From ^ Sev. ^ Finding ^ Verdict ^ Action ^ | ||
| + | | D1 | fable | MAJOR | The published '' | ||
| + | | D2 | fable, sonnet | MAJOR | ipapi.is no longer returns '' | ||
| + | | E1 | fable | MAJOR | The content page promises "the report script and its unedited output" | ||
| + | | P1 / F4 | fable, own 09-03 | MAJOR | §3's "11 figures unaccounted" | ||
| + | | P2 / F2 | fable, own 09-03 | MAJOR | The run log omitted four content-page revisions, including 2026-09-11, which added the cross-page '' | ||
| + | | P3 | fable | MAJOR | §8 said 88 unread below-threshold quotes; §6's 2026-09-21 refresh says 24 | CONFIRMED | **Fixed** in §8. | | ||
| + | | C1 | fable | MAJOR | "the older TorDNSEL service was retired in April 2020" is misstated: a DNS exit list still answers | PARTLY. The Tor Project' | ||
| + | | P6 / F5 | fable (PLAUSIBLE), | ||
| + | | N1 | fix author | MAJOR | 66.4% (per-target) compared against 29.9% (paper-level) | CONFIRMED from '' | ||
| + | | — | sonnet | MAJOR (" | ||
| + | | A1 | fable | MINOR | Report script' | ||
| + | | B1 | fable | MINOR | Chiapponi et al. do not " | ||
| + | | B4 | fable | MINOR (PLAUSIBLE) | Shavitt & Zilberman are characterised more strongly than their text supports: no country-accuracy-vs-claim measurement and no MaxMind-to-US default | CONFIRMED on the arXiv version (1005.5674v3, | ||
| + | | L1 | fable | MINOR (PLAUSIBLE) | CJEU //EDPS v SRB// quotation not verified verbatim | **Resolved**: | ||
| + | | E2 | fable | MINOR | "the big five clouds" | ||
| + | | E3 | fable | MINOR (PLAUSIBLE) | Three "we found no …" sentences have no recorded search behind them | CONFIRMED that no search protocol is recorded | **Recorded, not fixed** — §14.7. | | ||
| + | | E4 | fable | MINOR | For the stated reader, the only runnable artefact was broken | = D1/D2 | Fixed with D1/D2. | | ||
| + | | P4 / F3 | fable, own 09-03 | MINOR | §3 documented '' | ||
| + | | P5 | fable | MINOR | Section numbering skips §9 | CONFIRMED | Noted at the head of §14, numbering kept. | | ||
| + | | P7 | fable | MINOR | Nothing recorded a re-check of the vendor half since 2026-08-06 | CONFIRMED | **Fixed**: §14.6. | | ||
| + | | C2 | fable | MINOR | IP2Proxy' | ||
| + | | C3 | fable | MINOR | MaxMind dropped the " | ||
| + | | C4 | fable | MINOR | '' | ||
| + | | C5 | fable | MINOR | Seven cited vendor URLs redirect | CONFIRMED; all still reach the right content except the IPinfo one (C4) | **Not changed** apart from C4: a redirecting URL still resolves, and the content page links only two of the seven. | | ||
| + | | C6 | fable, sonnet | MINOR | DB-IP quote: " | ||
| + | | C7 | fable, sonnet | — | DigitalOcean CSV "404s intermittently" | ||
| + | | N2 | fix author | MINOR | Related Pages " | ||
| + | | F1 | own 09-03 | MAJOR | Related Pages marks '' | ||
| + | | F6 | own 09-03 | MINOR | '' | ||
| + | | — | sonnet | PLAUSIBLE | NetAcuity "no academic programme", | ||
| + | |||
| + | ==== 14.3 Revision ledger for the content page since this page's 2026-08-12 run log ==== | ||
| + | |||
| + | §11 records the 2026-08-12 refresh only. Every later revision of [[design: | ||
| + | |||
| + | ^ Rev ^ When (UTC) ^ What ^ Recorded in ^ | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | 1790116015 | 2026-09-22 | §14.2 fixes | §14 | | ||
| + | |||
| + | ==== 14.4 Was there a Fable review on 2026-08-12? ==== | ||
| + | |||
| + | The 2026-09-03 self-served pass (F5) said §10/§11 credit a Fable review that "did not report", | ||
| + | |||
| + | * ''/ | ||
| + | * Content-page revision '' | ||
| + | * The attempts recorded as not reporting are attempts at **this review item**; the one in the drain logs ran on 2026-09-03 (" | ||
| + | |||
| + | So §10's attribution is consistent with the record and is left standing. What it lacks is an artefact: neither the 2026-08-06 nor the 2026-08-12 review left a findings file that this run could find, so the attribution rests on those runs' own summaries. Session transcripts were not searched. This 2026-09-22 review is the first of this page whose findings are on disk. | ||
| + | |||
| + | ==== 14.5 classify_ips.py: | ||
| + | |||
| + | ipapi.is changed its keyless response on 1 September 2026 ('' | ||
| + | |||
| + | * ipapi.is is queried only when '' | ||
| + | * An extractor that meets an unexpected response shape now raises with the service, the address and the keys received, instead of a bare '' | ||
| + | * **Tested**: keyless on the five addresses (output on the content page); the keyed extractor against the full example response in ipapi.is' | ||
| + | |||
| + | What today' | ||
| + | |||
| + | Diff: | ||
| + | |||
| + | <code diff> | ||
| + | --- classify_ips_20260806.py 2026-09-22 22: | ||
| + | +++ classify_ips.py 2026-09-22 22: | ||
| + | @@ -12,8 +12,9 @@ | ||
| + | 2. OPERATOR-PUBLISHED PREFIXES (AWS, Google Cloud, Cloudflare). If the operator | ||
| + | says the prefix is theirs, it is theirs. Free, authoritative, | ||
| + | than any commercial " | ||
| + | - 3. GEOLOCATION ESTIMATES (four free services). These are inferences. The script | ||
| + | - | ||
| + | + 3. GEOLOCATION ESTIMATES (three free keyless services, four with an ipapi.is | ||
| + | + key in IPAPI_IS_KEY). These are inferences. The script prints them side by | ||
| + | + side and flags disagreement rather than picking one. | ||
| + | |||
| + | | ||
| + | | ||
| + | @@ -23,6 +24,7 @@ | ||
| + | | ||
| + | | ||
| + | | ||
| + | +import os | ||
| + | | ||
| + | | ||
| + | | ||
| + | @@ -48,13 +50,15 @@ | ||
| + | " | ||
| + | " | ||
| + | d.get(" | ||
| + | - # ipapi.is returns a REDUCED object (cc, flags, asn_org, no city) for keyless | ||
| + | - # queries about a third-party address, and the full object with a key. Handle | ||
| + | - # both rather than crashing on the free tier. | ||
| + | - " | ||
| + | - | ||
| + | - | ||
| + | } | ||
| + | +# ipapi.is carries the risk flags (is_datacenter, | ||
| + | +# keyless query returns neither the flags nor an ISO country code, so without a | ||
| + | +# (free) key it is skipped rather than half-used: https:// | ||
| + | +IPAPI_IS_KEY = os.environ.get(" | ||
| + | +if IPAPI_IS_KEY: | ||
| + | + GEO_SERVICES[" | ||
| + | + lambda d: (d[" | ||
| + | + | ||
| + | FLAGS = [" | ||
| + | |||
| + | |||
| + | @@ -132,7 +136,11 @@ | ||
| + | | ||
| + | | ||
| + | | ||
| + | - results[name] = extract(data) | ||
| + | + try: | ||
| + | + results[name] = extract(data) | ||
| + | + except KeyError as exc: | ||
| + | + raise RuntimeError(f" | ||
| + | + | ||
| + | | ||
| + | |||
| + | |||
| + | @@ -184,6 +192,8 @@ | ||
| + | | ||
| + | |||
| + | | ||
| + | + if not IPAPI_IS_KEY: | ||
| + | + print(" | ||
| + | | ||
| + | for ip in ips: | ||
| + | | ||
| + | </ | ||
| + | |||
| + | Output of the **2026-08-06** version of the script, as published on the content page until 2026-09-22. This is the run the risk-flag conclusion cites; it cannot be reproduced keylessly today. | ||
| + | |||
| + | < | ||
| + | layer 1 routing (Team Cymru bulk whois, 5/5 answered) | ||
| + | |||
| + | IP | ||
| + | 8.8.8.8 | ||
| + | 1.1.1.1 | ||
| + | 104.16.132.229 | ||
| + | 13.32.99.63 | ||
| + | 82.220.84.43 | ||
| + | |||
| + | layer 2 operator-published prefixes (AWS, Google Cloud, Cloudflare) | ||
| + | |||
| + | 11628 prefixes loaded | ||
| + | 8.8.8.8 | ||
| + | 1.1.1.1 | ||
| + | 104.16.132.229 | ||
| + | 13.32.99.63 | ||
| + | 82.220.84.43 | ||
| + | |||
| + | layer 3 geolocation estimates -- these are inferences, not facts | ||
| + | |||
| + | 8.8.8.8 | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | |||
| + | 1.1.1.1 | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | |||
| + | 104.16.132.229 | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | |||
| + | 13.32.99.63 | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | |||
| + | 82.220.84.43 | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | |||
| + | summary: 2 of 5 addresses had a cross-service country disagreement | ||
| + | </ | ||
| + | |||
| + | ==== 14.6 Vendor and URL re-check, 2026-09-22 ==== | ||
| + | |||
| + | Two independent passes fetched the vendor, licence and URL claims in //Which geolocation source to use in 2026//, //Ask the operator first// and // | ||
| + | |||
| + | ^ Claim ^ 2026-08-06 ^ 2026-09-22 ^ Kind ^ | ||
| + | | ipapi.is flags | keyless, 1,000/day | free key needed; keyless 30/day, no flags | rot (vendor change 2026-09-01) | | ||
| + | | MaxMind product names | GeoLite2, GeoIP2 | GeoLite, GeoIP (EULA of 12 February 2026: "' | ||
| + | | IP2Location LITE licence | "CC BY-SA 4.0" | CC BY-SA badge on the DB1 page; data-licence terms forbid redistribution and resale | stated two ways by the vendor | | ||
| + | | IP2Proxy LITE | eight categories | public proxies only; the eight (plus EPN) are the commercial edition' | ||
| + | | TorDNSEL | " | ||
| + | | IPinfo Privacy Detection docs | ''/ | ||
| + | | DB-IP quote | " | ||
| + | | DigitalOcean CSV | "404s intermittently" | ||
| + | |||
| + | ==== 14.7 What could not be established ==== | ||
| + | |||
| + | * **The three "we found no …" sentences** — no peer-reviewed prefix-granularity hosting classifier; no NetAcuity academic programme; nothing peer-reviewed applying an LLM to geolocation or host typing. None has a search protocol on this page. The LLM one rests partly on the corpus query in §13; the other two rest on the 2026-08-06 author' | ||
| + | * **Shavitt & Zilberman' | ||
| + | * **Whether the DigitalOcean CSV was ever intermittent.** Nothing on disk records the original observation. | ||
| + | * **A live keyed ipapi.is response** (§14.5). | ||
| + | * **The //EDPS v SRB// ECLI** ('' | ||
| + | |||
| + | ==== 14.8 The report script, the fold, and their unedited output ==== | ||
| + | |||
| + | The scripts exactly as in '' | ||
| + | |||
| + | <file javascript report_ip_classification.mjs> | ||
| + | // Every figure on design: | ||
| + | // | ||
| + | // node scripts/ | ||
| + | // node scripts/ | ||
| + | // | ||
| + | // Populations used here (each query names its own; "of <corpus size> papers" | ||
| + | // the answer): | ||
| + | // | ||
| + | // | ||
| + | // | ||
| + | // | ||
| + | // Free-text resource names are folded through scripts/ | ||
| + | // residue is printed at the bottom. | ||
| + | |||
| + | import { loadExtractions, | ||
| + | import { foldIpResource, | ||
| + | |||
| + | const WIKI = process.argv.includes(' | ||
| + | const T = (h, r) => (WIKI ? wikiTable(h, | ||
| + | const key = (p) => `${p.venue}/ | ||
| + | const head = (s) => console.log(`\n${WIKI ? '==== ' + s + ' ====' : '### ' + s}\n`); | ||
| + | |||
| + | const rows = loadExtractions(); | ||
| + | const crawled = rows.filter(POPULATIONS.crawled); | ||
| + | const measuredFrom = rows.filter(POPULATIONS.measuredFrom); | ||
| + | const ipTuples = (p) => p.classification.filter((t) => t.target === ' | ||
| + | const ipClassified = rows.filter((p) => ipTuples(p).length > 0); | ||
| + | |||
| + | console.log(`corpus | ||
| + | console.log(`crawled | ||
| + | console.log(`measuredFrom | ||
| + | console.log(`ipClassified | ||
| + | |||
| + | // ---------------------------------------------------------------- reach ---- | ||
| + | head(' | ||
| + | { | ||
| + | const byVenue = new Map(); | ||
| + | for (const p of ipClassified) byVenue.set(p.venue, | ||
| + | const venueTotal = new Map(); | ||
| + | for (const p of rows) venueTotal.set(p.venue, | ||
| + | console.log( | ||
| + | T( | ||
| + | [' | ||
| + | [...byVenue.entries()] | ||
| + | .sort((a, b) => b[1] - a[1]) | ||
| + | .map(([v, n]) => [v, n, venueTotal.get(v), | ||
| + | ) | ||
| + | ); | ||
| + | |||
| + | const buckets = [ | ||
| + | [' | ||
| + | [' | ||
| + | [' | ||
| + | [' | ||
| + | // 2025–2026 is provisional: | ||
| + | // incompletely selected. Labelled, not dropped. | ||
| + | [' | ||
| + | ]; | ||
| + | console.log(); | ||
| + | console.log( | ||
| + | T( | ||
| + | [' | ||
| + | buckets.map(([label, | ||
| + | const a = ipClassified.filter((p) => f(p.year)).length; | ||
| + | const b = rows.filter((p) => f(p.year)).length; | ||
| + | return [label, a, b, pct(a, b)]; | ||
| + | }) | ||
| + | ) | ||
| + | ); | ||
| + | } | ||
| + | |||
| + | // -------------------------------------------------- what method, enum'd ---- | ||
| + | head(`How the IP was classified (enum, papers of ${ipClassified.length})`); | ||
| + | { | ||
| + | const m = new Map(); | ||
| + | for (const p of ipClassified) | ||
| + | for (const t of ipTuples(p)) { | ||
| + | if (isSentinel(t.method)) continue; | ||
| + | if (!m.has(t.method)) m.set(t.method, | ||
| + | m.get(t.method).add(key(p)); | ||
| + | } | ||
| + | console.log( | ||
| + | T( | ||
| + | [' | ||
| + | [...m.entries()] | ||
| + | .sort((a, b) => b[1].size - a[1].size) | ||
| + | .map(([k, s]) => [k, s.size, pct(s.size, ipClassified.length)]) | ||
| + | ) | ||
| + | ); | ||
| + | } | ||
| + | |||
| + | // ------------------------------------------------------------ validation ---- | ||
| + | head(`Whether the IP classification was validated (papers of ${ipClassified.length})`); | ||
| + | { | ||
| + | const m = new Map(); | ||
| + | for (const p of ipClassified) | ||
| + | for (const t of ipTuples(p)) { | ||
| + | const v = t.validation ?? ' | ||
| + | if (!m.has(v)) m.set(v, new Set()); | ||
| + | m.get(v).add(key(p)); | ||
| + | } | ||
| + | console.log( | ||
| + | T( | ||
| + | [' | ||
| + | [...m.entries()] | ||
| + | .sort((a, b) => b[1].size - a[1].size) | ||
| + | .map(([k, s]) => [k, s.size, pct(s.size, ipClassified.length)]) | ||
| + | ) | ||
| + | ); | ||
| + | |||
| + | const gt = new Set(); | ||
| + | for (const p of ipClassified) | ||
| + | for (const t of ipTuples(p)) if (!isSentinel(t.groundTruthSource) && t.groundTruthSource) gt.add(key(p)); | ||
| + | console.log(`\nnames a ground-truth source: | ||
| + | |||
| + | // Papers whose *only* validation value is none-reported or not-applicable. | ||
| + | const weak = ipClassified.filter((p) => | ||
| + | ipTuples(p).every((t) => [' | ||
| + | ); | ||
| + | console.log(`no validation on any IP tuple: ${weak.length} / ${ipClassified.length} | ||
| + | } | ||
| + | |||
| + | // ------------------------------------------------- the named resources ----- | ||
| + | head(`Which resources, folded (papers of ${ipClassified.length})`); | ||
| + | { | ||
| + | const fam = new Map(); // family -> {task, set} | ||
| + | const residue = new Map(); // raw -> Set(paper) | ||
| + | for (const p of ipClassified) | ||
| + | for (const t of ipTuples(p)) { | ||
| + | if (isSentinel(t.resourceName) || !t.resourceName) continue; | ||
| + | const f = foldIpResource(t.resourceName); | ||
| + | if (!f) { | ||
| + | if (!residue.has(t.resourceName)) residue.set(t.resourceName, | ||
| + | residue.get(t.resourceName).add(key(p)); | ||
| + | continue; | ||
| + | } | ||
| + | const id = f.family; | ||
| + | if (!fam.has(id)) fam.set(id, { task: f.task, set: new Set() }); | ||
| + | fam.get(id).set.add(key(p)); | ||
| + | } | ||
| + | console.log( | ||
| + | T( | ||
| + | [' | ||
| + | [...fam.entries()] | ||
| + | .sort((a, b) => b[1].set.size - a[1].set.size) | ||
| + | .filter(([, v]) => v.set.size >= 2) | ||
| + | .map(([k, v]) => [k, TASK_LABEL[v.task], | ||
| + | ) | ||
| + | ); | ||
| + | const singles = [...fam.entries()].filter(([, | ||
| + | console.log(`\nfamilies named by exactly one paper: ${singles.length} (${singles.map(([k]) => k).join('; | ||
| + | console.log(`unfolded residue: ${residue.size} distinct strings`); | ||
| + | for (const [k, v] of residue) console.log(` | ||
| + | } | ||
| + | |||
| + | // ---------------------------------- MaxMind spelling count, the headline ---- | ||
| + | head(' | ||
| + | { | ||
| + | const spellings = new Set(); | ||
| + | const folded = new Set(); | ||
| + | const exact = new Map(); | ||
| + | const scope = []; | ||
| + | for (const p of rows) { | ||
| + | for (const t of p.classification) | ||
| + | if (t.target === ' | ||
| + | scope.push([p, | ||
| + | for (const t of p.vantage) | ||
| + | if (t.geolocationService && !isSentinel(t.geolocationService)) scope.push([p, | ||
| + | } | ||
| + | for (const [p, name] of scope) { | ||
| + | const f = foldIpResource(name); | ||
| + | if (f?.family !== ' | ||
| + | spellings.add(name); | ||
| + | folded.add(key(p)); | ||
| + | if (!exact.has(name)) exact.set(name, | ||
| + | exact.get(name).add(key(p)); | ||
| + | } | ||
| + | const best = [...exact.entries()].sort((a, | ||
| + | console.log(`distinct spellings of MaxMind: | ||
| + | console.log(`papers, | ||
| + | console.log(`papers under the commonest spelling (" | ||
| + | console.log(`undercount if you count exact strings: ${(100 * (1 - best[1].size / folded.size)).toFixed(0)}%`); | ||
| + | } | ||
| + | |||
| + | // -------------------------------------- who says which service they used ---- | ||
| + | head(' | ||
| + | { | ||
| + | const named = (pop) => { | ||
| + | const s = new Set(); | ||
| + | for (const p of pop) | ||
| + | for (const t of p.vantage) | ||
| + | if (t.geolocationService && !isSentinel(t.geolocationService)) s.add(key(p)); | ||
| + | return s; | ||
| + | }; | ||
| + | console.log( | ||
| + | T( | ||
| + | [' | ||
| + | [ | ||
| + | [' | ||
| + | [ | ||
| + | ' | ||
| + | measuredFrom.length, | ||
| + | named(measuredFrom).size, | ||
| + | pct(named(measuredFrom).size, | ||
| + | ], | ||
| + | ] | ||
| + | ) | ||
| + | ); | ||
| + | |||
| + | const fam = new Map(); | ||
| + | for (const p of rows) | ||
| + | for (const t of p.vantage) { | ||
| + | if (!t.geolocationService || isSentinel(t.geolocationService)) continue; | ||
| + | const f = foldIpResource(t.geolocationService); | ||
| + | const id = f ? f.family : `UNFOLDED: ${t.geolocationService}`; | ||
| + | if (!fam.has(id)) fam.set(id, new Set()); | ||
| + | fam.get(id).add(key(p)); | ||
| + | } | ||
| + | const total = new Set([...fam.values()].flatMap((s) => [...s])).size; | ||
| + | console.log(`\nof the ${total} papers naming one:`); | ||
| + | console.log( | ||
| + | T( | ||
| + | [' | ||
| + | [...fam.entries()] | ||
| + | .sort((a, b) => b[1].size - a[1].size) | ||
| + | .filter(([, s]) => s.size >= 2) | ||
| + | .map(([k, s]) => [k, s.size, pct(s.size, total)]) | ||
| + | ) | ||
| + | ); | ||
| + | console.log( | ||
| + | `named by one paper each: ${[...fam.entries()].filter(([, | ||
| + | ); | ||
| + | } | ||
| + | |||
| + | // ------------------------------------------------------- version stated ---- | ||
| + | head(' | ||
| + | { | ||
| + | // A geolocation database is versioned by date. The extraction does not carry a | ||
| + | // version field for classification resources, so this is a text proxy: does | ||
| + | // the evidence quote or the resource name mention a date, month or version? | ||
| + | const dated = / | ||
| + | const users = ipClassified.filter((p) => | ||
| + | ipTuples(p).some((t) => { | ||
| + | const f = t.resourceName ? foldIpResource(t.resourceName) : null; | ||
| + | return f && (f.task === ' | ||
| + | }) | ||
| + | ); | ||
| + | const withDate = users.filter((p) => | ||
| + | ipTuples(p).some((t) => dated.test(`${t.resourceName ?? '' | ||
| + | ); | ||
| + | console.log( | ||
| + | `papers using a third-party geo or routing dataset: ${users.length}` | ||
| + | ); | ||
| + | console.log( | ||
| + | ` ...whose evidence quote carries any date/ | ||
| + | ); | ||
| + | console.log(' | ||
| + | } | ||
| + | |||
| + | // ------------------------------------------- measured results, verbatim ---- | ||
| + | head(' | ||
| + | { | ||
| + | const re = | ||
| + | / | ||
| + | const seen = []; | ||
| + | for (const p of rows) | ||
| + | for (const t of p.detection) { | ||
| + | if (!t.prevalence) continue; | ||
| + | if (!re.test(`${t.phenomenon} ${t.technique} ${t.metric}`)) continue; | ||
| + | seen.push([p, | ||
| + | } | ||
| + | console.log(`${seen.length} prevalence-bearing detection tuples match the IP-classification regex`); | ||
| + | console.log(' | ||
| + | } | ||
| + | |||
| + | console.log(' | ||
| + | |||
| + | // ------------------------------ operator-published prefix lists, full text ---- | ||
| + | // Not a schema field: does anyone in this literature cite the cloud operators' | ||
| + | // own IP-range files? Full-text grep over every paper in the corpus. | ||
| + | { | ||
| + | const fs = await import(' | ||
| + | const path = await import(' | ||
| + | const { dataRoot } = await import(' | ||
| + | const re = | ||
| + | / | ||
| + | const hits = []; | ||
| + | for (const p of rows) { | ||
| + | const f = path.join(dataRoot(), | ||
| + | if (!fs.existsSync(f)) continue; | ||
| + | if (re.test(fs.readFileSync(f, | ||
| + | } | ||
| + | head(' | ||
| + | console.log(`${hits.length} of ${rows.length} papers`); | ||
| + | for (const h of hits) console.log(` | ||
| + | console.log(' | ||
| + | } | ||
| + | |||
| + | // ---------------------------------------------- cross-checking behaviour ---- | ||
| + | head(' | ||
| + | { | ||
| + | const famsFor = (p, task) => | ||
| + | new Set( | ||
| + | ipTuples(p) | ||
| + | .map((t) => t.resourceName) | ||
| + | .filter((n) => n && !isSentinel(n)) | ||
| + | .map(foldIpResource) | ||
| + | .filter((f) => f && f.task === task) | ||
| + | .map((f) => f.family) | ||
| + | ); | ||
| + | const geoUsers = ipClassified.filter((p) => famsFor(p, ' | ||
| + | const multi = geoUsers.filter((p) => famsFor(p, ' | ||
| + | console.log(`name >=1 geolocation source: | ||
| + | console.log(` | ||
| + | for (const p of multi.sort((a, | ||
| + | console.log(` | ||
| + | for (const task of [' | ||
| + | const n = ipClassified.filter((p) => famsFor(p, task).size > 0); | ||
| + | console.log(`name >=1 ' | ||
| + | if (task === ' | ||
| + | console.log(` | ||
| + | } | ||
| + | } | ||
| + | |||
| + | // --------------------------------------- size of the folded name universe ---- | ||
| + | head(' | ||
| + | { | ||
| + | const names = new Set(); | ||
| + | for (const p of rows) { | ||
| + | for (const t of p.classification) | ||
| + | if (t.target === ' | ||
| + | names.add(t.resourceName); | ||
| + | for (const t of p.vantage) | ||
| + | if (t.geolocationService && !isSentinel(t.geolocationService)) names.add(t.geolocationService); | ||
| + | } | ||
| + | const unmapped = [...names].filter((n) => !foldIpResource(n)); | ||
| + | console.log(`distinct strings: ${names.size}`); | ||
| + | console.log(`unmapped residue: ${unmapped.length}${unmapped.length ? ' -> ' + unmapped.join('; | ||
| + | } | ||
| + | |||
| + | // ----------------------------------------------- figures the page carried ---- | ||
| + | // Added 2026-08-12. Each of these was on design: | ||
| + | // script, so none of them could be re-derived when the corpus grew. If a figure | ||
| + | // is on the page it belongs here. | ||
| + | head(' | ||
| + | { | ||
| + | const imc = rows.filter((p) => p.venue === ' | ||
| + | const imcIp = ipClassified.filter((p) => p.venue === ' | ||
| + | console.log(`IMC: | ||
| + | |||
| + | // MaxMind, across BOTH fields the fold covers, vs within the IP-classifying set. | ||
| + | const mmAll = new Set(); | ||
| + | const mmIp = new Set(); | ||
| + | const mmSpellings = new Set(); | ||
| + | for (const p of rows) { | ||
| + | let hitIp = false; | ||
| + | let hit = false; | ||
| + | for (const t of p.classification) { | ||
| + | if (t.target !== ' | ||
| + | const f = foldIpResource(t.resourceName); | ||
| + | if (f && f.family === ' | ||
| + | } | ||
| + | for (const t of p.vantage) { | ||
| + | if (isSentinel(t.geolocationService)) continue; | ||
| + | const f = foldIpResource(t.geolocationService); | ||
| + | if (f && f.family === ' | ||
| + | } | ||
| + | if (hit) mmAll.add(key(p)); | ||
| + | if (hitIp) mmIp.add(key(p)); | ||
| + | } | ||
| + | console.log(`MaxMind: | ||
| + | `so ${mmAll.size - mmIp.size} name it only for their own vantage point. ${mmSpellings.size} distinct spellings.`); | ||
| + | |||
| + | // Crawling papers: classify an observed address vs geolocate their own vantage. | ||
| + | const crawlIp = crawled.filter((p) => ipTuples(p).length > 0).length; | ||
| + | const crawlGeo = crawled.filter((p) => p.vantage.some((t) => !isSentinel(t.geolocationService))).length; | ||
| + | const crawlEither = crawled.filter( | ||
| + | (p) => ipTuples(p).length > 0 || p.vantage.some((t) => !isSentinel(t.geolocationService)) | ||
| + | ).length; | ||
| + | console.log(`Crawling papers: ${crawlIp} classify an observed address, ${crawlGeo} geolocate their own vantage point, ` + | ||
| + | `union ${crawlEither} of ${crawled.length} (${pct(crawlEither, | ||
| + | } | ||
| + | </ | ||
| + | |||
| + | Output ('' | ||
| + | |||
| + | <file text report_ip_classification-output.txt> | ||
| + | corpus | ||
| + | crawled | ||
| + | measuredFrom | ||
| + | ipClassified | ||
| + | |||
| + | ### Papers classifying an IP address, by venue and period | ||
| + | |||
| + | Venue Papers classifying an IP Papers in corpus | ||
| + | ------- | ||
| + | IMC 124 | ||
| + | USENIX | ||
| + | NDSS | ||
| + | CCS 24 990 2.4% | ||
| + | WWW 24 843 2.8% | ||
| + | IEEE-SP | ||
| + | PETS | ||
| + | |||
| + | Period | ||
| + | ---------- | ||
| + | 2010–2013 | ||
| + | 2014–2017 | ||
| + | 2018–2021 | ||
| + | 2022–2024 | ||
| + | 2025–2026* | ||
| + | |||
| + | ### How the IP was classified (enum, papers of 295) | ||
| + | |||
| + | Method | ||
| + | ------------------- | ||
| + | third-party-service | ||
| + | curated-database | ||
| + | heuristic-rules | ||
| + | blocklist | ||
| + | regex-or-signature | ||
| + | other 9 3.1% | ||
| + | manual-labelling | ||
| + | supervised-ml | ||
| + | graph-analysis | ||
| + | dynamic-analysis | ||
| + | static-analysis | ||
| + | llm 1 0.3% | ||
| + | |||
| + | ### Whether the IP classification was validated (papers of 295) | ||
| + | |||
| + | Validation | ||
| + | -------------------------- | ||
| + | not-applicable | ||
| + | none-reported | ||
| + | manual-validation | ||
| + | comparison-to-other-method | ||
| + | held-out-test-set | ||
| + | cross-validation | ||
| + | |||
| + | names a ground-truth source: | ||
| + | no validation on any IP tuple: 196 / 295 (66.4%) | ||
| + | |||
| + | ### Which resources, folded (papers of 295) | ||
| + | |||
| + | Resource family | ||
| + | -------------------------------------------------------------------- | ||
| + | Home-grown heuristic or classifier | ||
| + | MaxMind | ||
| + | Other IP blocklists (DShield, FireHOL, CBL, AbuseIPDB, Honey Pot, …) Is it known-bad? | ||
| + | IPinfo | ||
| + | Router alias / router-to-AS (bdrmapIT, MAP-IT, MIDAR, Hoiho) | ||
| + | Team Cymru IP-to-ASN | ||
| + | Censys / Shodan / Nmap / Snort / Suricata | ||
| + | IP2Location | ||
| + | VirusTotal / Google Safe Browsing | ||
| + | RouteViews | ||
| + | WHOIS / IRR / RIR delegation files Whose network is it? 13 4.4% | ||
| + | Spamhaus | ||
| + | CAIDA datasets (prefix2as, AS2Org, ITDK) Whose network is it? 11 3.7% | ||
| + | Raw BGP feeds and IX data Whose network is it? 7 2.4% | ||
| + | RIPE RIS / RIPEstat / RIPE Atlas Whose network is it? 7 2.4% | ||
| + | Free geo-lookup APIs (freegeoip, ipstack, HostIP, IPInfoDB, …) Where is it? 6 2.0% | ||
| + | PeeringDB | ||
| + | NetAcuity (Digital Element) | ||
| + | RIPE IPmap Where is it? 5 1.7% | ||
| + | Fraud-score APIs (IPQualityScore, | ||
| + | ASdb Whose network is it? 5 1.7% | ||
| + | GreyNoise | ||
| + | Unnamed commercial geo database | ||
| + | Geocoding / positioning reference (GeoNames, Google, Skyhook, WiGLE) | ||
| + | Chinese geo databases (Chunzhen/ | ||
| + | MaxMind Anonymous IP / minFraud | ||
| + | Quova / Neustar | ||
| + | Chainalysis | ||
| + | IP2Proxy | ||
| + | pyasn / iptoasn.com | ||
| + | Spur What kind of host is it? 2 0.7% | ||
| + | |||
| + | families named by exactly one paper: 9 (ip-api.com; | ||
| + | unfolded residue: 0 distinct strings | ||
| + | |||
| + | ### How badly exact-string counting undercounts (MaxMind) | ||
| + | |||
| + | distinct spellings of MaxMind: | ||
| + | papers, folded: | ||
| + | papers under the commonest spelling (" | ||
| + | undercount if you count exact strings: 79% | ||
| + | |||
| + | ### Papers that geolocate their own vantage point | ||
| + | |||
| + | Population | ||
| + | ------------ | ||
| + | crawled | ||
| + | measuredFrom | ||
| + | |||
| + | of the 194 papers naming one: | ||
| + | Service family | ||
| + | -------------------------------------------------------------------- | ||
| + | MaxMind | ||
| + | IPinfo | ||
| + | Geocoding / positioning reference (GeoNames, Google, Skyhook, WiGLE) | ||
| + | IP2Location | ||
| + | Free geo-lookup APIs (freegeoip, ipstack, HostIP, IPInfoDB, …) 7 3.6% | ||
| + | ip-api.com | ||
| + | RIPE IPmap 7 3.6% | ||
| + | Unnamed commercial geo database | ||
| + | NetAcuity (Digital Element) | ||
| + | CDN / platform internal geo | ||
| + | Chinese geo databases (Chunzhen/ | ||
| + | Akamai EdgeScape | ||
| + | Team Cymru IP-to-ASN | ||
| + | Quova / Neustar | ||
| + | Home-grown heuristic or classifier | ||
| + | WHOIS / IRR / RIR delegation files 2 1.0% | ||
| + | named by one paper each: 4 families | ||
| + | |||
| + | ### Do the papers say which snapshot of the database they used? | ||
| + | |||
| + | papers using a third-party geo or routing dataset: 162 | ||
| + | ...whose evidence quote carries any date/ | ||
| + | (text proxy, not a schema field — treat as an upper bound) | ||
| + | |||
| + | ### Measured figures on geolocation-database accuracy (detection[].prevalence) | ||
| + | |||
| + | 630 prevalence-bearing detection tuples match the IP-classification regex | ||
| + | (full dump written to out/ | ||
| + | |||
| + | done. | ||
| + | |||
| + | ### Papers citing an operator-published cloud prefix-list URL (full-text grep) | ||
| + | |||
| + | 7 of 5859 papers | ||
| + | 2017 IMC large-scale-scanning-of-tcps-initial-window | ||
| + | 2019 IEEE-SP resident-evil-understanding-residential-ip-proxy-as-a-dark-service | ||
| + | 2020 CCS censored-planet-an-internet-wide-longitudinal-censorship-observatory | ||
| + | 2020 NDSS cdn-judo-breaking-the-cdn-dos-protection-with-itself | ||
| + | 2021 CCS warmonger-inflicting-denial-of-service-via-serverless-functions-in-the-cloud | ||
| + | 2023 USENIX dscope-a-cloud-native-internet-telescope | ||
| + | 2025 NDSS secure-ip-address-allocation-at-cloud-scale | ||
| + | (a floor: undercounts papers that used a list without citing its URL) | ||
| + | |||
| + | ### Papers naming more than one geolocation source | ||
| + | |||
| + | name >=1 geolocation source: | ||
| + | ...of which name > | ||
| + | 2010 IMC Eyeball ASes: from geography to connectivity. | ||
| + | 2017 IMC A look at router geolocation in public and commercial databases. | ||
| + | 2018 IMC An Empirical Analysis of the Commercial VPN Ecosystem. | ||
| + | 2018 IMC Tracing Cross Border Web Tracking. | ||
| + | 2022 CCS An Extensive Study of Residential Proxies in China. | ||
| + | 2022 IMC Are we ready for metaverse?: a measurement study of social virtual reality pla | ||
| + | 2022 USENIX Gossamer: Securely Measuring Password-based Logins | ||
| + | 2023 IEEE-SP IPvSeeYou: Exploiting Leaked Identifiers in IPv6 for Street-Level Geolocation. | ||
| + | 2023 IMC How to Operate a Meta-Telescope in your Spare Time. | ||
| + | 2023 IMC Replication: | ||
| + | 2024 IMC A First Look at Immersive Telepresence on Apple Vision Pro. | ||
| + | 2024 IMC Watching TV with the Second-Party: | ||
| + | 2025 NDSS Wallbleed: A Memory Disclosure Vulnerability in the Great Firewall of China | ||
| + | 2025 USENIX eSIMplicity or eSIMplification? | ||
| + | name >=1 ' | ||
| + | name >=1 ' | ||
| + | earliest 2018, 16 from 2019 or later | ||
| + | |||
| + | ### Free-text name universe that ip_fold.mjs covers | ||
| + | |||
| + | distinct strings: 443 | ||
| + | unmapped residue: 0 | ||
| + | |||
| + | ### Figures the page carries that were not previously printed | ||
| + | |||
| + | IMC: 124 of 638. Other six venues: 171 of 5221. | ||
| + | MaxMind: 134 papers across both fields, 65 of them inside the 295 IP-classifying papers, so 69 name it only for their own vantage point. 50 distinct spellings. | ||
| + | Crawling papers: 54 classify an observed address, 45 geolocate their own vantage point, union 76 of 1120 (6.8%). | ||
| + | </ | ||
| + | |||
| + | <file javascript ip_fold.mjs> | ||
| + | // Fold the free-text names that appear when a paper classifies an IP address. | ||
| + | // | ||
| + | // Scope: the union of | ||
| + | // | ||
| + | // | ||
| + | // 363 distinct strings in run1. MaxMind alone appears under 24 spellings | ||
| + | // (" | ||
| + | // MaxMind", | ||
| + | // single most-used resource in the field by a factor of about four. | ||
| + | // | ||
| + | // Rules, same as scripts/ | ||
| + | // * explicit ordered regex families, first match wins | ||
| + | // * specific products before the generic term they contain | ||
| + | // | ||
| + | // * everything unmatched is returned as residue and PRINTED, never dropped | ||
| + | // | ||
| + | // Each family also carries the *question* it answers, because " | ||
| + | // is five different measurements with five different error budgets. | ||
| + | |||
| + | /** @typedef {' | ||
| + | |||
| + | /** @type {{family: string, task: Task, re: RegExp}[]} */ | ||
| + | export const FAMILIES = [ | ||
| + | // ---- proxy / VPN / hosting detection: match before their parent vendors ---- | ||
| + | { family: ' | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { family: ' | ||
| + | { family: ' | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | |||
| + | // ---- geolocation databases and services ---- | ||
| + | { family: ' | ||
| + | { family: ' | ||
| + | { family: ' | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { family: ' | ||
| + | { family: 'Quova / Neustar', | ||
| + | { family: ' | ||
| + | { family: ' | ||
| + | { family: 'RIPE IPmap', | ||
| + | { | ||
| + | family: 'Free geo-lookup APIs (freegeoip, ipstack, HostIP, IPInfoDB, …)', | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { | ||
| + | family: 'CDN / platform internal geo', | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | |||
| + | // ---- routing / ASN ---- | ||
| + | { family: 'Team Cymru IP-to-ASN', | ||
| + | { family: ' | ||
| + | { | ||
| + | family: 'RIPE RIS / RIPEstat / RIPE Atlas', | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { family: 'CAIDA datasets (prefix2as, AS2Org, ITDK)', | ||
| + | { family: ' | ||
| + | { family: ' | ||
| + | { family: ' | ||
| + | { family: 'pyasn / iptoasn.com', | ||
| + | { | ||
| + | family: 'WHOIS / IRR / RIR delegation files', | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { | ||
| + | family: 'Raw BGP feeds and IX data', | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | |||
| + | // ---- reputation / abuse blocklists ---- | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { | ||
| + | family: 'Other IP blocklists (DShield, FireHOL, CBL, AbuseIPDB, Honey Pot, …)', | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { family: ' | ||
| + | // Criminal IP is an attack-surface search engine in the Censys/ | ||
| + | // marketed as threat intelligence. Filed under scanning, which is what it does. | ||
| + | { family: ' | ||
| + | // Published hosting/ | ||
| + | // Distinct from a fraud-score API: this is the provider' | ||
| + | { family: ' | ||
| + | // 2026-08-12: the first LLM in this field. One paper, GPT-4o. Kept as its own | ||
| + | // family rather than folded into ' | ||
| + | { family: 'LLM (GPT-4o)', | ||
| + | // Email-authentication records: not an IP classifier, but the extraction files | ||
| + | // them here because the unit of analysis is the sending IP. | ||
| + | { family: 'Email authentication (SPF/ | ||
| + | // Phone-number reference services, from one paper whose unit was a phone | ||
| + | // number rather than an IP. Non-IP reference data, like the geocoding row. | ||
| + | { family: ' | ||
| + | |||
| + | // ---- active scanning / host fingerprinting ---- | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | |||
| + | // ---- router topology: alias resolution and router-to-AS ---- | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | |||
| + | // ---- non-IP location reference data ---- | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | |||
| + | // ---- home-grown ---- | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | ]; | ||
| + | |||
| + | /** | ||
| + | * @param {string} name | ||
| + | * @returns {{family: string, task: Task} | null} | ||
| + | */ | ||
| + | export function foldIpResource(name) { | ||
| + | const s = String(name).trim(); | ||
| + | for (const f of FAMILIES) { | ||
| + | if (f.re.test(s)) return { family: f.family, task: f.task }; | ||
| + | } | ||
| + | return null; | ||
| + | } | ||
| + | |||
| + | export const TASK_LABEL = { | ||
| + | geolocation: | ||
| + | routing: 'Whose network is it?', | ||
| + | ' | ||
| + | reputation: 'Is it known-bad?', | ||
| + | topology: 'Is it a router, and which one?', | ||
| + | scanning: 'What is running on it?', | ||
| + | reference: ' | ||
| + | custom: ' | ||
| + | }; | ||
| + | </ | ||
| + | |||
| + | <file javascript maxmind_version.mjs> | ||
| + | // Do papers that use MaxMind say WHICH snapshot of MaxMind they used? | ||
| + | // | ||
| + | // A geolocation database is a moving target: MaxMind reissues GeoLite2 twice a | ||
| + | // week. "We used MaxMind" | ||
| + | // extraction schema has no version field for a classification resource, so this | ||
| + | // measures it directly against the full text: for every paper whose extraction | ||
| + | // names MaxMind, find every sentence in paper.cols.txt that mentions MaxMind and | ||
| + | // ask whether any of them carries a date, a month, or a version/ | ||
| + | // | ||
| + | // node scripts/ | ||
| + | // | ||
| + | // This is a generous test. A sentence saying "we crawled in March 2019 using | ||
| + | // MaxMind" | ||
| + | // Read the figure as an UPPER BOUND on how often the snapshot is identifiable. | ||
| + | |||
| + | import fs from ' | ||
| + | import path from ' | ||
| + | import { loadExtractions, | ||
| + | import { foldIpResource } from ' | ||
| + | |||
| + | const DUMP = process.argv.includes(' | ||
| + | const root = dataRoot(); | ||
| + | const rows = loadExtractions(); | ||
| + | |||
| + | const isMaxMind = (s) => s && !isSentinel(s) && foldIpResource(s)? | ||
| + | const users = rows.filter( | ||
| + | (p) => | ||
| + | p.classification.some((t) => t.target === ' | ||
| + | p.vantage.some((t) => isMaxMind(t.geolocationService)) | ||
| + | ); | ||
| + | |||
| + | // A month name, a year, or an explicit version/ | ||
| + | const DATED = | ||
| + | / | ||
| + | |||
| + | // A bibliography entry (" | ||
| + | // year but tells the reader nothing about which snapshot was queried. Drop | ||
| + | // sentences that look like reference-list entries before the strict count. | ||
| + | const BIBLIKE = / | ||
| + | |||
| + | let read = 0, | ||
| + | missing = 0, | ||
| + | mentioned = 0, | ||
| + | dated = 0, | ||
| + | datedStrict = 0; | ||
| + | const examples = []; | ||
| + | |||
| + | for (const p of users) { | ||
| + | const file = path.join(root, | ||
| + | if (!fs.existsSync(file)) { | ||
| + | missing += 1; | ||
| + | continue; | ||
| + | } | ||
| + | read += 1; | ||
| + | const text = fs.readFileSync(file, | ||
| + | // Split on sentence-ish boundaries; keep it crude, the unit is " | ||
| + | const sentences = text.split(/ | ||
| + | const hits = sentences.filter((s) => / | ||
| + | if (hits.length === 0) continue; // extraction says MaxMind, text does not — see note | ||
| + | mentioned += 1; | ||
| + | const withDate = hits.filter((s) => DATED.test(s)); | ||
| + | const strict = withDate.filter((s) => !BIBLIKE.test(s)); | ||
| + | if (strict.length > 0) datedStrict += 1; | ||
| + | if (withDate.length > 0) { | ||
| + | dated += 1; | ||
| + | if (strict.length > 0 && examples.length < 8) | ||
| + | examples.push([`${p.year}/ | ||
| + | } else if (DUMP) { | ||
| + | console.log(`UNDATED | ||
| + | } | ||
| + | } | ||
| + | |||
| + | console.log(`papers whose extraction names MaxMind: | ||
| + | console.log(` | ||
| + | console.log(` | ||
| + | console.log( | ||
| + | ` ...with a date/ | ||
| + | ); | ||
| + | console.log( | ||
| + | ` ...excluding sentences that are bibliography entries: | ||
| + | ); | ||
| + | console.log(' | ||
| + | for (const [k, s] of examples) console.log(` | ||
| + | </ | ||
| + | |||
| + | Output ('' | ||
| + | |||
| + | <file text maxmind_version-output.txt> | ||
| + | papers whose extraction names MaxMind: | ||
| + | full text available: | ||
| + | MaxMind/ | ||
| + | ...with a date/ | ||
| + | ...excluding sentences that are bibliography entries: | ||
| + | |||
| + | Still an upper bound: any year token in the sentence counts, including crawl dates. | ||
| + | |||
| + | 2011/ | ||
| + | IPv4 address space delegated to Egypt (as of January 24, 2011) and Libya (as of February 15, 2011) by AfriNIC (top half) as well as additional IPv4 address ranges associated with the two countries based on MaxMind GeoLite database (as of Ja | ||
| + | |||
| + | 2013/ | ||
| + | Using the database dated from February 1, 2011 (so number of HTTP requests for all LDNS across all their TTL that our analysis would reflect the GeoIP map at the time intervals. | ||
| + | |||
| + | 2014/ | ||
| + | We geolocalize each IP address in DIP v4 using the Maxmind GeoIP database.9 We then introduce, for each identified country, the cor- 6. | ||
| + | |||
| + | 2015/ | ||
| + | All IPs returned in each hop of the traceroute were We use a simplified version of this check when examining geo-located with the MaxMind GeoLite27 country databases. | ||
| + | |||
| + | 2014/ | ||
| + | GeoIP, 2013. | ||
| + | |||
| + | 2015/ | ||
| + | Since MaxMind updates the database regularly (to reflect changes in the address space), we use the databases produced on August 1, 2012 and August 16, 2013 for the 2012 census and 2013 census periods, respectively. | ||
| + | |||
| + | 2015/ | ||
| + | However, 49% reach the AS of the destination, | ||
| + | |||
| + | 2017/ | ||
| + | First, our recommendations 467 IMC '17, November 1-3, 2017, London, United Kingdom 0.0 0.2 0.4 0.6 0.8 1.0 ARIN (4761) APNIC (468) AFRINIC (58) LACNIC (38) RIPENCC (1523) CDF AFRINIC APNIC −4 −3 −2 −1 0 10 10 10 10 10 (a) MaxMind-Paid (41.2 | ||
| + | |||
| + | </ | ||
| + | |||
| + | ==== 14.9 Re-review of the fixes ==== | ||
| + | |||
| + | Both reviewers whose findings were acted on re-read the drafted fixes **before** they were saved, told again that their context might not be exhaustive, and wrote to '' | ||
| + | |||
| + | ^ ID ^ From ^ Sev. ^ Finding ^ Action ^ | ||
| + | | R1 | fable | MAJOR | The vendor-table date line said MaxMind dropped the " | ||
| + | | R2 | fable | MINOR | '' | ||
| + | | R3 | fable | MINOR | §14.6 said the two passes " | ||
| + | | R4 | fable | MINOR (PLAUSIBLE) | The IP2Location data-licence quote had no fetched copy in the run directory | **Fixed**: '' | ||
| + | | R5 | fable | MINOR | The flags box said "this page's 2026-08-06 run of the script below", | ||
| + | | R6 | fable | note | §14.2 says four unrecorded revisions where fable' | ||
| + | | — | sonnet | — | 11 changed external claims re-fetched and CONFIRMED (MaxMind EULA date and prices, IP2Location badge only on DB1, IP2Proxy LITE = '' | ||
| + | |||
| + | The fixes to R1–R5 were not re-reviewed a third time. | ||
| + | |||
| + | [[design: | ||
provenance/design/ip_classification.1789997452.txt.gz · Last modified: by karel.kubicek.claude
