provenance:design:ip_classification
Differences
This shows you the differences between two versions of the page.
| Next revision | Previous revision | ||
| provenance:design:ip_classification [2026/08/12 09:23] – Create provenance page for design:ip_classification: queries, denominators, the 18-string fold residue and where each went, two report-script bugs, quote verification, and the 4,322 -> 5,859 refresh diff. Partly reconstructed; marked as such. Authored by karel.kubicek.claude | provenance:design:ip_classification [2026/09/22 22:27] (current) – Add §14: the 2026-09-22 Fable review (findings on disk), Sonnet currency pass, settlement of the 2026-09-03 self-served pass, revision ledger, classify_ips.py change with the preserved 2026-08-06 output, vendor re-check, embedded report script/fold/output karel.kubicek.claude | ||
|---|---|---|---|
| Line 11: | Line 11: | ||
| ^ Item ^ Value ^ | ^ Item ^ Value ^ | ||
| | Content page | [[design: | | Content page | [[design: | ||
| - | | Report script | '' | + | | Report script | '' |
| | Folds | '' | | Folds | '' | ||
| | Supporting script | '' | | Supporting script | '' | ||
| - | | Quote verification | '' | + | | Quote verification | '' |
| | Data | '' | | Data | '' | ||
| | Refreshed | 2026-08-12 | | | Refreshed | 2026-08-12 | | ||
| + | | Reviewed | 2026-09-22, Fable plus a Sonnet currency pass — §14 | | ||
| ===== 2. Populations and denominators ===== | ===== 2. Populations and denominators ===== | ||
| Line 35: | Line 36: | ||
| <code bash> | <code bash> | ||
| cd / | cd / | ||
| - | node scripts/ | ||
| node scripts/ | node scripts/ | ||
| node scripts/ | node scripts/ | ||
| Line 45: | Line 45: | ||
| </ | </ | ||
| - | '' | + | The fold has no self-test of its own; its residue is printed at the foot of the report script' |
| + | |||
| + | Re-run 2026-09-22 on the corrected page, '' | ||
| + | |||
| + | ^ Figures ^ Where they come from ^ | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | |||
| + | Before 2026-09-22 this paragraph said 11, then was three edits behind at 16; the list above replaces it. '' | ||
| ===== 4. What the refresh changed ===== | ===== 4. What the refresh changed ===== | ||
| Line 115: | Line 131: | ||
| ===== 6. Quotes checked ===== | ===== 6. Quotes checked ===== | ||
| - | //Recorded, 2026-08-12.// | + | //Recorded, 2026-08-12. Re-run 2026-09-21 with the PDF fallback — see // |
| < | < | ||
| - | $ node scripts/ | + | $ node scripts/ |
| - | 484 quotes checked: 232 exact, 158 partial (>=60% of 5-word windows), | + | 484 quotes checked: 232 exact, 158 partial (>=60% of 5-word windows), |
| - | 94 below threshold, 0 with no full text on disk. | + | |
| </ | </ | ||
| - | **94 below threshold | + | **The old figure was 94 below threshold, and 67 of those 94 are a defect in the stored text rather than in the extraction** — below threshold against the rendering the extractor |
| + | |||
| + | Six of the then-94 | ||
| < | < | ||
| Line 131: | Line 148: | ||
| </ | </ | ||
| - | **Open, and stated as such on this page rather than on the content page:** the other 88 have not been read. A 19% below-threshold rate is higher than the '' | + | **Open, and stated as such on this page rather than on the content page: |
| // | // | ||
| Line 145: | Line 162: | ||
| **Rejected: | **Rejected: | ||
| + | |||
| + | * **DynamIPs** ({[padmanabhan2020_dynamips]}, | ||
| ===== 8. What could not be established ===== | ===== 8. What could not be established ===== | ||
| - | * **Whether the 88 unread below-threshold quotes check out.** See §6. | + | * **Whether the 24 unread below-threshold quotes check out.** See §6 and // |
| * **Free vs paid MaxMind.** The fold does not separate GeoLite2 from GeoIP2, and most papers do not say which they used. The page says so. Nothing in the extraction can close this; only reading the 134 papers can. | * **Free vs paid MaxMind.** The fold does not separate GeoLite2 from GeoIP2, and most papers do not say which they used. The page says so. Nothing in the extraction can close this; only reading the 134 papers can. | ||
| * **Whether "no validation" | * **Whether "no validation" | ||
| Line 154: | Line 173: | ||
| * **Criminal IP's family.** See §5.1. One paper, and it could reasonably go under //Is it known-bad?// | * **Criminal IP's family.** See §5.1. One paper, and it could reasonably go under //Is it known-bad?// | ||
| - | ===== 9. Run log ===== | + | ===== 10. Review pass, 2026-08-12 ===== |
| + | |||
| + | // | ||
| + | |||
| + | * '' | ||
| + | * The matcher was **substring**, | ||
| + | |||
| + | Both are fixed in '' | ||
| + | Fixed on this page's content page as a result: the **intro paragraph**, | ||
| + | |||
| + | ===== 11. Run log ===== | ||
| ^ ^ ^ | ^ ^ ^ | ||
| Line 161: | Line 190: | ||
| | Model | Claude Opus 5, no sub-agents used for this page | | | Model | Claude Opus 5, no sub-agents used for this page | | ||
| | Scope | Mechanical re-derivation. Prose, structure and method selection were not revisited; two sentences changed because the numbers no longer supported them ("has since halved" | | Scope | Mechanical re-derivation. Prose, structure and method selection were not revisited; two sentences changed because the numbers no longer supported them ("has since halved" | ||
| - | | Script changes | '' | + | | Script changes | '' |
| | Caveats deleted | "IEEE S&P is only 43% retrieved (paywall)" | | Caveats deleted | "IEEE S&P is only 43% retrieved (paywall)" | ||
| | Mistake caught in review | The hardcoded '' | | Mistake caught in review | The hardcoded '' | ||
| + | | Review | Reviewed by Claude Fable 5 on 2026-08-12 with the instruction that the summary might not be exhaustive. It found the windowed-guard defect in §10 and 4 stale figures on this page, one of them a surviving corpus-window statement. All fixes were applied and re-saved the same day. Later content-page revisions are in §14.3. | | ||
| + | |||
| + | |||
| + | ===== 12. DynamIPs wording verification, | ||
| + | |||
| + | //Recorded as it happened.// The content page cited {[padmanabhan2020_dynamips]} twice (the //Churn// and //IPv6 is different// paragraphs of "IP as an Identifier" | ||
| + | |||
| + | **Getting the paper.** | ||
| + | |||
| + | ^ Source ^ Result ^ | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | |||
| + | pypdf 6.16.2 extracted all 16 pages (95,238 characters) without trouble; the earlier "would not extract" | ||
| + | |||
| + | **Quotes checked.** '' | ||
| + | |||
| + | <file text dynamips_quote_check-output.txt> | ||
| + | using cached review_ip/ | ||
| + | sha256 ac51460ca1ca5a3b40fad31cde7d8229b7fed9dab8d56609e4a59ca2c23b06eb | ||
| + | pages 16 chars 95238 pypdf 6.16.2 | ||
| + | first line: DynamIPs: Analyzing address assignment practices in IPv4 and | ||
| + | |||
| + | [1] ws abstract | ||
| + | "IPv6 assignments have longer durations than IPv4 assignments—often remaining stable for months—thereby allowing the possibility of long-term fingerprinting of IPv6 subscribers" | ||
| + | …ajor CDN. Our investigation of temporal dynamics with these datasets shows that IPv6 assignments have longer durations than IPv4 assignments—often remaining stable for months—thereby allowing the possibility of long-term fingerprinting of IPv6 subscribers. Our analysis of spatial dynamics reveals IPv6 addressassignment patterns that … | ||
| + | |||
| + | [2] ws §1 contributions | ||
| + | "IPv6 prefixes delegated to residential subscribers can remain stable for months, permitting long-term use of IPv6 prefixes to identify individual subscribers (at the CPE granularity), | ||
| + | …ress assignments in IPv4 and IPv6 on over 3,000 dual-stack probes. We find that IPv6 prefixes delegated to residential subscribers can remain stable for months, permitting long-term use of IPv6 prefixes to identify individual subscribers (at the CPE granularity), | ||
| + | |||
| + | [3] ws §3 assignment durations | ||
| + | " | ||
| + | …v4 address durations tend to be shorter, particularly for DTAG, Orange, and BT. Well-defined modes—at 1 day (DTAG), 1.5 days (Proximus), 1 week (Orange), and 2 weeks (BT)—in IPv4 non dual-stack address durations suggest that ISPs renumber addresses periodically. This result, using 6 years' worth of "IP echo" data, is consistent with observ… | ||
| + | |||
| + | [4] ws §3 assignment durations | ||
| + | "we observe evidence of consistent periodic renumbering on 35 networks when considering non-dual-stack probes" | ||
| + | …or work that also noted periodic renumbering within these ISPs [ 34]. In total, we observe evidence of consistent periodic renumbering on 35 networks when considering non-dual-stack probes. DTAG appears to renumber IPv6 prefixes after 1-day durations as well but this … | ||
| + | |||
| + | [5] ws §3 assignment durations | ||
| + | " | ||
| + | …arentheses is the total assignment duration in years from all probes in the AS. renumbering every 24 hours in IPv6 mainly in the following German ASes: DTAG, Versatel (AS8881), Netcologne (AS8422), Telefonica DE (AS6805), and M-net (AS8767). We also observe consistent period renumbering with a 12-hour period in ANTEL (… | ||
| + | |||
| + | [6] ws §3 assignment durations | ||
| + | " | ||
| + | …6805), and M-net (AS8767). We also observe consistent period renumbering with a 12-hour period in ANTEL (AS6057) in Uruguay and with a 48-hour period in Global Village (AS18881) in Brazil. Long IPv6 /64 durations in most ASes suggest that a /64 can be used to identif… | ||
| + | |||
| + | [7] nows §3 evolution over time | ||
| + | "IPv6 durations have consistently been longer than IPv4 durations and address durations in dual-stack networks tend to be longer than durations in non-dual-stack IPv4 networks" | ||
| + | …(matched in mode nows: pypdf line-break artefact) …mefractionsperyear.Theyear-to-yeartrendsconfirmourinsightsfromtheoveralldataset: | ||
| + | |||
| + | [8] ws §3 related results | ||
| + | "IPv6 /64 prefixes tend to be stable for months and years in various ASNs, although we find evidence of periodic renumbering in a handful of ISPs" | ||
| + | …2014 to March 2015) [36]. Our results using the RIPE Atlas dataset confirm that IPv6 /64 prefixes tend to be stable for months and years in various ASNs, although we find evidence of periodic renumbering in a handful of ISPs. 4 IPV4-IPV6 INTERPLAY So far, we have studied temporal properties of IPv4 and … | ||
| + | |||
| + | [9] nows §7 conclusion | ||
| + | "IPv6 assignments typically last longer than IPv4 assignments and can persist for months in several large residential ISPs" | ||
| + | …(matched in mode nows: pypdf line-break artefact) …investigatetemporalandspatialdynamicsofIPv4andIPv6addressassignments.WefoundthatIPv6assignmentstypicallylastlongerthanIPv4assignmentsandcanpersistformonthsinseverallargeresidentialISPs.WestudiedspatialaspectsofIPv6addressesindetail, | ||
| + | |||
| + | [10] ws abstract (datasets) | ||
| + | "over 3,000 RIPE Atlas probes in dual-stack networks" | ||
| + | …mics. We present finegrained observations of dynamics using data collected from over 3,000 RIPE Atlas probes in dual-stack networks. RIPE Atlas probes in these networks report both their IPv4 and their IPv6 addr… | ||
| + | |||
| + | [11] ws abstract (datasets) | ||
| + | "32.7 billion IPv4 and IPv6 address associations observed by a major CDN" | ||
| + | …space. To corroborate and extend our findings, we also use a dataset containing 32.7 billion IPv4 and IPv6 address associations observed by a major CDN. Our investigation of temporal dynamics with these datasets shows that IPv6 ass… | ||
| + | |||
| + | 11 quotes, 11 present, 0 missing | ||
| + | </ | ||
| + | |||
| + | **What the paper actually says, against what the page said.** | ||
| + | |||
| + | ^ Page wording before ^ Verdict ^ Page wording now ^ | ||
| + | | "IPv6 assignments last //longer// than IPv4 ones, often remaining stable for months" | ||
| + | | "found the distribution spans orders of magnitude between ISPs — some reassign on a fixed daily cycle, others leave an address in place for months" | ||
| + | | "RIPE Atlas dual-stack probes plus 32.7 billion address associations observed by a CDN" | **Accurate.** "over 3,000 RIPE Atlas probes in dual-stack networks"; | ||
| + | |||
| + | **Judgement calls.** | ||
| + | |||
| + | * The "35 networks" | ||
| + | * Not published: the CDN-side figures (median association duration 61 days; 20% of associations lasting more than 143 of a possible 150 days; 75% of mobile associations lasting a day or less). They are about IPv4–IPv6 address // | ||
| + | * The 45% / 44% of probes that saw no change in over a year are **excluded** from the paper' | ||
| + | |||
| + | **Run.** | ||
| + | |||
| + | ^ ^ ^ | ||
| + | | Date | 2026-09-03 | | ||
| + | | Scope | One citation' | ||
| + | | Model | Claude Fable 5.1, no sub-agents | | ||
| + | | Content page change | Two sentences in "IP as an Identifier: Four Ways It Breaks"; | ||
| + | | Mistake caught | The first version of the checker re-joined every hyphen at a line break and so reported the " | ||
| + | | Review | One focused pass (Claude Sonnet, citations and quotes), given the flattened paper text, the two paragraphs and this section, told its context might not be exhaustive. It confirmed all quotes and every other attributed specific, and returned three findings, **all accepted**: (1) blocker — the page sentence presented the IPv6 12 h / 24 h / 48 h cycles as instances of the IPv4 "35 networks" | ||
| + | |||
| + | |||
| + | ===== 13. LLM-classification currency, 2026-09-03 ===== | ||
| + | |||
| + | //Recorded during the run.// Shared numbers, the script, its unedited output, the folds and the quote check are on **[[provenance: | ||
| + | |||
| + | ==== 13.1 A retracted sentence ==== | ||
| + | |||
| + | The [[design: | ||
| + | |||
| + | > LLMs have reached AS-to-organisation mapping {[selmo2025_borges]} but not IP classification. We found nothing peer-reviewed applying an LLM to geolocation, | ||
| + | |||
| + | The first two sentences hold. The clause in bold does not, and nothing on this site owned it: there was no script behind it, no paper cited for it, and [[privacy: | ||
| + | |||
| + | Measured per target, as the LLM share of the papers that classify that target at all: | ||
| + | |||
| + | ^ Target ^ LLM papers ^ Papers classifying it at all ^ Share ^ | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | |||
| + | So " | ||
| + | |||
| + | ==== 13.2 Two different papers, both described as "the one LLM paper" ==== | ||
| + | |||
| + | The page said both of these, five hundred lines apart: | ||
| + | |||
| + | * //" | ||
| + | * //"the '' | ||
| + | |||
| + | They are different papers, and a reader would reasonably merge them into one: | ||
| + | |||
| + | ^ ^ {[selmo2025_borges]} ^ {[schwartz2025_llmcloudhunter]} ^ | ||
| + | | Venue | IMC 2025 | TheWebConf 2025 | | ||
| + | | Model | **GPT-4o-mini**, | ||
| + | | '' | ||
| + | | In this page's method table? | **no** — its tuples are filed under '' | ||
| + | | What it does | few-shot extraction over PeeringDB notes and //aka// fields for sibling-ASN mapping | extracts IP indicators and user agents from threat-intelligence prose | | ||
| + | |||
| + | Both sentences are now explicit about which paper they mean, and the first says that Borges does not appear in the method table at all. '' | ||
| + | |||
| + | ==== 13.3 Quotes checked ==== | ||
| + | |||
| + | * '' | ||
| + | * '' | ||
| + | |||
| + | ==== 13.4 What could not be established ==== | ||
| + | |||
| + | * **Whether '' | ||
| + | * **Nothing else on this page was re-derived.** The 295-paper population, the geolocation figures and the fold residue are unchanged from the 2026-08-12 refresh (§4) and from §12; only the two LLM sentences were touched. | ||
| + | |||
| + | ===== Quote-check refresh, 2026-09-21 ===== | ||
| + | |||
| + | The 2026-09-04 '' | ||
| + | |||
| + | < | ||
| + | $ node scripts/ | ||
| + | 484 quotes checked: 232 exact, 158 partial (>=60% of 5-word windows), 67 rescued from the PDF, 27 below threshold in both renderings, 0 with no full text on disk. | ||
| + | </ | ||
| + | |||
| + | ^ Figure ^ Was ^ Is ^ Why ^ | ||
| + | | quotes checked | 484 | 484 | population unchanged — the corpus has not moved | | ||
| + | | exact | 232 | 232 | unchanged | | ||
| + | | partial (≥60% of 5-word windows) | 158 | 158 | unchanged | | ||
| + | | rescued from the PDF | — | **67** | new verdict; these were inside the old 94 | | ||
| + | | below threshold | **94** | **27** (in both renderings) | 94 = 67 + 27 exactly; nothing else moved | | ||
| + | | below-threshold rate | 19% | **5.6%** | 27 of 484 | | ||
| + | | unread below-threshold quotes | 88 | **24** | 3 of the 6 hand-read rows are still below threshold in both | | ||
| + | |||
| + | **What this does and does not say.** It does not say 67 extractions were wrong and are now right — the quotes were always in the papers. It says the //stored text// could not locate them and a second rendering of the same PDF can, so counting them as quote failures measured '' | ||
| + | |||
| + | **Scope of this edit.** §6 only. '' | ||
| + | |||
| + | ^ Item ^ Value ^ | ||
| + | | Date | 2026-09-21, unsupervised | | ||
| + | | Command | '' | ||
| + | | Artifact | '' | ||
| + | | Script changes | none — '' | ||
| + | | Reviewers | one '' | ||
| + | | Pages saved | this page only | | ||
| + | |||
| + | [[design: | ||
| + | |||
| + | ===== Markup sweep, 2026-09-17 ===== | ||
| + | |||
| + | Mechanical rendering repair only: a fresh live raw/XHTML export of 188 pages was checked with '' | ||
| + | |||
| + | ===== 14. Fable review and fixes, 2026-09-22 ===== | ||
| + | |||
| + | // | ||
| + | |||
| + | ^ Item ^ Value ^ | ||
| + | | Date | 2026-09-22, unsupervised | | ||
| + | | Revisions reviewed | content page rev '' | ||
| + | | Corpus | '' | ||
| + | | Reviewers | '' | ||
| + | | Author of the fixes | Claude Opus 5.5, which also settled the 2026-09-03 self-served pass (F1–F6) against the live revisions | | ||
| + | | Artefacts | '' | ||
| + | | Script changes | '' | ||
| + | | Pages saved | [[design: | ||
| + | |||
| + | §9 was never used; the numbering is kept so that revision summaries citing §10–§13 still resolve. | ||
| + | |||
| + | ==== 14.1 Corpus figures ==== | ||
| + | |||
| + | **No stale or mis-denominated corpus figure on the content page.** '' | ||
| + | |||
| + | One figure was **true but compared against the wrong base** (found by the fix author, not by a reviewer): "196 of 295 (66.4%) report no validation … against 29.9% across all 4,439 papers" | ||
| + | |||
| + | Guard caveat for the next run: the page's new '' | ||
| + | |||
| + | ==== 14.2 Findings and what was done with each ==== | ||
| + | |||
| + | Severity is the reviewer' | ||
| + | |||
| + | ^ ID ^ From ^ Sev. ^ Finding ^ Verdict ^ Action ^ | ||
| + | | D1 | fable | MAJOR | The published '' | ||
| + | | D2 | fable, sonnet | MAJOR | ipapi.is no longer returns '' | ||
| + | | E1 | fable | MAJOR | The content page promises "the report script and its unedited output" | ||
| + | | P1 / F4 | fable, own 09-03 | MAJOR | §3's "11 figures unaccounted" | ||
| + | | P2 / F2 | fable, own 09-03 | MAJOR | The run log omitted four content-page revisions, including 2026-09-11, which added the cross-page '' | ||
| + | | P3 | fable | MAJOR | §8 said 88 unread below-threshold quotes; §6's 2026-09-21 refresh says 24 | CONFIRMED | **Fixed** in §8. | | ||
| + | | C1 | fable | MAJOR | "the older TorDNSEL service was retired in April 2020" is misstated: a DNS exit list still answers | PARTLY. The Tor Project' | ||
| + | | P6 / F5 | fable (PLAUSIBLE), | ||
| + | | N1 | fix author | MAJOR | 66.4% (per-target) compared against 29.9% (paper-level) | CONFIRMED from '' | ||
| + | | — | sonnet | MAJOR (" | ||
| + | | A1 | fable | MINOR | Report script' | ||
| + | | B1 | fable | MINOR | Chiapponi et al. do not " | ||
| + | | B4 | fable | MINOR (PLAUSIBLE) | Shavitt & Zilberman are characterised more strongly than their text supports: no country-accuracy-vs-claim measurement and no MaxMind-to-US default | CONFIRMED on the arXiv version (1005.5674v3, | ||
| + | | L1 | fable | MINOR (PLAUSIBLE) | CJEU //EDPS v SRB// quotation not verified verbatim | **Resolved**: | ||
| + | | E2 | fable | MINOR | "the big five clouds" | ||
| + | | E3 | fable | MINOR (PLAUSIBLE) | Three "we found no …" sentences have no recorded search behind them | CONFIRMED that no search protocol is recorded | **Recorded, not fixed** — §14.7. | | ||
| + | | E4 | fable | MINOR | For the stated reader, the only runnable artefact was broken | = D1/D2 | Fixed with D1/D2. | | ||
| + | | P4 / F3 | fable, own 09-03 | MINOR | §3 documented '' | ||
| + | | P5 | fable | MINOR | Section numbering skips §9 | CONFIRMED | Noted at the head of §14, numbering kept. | | ||
| + | | P7 | fable | MINOR | Nothing recorded a re-check of the vendor half since 2026-08-06 | CONFIRMED | **Fixed**: §14.6. | | ||
| + | | C2 | fable | MINOR | IP2Proxy' | ||
| + | | C3 | fable | MINOR | MaxMind dropped the " | ||
| + | | C4 | fable | MINOR | '' | ||
| + | | C5 | fable | MINOR | Seven cited vendor URLs redirect | CONFIRMED; all still reach the right content except the IPinfo one (C4) | **Not changed** apart from C4: a redirecting URL still resolves, and the content page links only two of the seven. | | ||
| + | | C6 | fable, sonnet | MINOR | DB-IP quote: " | ||
| + | | C7 | fable, sonnet | — | DigitalOcean CSV "404s intermittently" | ||
| + | | N2 | fix author | MINOR | Related Pages " | ||
| + | | F1 | own 09-03 | MAJOR | Related Pages marks '' | ||
| + | | F6 | own 09-03 | MINOR | '' | ||
| + | | — | sonnet | PLAUSIBLE | NetAcuity "no academic programme", | ||
| + | |||
| + | ==== 14.3 Revision ledger for the content page since this page's 2026-08-12 run log ==== | ||
| + | |||
| + | §11 records the 2026-08-12 refresh only. Every later revision of [[design: | ||
| + | |||
| + | ^ Rev ^ When (UTC) ^ What ^ Recorded in ^ | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | 1790116015 | 2026-09-22 | §14.2 fixes | §14 | | ||
| + | |||
| + | ==== 14.4 Was there a Fable review on 2026-08-12? ==== | ||
| + | |||
| + | The 2026-09-03 self-served pass (F5) said §10/§11 credit a Fable review that "did not report", | ||
| + | |||
| + | * ''/ | ||
| + | * Content-page revision '' | ||
| + | * The attempts recorded as not reporting are attempts at **this review item**; the one in the drain logs ran on 2026-09-03 (" | ||
| + | |||
| + | So §10's attribution is consistent with the record and is left standing. What it lacks is an artefact: neither the 2026-08-06 nor the 2026-08-12 review left a findings file that this run could find, so the attribution rests on those runs' own summaries. Session transcripts were not searched. This 2026-09-22 review is the first of this page whose findings are on disk. | ||
| + | |||
| + | ==== 14.5 classify_ips.py: | ||
| + | |||
| + | ipapi.is changed its keyless response on 1 September 2026 ('' | ||
| + | |||
| + | * ipapi.is is queried only when '' | ||
| + | * An extractor that meets an unexpected response shape now raises with the service, the address and the keys received, instead of a bare '' | ||
| + | * **Tested**: keyless on the five addresses (output on the content page); the keyed extractor against the full example response in ipapi.is' | ||
| + | |||
| + | What today' | ||
| + | |||
| + | Diff: | ||
| + | |||
| + | <code diff> | ||
| + | --- classify_ips_20260806.py 2026-09-22 22: | ||
| + | +++ classify_ips.py 2026-09-22 22: | ||
| + | @@ -12,8 +12,9 @@ | ||
| + | 2. OPERATOR-PUBLISHED PREFIXES (AWS, Google Cloud, Cloudflare). If the operator | ||
| + | says the prefix is theirs, it is theirs. Free, authoritative, | ||
| + | than any commercial " | ||
| + | - 3. GEOLOCATION ESTIMATES (four free services). These are inferences. The script | ||
| + | - | ||
| + | + 3. GEOLOCATION ESTIMATES (three free keyless services, four with an ipapi.is | ||
| + | + key in IPAPI_IS_KEY). These are inferences. The script prints them side by | ||
| + | + side and flags disagreement rather than picking one. | ||
| + | |||
| + | | ||
| + | | ||
| + | @@ -23,6 +24,7 @@ | ||
| + | | ||
| + | | ||
| + | | ||
| + | +import os | ||
| + | | ||
| + | | ||
| + | | ||
| + | @@ -48,13 +50,15 @@ | ||
| + | " | ||
| + | " | ||
| + | d.get(" | ||
| + | - # ipapi.is returns a REDUCED object (cc, flags, asn_org, no city) for keyless | ||
| + | - # queries about a third-party address, and the full object with a key. Handle | ||
| + | - # both rather than crashing on the free tier. | ||
| + | - " | ||
| + | - | ||
| + | - | ||
| + | } | ||
| + | +# ipapi.is carries the risk flags (is_datacenter, | ||
| + | +# keyless query returns neither the flags nor an ISO country code, so without a | ||
| + | +# (free) key it is skipped rather than half-used: https:// | ||
| + | +IPAPI_IS_KEY = os.environ.get(" | ||
| + | +if IPAPI_IS_KEY: | ||
| + | + GEO_SERVICES[" | ||
| + | + lambda d: (d[" | ||
| + | + | ||
| + | FLAGS = [" | ||
| + | |||
| + | |||
| + | @@ -132,7 +136,11 @@ | ||
| + | | ||
| + | | ||
| + | | ||
| + | - results[name] = extract(data) | ||
| + | + try: | ||
| + | + results[name] = extract(data) | ||
| + | + except KeyError as exc: | ||
| + | + raise RuntimeError(f" | ||
| + | + | ||
| + | | ||
| + | |||
| + | |||
| + | @@ -184,6 +192,8 @@ | ||
| + | | ||
| + | |||
| + | | ||
| + | + if not IPAPI_IS_KEY: | ||
| + | + print(" | ||
| + | | ||
| + | for ip in ips: | ||
| + | | ||
| + | </ | ||
| + | |||
| + | Output of the **2026-08-06** version of the script, as published on the content page until 2026-09-22. This is the run the risk-flag conclusion cites; it cannot be reproduced keylessly today. | ||
| + | |||
| + | < | ||
| + | layer 1 routing (Team Cymru bulk whois, 5/5 answered) | ||
| + | |||
| + | IP | ||
| + | 8.8.8.8 | ||
| + | 1.1.1.1 | ||
| + | 104.16.132.229 | ||
| + | 13.32.99.63 | ||
| + | 82.220.84.43 | ||
| + | |||
| + | layer 2 operator-published prefixes (AWS, Google Cloud, Cloudflare) | ||
| + | |||
| + | 11628 prefixes loaded | ||
| + | 8.8.8.8 | ||
| + | 1.1.1.1 | ||
| + | 104.16.132.229 | ||
| + | 13.32.99.63 | ||
| + | 82.220.84.43 | ||
| + | |||
| + | layer 3 geolocation estimates -- these are inferences, not facts | ||
| + | |||
| + | 8.8.8.8 | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | |||
| + | 1.1.1.1 | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | |||
| + | 104.16.132.229 | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | |||
| + | 13.32.99.63 | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | |||
| + | 82.220.84.43 | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | | ||
| + | |||
| + | summary: 2 of 5 addresses had a cross-service country disagreement | ||
| + | </ | ||
| + | |||
| + | ==== 14.6 Vendor and URL re-check, 2026-09-22 ==== | ||
| + | |||
| + | Two independent passes fetched the vendor, licence and URL claims in //Which geolocation source to use in 2026//, //Ask the operator first// and // | ||
| + | |||
| + | ^ Claim ^ 2026-08-06 ^ 2026-09-22 ^ Kind ^ | ||
| + | | ipapi.is flags | keyless, 1,000/day | free key needed; keyless 30/day, no flags | rot (vendor change 2026-09-01) | | ||
| + | | MaxMind product names | GeoLite2, GeoIP2 | GeoLite, GeoIP (EULA of 12 February 2026: "' | ||
| + | | IP2Location LITE licence | "CC BY-SA 4.0" | CC BY-SA badge on the DB1 page; data-licence terms forbid redistribution and resale | stated two ways by the vendor | | ||
| + | | IP2Proxy LITE | eight categories | public proxies only; the eight (plus EPN) are the commercial edition' | ||
| + | | TorDNSEL | " | ||
| + | | IPinfo Privacy Detection docs | ''/ | ||
| + | | DB-IP quote | " | ||
| + | | DigitalOcean CSV | "404s intermittently" | ||
| + | |||
| + | ==== 14.7 What could not be established ==== | ||
| + | |||
| + | * **The three "we found no …" sentences** — no peer-reviewed prefix-granularity hosting classifier; no NetAcuity academic programme; nothing peer-reviewed applying an LLM to geolocation or host typing. None has a search protocol on this page. The LLM one rests partly on the corpus query in §13; the other two rest on the 2026-08-06 author' | ||
| + | * **Shavitt & Zilberman' | ||
| + | * **Whether the DigitalOcean CSV was ever intermittent.** Nothing on disk records the original observation. | ||
| + | * **A live keyed ipapi.is response** (§14.5). | ||
| + | * **The //EDPS v SRB// ECLI** ('' | ||
| + | |||
| + | ==== 14.8 The report script, the fold, and their unedited output ==== | ||
| + | |||
| + | The scripts exactly as in '' | ||
| + | |||
| + | <file javascript report_ip_classification.mjs> | ||
| + | // Every figure on design: | ||
| + | // | ||
| + | // node scripts/ | ||
| + | // node scripts/ | ||
| + | // | ||
| + | // Populations used here (each query names its own; "of <corpus size> papers" | ||
| + | // the answer): | ||
| + | // | ||
| + | // | ||
| + | // | ||
| + | // | ||
| + | // Free-text resource names are folded through scripts/ | ||
| + | // residue is printed at the bottom. | ||
| + | |||
| + | import { loadExtractions, | ||
| + | import { foldIpResource, | ||
| + | |||
| + | const WIKI = process.argv.includes(' | ||
| + | const T = (h, r) => (WIKI ? wikiTable(h, | ||
| + | const key = (p) => `${p.venue}/ | ||
| + | const head = (s) => console.log(`\n${WIKI ? '==== ' + s + ' ====' : '### ' + s}\n`); | ||
| + | |||
| + | const rows = loadExtractions(); | ||
| + | const crawled = rows.filter(POPULATIONS.crawled); | ||
| + | const measuredFrom = rows.filter(POPULATIONS.measuredFrom); | ||
| + | const ipTuples = (p) => p.classification.filter((t) => t.target === ' | ||
| + | const ipClassified = rows.filter((p) => ipTuples(p).length > 0); | ||
| + | |||
| + | console.log(`corpus | ||
| + | console.log(`crawled | ||
| + | console.log(`measuredFrom | ||
| + | console.log(`ipClassified | ||
| + | |||
| + | // ---------------------------------------------------------------- reach ---- | ||
| + | head(' | ||
| + | { | ||
| + | const byVenue = new Map(); | ||
| + | for (const p of ipClassified) byVenue.set(p.venue, | ||
| + | const venueTotal = new Map(); | ||
| + | for (const p of rows) venueTotal.set(p.venue, | ||
| + | console.log( | ||
| + | T( | ||
| + | [' | ||
| + | [...byVenue.entries()] | ||
| + | .sort((a, b) => b[1] - a[1]) | ||
| + | .map(([v, n]) => [v, n, venueTotal.get(v), | ||
| + | ) | ||
| + | ); | ||
| + | |||
| + | const buckets = [ | ||
| + | [' | ||
| + | [' | ||
| + | [' | ||
| + | [' | ||
| + | // 2025–2026 is provisional: | ||
| + | // incompletely selected. Labelled, not dropped. | ||
| + | [' | ||
| + | ]; | ||
| + | console.log(); | ||
| + | console.log( | ||
| + | T( | ||
| + | [' | ||
| + | buckets.map(([label, | ||
| + | const a = ipClassified.filter((p) => f(p.year)).length; | ||
| + | const b = rows.filter((p) => f(p.year)).length; | ||
| + | return [label, a, b, pct(a, b)]; | ||
| + | }) | ||
| + | ) | ||
| + | ); | ||
| + | } | ||
| + | |||
| + | // -------------------------------------------------- what method, enum'd ---- | ||
| + | head(`How the IP was classified (enum, papers of ${ipClassified.length})`); | ||
| + | { | ||
| + | const m = new Map(); | ||
| + | for (const p of ipClassified) | ||
| + | for (const t of ipTuples(p)) { | ||
| + | if (isSentinel(t.method)) continue; | ||
| + | if (!m.has(t.method)) m.set(t.method, | ||
| + | m.get(t.method).add(key(p)); | ||
| + | } | ||
| + | console.log( | ||
| + | T( | ||
| + | [' | ||
| + | [...m.entries()] | ||
| + | .sort((a, b) => b[1].size - a[1].size) | ||
| + | .map(([k, s]) => [k, s.size, pct(s.size, ipClassified.length)]) | ||
| + | ) | ||
| + | ); | ||
| + | } | ||
| + | |||
| + | // ------------------------------------------------------------ validation ---- | ||
| + | head(`Whether the IP classification was validated (papers of ${ipClassified.length})`); | ||
| + | { | ||
| + | const m = new Map(); | ||
| + | for (const p of ipClassified) | ||
| + | for (const t of ipTuples(p)) { | ||
| + | const v = t.validation ?? ' | ||
| + | if (!m.has(v)) m.set(v, new Set()); | ||
| + | m.get(v).add(key(p)); | ||
| + | } | ||
| + | console.log( | ||
| + | T( | ||
| + | [' | ||
| + | [...m.entries()] | ||
| + | .sort((a, b) => b[1].size - a[1].size) | ||
| + | .map(([k, s]) => [k, s.size, pct(s.size, ipClassified.length)]) | ||
| + | ) | ||
| + | ); | ||
| + | |||
| + | const gt = new Set(); | ||
| + | for (const p of ipClassified) | ||
| + | for (const t of ipTuples(p)) if (!isSentinel(t.groundTruthSource) && t.groundTruthSource) gt.add(key(p)); | ||
| + | console.log(`\nnames a ground-truth source: | ||
| + | |||
| + | // Papers whose *only* validation value is none-reported or not-applicable. | ||
| + | const weak = ipClassified.filter((p) => | ||
| + | ipTuples(p).every((t) => [' | ||
| + | ); | ||
| + | console.log(`no validation on any IP tuple: ${weak.length} / ${ipClassified.length} | ||
| + | } | ||
| + | |||
| + | // ------------------------------------------------- the named resources ----- | ||
| + | head(`Which resources, folded (papers of ${ipClassified.length})`); | ||
| + | { | ||
| + | const fam = new Map(); // family -> {task, set} | ||
| + | const residue = new Map(); // raw -> Set(paper) | ||
| + | for (const p of ipClassified) | ||
| + | for (const t of ipTuples(p)) { | ||
| + | if (isSentinel(t.resourceName) || !t.resourceName) continue; | ||
| + | const f = foldIpResource(t.resourceName); | ||
| + | if (!f) { | ||
| + | if (!residue.has(t.resourceName)) residue.set(t.resourceName, | ||
| + | residue.get(t.resourceName).add(key(p)); | ||
| + | continue; | ||
| + | } | ||
| + | const id = f.family; | ||
| + | if (!fam.has(id)) fam.set(id, { task: f.task, set: new Set() }); | ||
| + | fam.get(id).set.add(key(p)); | ||
| + | } | ||
| + | console.log( | ||
| + | T( | ||
| + | [' | ||
| + | [...fam.entries()] | ||
| + | .sort((a, b) => b[1].set.size - a[1].set.size) | ||
| + | .filter(([, v]) => v.set.size >= 2) | ||
| + | .map(([k, v]) => [k, TASK_LABEL[v.task], | ||
| + | ) | ||
| + | ); | ||
| + | const singles = [...fam.entries()].filter(([, | ||
| + | console.log(`\nfamilies named by exactly one paper: ${singles.length} (${singles.map(([k]) => k).join('; | ||
| + | console.log(`unfolded residue: ${residue.size} distinct strings`); | ||
| + | for (const [k, v] of residue) console.log(` | ||
| + | } | ||
| + | |||
| + | // ---------------------------------- MaxMind spelling count, the headline ---- | ||
| + | head(' | ||
| + | { | ||
| + | const spellings = new Set(); | ||
| + | const folded = new Set(); | ||
| + | const exact = new Map(); | ||
| + | const scope = []; | ||
| + | for (const p of rows) { | ||
| + | for (const t of p.classification) | ||
| + | if (t.target === ' | ||
| + | scope.push([p, | ||
| + | for (const t of p.vantage) | ||
| + | if (t.geolocationService && !isSentinel(t.geolocationService)) scope.push([p, | ||
| + | } | ||
| + | for (const [p, name] of scope) { | ||
| + | const f = foldIpResource(name); | ||
| + | if (f?.family !== ' | ||
| + | spellings.add(name); | ||
| + | folded.add(key(p)); | ||
| + | if (!exact.has(name)) exact.set(name, | ||
| + | exact.get(name).add(key(p)); | ||
| + | } | ||
| + | const best = [...exact.entries()].sort((a, | ||
| + | console.log(`distinct spellings of MaxMind: | ||
| + | console.log(`papers, | ||
| + | console.log(`papers under the commonest spelling (" | ||
| + | console.log(`undercount if you count exact strings: ${(100 * (1 - best[1].size / folded.size)).toFixed(0)}%`); | ||
| + | } | ||
| + | |||
| + | // -------------------------------------- who says which service they used ---- | ||
| + | head(' | ||
| + | { | ||
| + | const named = (pop) => { | ||
| + | const s = new Set(); | ||
| + | for (const p of pop) | ||
| + | for (const t of p.vantage) | ||
| + | if (t.geolocationService && !isSentinel(t.geolocationService)) s.add(key(p)); | ||
| + | return s; | ||
| + | }; | ||
| + | console.log( | ||
| + | T( | ||
| + | [' | ||
| + | [ | ||
| + | [' | ||
| + | [ | ||
| + | ' | ||
| + | measuredFrom.length, | ||
| + | named(measuredFrom).size, | ||
| + | pct(named(measuredFrom).size, | ||
| + | ], | ||
| + | ] | ||
| + | ) | ||
| + | ); | ||
| + | |||
| + | const fam = new Map(); | ||
| + | for (const p of rows) | ||
| + | for (const t of p.vantage) { | ||
| + | if (!t.geolocationService || isSentinel(t.geolocationService)) continue; | ||
| + | const f = foldIpResource(t.geolocationService); | ||
| + | const id = f ? f.family : `UNFOLDED: ${t.geolocationService}`; | ||
| + | if (!fam.has(id)) fam.set(id, new Set()); | ||
| + | fam.get(id).add(key(p)); | ||
| + | } | ||
| + | const total = new Set([...fam.values()].flatMap((s) => [...s])).size; | ||
| + | console.log(`\nof the ${total} papers naming one:`); | ||
| + | console.log( | ||
| + | T( | ||
| + | [' | ||
| + | [...fam.entries()] | ||
| + | .sort((a, b) => b[1].size - a[1].size) | ||
| + | .filter(([, s]) => s.size >= 2) | ||
| + | .map(([k, s]) => [k, s.size, pct(s.size, total)]) | ||
| + | ) | ||
| + | ); | ||
| + | console.log( | ||
| + | `named by one paper each: ${[...fam.entries()].filter(([, | ||
| + | ); | ||
| + | } | ||
| + | |||
| + | // ------------------------------------------------------- version stated ---- | ||
| + | head(' | ||
| + | { | ||
| + | // A geolocation database is versioned by date. The extraction does not carry a | ||
| + | // version field for classification resources, so this is a text proxy: does | ||
| + | // the evidence quote or the resource name mention a date, month or version? | ||
| + | const dated = / | ||
| + | const users = ipClassified.filter((p) => | ||
| + | ipTuples(p).some((t) => { | ||
| + | const f = t.resourceName ? foldIpResource(t.resourceName) : null; | ||
| + | return f && (f.task === ' | ||
| + | }) | ||
| + | ); | ||
| + | const withDate = users.filter((p) => | ||
| + | ipTuples(p).some((t) => dated.test(`${t.resourceName ?? '' | ||
| + | ); | ||
| + | console.log( | ||
| + | `papers using a third-party geo or routing dataset: ${users.length}` | ||
| + | ); | ||
| + | console.log( | ||
| + | ` ...whose evidence quote carries any date/ | ||
| + | ); | ||
| + | console.log(' | ||
| + | } | ||
| + | |||
| + | // ------------------------------------------- measured results, verbatim ---- | ||
| + | head(' | ||
| + | { | ||
| + | const re = | ||
| + | / | ||
| + | const seen = []; | ||
| + | for (const p of rows) | ||
| + | for (const t of p.detection) { | ||
| + | if (!t.prevalence) continue; | ||
| + | if (!re.test(`${t.phenomenon} ${t.technique} ${t.metric}`)) continue; | ||
| + | seen.push([p, | ||
| + | } | ||
| + | console.log(`${seen.length} prevalence-bearing detection tuples match the IP-classification regex`); | ||
| + | console.log(' | ||
| + | } | ||
| + | |||
| + | console.log(' | ||
| + | |||
| + | // ------------------------------ operator-published prefix lists, full text ---- | ||
| + | // Not a schema field: does anyone in this literature cite the cloud operators' | ||
| + | // own IP-range files? Full-text grep over every paper in the corpus. | ||
| + | { | ||
| + | const fs = await import(' | ||
| + | const path = await import(' | ||
| + | const { dataRoot } = await import(' | ||
| + | const re = | ||
| + | / | ||
| + | const hits = []; | ||
| + | for (const p of rows) { | ||
| + | const f = path.join(dataRoot(), | ||
| + | if (!fs.existsSync(f)) continue; | ||
| + | if (re.test(fs.readFileSync(f, | ||
| + | } | ||
| + | head(' | ||
| + | console.log(`${hits.length} of ${rows.length} papers`); | ||
| + | for (const h of hits) console.log(` | ||
| + | console.log(' | ||
| + | } | ||
| + | |||
| + | // ---------------------------------------------- cross-checking behaviour ---- | ||
| + | head(' | ||
| + | { | ||
| + | const famsFor = (p, task) => | ||
| + | new Set( | ||
| + | ipTuples(p) | ||
| + | .map((t) => t.resourceName) | ||
| + | .filter((n) => n && !isSentinel(n)) | ||
| + | .map(foldIpResource) | ||
| + | .filter((f) => f && f.task === task) | ||
| + | .map((f) => f.family) | ||
| + | ); | ||
| + | const geoUsers = ipClassified.filter((p) => famsFor(p, ' | ||
| + | const multi = geoUsers.filter((p) => famsFor(p, ' | ||
| + | console.log(`name >=1 geolocation source: | ||
| + | console.log(` | ||
| + | for (const p of multi.sort((a, | ||
| + | console.log(` | ||
| + | for (const task of [' | ||
| + | const n = ipClassified.filter((p) => famsFor(p, task).size > 0); | ||
| + | console.log(`name >=1 ' | ||
| + | if (task === ' | ||
| + | console.log(` | ||
| + | } | ||
| + | } | ||
| + | |||
| + | // --------------------------------------- size of the folded name universe ---- | ||
| + | head(' | ||
| + | { | ||
| + | const names = new Set(); | ||
| + | for (const p of rows) { | ||
| + | for (const t of p.classification) | ||
| + | if (t.target === ' | ||
| + | names.add(t.resourceName); | ||
| + | for (const t of p.vantage) | ||
| + | if (t.geolocationService && !isSentinel(t.geolocationService)) names.add(t.geolocationService); | ||
| + | } | ||
| + | const unmapped = [...names].filter((n) => !foldIpResource(n)); | ||
| + | console.log(`distinct strings: ${names.size}`); | ||
| + | console.log(`unmapped residue: ${unmapped.length}${unmapped.length ? ' -> ' + unmapped.join('; | ||
| + | } | ||
| + | |||
| + | // ----------------------------------------------- figures the page carried ---- | ||
| + | // Added 2026-08-12. Each of these was on design: | ||
| + | // script, so none of them could be re-derived when the corpus grew. If a figure | ||
| + | // is on the page it belongs here. | ||
| + | head(' | ||
| + | { | ||
| + | const imc = rows.filter((p) => p.venue === ' | ||
| + | const imcIp = ipClassified.filter((p) => p.venue === ' | ||
| + | console.log(`IMC: | ||
| + | |||
| + | // MaxMind, across BOTH fields the fold covers, vs within the IP-classifying set. | ||
| + | const mmAll = new Set(); | ||
| + | const mmIp = new Set(); | ||
| + | const mmSpellings = new Set(); | ||
| + | for (const p of rows) { | ||
| + | let hitIp = false; | ||
| + | let hit = false; | ||
| + | for (const t of p.classification) { | ||
| + | if (t.target !== ' | ||
| + | const f = foldIpResource(t.resourceName); | ||
| + | if (f && f.family === ' | ||
| + | } | ||
| + | for (const t of p.vantage) { | ||
| + | if (isSentinel(t.geolocationService)) continue; | ||
| + | const f = foldIpResource(t.geolocationService); | ||
| + | if (f && f.family === ' | ||
| + | } | ||
| + | if (hit) mmAll.add(key(p)); | ||
| + | if (hitIp) mmIp.add(key(p)); | ||
| + | } | ||
| + | console.log(`MaxMind: | ||
| + | `so ${mmAll.size - mmIp.size} name it only for their own vantage point. ${mmSpellings.size} distinct spellings.`); | ||
| + | |||
| + | // Crawling papers: classify an observed address vs geolocate their own vantage. | ||
| + | const crawlIp = crawled.filter((p) => ipTuples(p).length > 0).length; | ||
| + | const crawlGeo = crawled.filter((p) => p.vantage.some((t) => !isSentinel(t.geolocationService))).length; | ||
| + | const crawlEither = crawled.filter( | ||
| + | (p) => ipTuples(p).length > 0 || p.vantage.some((t) => !isSentinel(t.geolocationService)) | ||
| + | ).length; | ||
| + | console.log(`Crawling papers: ${crawlIp} classify an observed address, ${crawlGeo} geolocate their own vantage point, ` + | ||
| + | `union ${crawlEither} of ${crawled.length} (${pct(crawlEither, | ||
| + | } | ||
| + | </ | ||
| + | |||
| + | Output ('' | ||
| + | |||
| + | <file text report_ip_classification-output.txt> | ||
| + | corpus | ||
| + | crawled | ||
| + | measuredFrom | ||
| + | ipClassified | ||
| + | |||
| + | ### Papers classifying an IP address, by venue and period | ||
| + | |||
| + | Venue Papers classifying an IP Papers in corpus | ||
| + | ------- | ||
| + | IMC 124 | ||
| + | USENIX | ||
| + | NDSS | ||
| + | CCS 24 990 2.4% | ||
| + | WWW 24 843 2.8% | ||
| + | IEEE-SP | ||
| + | PETS | ||
| + | |||
| + | Period | ||
| + | ---------- | ||
| + | 2010–2013 | ||
| + | 2014–2017 | ||
| + | 2018–2021 | ||
| + | 2022–2024 | ||
| + | 2025–2026* | ||
| + | |||
| + | ### How the IP was classified (enum, papers of 295) | ||
| + | |||
| + | Method | ||
| + | ------------------- | ||
| + | third-party-service | ||
| + | curated-database | ||
| + | heuristic-rules | ||
| + | blocklist | ||
| + | regex-or-signature | ||
| + | other 9 3.1% | ||
| + | manual-labelling | ||
| + | supervised-ml | ||
| + | graph-analysis | ||
| + | dynamic-analysis | ||
| + | static-analysis | ||
| + | llm 1 0.3% | ||
| + | |||
| + | ### Whether the IP classification was validated (papers of 295) | ||
| + | |||
| + | Validation | ||
| + | -------------------------- | ||
| + | not-applicable | ||
| + | none-reported | ||
| + | manual-validation | ||
| + | comparison-to-other-method | ||
| + | held-out-test-set | ||
| + | cross-validation | ||
| + | |||
| + | names a ground-truth source: | ||
| + | no validation on any IP tuple: 196 / 295 (66.4%) | ||
| + | |||
| + | ### Which resources, folded (papers of 295) | ||
| + | |||
| + | Resource family | ||
| + | -------------------------------------------------------------------- | ||
| + | Home-grown heuristic or classifier | ||
| + | MaxMind | ||
| + | Other IP blocklists (DShield, FireHOL, CBL, AbuseIPDB, Honey Pot, …) Is it known-bad? | ||
| + | IPinfo | ||
| + | Router alias / router-to-AS (bdrmapIT, MAP-IT, MIDAR, Hoiho) | ||
| + | Team Cymru IP-to-ASN | ||
| + | Censys / Shodan / Nmap / Snort / Suricata | ||
| + | IP2Location | ||
| + | VirusTotal / Google Safe Browsing | ||
| + | RouteViews | ||
| + | WHOIS / IRR / RIR delegation files Whose network is it? 13 4.4% | ||
| + | Spamhaus | ||
| + | CAIDA datasets (prefix2as, AS2Org, ITDK) Whose network is it? 11 3.7% | ||
| + | Raw BGP feeds and IX data Whose network is it? 7 2.4% | ||
| + | RIPE RIS / RIPEstat / RIPE Atlas Whose network is it? 7 2.4% | ||
| + | Free geo-lookup APIs (freegeoip, ipstack, HostIP, IPInfoDB, …) Where is it? 6 2.0% | ||
| + | PeeringDB | ||
| + | NetAcuity (Digital Element) | ||
| + | RIPE IPmap Where is it? 5 1.7% | ||
| + | Fraud-score APIs (IPQualityScore, | ||
| + | ASdb Whose network is it? 5 1.7% | ||
| + | GreyNoise | ||
| + | Unnamed commercial geo database | ||
| + | Geocoding / positioning reference (GeoNames, Google, Skyhook, WiGLE) | ||
| + | Chinese geo databases (Chunzhen/ | ||
| + | MaxMind Anonymous IP / minFraud | ||
| + | Quova / Neustar | ||
| + | Chainalysis | ||
| + | IP2Proxy | ||
| + | pyasn / iptoasn.com | ||
| + | Spur What kind of host is it? 2 0.7% | ||
| + | |||
| + | families named by exactly one paper: 9 (ip-api.com; | ||
| + | unfolded residue: 0 distinct strings | ||
| + | |||
| + | ### How badly exact-string counting undercounts (MaxMind) | ||
| + | |||
| + | distinct spellings of MaxMind: | ||
| + | papers, folded: | ||
| + | papers under the commonest spelling (" | ||
| + | undercount if you count exact strings: 79% | ||
| + | |||
| + | ### Papers that geolocate their own vantage point | ||
| + | |||
| + | Population | ||
| + | ------------ | ||
| + | crawled | ||
| + | measuredFrom | ||
| + | |||
| + | of the 194 papers naming one: | ||
| + | Service family | ||
| + | -------------------------------------------------------------------- | ||
| + | MaxMind | ||
| + | IPinfo | ||
| + | Geocoding / positioning reference (GeoNames, Google, Skyhook, WiGLE) | ||
| + | IP2Location | ||
| + | Free geo-lookup APIs (freegeoip, ipstack, HostIP, IPInfoDB, …) 7 3.6% | ||
| + | ip-api.com | ||
| + | RIPE IPmap 7 3.6% | ||
| + | Unnamed commercial geo database | ||
| + | NetAcuity (Digital Element) | ||
| + | CDN / platform internal geo | ||
| + | Chinese geo databases (Chunzhen/ | ||
| + | Akamai EdgeScape | ||
| + | Team Cymru IP-to-ASN | ||
| + | Quova / Neustar | ||
| + | Home-grown heuristic or classifier | ||
| + | WHOIS / IRR / RIR delegation files 2 1.0% | ||
| + | named by one paper each: 4 families | ||
| + | |||
| + | ### Do the papers say which snapshot of the database they used? | ||
| + | |||
| + | papers using a third-party geo or routing dataset: 162 | ||
| + | ...whose evidence quote carries any date/ | ||
| + | (text proxy, not a schema field — treat as an upper bound) | ||
| + | |||
| + | ### Measured figures on geolocation-database accuracy (detection[].prevalence) | ||
| + | |||
| + | 630 prevalence-bearing detection tuples match the IP-classification regex | ||
| + | (full dump written to out/ | ||
| + | |||
| + | done. | ||
| + | |||
| + | ### Papers citing an operator-published cloud prefix-list URL (full-text grep) | ||
| + | |||
| + | 7 of 5859 papers | ||
| + | 2017 IMC large-scale-scanning-of-tcps-initial-window | ||
| + | 2019 IEEE-SP resident-evil-understanding-residential-ip-proxy-as-a-dark-service | ||
| + | 2020 CCS censored-planet-an-internet-wide-longitudinal-censorship-observatory | ||
| + | 2020 NDSS cdn-judo-breaking-the-cdn-dos-protection-with-itself | ||
| + | 2021 CCS warmonger-inflicting-denial-of-service-via-serverless-functions-in-the-cloud | ||
| + | 2023 USENIX dscope-a-cloud-native-internet-telescope | ||
| + | 2025 NDSS secure-ip-address-allocation-at-cloud-scale | ||
| + | (a floor: undercounts papers that used a list without citing its URL) | ||
| + | |||
| + | ### Papers naming more than one geolocation source | ||
| + | |||
| + | name >=1 geolocation source: | ||
| + | ...of which name > | ||
| + | 2010 IMC Eyeball ASes: from geography to connectivity. | ||
| + | 2017 IMC A look at router geolocation in public and commercial databases. | ||
| + | 2018 IMC An Empirical Analysis of the Commercial VPN Ecosystem. | ||
| + | 2018 IMC Tracing Cross Border Web Tracking. | ||
| + | 2022 CCS An Extensive Study of Residential Proxies in China. | ||
| + | 2022 IMC Are we ready for metaverse?: a measurement study of social virtual reality pla | ||
| + | 2022 USENIX Gossamer: Securely Measuring Password-based Logins | ||
| + | 2023 IEEE-SP IPvSeeYou: Exploiting Leaked Identifiers in IPv6 for Street-Level Geolocation. | ||
| + | 2023 IMC How to Operate a Meta-Telescope in your Spare Time. | ||
| + | 2023 IMC Replication: | ||
| + | 2024 IMC A First Look at Immersive Telepresence on Apple Vision Pro. | ||
| + | 2024 IMC Watching TV with the Second-Party: | ||
| + | 2025 NDSS Wallbleed: A Memory Disclosure Vulnerability in the Great Firewall of China | ||
| + | 2025 USENIX eSIMplicity or eSIMplification? | ||
| + | name >=1 ' | ||
| + | name >=1 ' | ||
| + | earliest 2018, 16 from 2019 or later | ||
| + | |||
| + | ### Free-text name universe that ip_fold.mjs covers | ||
| + | |||
| + | distinct strings: 443 | ||
| + | unmapped residue: 0 | ||
| + | |||
| + | ### Figures the page carries that were not previously printed | ||
| + | |||
| + | IMC: 124 of 638. Other six venues: 171 of 5221. | ||
| + | MaxMind: 134 papers across both fields, 65 of them inside the 295 IP-classifying papers, so 69 name it only for their own vantage point. 50 distinct spellings. | ||
| + | Crawling papers: 54 classify an observed address, 45 geolocate their own vantage point, union 76 of 1120 (6.8%). | ||
| + | </ | ||
| + | |||
| + | <file javascript ip_fold.mjs> | ||
| + | // Fold the free-text names that appear when a paper classifies an IP address. | ||
| + | // | ||
| + | // Scope: the union of | ||
| + | // | ||
| + | // | ||
| + | // 363 distinct strings in run1. MaxMind alone appears under 24 spellings | ||
| + | // (" | ||
| + | // MaxMind", | ||
| + | // single most-used resource in the field by a factor of about four. | ||
| + | // | ||
| + | // Rules, same as scripts/ | ||
| + | // * explicit ordered regex families, first match wins | ||
| + | // * specific products before the generic term they contain | ||
| + | // | ||
| + | // * everything unmatched is returned as residue and PRINTED, never dropped | ||
| + | // | ||
| + | // Each family also carries the *question* it answers, because " | ||
| + | // is five different measurements with five different error budgets. | ||
| + | |||
| + | /** @typedef {' | ||
| + | |||
| + | /** @type {{family: string, task: Task, re: RegExp}[]} */ | ||
| + | export const FAMILIES = [ | ||
| + | // ---- proxy / VPN / hosting detection: match before their parent vendors ---- | ||
| + | { family: ' | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { family: ' | ||
| + | { family: ' | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | |||
| + | // ---- geolocation databases and services ---- | ||
| + | { family: ' | ||
| + | { family: ' | ||
| + | { family: ' | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { family: ' | ||
| + | { family: 'Quova / Neustar', | ||
| + | { family: ' | ||
| + | { family: ' | ||
| + | { family: 'RIPE IPmap', | ||
| + | { | ||
| + | family: 'Free geo-lookup APIs (freegeoip, ipstack, HostIP, IPInfoDB, …)', | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { | ||
| + | family: 'CDN / platform internal geo', | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | |||
| + | // ---- routing / ASN ---- | ||
| + | { family: 'Team Cymru IP-to-ASN', | ||
| + | { family: ' | ||
| + | { | ||
| + | family: 'RIPE RIS / RIPEstat / RIPE Atlas', | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { family: 'CAIDA datasets (prefix2as, AS2Org, ITDK)', | ||
| + | { family: ' | ||
| + | { family: ' | ||
| + | { family: ' | ||
| + | { family: 'pyasn / iptoasn.com', | ||
| + | { | ||
| + | family: 'WHOIS / IRR / RIR delegation files', | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { | ||
| + | family: 'Raw BGP feeds and IX data', | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | |||
| + | // ---- reputation / abuse blocklists ---- | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { | ||
| + | family: 'Other IP blocklists (DShield, FireHOL, CBL, AbuseIPDB, Honey Pot, …)', | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | { family: ' | ||
| + | // Criminal IP is an attack-surface search engine in the Censys/ | ||
| + | // marketed as threat intelligence. Filed under scanning, which is what it does. | ||
| + | { family: ' | ||
| + | // Published hosting/ | ||
| + | // Distinct from a fraud-score API: this is the provider' | ||
| + | { family: ' | ||
| + | // 2026-08-12: the first LLM in this field. One paper, GPT-4o. Kept as its own | ||
| + | // family rather than folded into ' | ||
| + | { family: 'LLM (GPT-4o)', | ||
| + | // Email-authentication records: not an IP classifier, but the extraction files | ||
| + | // them here because the unit of analysis is the sending IP. | ||
| + | { family: 'Email authentication (SPF/ | ||
| + | // Phone-number reference services, from one paper whose unit was a phone | ||
| + | // number rather than an IP. Non-IP reference data, like the geocoding row. | ||
| + | { family: ' | ||
| + | |||
| + | // ---- active scanning / host fingerprinting ---- | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | |||
| + | // ---- router topology: alias resolution and router-to-AS ---- | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | |||
| + | // ---- non-IP location reference data ---- | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | |||
| + | // ---- home-grown ---- | ||
| + | { | ||
| + | family: ' | ||
| + | task: ' | ||
| + | re: / | ||
| + | }, | ||
| + | ]; | ||
| + | |||
| + | /** | ||
| + | * @param {string} name | ||
| + | * @returns {{family: string, task: Task} | null} | ||
| + | */ | ||
| + | export function foldIpResource(name) { | ||
| + | const s = String(name).trim(); | ||
| + | for (const f of FAMILIES) { | ||
| + | if (f.re.test(s)) return { family: f.family, task: f.task }; | ||
| + | } | ||
| + | return null; | ||
| + | } | ||
| + | |||
| + | export const TASK_LABEL = { | ||
| + | geolocation: | ||
| + | routing: 'Whose network is it?', | ||
| + | ' | ||
| + | reputation: 'Is it known-bad?', | ||
| + | topology: 'Is it a router, and which one?', | ||
| + | scanning: 'What is running on it?', | ||
| + | reference: ' | ||
| + | custom: ' | ||
| + | }; | ||
| + | </ | ||
| + | |||
| + | <file javascript maxmind_version.mjs> | ||
| + | // Do papers that use MaxMind say WHICH snapshot of MaxMind they used? | ||
| + | // | ||
| + | // A geolocation database is a moving target: MaxMind reissues GeoLite2 twice a | ||
| + | // week. "We used MaxMind" | ||
| + | // extraction schema has no version field for a classification resource, so this | ||
| + | // measures it directly against the full text: for every paper whose extraction | ||
| + | // names MaxMind, find every sentence in paper.cols.txt that mentions MaxMind and | ||
| + | // ask whether any of them carries a date, a month, or a version/ | ||
| + | // | ||
| + | // node scripts/ | ||
| + | // | ||
| + | // This is a generous test. A sentence saying "we crawled in March 2019 using | ||
| + | // MaxMind" | ||
| + | // Read the figure as an UPPER BOUND on how often the snapshot is identifiable. | ||
| + | |||
| + | import fs from ' | ||
| + | import path from ' | ||
| + | import { loadExtractions, | ||
| + | import { foldIpResource } from ' | ||
| + | |||
| + | const DUMP = process.argv.includes(' | ||
| + | const root = dataRoot(); | ||
| + | const rows = loadExtractions(); | ||
| + | |||
| + | const isMaxMind = (s) => s && !isSentinel(s) && foldIpResource(s)? | ||
| + | const users = rows.filter( | ||
| + | (p) => | ||
| + | p.classification.some((t) => t.target === ' | ||
| + | p.vantage.some((t) => isMaxMind(t.geolocationService)) | ||
| + | ); | ||
| + | |||
| + | // A month name, a year, or an explicit version/ | ||
| + | const DATED = | ||
| + | / | ||
| + | |||
| + | // A bibliography entry (" | ||
| + | // year but tells the reader nothing about which snapshot was queried. Drop | ||
| + | // sentences that look like reference-list entries before the strict count. | ||
| + | const BIBLIKE = / | ||
| + | |||
| + | let read = 0, | ||
| + | missing = 0, | ||
| + | mentioned = 0, | ||
| + | dated = 0, | ||
| + | datedStrict = 0; | ||
| + | const examples = []; | ||
| + | |||
| + | for (const p of users) { | ||
| + | const file = path.join(root, | ||
| + | if (!fs.existsSync(file)) { | ||
| + | missing += 1; | ||
| + | continue; | ||
| + | } | ||
| + | read += 1; | ||
| + | const text = fs.readFileSync(file, | ||
| + | // Split on sentence-ish boundaries; keep it crude, the unit is " | ||
| + | const sentences = text.split(/ | ||
| + | const hits = sentences.filter((s) => / | ||
| + | if (hits.length === 0) continue; // extraction says MaxMind, text does not — see note | ||
| + | mentioned += 1; | ||
| + | const withDate = hits.filter((s) => DATED.test(s)); | ||
| + | const strict = withDate.filter((s) => !BIBLIKE.test(s)); | ||
| + | if (strict.length > 0) datedStrict += 1; | ||
| + | if (withDate.length > 0) { | ||
| + | dated += 1; | ||
| + | if (strict.length > 0 && examples.length < 8) | ||
| + | examples.push([`${p.year}/ | ||
| + | } else if (DUMP) { | ||
| + | console.log(`UNDATED | ||
| + | } | ||
| + | } | ||
| + | |||
| + | console.log(`papers whose extraction names MaxMind: | ||
| + | console.log(` | ||
| + | console.log(` | ||
| + | console.log( | ||
| + | ` ...with a date/ | ||
| + | ); | ||
| + | console.log( | ||
| + | ` ...excluding sentences that are bibliography entries: | ||
| + | ); | ||
| + | console.log(' | ||
| + | for (const [k, s] of examples) console.log(` | ||
| + | </ | ||
| + | |||
| + | Output ('' | ||
| + | |||
| + | <file text maxmind_version-output.txt> | ||
| + | papers whose extraction names MaxMind: | ||
| + | full text available: | ||
| + | MaxMind/ | ||
| + | ...with a date/ | ||
| + | ...excluding sentences that are bibliography entries: | ||
| + | |||
| + | Still an upper bound: any year token in the sentence counts, including crawl dates. | ||
| + | |||
| + | 2011/ | ||
| + | IPv4 address space delegated to Egypt (as of January 24, 2011) and Libya (as of February 15, 2011) by AfriNIC (top half) as well as additional IPv4 address ranges associated with the two countries based on MaxMind GeoLite database (as of Ja | ||
| + | |||
| + | 2013/ | ||
| + | Using the database dated from February 1, 2011 (so number of HTTP requests for all LDNS across all their TTL that our analysis would reflect the GeoIP map at the time intervals. | ||
| + | |||
| + | 2014/ | ||
| + | We geolocalize each IP address in DIP v4 using the Maxmind GeoIP database.9 We then introduce, for each identified country, the cor- 6. | ||
| + | |||
| + | 2015/ | ||
| + | All IPs returned in each hop of the traceroute were We use a simplified version of this check when examining geo-located with the MaxMind GeoLite27 country databases. | ||
| + | |||
| + | 2014/ | ||
| + | GeoIP, 2013. | ||
| + | |||
| + | 2015/ | ||
| + | Since MaxMind updates the database regularly (to reflect changes in the address space), we use the databases produced on August 1, 2012 and August 16, 2013 for the 2012 census and 2013 census periods, respectively. | ||
| + | |||
| + | 2015/ | ||
| + | However, 49% reach the AS of the destination, | ||
| + | |||
| + | 2017/ | ||
| + | First, our recommendations 467 IMC '17, November 1-3, 2017, London, United Kingdom 0.0 0.2 0.4 0.6 0.8 1.0 ARIN (4761) APNIC (468) AFRINIC (58) LACNIC (38) RIPENCC (1523) CDF AFRINIC APNIC −4 −3 −2 −1 0 10 10 10 10 10 (a) MaxMind-Paid (41.2 | ||
| + | |||
| + | </ | ||
| + | |||
| + | ==== 14.9 Re-review of the fixes ==== | ||
| + | |||
| + | Both reviewers whose findings were acted on re-read the drafted fixes **before** they were saved, told again that their context might not be exhaustive, and wrote to '' | ||
| + | |||
| + | ^ ID ^ From ^ Sev. ^ Finding ^ Action ^ | ||
| + | | R1 | fable | MAJOR | The vendor-table date line said MaxMind dropped the " | ||
| + | | R2 | fable | MINOR | '' | ||
| + | | R3 | fable | MINOR | §14.6 said the two passes " | ||
| + | | R4 | fable | MINOR (PLAUSIBLE) | The IP2Location data-licence quote had no fetched copy in the run directory | **Fixed**: '' | ||
| + | | R5 | fable | MINOR | The flags box said "this page's 2026-08-06 run of the script below", | ||
| + | | R6 | fable | note | §14.2 says four unrecorded revisions where fable' | ||
| + | | — | sonnet | — | 11 changed external claims re-fetched and CONFIRMED (MaxMind EULA date and prices, IP2Location badge only on DB1, IP2Proxy LITE = '' | ||
| + | |||
| + | The fixes to R1–R5 were not re-reviewed a third time. | ||
| [[design: | [[design: | ||
provenance/design/ip_classification.1786526634.txt.gz · Last modified: by karel.kubicek.claude
