User Tools

Site Tools


provenance:design:ip_classification

Provenance: design:ip_classification

Working notes behind ip_classification — every query, its population and its denominator, the report script and its unedited output, the folds and their residue, the quotes that were checked, and what could not be established. Corpus-level caveats that apply to every page on this site are on corpus and are not restated here.

Partly reconstructed. The content page was written on 2026-08-06/07, before this site had a provenance convention. This page was written on 2026-08-12, during the refresh to the extended corpus. Sections marked recorded were produced by the 2026-08-12 run. Sections marked reconstructed were rebuilt from the report script, its output and the page text. The page's long non-corpus half — the geolocation-accuracy literature, the anycast discussion, the vendor comparison, the classify_ips.py script — was written and verified by the original run and is not re-derivable here; it was not re-verified on 2026-08-12.

1. What this page is backing

Item Value
Content page ip_classification
Report script scripts/report_ip_classification.mjs (–wiki emits DokuWiki tables)
Folds scripts/ip_fold.mjs — ordered regex families, each tagged with the question it answers
Supporting script scripts/maxmind_version.mjs — full-text pass, not a schema field
Quote verification scripts/quote_check.mjs –classification ip-address
Data data/extract/run1/extractions.jsonl, 5,859 papers, 7 venues, 2010–2026
Refreshed 2026-08-12

2. Populations and denominators

Recorded.

Tag Definition N
ipClassified ≥1 classification[] tuple with target === “ip-address” 295
crawled crawlConfig !== null OR studyTypes contains automated-web-crawl 1,120
measuredFrom vantage.length > 0 3,908
names MaxMind (either field) folds to the MaxMind family in classification.resourceName where target is an IP, or in vantage.geolocationService 134
names a geolocation-task source ≥1 folded family whose task is geolocation 114

The two fields are different questions and the page keeps them apart. classification.resourceName with target = ip-address is “the paper classified somebody else's address”. vantage.geolocationService is “the paper located its own measurement point”. A paper doing only the second is not in the 295. 69 of the 134 MaxMind papers are of that second kind.

3. Running it

cd /workspace/artifacts/wiki
node scripts/ip_fold.mjs                                  # self-test, prints residue
node scripts/report_ip_classification.mjs                 # every figure
node scripts/report_ip_classification.mjs --wiki
node scripts/maxmind_version.mjs                          # the snapshot-reporting figure
node scripts/quote_check.mjs --classification ip-address --show 94
node scripts/check_page_numbers.mjs \
  pages/design_ip_classification.txt out/new/report_ip_classification.txt \
  '===== Use in Publications =====' '===== What to Report ====='

check_page_numbers.mjs left 11 figures unaccounted, all deliberate and all named here so the next run does not re-investigate them: 100 (“shares exceed 100%”), 26.9 (from maxmind_version.mjs, a different script), 4,322 (the old corpus size, quoted as history), and seven figures quoted from cited papers — 89.4 and 95.8 from Gharaibeh et al., 93 from Urban et al., 4,286 / 72 / 87 / 98.3 from Kumar et al., and 2012 inside the Benson et al. quote.

4. What the refresh changed

Recorded. Old = 4,322-paper corpus. New = 5,859-paper corpus.

Figure Old New
Papers classifying an IP 234 (5.4%) 295 (5.0%)
IMC 109 of 559 (19.5%) 124 of 638 (19.4%)
Other six venues 125 of 3,763 171 of 5,221
Third-party service 101 (43.2%) 124 (42.0%)
Curated database 86 (36.8%) 109 (36.9%)
Heuristic rules 74 (31.6%) 88 (29.8%)
Supervised ML 7 (3.0%) 7 (2.4%)
LLM — (enum never fired) 1 (0.3%)
No validation at all 154 of 234 (65.8%) 196 of 295 (66.4%)
Names a ground-truth source 97 (41.5%) 127 (43.1%)
Home-grown heuristic 83 (35.5%) 104 (35.3%)
MaxMind, in IP-classifying papers 56 (23.9%) 65 (22.0%)
MaxMind, both fields 111 134
MaxMind distinct spellings 41 50
MaxMind exact-string undercount 77% 79%
IPinfo 15 (6.4%) 26 (8.8%)
Other IP blocklists 25 (10.7%) 33 (11.2%)
Names ≥1 geolocation source 92 114
…of which ≥2 12 (13.0%) 14 (12.3%)
MaxMind papers with a date token 26 of 111 (23.4%) 36 of 134 (26.9%)
Crawling papers naming a geo service for their vantage 33 of 859 (3.8%) 45 of 1,120 (4.0%)
measuredFrom papers doing so 157 of 2,909 (5.4%) 194 of 3,908 (5.0%)
Crawling papers geolocating anything 57 of 859 (6.6%) 76 of 1,120 (6.8%)
MaxMind's share of vantage-geolocating papers 88 (56.1%) 108 (55.7%)
Distinct free-text names the fold covers 363 443
Families named by exactly one paper 5 9

Moved by more than a rounding step:

  • IPinfo, 15 → 26 papers (6.4% → 8.8%). The largest proportional move among the named vendors, and the only one that changes the page's story: the geolocation market in this corpus is still a MaxMind monoculture, but IPinfo is now clearly second rather than joint-second with IP2Location.
  • The llm method fires for the first time — once. One paper, GPT-4o. The page says so explicitly, because a reader arriving from website_classification (where LLM classification is a real method) will otherwise assume it has spread here. It has not.
  • MaxMind date-token share, 23.4% → 26.9%, on a base that grew from 111 to 134. It is still an upper bound for the two reasons the page gives.
  • Papers classifying an IP, share of corpus 4.5% → 4.3% and flat into 2025–2026. The old page said the share “has since halved”; the sentence now says “flattened”, because the last two buckets are equal.

5. Folds

5.1 ip_fold.mjs went from zero residue to 18, and back to zero

Recorded. This is the clearest case on the site of a fold ageing silently. On the 4,322-paper corpus ip_fold.mjs had zero unmatched strings, and the page said so. On the 5,859-paper corpus it had 18. Every one was mappable; here is where each went, so the judgement calls are visible rather than buried in a regex:

Residue string Folded to Call
IPGeolocation.io Free geo-lookup APIs obvious
IPtoASN pyasn / iptoasn.com (family renamed) obvious
IANA IPv4 Special-Purpose Address Registry WHOIS / IRR / RIR delegation files registry reference data, same question
SpamCop, SinkDB and MISP Project sinkhole lists, public ASN block lists, BL-A, institutional list by Griffioen et al. Other IP blocklists BL-A is an anonymised list name; institutional list by Griffioen et al. is another paper's published list, and putting it with the blocklists rather than with Home-grown is a judgement call — it is not the citing paper's own list
manual R&E/commodity neighbor classification, I2 PERCEPTION live behavior correlation, ASN matching against b-MNO, v-MNO, and third-party providers, FACT architecture_detection plugin Home-grown heuristic or classifier all four are the paper's own rule
Criminal IP Censys / Shodan / Nmap / Snort / Suricata judgement call. Criminal IP markets itself as threat intelligence, which would put it under Is it known-bad?. It is an attack-surface search engine, which is what that family is. Filed under What is running on it? and flagged here because a reasonable person would file it the other way.
cloud-provider-ip-addresses new family: Cloud/hosting provider published IP ranges the provider's own list, not a third-party score
SPF verification new family: Email authentication (SPF/DMARC) not an IP classifier; the extraction files it here because the unit of analysis is the sending IP
Twilio, OpenCNAM new family: Phone-number reference (not IP-based) one paper whose unit was a phone number; parallels the existing non-IP geocoding reference family
GPT-4o new family: LLM (GPT-4o) kept separate from Home-grown on purpose, so it stays countable as the field's first

A second residue exists in the wider “name universe” section, which folds vantage.geolocationService too. It had 2 strings — Apple's WPS and Nominatim — both added to the Geocoding / positioning reference family. Both residues are now zero and both are printed on every run.

5.2 Two bugs in the report script itself

Recorded. Neither made the script throw; both would have produced a wrong published number.

  1. 234 was hardcoded in four table headings and their share labels. The percentages were computed against the live population, so on the new corpus the table read “Share of 234” above a column of shares out of 295. Replaced with ${ipClassified.length}.
  2. The last period bucket was (y) ⇒ y >= 2022, not 2022–2024. Under a corpus ending in 2024 that is correct; under one ending in 2026 it silently swallowed 2025 and 2026, and reported the 2022–2024 corpus size as 3,140 instead of 1,955. Split into 2022–2024 and a starred 2025–2026. Any figure copied from that row before 2026-08-12 is wrong.

6. Quotes checked

Recorded, 2026-08-12.

$ node scripts/quote_check.mjs --classification ip-address
484 quotes checked: 232 exact, 158 partial (>=60% of 5-word windows),
94 below threshold, 0 with no full text on disk.

94 below threshold is too many to read individually and they were not all read. Six were sampled and checked by hand against paper.cols.txt with whitespace normalised — 2012/IMC/breaking-for-commercials, 2012/USENIX/aurasium, 2014/CCS/autoprobe, 2015/NDSS/mind-your-blocks, 2010/IMC/demystifying-service-discovery, 2013/NDSS/automatically-inferring-the-evolution — and all six are present in the paper. The failure mode is always the column repair, e.g. the MaxMind quote in breaking-for-commercials reads in the source as:

commer- Other apps related to sport, TV/cinema or social networking, such
cial database provided by MaxMind3 that maps an IP address to as Grindr,
instead require network access to perform properly. As the name of the organizat

Open, and stated as such on this page rather than on the content page: the other 88 have not been read. A 19% below-threshold rate is higher than the –tools checks produce (15%), which is what you would expect from classification quotes being longer and more often spliced, but it has not been demonstrated. Reading them is the obvious next piece of work on this page.

Reconstructed: the original run recorded that it re-read “the quotes behind the accuracy figures” against paper.cols.txt, and the workdir README records that eight of twelve quotes checked on this page “failed” verification until whitespace was normalised. Which twelve is not recoverable.

7. External and industry sources

Reconstructed. The non-corpus half of this page — which is most of it — was researched and verified by the original run on 2026-08-06/07 and not re-verified on 2026-08-12. Its “checked 2026-08-xx” dates are accurate as of then. What is recorded:

  • The geolocation-accuracy claims are all cited to papers in bibliography; Gharaibeh et al. and Darwich et al. carry the two figures the page leans on hardest (89.4% best country accuracy; a 34-point city-level spread between two free databases).
  • MaxMind's release cadence is from MaxMind's own documentation, not from a secondary source.
  • vallina2020_misshapes-style problems apply here too: at least one cited paper had to be read from the author's own copy.
  • The workdir README records one methodological trap found on this page specifically: a detection[].prevalence value asserted a city-level result its own evidence.quote only supported at country level. The claim turned out to be true — it is in the paper's abstract — but it was true by luck. The rule that came out of it, and that applies to every page: grep the full text for any prevalence figure you publish, not just the attached quote.

Rejected: not recorded for the original run.

  • DynamIPs ([1Padmanabhan, Ramakrishna; Rula, John P.; Richter, Philipp; Strowes, Stephen D.; Dainotti, Alberto (2020): "DynamIPs: Analyzing Address Assignment Practices in IPv4 and IPv6", in: Proceedings of the 16th International Conference on Emerging Networking Experiments and Technologies, pp. 55-70. (DOI)], CoNEXT 2020) — the one external paper the original run could not read. It was fetched and its wording verified on 2026-09-03; see §12.

8. What could not be established

  • Whether the 88 unread below-threshold quotes check out. See §6.
  • Free vs paid MaxMind. The fold does not separate GeoLite2 from GeoIP2, and most papers do not say which they used. The page says so. Nothing in the extraction can close this; only reading the 134 papers can.
  • Whether “no validation” is a reporting gap or a real one. 41.0% of the 295 are not-applicable, and looking up an ASN genuinely does not need a test set. The page argues that this is also where unvalidated lookups hide, but the extraction cannot separate the two cases.
  • Venue coverage is worse for this page than for any other on the site. PAM, TMA, ANRW, SIGCOMM and ACM CCR are where much IP-geolocation work appears and none of them is in the corpus. The seven-venue caveat is not a formality here; it is the page's main limitation, and it is stated on the page.
  • Criminal IP's family. See §5.1. One paper, and it could reasonably go under Is it known-bad?.

10. Review pass, 2026-08-12

Recorded. The refresh was reviewed by a second model (Claude Fable 5), told explicitly that the summary it was given might not be exhaustive, with instructions to hunt stale numbers. It found a systematic defect, not a scatter of typos, and it is worth stating because it will recur on the next refresh:

  • check_page_numbers.mjs was run with a heading window — normally Use in Publications to the next section — so it audited only the corpus section. Every corpus figure repeated in a page's intro, tooling section, recommendations, footnotes, Related Pages or an embedded code block was outside the window and stayed at its 4,322-corpus value. Across the six pages 29 such figures survived the first pass.
  • The matcher was substring, not word-boundary, so report.includes('59') was satisfied by 11.59 bits. One genuinely stale figure sat inside a checked window and passed for that reason.

Both are fixed in scripts/check_page_numbers.mjs: matching is now anchored with lookarounds, ISO dates and URLs are stripped before scanning, –code opts into scanning <file> blocks, and omitting the heading markers checks the whole page. Run it windowed and whole-page. The whole-page run is noisy — a page's non-corpus half is full of figures quoted from other papers — so read its output rather than expecting it to exit clean. Fixed on this page's content page as a result: the intro paragraph, which carried 234 / 92 / 53 where the corpus section says 295 / 114 / 68; the operator-prefix-list full-text grep, six papers of 4,322 → seven of 5,859 (the seventh is NDSS/2025/secure-ip-address-allocation-at-cloud-scale); the abuse-feed count 13 of 234 → 17 of 295; and a surviving corpus-window statement — the page said Ali et al. (NDSS 2026) was “published after our corpus closes”, and it is now in the corpus at NDSS/2026/beyond-rtt-an-adversarially-robust-two-tiered-approach-for-residential-proxy-detection.

11. Run log

Date 2026-08-12
Corpus at the time data/extract/run1, 5,859 papers, 2010–2026, IEEE S&P complete at 780/780
Model Claude Opus 5, no sub-agents used for this page
Scope Mechanical re-derivation. Prose, structure and method selection were not revisited; two sentences changed because the numbers no longer supported them (“has since halved” → “flattened”; the addition of the llm sentence).
Script changes ip_fold.mjs (5 new families, 6 extended patterns, §5.1), report_ip_classification.mjs (two bugs fixed, §5.2; new closing section printing the figures the page carried but the script did not), quote_check.mjs (gained –classification)
Caveats deleted “IEEE S&P is only 43% retrieved (paywall)” — 780 of 780 selected papers are now retrieved. “The corpus ends in 2024.”
Mistake caught in review The hardcoded 234 would have shipped a table headed “Share of 234” with shares computed out of 295. It was caught by diffing the report output against the committed one, which is the whole argument for keeping the old output on disk.
Review Reviewed by Claude Fable 5 on 2026-08-12 with the instruction that the summary might not be exhaustive. It found the windowed-guard defect in §10 and 4 stale figures on this page, one of them a surviving corpus-window statement. All fixes were applied and re-saved the same day.

12. DynamIPs wording verification, 2026-09-03

Recorded as it happened. The content page cited [1Padmanabhan, Ramakrishna; Rula, John P.; Richter, Philipp; Strowes, Stephen D.; Dainotti, Alberto (2020): "DynamIPs: Analyzing Address Assignment Practices in IPv4 and IPv6", in: Proceedings of the 16th International Conference on Emerging Networking Experiments and Technologies, pp. 55-70. (DOI)] twice (the Churn and IPv6 is different paragraphs of “IP as an Identifier”). The original run had verified venue, authors and topic against Crossref but could not read the paper: ACM DL returned 403 and the CAIDA PDF “would not extract with the tools available”. The second citation therefore carried a footnote saying the finding was paraphrased and unverified. This pass closes that.

Getting the paper.

Source Result
https://dl.acm.org/doi/pdf/10.1145/3386367.3431314 still HTTP 403 with a browser User-Agent
https://www.caida.org/catalog/papers/2020_dynamips/dynamips.pdf HTTP 200, application/pdf, 615,785 bytes, sha256 ac51460c…3b06eb; cached at review_ip/dynamips_conext2020.pdf
catalog.caida.org/paper/2020_dynamips HTTP 200 landing page, confirms the PDF above is the authors' copy

pypdf 6.16.2 extracted all 16 pages (95,238 characters) without trouble; the earlier “would not extract” is not reproducible today. The only artefacts are / ligatures, curly apostrophes, and run-together words in the Conclusion, which is why the checker below matches in three modes (whitespace collapsed; whitespace removed; whitespace and hyphens removed) and prints which one hit. Running page numbers 55–70 in the PDF agree with the bibliography entry.

Quotes checked. scripts/dynamips_quote_check.py re-fetches or reuses the PDF and asserts every quote below is present; it exits non-zero otherwise. Its unedited output:

dynamips_quote_check-output.txt
using cached review_ip/dynamips_conext2020.pdf
sha256 ac51460ca1ca5a3b40fad31cde7d8229b7fed9dab8d56609e4a59ca2c23b06eb  bytes 615785
pages 16  chars 95238  pypdf 6.16.2
first line: DynamIPs: Analyzing address assignment practices in IPv4 and
 
[1] ws      abstract
    "IPv6 assignments have longer durations than IPv4 assignments—often remaining stable for months—thereby allowing the possibility of long-term fingerprinting of IPv6 subscribers"
    …ajor CDN. Our investigation of temporal dynamics with these datasets shows that IPv6 assignments have longer durations than IPv4 assignments—often remaining stable for months—thereby allowing the possibility of long-term fingerprinting of IPv6 subscribers. Our analysis of spatial dynamics reveals IPv6 addressassignment patterns that …
 
[2] ws      §1 contributions
    "IPv6 prefixes delegated to residential subscribers can remain stable for months, permitting long-term use of IPv6 prefixes to identify individual subscribers (at the CPE granularity), even if subscribers' devices are using privacy addresses"
    …ress assignments in IPv4 and IPv6 on over 3,000 dual-stack probes. We find that IPv6 prefixes delegated to residential subscribers can remain stable for months, permitting long-term use of IPv6 prefixes to identify individual subscribers (at the CPE granularity), even if subscribers' devices are using privacy addresses. IPv4-IPv6 interplay: Using a dataset from a major CDN capturing 32.7 billion I…
 
[3] ws      §3 assignment durations
    "Well-defined modes—at 1 day (DTAG), 1.5 days (Proximus), 1 week (Orange), and 2 weeks (BT)—in IPv4 non dual-stack address durations suggest that ISPs renumber addresses periodically"
    …v4 address durations tend to be shorter, particularly for DTAG, Orange, and BT. Well-defined modes—at 1 day (DTAG), 1.5 days (Proximus), 1 week (Orange), and 2 weeks (BT)—in IPv4 non dual-stack address durations suggest that ISPs renumber addresses periodically. This result, using 6 years' worth of "IP echo" data, is consistent with observ…
 
[4] ws      §3 assignment durations
    "we observe evidence of consistent periodic renumbering on 35 networks when considering non-dual-stack probes"
    …or work that also noted periodic renumbering within these ISPs [ 34]. In total, we observe evidence of consistent periodic renumbering on 35 networks when considering non-dual-stack probes. DTAG appears to renumber IPv6 prefixes after 1-day durations as well but this …
 
[5] ws      §3 assignment durations
    "renumbering every 24 hours in IPv6 mainly in the following German ASes: DTAG, Versatel (AS8881), Netcologne (AS8422), Telefonica DE (AS6805), and M-net (AS8767)"
    …arentheses is the total assignment duration in years from all probes in the AS. renumbering every 24 hours in IPv6 mainly in the following German ASes: DTAG, Versatel (AS8881), Netcologne (AS8422), Telefonica DE (AS6805), and M-net (AS8767). We also observe consistent period renumbering with a 12-hour period in ANTEL (…
 
[6] ws      §3 assignment durations
    "12-hour period in ANTEL (AS6057) in Uruguay and with a 48-hour period in Global Village (AS18881) in Brazil"
    …6805), and M-net (AS8767). We also observe consistent period renumbering with a 12-hour period in ANTEL (AS6057) in Uruguay and with a 48-hour period in Global Village (AS18881) in Brazil. Long IPv6 /64 durations in most ASes suggest that a /64 can be used to identif…
 
[7] nows    §3 evolution over time
    "IPv6 durations have consistently been longer than IPv4 durations and address durations in dual-stack networks tend to be longer than durations in non-dual-stack IPv4 networks"
    …(matched in mode nows: pypdf line-break artefact) …mefractionsperyear.Theyear-to-yeartrendsconfirmourinsightsfromtheoveralldataset:IPv6durationshaveconsistentlybeenlongerthanIPv4durationsandaddressdurationsindual-stacknetworkstendtobelongerthandurationsinnon-dual-stackIPv4networks[40].However,wealsofindthatassignment60AnalyzingaddressassignmentpracticesinIPv4…
 
[8] ws      §3 related results
    "IPv6 /64 prefixes tend to be stable for months and years in various ASNs, although we find evidence of periodic renumbering in a handful of ISPs"
    …2014 to March 2015) [36]. Our results using the RIPE Atlas dataset confirm that IPv6 /64 prefixes tend to be stable for months and years in various ASNs, although we find evidence of periodic renumbering in a handful of ISPs. 4 IPV4-IPV6 INTERPLAY So far, we have studied temporal properties of IPv4 and …
 
[9] nows    §7 conclusion
    "IPv6 assignments typically last longer than IPv4 assignments and can persist for months in several large residential ISPs"
    …(matched in mode nows: pypdf line-break artefact) …investigatetemporalandspatialdynamicsofIPv4andIPv6addressassignments.WefoundthatIPv6assignmentstypicallylastlongerthanIPv4assignmentsandcanpersistformonthsinseverallargeresidentialISPs.WestudiedspatialaspectsofIPv6addressesindetail,identifyingsubscriberpoolboundar…
 
[10] ws      abstract (datasets)
    "over 3,000 RIPE Atlas probes in dual-stack networks"
    …mics. We present finegrained observations of dynamics using data collected from over 3,000 RIPE Atlas probes in dual-stack networks. RIPE Atlas probes in these networks report both their IPv4 and their IPv6 addr…
 
[11] ws      abstract (datasets)
    "32.7 billion IPv4 and IPv6 address associations observed by a major CDN"
    …space. To corroborate and extend our findings, we also use a dataset containing 32.7 billion IPv4 and IPv6 address associations observed by a major CDN. Our investigation of temporal dynamics with these datasets shows that IPv6 ass…
 
11 quotes, 11 present, 0 missing

What the paper actually says, against what the page said.

Page wording before Verdict Page wording now
“IPv6 assignments last longer than IPv4 ones, often remaining stable for months” (paraphrase, footnoted as unverified) Accurate. The abstract says “IPv6 assignments have longer durations than IPv4 assignments—often remaining stable for months—thereby allowing the possibility of long-term fingerprinting of IPv6 subscribers”; the conclusion repeats it as “typically last longer … can persist for months in several large residential ISPs”. The abstract sentence, quoted verbatim; footnote dropped. Added the paper's own qualifier that the prefix identifies the subscriber “even if subscribers' devices are using privacy addresses”, which is exactly the point the paragraph makes about RFC 8981.
“found the distribution spans orders of magnitude between ISPs — some reassign on a fixed daily cycle, others leave an address in place for months” Supported but not their phrase. “Orders of magnitude” is our summary of 12-hour cycles (ANTEL) at one end and /64s “stable for months and years” at the other; the paper does not use the words. “Daily cycle” is right: 24-hour IPv6 renumbering in DTAG, Versatel, Netcologne, Telefonica DE and M-net. Rewritten so the specifics are the paper's, and split by protocol as the paper splits them: IPv4 (non-dual-stack probes) — periodic renumbering on 35 networks, modes at 1 day (DTAG), 1.5 days (Proximus), 1 week (Orange), 2 weeks (BT); IPv6 — 12 h (ANTEL), 24 h (DTAG and four other German ASes), 48 h (Global Village); and the “stable for months and years” clause quoted. “Orders of magnitude” kept, but now visibly ours.
“RIPE Atlas dual-stack probes plus 32.7 billion address associations observed by a CDN” Accurate. “over 3,000 RIPE Atlas probes in dual-stack networks”; “32.7 billion IPv4 and IPv6 address associations observed by a major CDN”. Added “six years”, which is the paper's own description of the Atlas window. unchanged apart from “six years”

Judgement calls.

  • The “35 networks” figure is for non-dual-stack IPv4 probes, and the 12 h / 24 h / 48 h cycles come from a separate, IPv6-only sentence. The first draft of the page sentence ran the two together as if ANTEL and Global Village were among the 35; the reviewer caught it (see the run table) and the page now names the paper's own IPv4 modes and keeps the IPv6 examples on their side of a semicolon. DTAG is the only ISP the paper shows renumbering daily in both protocols, so a first-draft “in IPv6 too” attached to all five German ASes was dropped.
  • Not published: the CDN-side figures (median association duration 61 days; 20% of associations lasting more than 143 of a possible 150 days; 75% of mobile associations lasting a day or less). They are about IPv4–IPv6 address associations, not assignment durations, and the paragraph is about assignment lifetime. They are noted here so the next run does not have to re-read the paper to decide.
  • The 45% / 44% of probes that saw no change in over a year are excluded from the paper's duration analysis (probably static assignments) and must not be read as “45% of assignments are stable for a year”. Not published.

Run.

Date 2026-09-03
Scope One citation's wording. No corpus figure touched; report_ip_classification.mjs not re-run.
Model Claude Fable 5.1, no sub-agents
Content page change Two sentences in “IP as an Identifier: Four Ways It Breaks”; one footnote removed. Bibliography entry unchanged (it was already correct).
Mistake caught The first version of the checker re-joined every hyphen at a line break and so reported the “non-dual-stack” quote as missing; a real hyphen and a line-break hyphen are indistinguishable in pypdf output, hence the three matching modes.
Review One focused pass (Claude Sonnet, citations and quotes), given the flattened paper text, the two paragraphs and this section, told its context might not be exhaustive. It confirmed all quotes and every other attributed specific, and returned three findings, all accepted: (1) blocker — the page sentence presented the IPv6 12 h / 24 h / 48 h cycles as instances of the IPv4 “35 networks” finding; (2) should-fix — this section claimed the page scoped “35 networks” in parentheses when the delivered sentence did not; (3) nit — “in IPv6 too” was demonstrable only for DTAG. All three fixed before publishing; the IPv4 modes sentence was added to the checker (quote 3) at the same time.

13. LLM-classification currency, 2026-09-03

Recorded during the run. Shared numbers, the script, its unedited output, the folds and the quote check are on website_classification §12. This section records only what is specific to this page — including the one sentence on it that was simply wrong.

13.1 A retracted sentence

The Open Questions bullet read:

LLMs have reached AS-to-organisation mapping [2Selmo, Carlos; Carisimo, Esteban; Bustamante, Fabián E.; Alvarez-Hamelin, J. Ignacio (2025): "Learning AS-to-Organization Mappings with Borges", in: Proceedings of the 2025 ACM Internet Measurement Conference, pp. 120-133. (DOI)] but not IP classification. We found nothing peer-reviewed applying an LLM to geolocation, host typing or residential/VPN/datacenter labelling as of August 2026 — unlike cookie and policy classification, where LLM methods are now routine.

The first two sentences hold. The clause in bold does not, and nothing on this site owned it: there was no script behind it, no paper cited for it, and cookies — the page it is a claim about — does not mention LLMs at all. It was a plausible aside that no reviewer brief had a reason to check.

Measured per target, as the LLM share of the papers that classify that target at all:

Target LLM papers Papers classifying it at all Share
privacy-policy 12 102 11.8%
cookie 1 53 1.9%
ip-address 1 295 0.3%

So “routine” is defensible for neither. privacy-policy is the highest share of any target in the corpus and is the only one that comes close; cookie is a single TheWebConf 2025 paper, [3Chen, Baiqi; Lyu, Jiawei; Wu, Tingmin; Chhetri, Mohan Baruwal; Bai, Guangdong (2025): "Semantics-Aware Cookie Purpose Compliance", in: Proceedings of the ACM Web Conference. (DOI)]. And the comparison the old sentence was making — that IP classification is unusually untouched — is weaker than it claimed: at 1 of 295 it is low, but web-request is 1 of 258 and javascript and fingerprinting-script are at zero. The page now states the measured shares, says the old clause was wrong, and links the full table.

13.2 Two different papers, both described as "the one LLM paper"

The page said both of these, five hundred lines apart:

  • “[2Selmo, Carlos; Carisimo, Esteban; Bustamante, Fabián E.; Alvarez-Hamelin, J. Ignacio (2025): "Learning AS-to-Organization Mappings with Borges", in: Proceedings of the 2025 ACM Internet Measurement Conference, pp. 120-133. (DOI)] … reports a 7% improvement in sibling-ASN identification … That is, as of 2026, the one place in IP classification where an LLM method has cleared peer review
  • “the llm method fires exactly once — one paper, GPT-4o, in the 2025–2026 window”

They are different papers, and a reader would reasonably merge them into one:

[2Selmo, Carlos; Carisimo, Esteban; Bustamante, Fabián E.; Alvarez-Hamelin, J. Ignacio (2025): "Learning AS-to-Organization Mappings with Borges", in: Proceedings of the 2025 ACM Internet Measurement Conference, pp. 120-133. (DOI)] [4Schwartz, Yuval; Ben-Shimol, Lavi; Mimran, Dudu; Elovici, Yuval; Shabtai, Asaf (2025): "LLMCloudHunter: Harnessing LLMs for Automated Extraction of Detection Rules from Cloud-Based CTI", in: Proceedings of the ACM Web Conference. (DOI)]
Venue IMC 2025 TheWebConf 2025
Model GPT-4o-mini, temperature 0 GPT-4o
classification.target other ip-address
In this page's method table? no — its tuples are filed under other yes, it is the single llm row
What it does few-shot extraction over PeeringDB notes and aka fields for sibling-ASN mapping extracts IP indicators and user agents from threat-intelligence prose

Both sentences are now explicit about which paper they mean, and the first says that Borges does not appear in the method table at all. schwartz2025_llmcloudhunter was added to bibliography in this run so the second sentence can name its paper; it was generated by scripts/bibgen.mjs from the venue index (DOI 10.1145/3696410.3714798, OpenAlex metadata).

13.3 Quotes checked

  • selmo2025_borges“utilizing OpenAI's GPT-4o-mini [40] with a temperature set to 0 and a Top P probability mass of 1”, located in data/fulltext/2025/IMC/learning-as-to-organization-mappings-with-borges/paper.cols.txt on 2026-09-03. This is the source for the GPT-4o-mini correction; the page previously named no model here and the sentence five hundred lines later named GPT-4o, which is how the two papers got conflated.
  • schwartz2025_llmcloudhunter — its ip-address tuple's evidence quote passes the shared quote check at the PASS-ELID tier: the extractor wrote “This component … parses OSCTIs to identify and extract IoCs, notably IP addresses and user agents pertinent to AWS CloudTrail logs”, and both fragments either side of the elision are present. The elided middle is unverified, which is why the page describes what the paper does rather than quoting it.

13.4 What could not be established

  • Whether other hides an LLM IP-classification paper — probed, and Borges is the reason it had to be. Borges is the proof that the bucket can hide one: an LLM paper squarely about AS-and-organisation mapping sits under other and therefore outside this page's method table. 116 of the 177 corpus LLM papers are in that bucket. other does carry a free-text targetDetail, stated on all 157 such tuples, and report_llm_currency.mjs now probes it; Borges's reads “favicon and associated final-URL groups”, which is exactly the kind of string a keyword probe for IP, geolocation, ASN would miss. So treat the probe as evidence about scripts and website topics, where the strings are unambiguous, and as weak evidence here. The page's “fires exactly once” is a statement about the ip-address target, not about the topic, and it now says so. Full probe output on website_classification §12.9.
  • Nothing else on this page was re-derived. The 295-paper population, the geolocation figures and the fold residue are unchanged from the 2026-08-12 refresh (§4) and from §12; only the two LLM sentences were touched.

← back to the content page · corpus-level provenance

provenance/design/ip_classification.txt · Last modified: by karel.kubicek.claude