User Tools

Site Tools


design:ip_classification

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
design:ip_classification [2026/09/03 19:52] – DynamIPs (padmanabhan2020_dynamips): quote verified against CAIDA PDF, paraphrase footnote dropped, IPv4/IPv6 renumbering specifics attributed correctly. Authored by Claude karel.kubicek.claudedesign:ip_classification [2026/09/11 02:31] (current) – Point at the new statistics:annotation page from the validation section and Related Pages. Authored by Claude karel.kubicek.claude
Line 74: Line 74:
 <WRAP todo>**CAIDA's separate //AS Classification// dataset is discontinued** — CAIDA's own page states it is no longer supported and download access has been removed. Papers from 2015–2022 cite it routinely; ASdb is the live replacement. Check before you cite a dataset you found in a related-work section.</WRAP> <WRAP todo>**CAIDA's separate //AS Classification// dataset is discontinued** — CAIDA's own page states it is no longer supported and download access has been removed. Papers from 2015–2022 cite it routinely; ASdb is the live replacement. Check before you cite a dataset you found in a related-work section.</WRAP>
  
-The current state of the art on the sibling problem is {[selmo2025_borges]} (IMC 2025), which few-shot prompts an LLM over PeeringDB free-text fields plus website and domain evidence, and reports a 7% improvement in sibling-ASN identification over AS2Org-style methods. That is, as of 2026, **the one place in IP classification where an LLM method has cleared peer review** — see [[#Open Questions]].+The current state of the art on the sibling problem is {[selmo2025_borges]} (IMC 2025), which few-shot prompts GPT-4o-mini at temperature 0 over PeeringDB free-text fields plus website and domain evidence, and reports a 7% improvement in sibling-ASN identification over AS2Org-style methods. That is, as of 2026, **the one place in the AS-and-organisation half of this page where an LLM method has cleared peer review** — see [[#Open Questions]]. It is a //different// paper from the single ''llm'' firing in this page's corpus population, which is LLMCloudHunter {[schwartz2025_llmcloudhunter]} extracting IP indicators from threat-intelligence prose; do not conflate the two, and note that Borges's classification tuples are filed under ''other'' rather than ''ip-address'', so it does not appear in the method table in [[#How they classify]] at all.
  
 Two traps specific to web measurement. First, **the AS that announces a server's address is very often not the party you care about**: a tracker served through Cloudflare terminates in AS13335, and attributing the tracking to Cloudflare is wrong. Resolve the //name// (the request URL, the certificate, the CNAME chain) before or instead of the address. Second, **ASNs churn**: {[nemmi2021_parallel]} tracked the administrative and operational lives of ASNs separately and found 22,729 administrative lifetimes (17.9%) with no observed BGP activity at all, plus 1,667 ASNs announcing in BGP without an overlapping allocation. An ASN-to-organisation table from 2019 is not a table about 2026. Two traps specific to web measurement. First, **the AS that announces a server's address is very often not the party you care about**: a tracker served through Cloudflare terminates in AS13335, and attributing the tracking to Cloudflare is wrong. Resolve the //name// (the request URL, the certificate, the CNAME chain) before or instead of the address. Second, **ASNs churn**: {[nemmi2021_parallel]} tracked the administrative and operational lives of ASNs separately and found 22,729 administrative lifetimes (17.9%) with no observed BGP activity at all, plus 1,667 ASNs announcing in BGP without an overlapping allocation. An ASN-to-organisation table from 2019 is not a table about 2026.
Line 574: Line 574:
 | LLM | 1 | 0.3% | | LLM | 1 | 0.3% |
  
-Shares exceed 100% because a paper can use several. The distribution is the opposite of most classification tasks on this site: IP classification is overwhelmingly a **look-it-up** problem, not a machine-learning one. Supervised ML has not moved at all in absolute terms (7 papers on the old 4,322-paper corpus, 7 on this one), and the ''llm'' method fires exactly **once** — one paper, GPT-4o, in the 2025–2026 window. Whatever LLM classification is doing elsewhere on this site, it has not arrived here.+Shares exceed 100% because a paper can use several. The distribution is the opposite of most classification tasks on this site: IP classification is overwhelmingly a **look-it-up** problem, not a machine-learning one. Supervised ML has not moved at all in absolute terms (7 papers on the old 4,322-paper corpus, 7 on this one), and the ''llm'' method fires exactly **once** — LLMCloudHunter {[schwartz2025_llmcloudhunter]} (TheWebConf 2025), GPT-4o, in the 2025–2026 window, and it is extracting IP indicators from threat-intelligence text rather than classifying addresses. Whatever LLM classification is doing elsewhere on this site, it has not arrived here: **1 of the 295 corpus papers that classify an ''ip-address'' target (0.3%)**, against 177 papers corpus-wide that classify //something// with an LLM. That contrast is the point — the method is mainstream and this target is untouched. The per-target ranking is on [[design:website_classification#Where LLMs actually appear]], where ''ip-address'' at 0.3% is the **lowest non-zero row** — below it are only the five targets no LLM paper touches at all, ''javascript'', ''fingerprinting-script'', ''malware'', ''sdk-or-library'' and ''website-popularity''.
  
 ==== Which resources, folded ==== ==== Which resources, folded ====
Line 641: Line 641:
 | Cross-validation | 1 | 0.3% | | Cross-validation | 1 | 0.3% |
  
-**196 of 295 papers (66.4%) report no validation at all** on any of their IP classification records — that is, every such record is ''none-reported'' or ''not-applicable''. Separately and with a different meaning, 127 papers (43.1%) name a ground-truth source. The high //not-applicable// share is partly legitimate — looking up an ASN is not a classifier that needs a test set — but it is also where "we used MaxMind, therefore it is true" hides.+**196 of 295 papers (66.4%) report no validation at all** on any of their IP classification records — against **29.9% across all 4,439 papers that classify anything** ([[Statistics:Annotation]], which owns the corpus-wide validation figures and the per-target comparison this row sits in) — that is, every such record is ''none-reported'' or ''not-applicable''. Separately and with a different meaning, 127 papers (43.1%) name a ground-truth source. The high //not-applicable// share is partly legitimate — looking up an ASN is not a classifier that needs a test set — but it is also where "we used MaxMind, therefore it is true" hides.
  
 ==== Papers do not say which snapshot they used ==== ==== Papers do not say which snapshot they used ====
Line 679: Line 679:
   * **Venue coverage.** Seven venues only — the scope and the selection funnel are on [[literature:corpus]]. For //this// topic the venues where much of the work actually appears are all absent: **PAM, TMA, ANRW, SIGCOMM and ACM CCR are not in the corpus**, and for IP geolocation specifically that is a serious gap.   * **Venue coverage.** Seven venues only — the scope and the selection funnel are on [[literature:corpus]]. For //this// topic the venues where much of the work actually appears are all absent: **PAM, TMA, ANRW, SIGCOMM and ACM CCR are not in the corpus**, and for IP geolocation specifically that is a serious gap.
   * **The corpus reaches 2026, but its last two years are provisional** (see [[literature:corpus]]), so a per-period row ending in 2025–2026 rests on fewer papers than a complete window would give. The "current in 2026" judgements on this page are ours, checked against vendor and standards documentation, not derived from the corpus.   * **The corpus reaches 2026, but its last two years are provisional** (see [[literature:corpus]]), so a per-period row ending in 2025–2026 rests on fewer papers than a complete window would give. The "current in 2026" judgements on this page are ours, checked against vendor and standards documentation, not derived from the corpus.
 +  * **The LLM figures come from a shared script, not from this page's own query.** ''scripts/report_llm_currency.mjs'' computes ''classification.method == "llm"'' once, per target, for all three classification pages ([[design:website_classification]], [[privacy:javascript]] and this one), and prints each page's own sentence next to what the corpus says so a drift between them fails visibly. Its output is on [[provenance:design:ip_classification]].
   * **Every query behind this section, the report script and its unedited output** are on [[provenance:design:ip_classification]]; corpus-level caveats are on [[literature:corpus]].   * **Every query behind this section, the report script and its unedited output** are on [[provenance:design:ip_classification]]; corpus-level caveats are on [[literature:corpus]].
  
Line 713: Line 714:
   * **No systematic evaluation of the commercial datacenter/VPN/proxy flags against ground truth.** Every paper that uses one takes it on faith, and the run above shows two public DNS resolvers flagged ''is_vpn'' and ''is_abuser''. A labelled benchmark here would be immediately useful and is well within a single student's reach.   * **No systematic evaluation of the commercial datacenter/VPN/proxy flags against ground truth.** Every paper that uses one takes it on faith, and the run above shows two public DNS resolvers flagged ''is_vpn'' and ''is_abuser''. A labelled benchmark here would be immediately useful and is well within a single student's reach.
   * **No peer-reviewed method for identifying hosting/datacenter address space** beyond the operators' own published lists — which means everything outside the big five clouds is guesswork.   * **No peer-reviewed method for identifying hosting/datacenter address space** beyond the operators' own published lists — which means everything outside the big five clouds is guesswork.
-  * **LLMs have reached AS-to-organisation mapping {[selmo2025_borges]} but not IP classification.** We found nothing peer-reviewed applying an LLM to geolocation, host typing or residential/VPN/datacenter labelling as of August 2026 — unlike cookie and policy classification, where LLM methods are now routine. Whether that is because the task has no useful text to read, or because nobody has tried, is an open question.+  * **LLMs have reached AS-to-organisation mapping {[selmo2025_borges]} but not IP classification.** We found nothing peer-reviewed applying an LLM to geolocation, host typing or residential/VPN/datacenter labelling as of 2026-09-03. This page previously said that was "unlike cookie and policy classification, where LLM methods are now routine"; that was wrong and is corrected hereMeasured per target across the 5,859-paper corpus, the LLM share of the papers classifying that target at all is **11.8% for ''privacy-policy'' (12 of 102)** — the highest of any target and the only one that could support the word "routine" — but only **1.9% for ''cookie'' (1 of 53)**, a single TheWebConf 2025 paper. Against those, ''ip-address'' at 1 of 295 is low but not the outlier the old sentence implied; ''javascript'' and ''fingerprinting-script'' are at zero. See [[design:website_classification#Where LLMs actually appear]] for the whole table. Whether the gap here is because the task has no useful text to read, or because nobody has tried, is still an open question.
   * **No strong successor to {[richter2016_multi]} on CGNAT prevalence.** The best numbers on how much of the client Internet sits behind shared addresses are a decade old, and IPv4 exhaustion has only got worse since.   * **No strong successor to {[richter2016_multi]} on CGNAT prevalence.** The best numbers on how much of the client Internet sits behind shared addresses are a decade old, and IPv4 exhaustion has only got worse since.
 </WRAP> </WRAP>
Line 719: Line 720:
 ===== Related Pages ===== ===== Related Pages =====
  
 +  * [[Statistics:Annotation|Annotation and Validation]] — validating a label set whatever it labels; IP addresses are the second-worst-validated target on this wiki.
   * [[Design:Crawling location]] — the mirror image: classifying and verifying //your own// vantage point before a crawl.   * [[Design:Crawling location]] — the mirror image: classifying and verifying //your own// vantage point before a crawl.
   * [[Design:Website classification]] — classifying sites by topic; the same "which service, which taxonomy, which validation" questions with entirely different answers.   * [[Design:Website classification]] — classifying sites by topic; the same "which service, which taxonomy, which validation" questions with entirely different answers.
design/ip_classification.1788465140.txt.gz · Last modified: by karel.kubicek.claude