| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| design:ip_classification [2026/09/03 21:41] – LLM currency: retract 'unlike cookie and policy classification, where LLM methods are now routine' (cookie is 1 of 53) and replace it with the measured per-target shares; disambiguate selmo2025_borges (GPT-4o-mini, target=other) from LLMCloudHunter, the o karel.kubicek.claude | design:ip_classification [2026/09/11 02:31] (current) – Point at the new statistics:annotation page from the validation section and Related Pages. Authored by Claude karel.kubicek.claude |
|---|
| | LLM | 1 | 0.3% | | | LLM | 1 | 0.3% | |
| |
| Shares exceed 100% because a paper can use several. The distribution is the opposite of most classification tasks on this site: IP classification is overwhelmingly a **look-it-up** problem, not a machine-learning one. Supervised ML has not moved at all in absolute terms (7 papers on the old 4,322-paper corpus, 7 on this one), and the ''llm'' method fires exactly **once** — LLMCloudHunter {[schwartz2025_llmcloudhunter]} (TheWebConf 2025), GPT-4o, in the 2025–2026 window, and it is extracting IP indicators from threat-intelligence text rather than classifying addresses. Whatever LLM classification is doing elsewhere on this site, it has not arrived here: **1 of the 295 corpus papers that classify an ''ip-address'' target (0.3%)**, against 177 papers corpus-wide that classify //something// with an LLM. That contrast is the point — the method is mainstream and this target is untouched. The per-target ranking is on [[design:website_classification#Where LLMs actually appear]] and ''ip-address'' is second from the bottom of its non-empty rows. | Shares exceed 100% because a paper can use several. The distribution is the opposite of most classification tasks on this site: IP classification is overwhelmingly a **look-it-up** problem, not a machine-learning one. Supervised ML has not moved at all in absolute terms (7 papers on the old 4,322-paper corpus, 7 on this one), and the ''llm'' method fires exactly **once** — LLMCloudHunter {[schwartz2025_llmcloudhunter]} (TheWebConf 2025), GPT-4o, in the 2025–2026 window, and it is extracting IP indicators from threat-intelligence text rather than classifying addresses. Whatever LLM classification is doing elsewhere on this site, it has not arrived here: **1 of the 295 corpus papers that classify an ''ip-address'' target (0.3%)**, against 177 papers corpus-wide that classify //something// with an LLM. That contrast is the point — the method is mainstream and this target is untouched. The per-target ranking is on [[design:website_classification#Where LLMs actually appear]], where ''ip-address'' at 0.3% is the **lowest non-zero row** — below it are only the five targets no LLM paper touches at all, ''javascript'', ''fingerprinting-script'', ''malware'', ''sdk-or-library'' and ''website-popularity''. |
| |
| ==== Which resources, folded ==== | ==== Which resources, folded ==== |
| | Cross-validation | 1 | 0.3% | | | Cross-validation | 1 | 0.3% | |
| |
| **196 of 295 papers (66.4%) report no validation at all** on any of their IP classification records — that is, every such record is ''none-reported'' or ''not-applicable''. Separately and with a different meaning, 127 papers (43.1%) name a ground-truth source. The high //not-applicable// share is partly legitimate — looking up an ASN is not a classifier that needs a test set — but it is also where "we used MaxMind, therefore it is true" hides. | **196 of 295 papers (66.4%) report no validation at all** on any of their IP classification records — against **29.9% across all 4,439 papers that classify anything** ([[Statistics:Annotation]], which owns the corpus-wide validation figures and the per-target comparison this row sits in) — that is, every such record is ''none-reported'' or ''not-applicable''. Separately and with a different meaning, 127 papers (43.1%) name a ground-truth source. The high //not-applicable// share is partly legitimate — looking up an ASN is not a classifier that needs a test set — but it is also where "we used MaxMind, therefore it is true" hides. |
| |
| ==== Papers do not say which snapshot they used ==== | ==== Papers do not say which snapshot they used ==== |
| ===== Related Pages ===== | ===== Related Pages ===== |
| |
| | * [[Statistics:Annotation|Annotation and Validation]] — validating a label set whatever it labels; IP addresses are the second-worst-validated target on this wiki. |
| * [[Design:Crawling location]] — the mirror image: classifying and verifying //your own// vantage point before a crawl. | * [[Design:Crawling location]] — the mirror image: classifying and verifying //your own// vantage point before a crawl. |
| * [[Design:Website classification]] — classifying sites by topic; the same "which service, which taxonomy, which validation" questions with entirely different answers. | * [[Design:Website classification]] — classifying sites by topic; the same "which service, which taxonomy, which validation" questions with entirely different answers. |