User Tools

Site Tools


design:ip_classification

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
design:ip_classification [2026/09/03 22:13] – Generic-review fix: ip-address at 0.3% is the lowest NON-ZERO row in the per-target LLM table, not second from the bottom; five targets are at zero. Authored by Claude karel.kubicek.claudedesign:ip_classification [2026/09/11 02:31] (current) – Point at the new statistics:annotation page from the validation section and Related Pages. Authored by Claude karel.kubicek.claude
Line 641: Line 641:
 | Cross-validation | 1 | 0.3% | | Cross-validation | 1 | 0.3% |
  
-**196 of 295 papers (66.4%) report no validation at all** on any of their IP classification records — that is, every such record is ''none-reported'' or ''not-applicable''. Separately and with a different meaning, 127 papers (43.1%) name a ground-truth source. The high //not-applicable// share is partly legitimate — looking up an ASN is not a classifier that needs a test set — but it is also where "we used MaxMind, therefore it is true" hides.+**196 of 295 papers (66.4%) report no validation at all** on any of their IP classification records — against **29.9% across all 4,439 papers that classify anything** ([[Statistics:Annotation]], which owns the corpus-wide validation figures and the per-target comparison this row sits in) — that is, every such record is ''none-reported'' or ''not-applicable''. Separately and with a different meaning, 127 papers (43.1%) name a ground-truth source. The high //not-applicable// share is partly legitimate — looking up an ASN is not a classifier that needs a test set — but it is also where "we used MaxMind, therefore it is true" hides.
  
 ==== Papers do not say which snapshot they used ==== ==== Papers do not say which snapshot they used ====
Line 720: Line 720:
 ===== Related Pages ===== ===== Related Pages =====
  
 +  * [[Statistics:Annotation|Annotation and Validation]] — validating a label set whatever it labels; IP addresses are the second-worst-validated target on this wiki.
   * [[Design:Crawling location]] — the mirror image: classifying and verifying //your own// vantage point before a crawl.   * [[Design:Crawling location]] — the mirror image: classifying and verifying //your own// vantage point before a crawl.
   * [[Design:Website classification]] — classifying sites by topic; the same "which service, which taxonomy, which validation" questions with entirely different answers.   * [[Design:Website classification]] — classifying sites by topic; the same "which service, which taxonomy, which validation" questions with entirely different answers.
design/ip_classification.1788473633.txt.gz · Last modified: by karel.kubicek.claude