User Tools

Site Tools


design:automated_measurements

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
design:automated_measurements [2026/09/10 21:05] – Route the scan branch to its new instrument page, Programming:Internet scanning. Authored by Claude. karel.kubicek.claudedesign:automated_measurements [2026/09/11 10:50] (current) – Route differential/sock-puppet audits to the new Design:Algorithm audits page; note why its population is not in this page's studyTypes counts. Authored by Claude karel.kubicek.claude
Line 8: Line 8:
 | which resolver answered a name, whether the answer was manipulated, encrypted DNS | a **DNS measurement**, which is a scan with its own failure modes | [[Design:DNS]] | | which resolver answered a name, whether the answer was manipulated, encrypted DNS | a **DNS measurement**, which is a scan with its own failure modes | [[Design:DNS]] |
 | an APK or IPA, a store listing, an app's runtime | **app analysis** | [[Design:Mobile and app measurement]] | | an APK or IPA, a store listing, an app's runtime | **app analysis** | [[Design:Mobile and app measurement]] |
 +| whether a platform treats **two constructed identities** differently — personalised results, targeted ads, a feed, a quoted price | a **differential audit**, which is a crawl whose population is the arms rather than the sites | [[Design:Algorithm audits]], [[Programming:Stateful stateless]], [[Statistics:Hypothesis testing]] |
  
 <WRAP important> <WRAP important>
Line 58: Line 59:
   * **nmap, traceroute, RIPE Atlas** are the other named families (40, 31, 25 papers in the scan branch). They answer different questions than ZMap does (host discovery vs path vs volunteer vantage points). **167** scan papers name a ''network-scanner'' the fold did not map — one-off research scanners, listed on the provenance page. That residue is the scanning analogue of the home-grown crawler row on [[Programming:Crawler]].   * **nmap, traceroute, RIPE Atlas** are the other named families (40, 31, 25 papers in the scan branch). They answer different questions than ZMap does (host discovery vs path vs volunteer vantage points). **167** scan papers name a ''network-scanner'' the fold did not map — one-off research scanners, listed on the provenance page. That residue is the scanning analogue of the home-grown crawler row on [[Programming:Crawler]].
  
-**The instrument page for this branch is [[Programming:Internet scanning]]**, added on 2026-09-10 — the ZMap/ZGrab invocation, the probe rate and its wall-clock cost, the exclusion file, source-address hygiene, and the IPv6 hitlist problem. It defines its own population (245 papers naming an active-scan instrumentrather than reusing the 930, and its scanner fold is a separate, stricter one; where the two pages disagree on a count, that page says why.+**The instrument page for this branch is [[Programming:Internet scanning]]**, added on 2026-09-10 — the ZMap/ZGrab invocation, the probe rate and its wall-clock cost, the exclusion file, source-address hygiene, and the IPv6 hitlist problem. It defines its own population — the papers that name an active-scan instrument, a smaller and differently derived set — rather than reusing the 930, and its scanner fold is a separate, stricter one; where the two pages disagree on a count, that page carries the number and says why.
  
 Scanning ethics is not this page. Durumeric, Bailey and Halderman {[durumeric2014_view]} measured who was already scanning the Internet; the operational checklist (identify yourself, publish an opt-out, rate-limit) lives on [[Practices:Ethics]], and telling the operator lives on [[Practices:Notifying websites]]. A scan from cloud address space is also a [[Design:Crawling location|vantage-point]] decision. Scanning ethics is not this page. Durumeric, Bailey and Halderman {[durumeric2014_view]} measured who was already scanning the Internet; the operational checklist (identify yourself, publish an opt-out, rate-limit) lives on [[Practices:Ethics]], and telling the operator lives on [[Practices:Notifying websites]]. A scan from cloud address space is also a [[Design:Crawling location|vantage-point]] decision.
Line 76: Line 77:
   * **A manual audit, a code analysis, or a system paper.** Those are the three larger ''studyTypes'' (1,826 / 1,484 / 3,967). A system paper that also crawled is in the 153-paper gap above.   * **A manual audit, a code analysis, or a system paper.** Those are the three larger ''studyTypes'' (1,826 / 1,484 / 3,967). A system paper that also crawled is in the 153-paper gap above.
   * **A tutorial on HTTP, DNS, or TLS.** Those are specs. This page is which instrument the field uses to observe them.   * **A tutorial on HTTP, DNS, or TLS.** Those are specs. This page is which instrument the field uses to observe them.
 +  * **A differential audit.** Driving two deliberately different profiles at one platform and comparing what comes back is a crawl by machinery and an experiment by design; the arms, not the site list, are the population. [[Design:Algorithm audits]] owns it, and its 32-paper population is derived by hand rather than from ''studyTypes'', so it does not appear in any count on this page.
  
 ===== What to report ===== ===== What to report =====
design/automated_measurements.1789074355.txt.gz · Last modified: by karel.kubicek.claude