This is an old revision of the document!
Table of Contents
Roadmap
What this wiki has decided to write next, and what it has decided not to write. Both halves are here on purpose: on a site where every page names its own population, “we looked and there is not enough literature yet” is a finding, and it should be recorded once rather than rediscovered by every gap pass.
Paper counts below are candidate sets from the publication corpus — title-and-summary probes over 5,859 extracted papers, so they are floors, not populations. The page that gets written derives its own.
Queued
The links below are red on purpose. These pages are scoped and queued but not written; the id is fixed now so later work lands in the right place. scripts/sitemap.mjs checks that every promised-but-missing page on the wiki appears in this table: a red link anywhere else is a defect, and a row here that has turned blue must be removed.
| Planned page | Why, and the evidence |
|---|---|
| Ad archives | Ad-transparency archives as an instrument. 69 papers; the two-sided error problem, and what a denominator means when the archive holds only ads that ran. |
| Email authentication | SPF, DKIM, DMARC, DANE, MTA-STS deployment. 31 hand-mapped papers, scoped out of Email tracking because it is not tracking. |
| Privacy policies and terms | Policy text as data you collect and label. 121 papers, and the tool lineage from Polisis (2018) to PoliGraph (2023) with dates, because a fresh student reaches for the 2018 tool. |
| Internet scanning | The scan branch of Automated measurements has no instrument page. 930 papers probe the network, 143 name ZMap/ZGrab/masscan/nmap; rate limits, blocklists, abuse contacts, and why IPv6 is not enumerable. |
| Blocking and geodifference | What your crawler was not allowed to see: geoblocking, GDPR walls, censorship. 74 papers. The claim needs a control vantage and a definition of “blocked” that survives a soft 200. |
| Ownership resolution | Attributing a third-party domain to a company. The material exists, but under a dead tool's name, where nobody looking for it will search. |
| Annotation and validation | Labelling data with humans or an LLM, and validating the result. 42.4% of the 4,439 papers that classified anything report no validation. LLM-as-annotator is 30 papers, 27 of them 2024–2026 — the fastest-growing method here, and currently buried inside Website classification. |
| Connected TV | Smart TVs, sticks and set-top boxes as targets. 16 papers, only one on the web platform — which is the point: none of the web instrumentation transfers. |
| Age assurance | Age verification and age gates as deployed. 11 papers, thin on purpose: the policy surface moves far faster than the literature, so expect method and caveat rather than figures. |
| Data subject rights | Access, deletion and opt-out requests as a measurement: DSARs, Global Privacy Control, “do not sell”. 9 papers, 6 of them 2024–2026. Note the id — Requests is web requests. |
Assessed, and not written
Recorded so the next gap pass does not re-propose them. Counts are candidate sets from the corpus, not populations.
| Proposal | Disposition |
|---|---|
| Anti-adblock and adblock walls | 3 papers. A section on Filter Lists, not a page. |
AI crawlers, robots.txt, training opt-out | 4 papers, 3 of them 2024 or later: the corpus lags the field by about two years. Extend Crawler Detection, which covers your crawler being noticed, to cover measuring who crawls you. |
| WebRTC leaks | 6 papers. Below the line — and the line matters, or WebAssembly and every other thinly-published web technology follows. |
| SSO, WebAuthn and MFA deployment | 46 papers unaudited, and the probe catches attacks on OAuth as well as deployment measurement. Gated on that audit; dropped if the deployment slice is under ~20. |
| Sock-puppet and personalization audits | 20 papers, the smallest live proposal. Likely a section on Stateful Stateless rather than a page. |
| Per-company platform pages (Facebook, Twitter, TikTok, Amazon) | Assessed against the corpus and deliberately not written; Platforms carries the reasoning and the counts. |
Adding to this page
A proposal earns a row in Queued when it has an id, a population that can be re-derived from the corpus, and a reason it is not a section on an existing page. It earns a row in Assessed when that check fails — with the count that made it fail. Both are recorded in roadmap when a decision changes.
