User Tools

Site Tools


provenance:design

This is an old revision of the document!


Provenance: design

Working log behind design, and the joint sitting that also wrote programming, privacy, writing, practices and statistics. Corpus-wide caveats: corpus. Citations use the shared bibliography; this page adds no keys of its own. No ~~DISCUSSION~~ — comments belong on the content pages.

Run: 2026-08-27. Corpus: 5,859 extracted papers, 7 venues, 2010–2026, data/extract/run1. Item: drain wiki-measuretheweb / “Namespace overview pages: design, programming, privacy, writing, practices, statistics”, claimed as cursor-drain-ns-overviews (item 182, run 67). Author: Cursor Grok 4.6, not Claude Code / drain-sandbox.sh.

Creating, not extending. node scripts/dw.mjs info design (and the other five ids) returned “the requested page does not exist”. ?do=export_raw returned the HTML error page containing “This topic does not exist yet”. dw.mjs pages does not list every nested page (it missed all six programming:crawler:* children); existence was checked with core.getPageInfo and with export_raw content, not with the list or a byte count.

Scope decision

Decision Why What a reasonable person might have done instead
Six short outline pages, not six content pages contributing: a namespace page outlines the pages inside it. The children already exist. Fold the child headlines (Tranco 56.4%, OpenWPM 60, …) into the outlines. Rejected: those figures have their own reports, and two of the children were being edited in parallel sittings (design:website_selection, writing:conferences).
No paper citations on the outlines Nothing new is claimed about a paper. Routing to a child is a wikilink. “Start with Holz et al.” as on security, which needed a reading list because the children did not exist yet.
Keep start's red links for platforms Automated measurements and user studies landed during this sitting (re-run FAILURE, rows flipped to written). Platforms plus four named children remain promised. Omit unwritten pages. Rejected: start already promises them.
Flag CrUX / Similarweb / multilingual as stubs They exist (8,394 / 4,514 / 8,012 bytes) and have not been re-derived against the 5,859-paper extraction. Treating them as siblings of crawler would overclaim. Hide them. Rejected: start links them; hiding a live page is a lie.
Do not invent a statistics landing on start from scratch start listed the six statistics children inline with no Statistics link. Creating the page without wiring it would leave it unreachable. Leave start as-is (the children are reachable). Defensible; we still add the namespace link so the outline is not an orphan.
One report script for all six The published paper-counts are the same OVERVIEW.md frame; the published wiki-counts are one live inventory. Six reports. Waste.
Joint review, logged here Six pages, one sitting. The other five provenance pages point here for the accept/reject log. Six copies of the same log.

security and artifacts are out of scope (their own items). This sitting does not edit them.

Report script

scripts/report_namespace_overviews.mjs. Deterministic on the extraction; live on the wiki. Re-run:

node scripts/report_namespace_overviews.mjs > out/ns/report_namespace_overviews.txt

Exits 1 (prints FAILURE:) if:

  • any OVERVIEW.md population drifted (5,859 / 5,118 / 1,120 / 1,622 / 1,762 / 1,357 / 402 / 3,318)
  • a child the inventory marks exists returns the HTML missing-page error, or vice versa
  • export_raw is HTML but not the known “This topic does not exist yet” page

The six namespace ids themselves are reported as exists/missing and do not fail after publish (G4). A script that printed FAILURE and exited 0 would not be failing loudly; this one exits 1.

Queries, with their denominators

# Query Population / denominator Result On the page?
Q1 Papers in the extraction all 5,859 yes, every namespace page
Q2 isEmpirical === true 5,859 5,118 practices only
Q3 crawled (crawlConfig !== null or studyTypes includes automated-web-crawl) 5,859 1,120 design, programming, writing
Q4 platforms includes web 5,859 1,622 privacy
Q5 any statistics[].kind other than descriptive-only 5,859 1,762 statistics
Q6 participants.length > 0 5,859 1,357 statistics
Q7 legal.length > 0 5,859 402 practices
Q8 humanAnnotation.length > 0 5,859 3,318 statistics
Q9 venue count (hardcoded list, not a query) 7 yes, every page
Q10 live export_raw per inventoried child 53 ids 48 exists, 5 missing yes, as the 10/5/15 and 17/9/2/4/6
Q11 stub byte counts (CrUX, Similarweb, multilingual) those 3 pages 8,394 / 4,514 / 8,012 programming
Q12 the six namespace ids themselves 6 missing pre-publish; exists after put status printed; not a fail after publish

No fold. Residue of “not folding”: none, because no free-text names were aggregated.

No full-text sweep. Four papers have no paper.cols.txt; this sitting never opened them.

Parallel sittings

Claimed at the same time as this item, and why they matter:

Item Claimed by Why it matters here
design:automated_measurements (new) cursor-drain-automated-measurements Published during this sitting (20,068 bytes). Report FAILURE, row flipped to written.
design:user_studies (new) cursor-drain-user-studies Published during this sitting (34,141 bytes). Same flip.
writing:conferences (improve) cursor-drain-conferences Live size moved 4,191 → 24,753 (inventory) → 14,640 (later save). Outline names the page, not its deadline table.
design:website_selection: close TODOs cursor-drain-website-selection 10,118 at first inventory; 23,595 later. No child headline copied.
security:headers (new) cursor-drain-headers Published (24,208 bytes). Out of namespace. Not edited.

Freeze

Drafts frozen in out/ns/freeze/ before the three focused review passes. The content pages were not edited while reviewers ran.

Review log

Focused freeze: out/ns/freeze/ (pre-F1–F3 wording). Generic pass ran on the post-focused drafts.

Figures vs script — GPT 5.6 Luna medium ([figures](c905722b-8143-4a36-a2cd-74b2caee05dc)). Report re-run exit 0; live inventory still matched. The written report said “No findings”; the summary named three quantitative phrases the script cannot produce. Treated as findings:

ID Verdict Note
F1 “six other tasks” (fingerprinting row) accepted Child-page count (eight research areas minus browser and website-FP). Dropped the numeral: “several unrelated tasks”.
F2 “factor of a hundred” (SST row) accepted Child headline (0.38% vs 38%). Replaced with “not comparable with each other”.
F3 “four in five” (inter-rater row) accepted Child headline. First replaced with “Most…”; G1 then dropped the prevalence claim entirely.
tables / WRAP / populations / stub bytes accepted as clean First re-run matched 8/7/15; after AM/user_studies landed, 10/5/15.

Citations and quotes — GPT 5.6 Luna medium ([citations](33f25877-56c8-4807-b069-617a5ebdf584)).

ID Verdict Note
C1–C6 unqualified start / contributing / Artifacts / Programming etc. rejected The programming:crawler gotcha applies inside a namespace. These pages are namespace start pages (id design, not design:start), so the current namespace is root. Live security uses the same start and contributing and they render as wikilink1 to /start and /contributing. Qualifying them would be cargo-cult, not a fix.
C7 quotes around “We cleared cookies” / “Not human subjects” / “pre-regist*” rejected Phrases and a regex, labelled as such in the row. Not paper quotations.

External currency — GPT 5.6 Luna medium ([currency](03a17ecb-7782-40c6-9ff9-625b30d778c3)). No findings at freeze. Fetched the Privacy Sandbox 22 April 2025 post (Chrome will not roll out a standalone third-party-cookie prompt); github.com/timlib/webXray 404; sample of written children exist. At freeze, design:automated_measurements was still the HTML missing-page fingerprint; it landed later (see Parallel sittings).

Generic — GPT 5.6 Sol medium ([generic](0f26a36d-9359-40b9-845a-2a71738c7ebb)), standing in for fable. No checklist.

ID Verdict Note
G1 child headlines on the outlines accepted Purpose column rewritten as questions/scopes, not findings (IMC crawl-rate, OpenWPM defaults, PanoptiChrome 0-use, “almost never”, ePrivacy-as-conclusion, …).
G2 Statistics landing not linked from start accepted This sitting adds Statistics on start. Provenance claimed it before the put; the put is the last step.
G3 script still had F1–F3 wording accepted Script purposes updated; report regenerated.
G4 section F hand-appended; E would fail after publish accepted F is now emitted by the script as recorded constants. E prints exists/missing and does not fail after publish.
G5 programming provenance overstated the stub-length fail accepted Bytes are printed for review; only exists/missing fails.
G6 provenance claimed “no child headlines” while drafts had them accepted G1 made the claim true; leftover sentences narrowed.
G7 parallel-sitting / drain-item prose on public outlines accepted Moved to this log. Public pages say written vs promised.
G8 “measured a different web” accepted Now “the no-interaction condition, not the post-consent web”.
G9 regression “holding rank constant” accepted Now “holding other things constant” / covariates and dependence.
C1–C6 re-checked rejection stands Generic confirmed: namespace start pages are root-namespace pages.

Published

2026-08-27. dw.mjs put (JSON-RPC, summaries include “Authored by Claude”):

  • new pages, no –if-rev: design, programming, privacy, writing, practices, statistics, and the six provenance:* logs
  • start with –if-rev 1787840313 → rev 1787840590 (Statistics wired; security children no longer promised as unwritten)

Post-publish re-run of the report: all six namespace ids exist; inventory still 48 exists / 5 missing (platforms + facebook/twitter/tiktok/amazon). Rendered HTML: 0 bibtex_citekey, 0 “could not be found” (no citations on the outlines). Red links on design and start are exactly those five platform pages. ?do=export_raw on each published id starts with ======, not the HTML error page. No bibliography append, no bibtex purge.

Unedited report output

report_namespace_overviews.txt
========================================================================
A. CORPUS FRAME (OVERVIEW.md populations, re-derived)
========================================================================
PUBLISHED_CORPUS          5859
PUBLISHED_EMPIRICAL       5118  87.4%
PUBLISHED_CRAWLED         1120  19.1%
PUBLISHED_WEB             1622  27.7%
PUBLISHED_INFERENTIAL     1762  30.1%
PUBLISHED_HUMAN_SUBJECTS  1357  23.2%
PUBLISHED_LEGAL           402  6.9%
PUBLISHED_ANNOTATED       3318  56.6%
 
Venues: CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P.
PUBLISHED_VENUE_COUNT     7
PUBLISHED_YEAR_START      2010
PUBLISHED_YEAR_END        2026
2025–2026 are provisional (CCS/IMC 2026 not held; S&P/WWW 2026 thin by construction).
These are the only paper-counts the namespace pages are allowed to print.
Child-page headline figures belong on the children.
 
========================================================================
B. LIVE CHILD INVENTORY
========================================================================
exists  = export_raw is wiki source, not the HTML "does not exist yet" error page.
missing = that HTML error page. Status code and byte count are not an existence test.
 
PUBLISHED_NS_COUNT        6
Namespace    Children in inventory  Live exists  Live missing (promised)
-----------  ---------------------  -----------  -----------------------
design       15                     10           5
programming  17                     17           0
privacy      9                      9            0
writing      2                      2            0
practices    4                      4            0
statistics   6                      6            0
 
PUBLISHED_EXISTS_TOTAL    48
PUBLISHED_MISSING_TOTAL   5
PUBLISHED_INVENTORY_TOTAL 53
PUBLISHED_DESIGN_EXISTS   10
PUBLISHED_DESIGN_MISSING  5
PUBLISHED_DESIGN_CHILDREN 15
PUBLISHED_PROGRAMMING_EXISTS   17
PUBLISHED_PROGRAMMING_MISSING  0
PUBLISHED_PROGRAMMING_CHILDREN 17
PUBLISHED_PRIVACY_EXISTS   9
PUBLISHED_PRIVACY_MISSING  0
PUBLISHED_PRIVACY_CHILDREN 9
PUBLISHED_WRITING_EXISTS   2
PUBLISHED_WRITING_MISSING  0
PUBLISHED_WRITING_CHILDREN 2
PUBLISHED_PRACTICES_EXISTS   4
PUBLISHED_PRACTICES_MISSING  0
PUBLISHED_PRACTICES_CHILDREN 4
PUBLISHED_STATISTICS_EXISTS   6
PUBLISHED_STATISTICS_MISSING  0
PUBLISHED_STATISTICS_CHILDREN 6
 
PUBLISHED_STUB_COUNT      3
Stubs (exist, short, not corpus-backed against the 2010–2026 extraction):
  programming:crux  8394 bytes
  programming:similarweb  4514 bytes
  programming:multilingual_support  8012 bytes
 
========================================================================
C. PER-NAMESPACE CHILD TABLE
========================================================================
 
--- design ---
Page                                   Live     Bytes   What a student needs it for
-------------------------------------  -------  ------  ---------------------------------------------------------------------------------------------------
[[Design:Website selection]]           exists   23505   Which list substitutes for the web (Tranco, CrUX, Alexa residue).
[[Design:Sampling]]                    exists   52327   How you draw from a list you already chose (top-n vs stratified, size, version).
[[Design:Website classification]]      exists   105360  Topic / industry / company labels — not popularity.
[[Design:IP classification]]           exists   74729   Turning an observed address into a defendable claim (ASN, geo, network type).
[[Design:Crawling location]]           exists   30657   The vantage point you control: country, datacenter vs residential, bot treatment.
[[Design:Archives]]                    exists   41808   Live crawl vs Wayback / Common Crawl — what an archive can and cannot answer.
[[Design:Longitudinal]]                exists   44652   Pinning list, browser, vantage and classifier so wave two is comparable to wave one.
[[Design:Mobile and app measurement]]  exists   56974   Store scraping, static vs dynamic, and whether pinning hid the traffic.
[[Design:Automated measurements]]      exists   19966   Orient between crawling, scanning and app analysis.
[[Design:User studies]]                exists   33966   Participants vs annotators. Crowdworkers labelling data are annotation, not a user study.
[[Design:Platforms]]                   missing  10988   What changes when you measure one large platform rather than a sample of the web. Later drain item.
[[Design:Platforms:Facebook]]          missing  11184   Promised by start. The platforms parent item will keep or drop this red link.
[[Design:Platforms:Twitter]]           missing  11166   Promised by start. Same as Facebook: parent item decides.
[[Design:Platforms:TikTok]]            missing  11148   Promised by start. Same as Facebook: parent item decides.
[[Design:Platforms:Amazon]]            missing  11148   Promised by start. Same as Facebook: parent item decides.
 
--- programming ---
Page                                             Live    Bytes  What a student needs it for
-----------------------------------------------  ------  -----  ---------------------------------------------------------------------------------------------------------
[[Programming:Crawler]]                          exists  52617  Browser vs control channel. Selenium / Puppeteer / Playwright / CDP, then the specialised crawlers.
[[Programming:Crawler:OpenWPM]]                  exists  61555  Unbranded Firefox + privileged WebExtension. Pin the version; cookie_instrument is the only default.
[[Programming:Crawler:Tracker Radar Collector]]  exists  21117  DuckDuckGo Chromium/CDP collector. Not the Tracker Radar dataset, wiki, entity map or detector.
[[Programming:Crawler:PageGraph]]                exists  21510  Brave page-execution recorder. Attribution (which actor caused this request), not a crawler library.
[[Programming:Crawler:webXray]]                  exists  53026  Domain-to-company ownership lists, and what remains of the tool.
[[Programming:Crawler:Foxhound]]                 exists  68958  Firefox fork that taints strings. A flow is not a vulnerability.
[[Programming:Crawler:PanoptiChrome]]            exists  33565  V8 object taint in Chromium. When it is the instrument and when it is not.
[[Programming:Interaction]]                      exists  56388  What the crawler does after page.goto: depth, scroll, click. A website is not a page.
[[Programming:Registration]]                     exists  37842  Login, account creation, newsletter subscribe. Logged-out is the default crawl.
[[Programming:Stateful stateless]]               exists  60504  Whether the profile survives between visits. "We cleared cookies" is not a reset.
[[Programming:Traffic files]]                    exists  45656  HAR vs pcap vs proxy vs CDP. The three taps are not nested.
[[Programming:Docker]]                           exists  24157  Pin the image digest and size /dev/shm. A tag is not a pin.
[[Programming:Tranco]]                           exists  27540  Tranco API and construction. Cite the list id, not "the Tranco top 1M".
[[Programming:Cloudflare Radar]]                 exists  23850  Radar Domain Rankings. A bucket is not a rank; only the top 100 is ordered.
[[Programming:CrUX]]                             exists  8394   Chrome UX Report access (BigQuery). Short instrument note, not corpus-backed.
[[Programming:Similarweb]]                       exists  4514   Similarweb extension/API note. Short stub, not corpus-backed.
[[Programming:Multilingual support]]             exists  8012   Language detection and translation in a crawl. Short, not yet refreshed against the 2010–2026 extraction.
 
--- privacy ---
Page                              Live    Bytes  What a student needs it for
--------------------------------  ------  -----  -------------------------------------------------------------------------------------------------------------------
[[Privacy:Requests]]              exists  78819  Classifying requests (filter lists vs learned). The list is usually both instrument and ground truth.
[[Privacy:Cookies]]               exists  19875  Classifying cookies. First-party is not always benign; third-party cookies were not discontinued.
[[Privacy:Fingerprinting]]        exists  45596  Measuring browser fingerprinting. The word also names unrelated tasks (website / traffic fingerprinting, …).
[[Privacy:JavaScript]]            exists  58197  Observing and labelling scripts. The unit of analysis (domain vs URL vs content) is the load-bearing choice.
[[Privacy:Server side tracking]]  exists  46745  Measuring tracking that never sends a request to the tracker.
[[Privacy:Cookie syncing]]        exists  49815  Detecting ID sync from a crawl. The identifier heuristic is the measurement.
[[Privacy:Consent]]               exists  67169  What the crawler does with the banner. That choice is the treatment, not a config detail.
[[Privacy:TCF consent strings]]   exists  55858  TC string and Google Additional Consent as artefacts. A decode is not a user choice.
[[Privacy:Darkpatterns]]          exists  15075  Deceptive patterns as the thing you measure, and as a confounder of the crawler. HCI venues are not in this corpus.
 
--- writing ---
Page                           Live    Bytes  What a student needs it for
-----------------------------  ------  -----  ---------------------------------------------------------------------------------------------------
[[Writing:Conferences]]        exists  14552  Where to send a web-measurement paper, among these seven venues and the ones the corpus is missing.
[[Writing:Literature review]]  exists  27896  The keyword-derived denominator trap, measured on the SoKs in this corpus.
 
--- practices ---
Page                              Live    Bytes  What a student needs it for
--------------------------------  ------  -----  -------------------------------------------------------------------------------------------------
[[Practices:Ethics]]              exists  71251  "Not human subjects" is a paperwork finding, not a harm finding. What the field actually reports.
[[Practices:Notifying websites]]  exists  58860  Telling operators at crawl scale. Hold back a control arm.
[[Practices:Legal enforcement]]   exists  51641  Which authority is competent. A cookie finding is ePrivacy, not the GDPR one-stop-shop.
[[Practices:Public relations]]    exists  55718  The number and its denominator have to travel in the same sentence.
 
--- statistics ---
Page                                  Live    Bytes  What a student needs it for
------------------------------------  ------  -----  --------------------------------------------------------------------------------------------------------------
[[Statistics:Hypothesis testing]]     exists  63563  Which tests this field uses. Sites are not independent; clustered errors are almost never used.
[[Statistics:Pvalue corrections]]     exists  62373  Multiplicity as a property of the crawl, not a decision. Correction tracks participants, not hypothesis count.
[[Statistics:Regression]]             exists  72483  Counts, clustered rows, and classifiers-that-look-like-regressions in this literature.
[[Statistics:Biases]]                 exists  49101  Selection, survivorship, vantage and denominator bias with the effect sizes this field has measured.
[[Statistics:Interrater agreement]]   exists  40942  Reliability of hand-coded ground truth, and what to report.
[[Statistics:Study preregistration]]  exists  50827  Almost unheard of for crawls. "pre-regist*" is usually a domain, not a study.
 
========================================================================
D. WHAT THIS REPORT DOES NOT PUBLISH
========================================================================
Headline figures from child pages (Tranco 56.4%, OpenWPM 60 papers, …).
Those have their own report scripts. Copying them here would go stale when
a parallel sitting refreshes the child (website_selection and conferences
are claimed in other Cursor sittings as of 2026-08-27).
Full-text sweeps: none. No fold, no residue.
security and artifacts are out of scope (their own drain items).
 
========================================================================
E. LIVE CHECK OF THE SIX NAMESPACE PAGES THEMSELVES
========================================================================
Before this sitting publishes them they are missing. After publish a re-run
will show exists; that is not a FAILURE. Clobbering is prevented by dw.mjs
--if-rev on later edits, not by this check forever requiring absence.
design: exists. export_raw 4375 bytes.
programming: exists. export_raw 4691 bytes.
privacy: exists. export_raw 3997 bytes.
writing: exists. export_raw 1980 bytes.
practices: exists. export_raw 2941 bytes.
statistics: exists. export_raw 3569 bytes.
 
========================================================================
F. SITTING-START OBSERVATIONS (recorded constants, not live-fetched)
========================================================================
These are contemporaneous notes from 2026-08-27, written into the script
so the report file is generated, not hand-appended. They are not claims
about a later wiki state.
PUBLISHED_CONFERENCES_BYTES_AT_SITTING_START  4191
LIVE_CONFERENCES_BYTES_AT_INVENTORY           24753
LIVE_WEBSITE_SELECTION_BYTES_AT_INVENTORY     10118
CHROME_TP_COOKIE_REVERSAL_DAY                 22
CHROME_TP_COOKIE_REVERSAL_YEAR                2025
CHROME_PRIVACY_SANDBOX_POST                   22 April 2025
REFUSED_CHILD_HEADLINE_EXAMPLE                Tranco 56.4%; OpenWPM 60 papers
 
PUBLISHED_SNAPSHOT_DATE   2026-08-27
OK — inventory matches live wiki; corpus frame matches OVERVIEW.md.
provenance/design.1787840722.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki