User Tools

Site Tools


provenance:design:sampling

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
provenance:design:sampling [2026/08/12 21:40] – Record the generic review pass (10 findings, all accepted), the two corrections applied to design:website_selection, and the missing spot-check row. Authored by Claude. karel.kubicek.claudeprovenance:design:sampling [2026/09/21 14:58] (current) – Amendment 2026-09-21: record the crawl-depth scope note against programming:interaction (same 155 numerator, 680 vs 417 denominators), the roster comparison showing the 155 sets are NOT identical (149 overlap, unpublished), the refreshed Hispar/HTTP Archi karel.kubicek.claude
Line 11: Line 11:
 | Report script | ''scripts/report_sampling.mjs'' | | Report script | ''scripts/report_sampling.mjs'' |
 | Fold it depends on | ''scripts/sample_fold.mjs'' (sampling-frame families) | | Fold it depends on | ''scripts/sample_fold.mjs'' (sampling-frame families) |
-| Published code | ''pages/stratified_sample.py'', embedded on the content page as a downloadable ''<file>'' |+| Published code | ''pages/stratified_sample.py'', embedded on the content page as a downloadable ''%%<file>%%'' |
 | Data | ''data/extract/run1/extractions.jsonl'', 5,859 papers, 7 venues, 2010–2026 | | Data | ''data/extract/run1/extractions.jsonl'', 5,859 papers, 7 venues, 2010–2026 |
 | Bibliography additions | 6 new entries, ''pages/bib_additions_sampling.bib'' | | Bibliography additions | 6 new entries, ''pages/bib_additions_sampling.bib'' |
Line 86: Line 86:
 | ''60'' | the prose "stalled around 60%", rounding 61.6% and 60.8% in the table above it | | ''60'' | the prose "stalled around 60%", rounding 61.6% and 60.8% in the table above it |
  
-The whole-page run (''--code'') returns 32, all of them quotes from cited papers, constants inside the published Python, or values in that script's real JSON output. Both were read rather than assumed clean.+The whole-page run (''%%--code%%'') returns 32, all of them quotes from cited papers, constants inside the published Python, or values in that script's real JSON output. Both were read rather than assumed clean.
  
 One figure passes the guard **by coincidence** and is worth flagging: the page says Alexa appears "and 350 other ways" after naming five spellings, which is 355 − 5 and is arithmetic, not a reported figure. It matched a ''350'' elsewhere in the report (a year-bucket count). One figure passes the guard **by coincidence** and is worth flagging: the page says Alexa appears "and 350 other ways" after naming five spellings, which is 355 − 5 and is arithmetic, not a reported figure. It matched a ''350'' elsewhere in the report (a year-bucket count).
Line 209: Line 209:
 ''pages/stratified_sample.py'' is embedded on the content page and was **run**, not asserted. Both sources exercised on 2026-08-12: ''pages/stratified_sample.py'' is embedded on the content page and was **run**, not asserted. Both sources exercised on 2026-08-12:
  
-  * ''--source tranco --per-stratum 200 --seed 20260812'' → list ''645KX'', 1,000,000 entries, 800 sites drawn, frame SHA-256 ''feab56e5…''. The JSON in the page's ''<code>'' block is that run verbatim, with the four ''strata'' objects reflowed to one line each — which the page says. +  * ''%%--source%% tranco %%--per-stratum%% 200 %%--seed%% 20260812'' → list ''645KX'', 1,000,000 entries, 800 sites drawn, frame SHA-256 ''feab56e5…''. The JSON in the page's ''%%<code>%%'' block is that run verbatim, with the four ''strata'' objects reflowed to one line each — which the page says. 
-  * ''--source crux --per-stratum 200 --seed 20260812'' → 1,000,000 entries, 800 drawn, strata populations 1,000 / 9,000 / 90,000 / 900,000, frame SHA-256 ''5a51e5dc…''.+  * ''%%--source%% crux %%--per-stratum%% 200 %%--seed%% 20260812'' → 1,000,000 entries, 800 drawn, strata populations 1,000 / 9,000 / 90,000 / 900,000, frame SHA-256 ''5a51e5dc…''.
  
 The embedded copy was diffed against the file byte-for-byte after embedding (identical). One bug was found and fixed by running it: ''/download/<id>/<n>'' serves **bare CSV**, not the zip that ''/top-1m.csv.zip'' serves, so the first version died with ''BadZipFile''. It now sniffs the ''PK'' magic. The lesson is the site's standing one — run the thing before publishing it. The embedded copy was diffed against the file byte-for-byte after embedding (identical). One bug was found and fixed by running it: ''/download/<id>/<n>'' serves **bare CSV**, not the zip that ''/top-1m.csv.zip'' serves, so the first version died with ''BadZipFile''. It now sniffs the ''PK'' magic. The lesson is the site's standing one — run the thing before publishing it.
Line 292: Line 292:
 | Mistakes caught in review of my own work | (a) the "impossible Alexa version" finding, retracted after hand-checking — §6; (b) an unverifiable ''prevalence'' figure from Zeber et al., dropped — §6; (c) the published sampler crashed on Tranco's real response format — §8; (d) a first draft asserted "coverage error dominates sampling error" flatly, now a footnoted judgement with its reasoning shown. | | Mistakes caught in review of my own work | (a) the "impossible Alexa version" finding, retracted after hand-checking — §6; (b) an unverifiable ''prevalence'' figure from Zeber et al., dropped — §6; (c) the published sampler crashed on Tranco's real response format — §8; (d) a first draft asserted "coverage error dominates sampling error" flatly, now a footnoted judgement with its reasoning shown. |
 | Discussion block | None on this page, following the convention set by the other ''provenance:'' pages — comments belong on the content page. Checked against the published ''provenance:design:crawling_location'', which likewise carries no ''<bibtex bibliography>'' block, so the citekeys here render as markers without a reference list. That is the existing convention, not an omission. | | Discussion block | None on this page, following the convention set by the other ''provenance:'' pages — comments belong on the content page. Checked against the published ''provenance:design:crawling_location'', which likewise carries no ''<bibtex bibliography>'' block, so the citekeys here render as markers without a reference list. That is the existing convention, not an omission. |
-| Publication order | ''literature:bibliography'' (6 entries) → ''design:sampling'' → ''provenance:design:sampling'', all 2026-08-12. Rendering verified afterwards: 32 inline ''bibtex_citekey'' markers and a ''bibtex_references'' list on the content page, 0 unresolved keys, 3 red links (''statistics:biases'', ''statistics:hypothesis_testing'', ''artifacts'' — all pages ''start'' already promises), and the published ''<file>'' round-trips byte-identical to ''pages/stratified_sample.py'' apart from a stripped trailing newline. |+| Publication order | ''literature:bibliography'' (6 entries) → ''design:sampling'' → ''provenance:design:sampling'', all 2026-08-12. Rendering verified afterwards: 32 inline ''bibtex_citekey'' markers and a ''bibtex_references'' list on the content page, 0 unresolved keys, 3 red links (''statistics:biases'', ''statistics:hypothesis_testing'', ''artifacts'' — all pages ''start'' already promises), and the published ''%%<file>%%'' round-trips byte-identical to ''pages/stratified_sample.py'' apart from a stripped trailing newline. |
 | Published before Pass D returned | **Yes, deliberately.** Three focused passes had returned and their findings were applied; the generic pass typically returns prose and structure findings, which are a second revision rather than a blocker. Given the run had already been interrupted several times, shipping a reviewed page and revising it was judged better than risking an unpublished one. Whatever Pass D finds is applied as a follow-up edit and recorded above. A reader comparing revisions should know the first published revision predates one of the four reviews. | | Published before Pass D returned | **Yes, deliberately.** Three focused passes had returned and their findings were applied; the generic pass typically returns prose and structure findings, which are a second revision rather than a blocker. Given the run had already been interrupted several times, shipping a reviewed page and revising it was judged better than risking an unpublished one. Whatever Pass D finds is applied as a follow-up edit and recorded above. A reader comparing revisions should know the first published revision predates one of the four reviews. |
  
 [[design:sampling|← back to the content page]] · [[literature:corpus|corpus-level provenance]] [[design:sampling|← back to the content page]] · [[literature:corpus|corpus-level provenance]]
 +
 +===== Markup sweep, 2026-09-17 =====
 +
 +Mechanical rendering repair only: a fresh live raw/XHTML export of 188 pages was checked with ''check_wrap.mjs'' and ''check_typography.mjs''. Affected plugin tags, CLI flags and heading markup were repaired; no figures or substantive prose were changed. The resulting source and rendered DOM were re-checked after saving.
 +
 +===== Amendment, 2026-09-21: crawl-depth figures scoped against programming:interaction, and the Hispar box refreshed =====
 +
 +**What was wrong.** Nothing on this page was a wrong number. Two things were wrong for a reader who reads two pages.
 +
 +  * **The depth figures had no scope note.** This page publishes **22.8% of 680** landing-page-only and **15.0%** stating a subpage count; [[programming:interaction]] publishes **37.2% of 417** on the site-depth axis and **12.1% of 857** giving a subpage number. Four figures, two pages, no sentence anywhere telling a reader why they differ.
 +  * **The Hispar box was stale.** Checked 2026-08-12, it said "no maintained equivalent list was found", and the //Open Questions// bullet said the 2020 landing-page result "stands unaddressed, with no tooling to address it". Both had been overtaken by [[programming:interaction]], which documents HTTP Archive's one-secondary-page-per-site crawl (April 2022 onwards, ''is_root_page'' / ''root_page'' columns, ''MAX_DEPTH = 1'' / ''MAX_BREADTH = 1'', the first same-origin link) and measures what that rule misses.
 +
 +**The commands and their real output.**
 +
 +<code>
 +$ node scripts/report_sampling.mjs
 +page population that also crawled: 680
 +Interaction depth      Papers  Share of 680
 +single-target-page     222     32.6%
 +landing-page-only      155     22.8%
 +landing-plus-subpages  128     18.8%
 +deep-crawl             92      13.5%
 +not-stated             69      10.1%
 +no crawlConfig record  14      2.1%
 +states how many subpages per site: 102  15.0%
 +  median 10, max 2000
 +goes past the landing page: 220  32.4%
 +papers with unit `websites` that crawled: 492
 +  landing page only:                      122  24.8%
 +
 +$ node scripts/report_interaction.mjs
 +webCrawled  (crawled AND platforms includes 'web')         857   <- this page's denominator
 +  denominator: 417 papers (NOT 857; single-target-page and not-stated are excluded)
 +landing-page-only      155     37.2%
 +  denominator: 857 web crawls; 104 (12.1%) give a number.
 +  of the 417 on the site-depth axis, 102 (24.5%) give a number.
 +</code>
 +
 +**The reconciliation, and why it is a scope note and not a merge.** The two pages report the **same raw count of landing-page-only papers, 155**, and divide it by different denominators: 680 here (every paper with a web-unit population that also crawled, keeping ''single-target-page'', ''not-stated'' and the no-record papers in the denominator) against 417 there (only the web crawls that put a value on the site-depth axis). That is the whole of the 22.8%-against-37.2% gap, and it is now stated on this page in those terms.
 +
 +**The matching 155 is a coincidence of count, not a shared set, and the page says so.** Before writing the sentence, the two rosters were compared directly, because "the same 155 papers" would have been the natural and wrong thing to write:
 +
 +<code>
 +interaction landing-only: 155     sampling landing-only: 155
 +intersection: 149
 +in interaction not sampling: 6    in sampling not interaction: 6
 +</code>
 +
 +The six each way are the two population definitions, not an error: this page needs a ''population[].unit'' in {websites, domains, web-pages}, that page needs ''platforms'' to include ''web''. Traffic-analysis and censorship papers sit in one and not the other. **No overlap figure was published** — it is not produced by either page's report script, and publishing it would put a number on a page that no script regenerates. It is recorded here instead, which is what this page is for, so that a later run does not "reconcile" two counts that only look identical.
 +
 +**What changed on the page.**
 +
 +  - The //One page per site is the norm// summary gained a paragraph stating the 680-against-417 scope difference explicitly, naming the 222 + 69 + 14 rows that this page keeps and that page excludes, and warning that the 155-paper sets are not identical.
 +  - The **15.0%** sentence gained its numerator (102 of 680) and a pointer to the 12.1% (104 of 857) with its own denominator named.
 +  - The Hispar ''%%<WRAP todo>%%'' box was rewritten: Hispar is still gone (NXDOMAIN, 2026-08-12, unchanged), but the box now records HTTP Archive as the one maintained source of internal pages, states the selection rule that makes it a by-product rather than a sample, and points at the measured demonstration on [[programming:interaction]]. The claim's footnotes and their check dates live on that page and were **not** copied here — one number, one page.
 +  - The //Open Questions// bullet was rewritten the same way: the gap is now "no maintained internal-page //list//", not "no tooling".
 +
 +**Rendered-DOM check.** Page body after the table of contents: list items 40, headings 33, tables 13, ''%%<pre>%%'' blocks 2 — **all four identical before and after**, so nothing was swallowed by the rewritten ''%%<WRAP>%%'' box or by the wrapped //Open Questions// bullet. **0** ''wikilink2'' red links. All three cross-page anchors resolve against real heading ids on ''programming:interaction'': ''depth_has_more_than_three_positions'', ''how_deep_the_field_actually_goes'', ''the_one_page_list_that_included_internal_pages_is_gone''.
 +
 +**Finding rejected.** The first draft of the scope note compared this page's **24.8% of 492** against the 37.2%, because that is the pairing the task named. It was dropped: 24.8% is the ''unit == websites'' cut (122 of 492) and shares neither numerator nor denominator with anything on the other page, so pairing them would have invented a comparison. The comparable pair is 22.8% of 680 against 37.2% of 417, which share the numerator 155. The 24.8% sentence is left exactly as it was.
 +
 +**Reviewers.** One ''sonnet'' figures-against-script pass over both edits and the sections either side of them. No citations pass: no ''{[citekey]}'' and no quoted claim was touched — ''{[aqeel2020_landing]}'' and the Aqeel quote in the same subsection are unchanged.
  
provenance/design/sampling.1786570821.txt.gz · Last modified: by karel.kubicek.claude