| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| provenance:design:sampling [2026/08/12 21:40] – Record the generic review pass (10 findings, all accepted), the two corrections applied to design:website_selection, and the missing spot-check row. Authored by Claude. karel.kubicek.claude | provenance:design:sampling [2026/09/21 14:58] (current) – Amendment 2026-09-21: record the crawl-depth scope note against programming:interaction (same 155 numerator, 680 vs 417 denominators), the roster comparison showing the 155 sets are NOT identical (149 overlap, unpublished), the refreshed Hispar/HTTP Archi karel.kubicek.claude |
|---|
| | Report script | ''scripts/report_sampling.mjs'' | | | Report script | ''scripts/report_sampling.mjs'' | |
| | Fold it depends on | ''scripts/sample_fold.mjs'' (sampling-frame families) | | | Fold it depends on | ''scripts/sample_fold.mjs'' (sampling-frame families) | |
| | Published code | ''pages/stratified_sample.py'', embedded on the content page as a downloadable ''<file>'' | | | Published code | ''pages/stratified_sample.py'', embedded on the content page as a downloadable ''%%<file>%%'' | |
| | Data | ''data/extract/run1/extractions.jsonl'', 5,859 papers, 7 venues, 2010–2026 | | | Data | ''data/extract/run1/extractions.jsonl'', 5,859 papers, 7 venues, 2010–2026 | |
| | Bibliography additions | 6 new entries, ''pages/bib_additions_sampling.bib'' | | | Bibliography additions | 6 new entries, ''pages/bib_additions_sampling.bib'' | |
| | ''60'' | the prose "stalled around 60%", rounding 61.6% and 60.8% in the table above it | | | ''60'' | the prose "stalled around 60%", rounding 61.6% and 60.8% in the table above it | |
| |
| The whole-page run (''--code'') returns 32, all of them quotes from cited papers, constants inside the published Python, or values in that script's real JSON output. Both were read rather than assumed clean. | The whole-page run (''%%--code%%'') returns 32, all of them quotes from cited papers, constants inside the published Python, or values in that script's real JSON output. Both were read rather than assumed clean. |
| |
| One figure passes the guard **by coincidence** and is worth flagging: the page says Alexa appears "and 350 other ways" after naming five spellings, which is 355 − 5 and is arithmetic, not a reported figure. It matched a ''350'' elsewhere in the report (a year-bucket count). | One figure passes the guard **by coincidence** and is worth flagging: the page says Alexa appears "and 350 other ways" after naming five spellings, which is 355 − 5 and is arithmetic, not a reported figure. It matched a ''350'' elsewhere in the report (a year-bucket count). |
| ''pages/stratified_sample.py'' is embedded on the content page and was **run**, not asserted. Both sources exercised on 2026-08-12: | ''pages/stratified_sample.py'' is embedded on the content page and was **run**, not asserted. Both sources exercised on 2026-08-12: |
| |
| * ''--source tranco --per-stratum 200 --seed 20260812'' → list ''645KX'', 1,000,000 entries, 800 sites drawn, frame SHA-256 ''feab56e5…''. The JSON in the page's ''<code>'' block is that run verbatim, with the four ''strata'' objects reflowed to one line each — which the page says. | * ''%%--source%% tranco %%--per-stratum%% 200 %%--seed%% 20260812'' → list ''645KX'', 1,000,000 entries, 800 sites drawn, frame SHA-256 ''feab56e5…''. The JSON in the page's ''%%<code>%%'' block is that run verbatim, with the four ''strata'' objects reflowed to one line each — which the page says. |
| * ''--source crux --per-stratum 200 --seed 20260812'' → 1,000,000 entries, 800 drawn, strata populations 1,000 / 9,000 / 90,000 / 900,000, frame SHA-256 ''5a51e5dc…''. | * ''%%--source%% crux %%--per-stratum%% 200 %%--seed%% 20260812'' → 1,000,000 entries, 800 drawn, strata populations 1,000 / 9,000 / 90,000 / 900,000, frame SHA-256 ''5a51e5dc…''. |
| |
| The embedded copy was diffed against the file byte-for-byte after embedding (identical). One bug was found and fixed by running it: ''/download/<id>/<n>'' serves **bare CSV**, not the zip that ''/top-1m.csv.zip'' serves, so the first version died with ''BadZipFile''. It now sniffs the ''PK'' magic. The lesson is the site's standing one — run the thing before publishing it. | The embedded copy was diffed against the file byte-for-byte after embedding (identical). One bug was found and fixed by running it: ''/download/<id>/<n>'' serves **bare CSV**, not the zip that ''/top-1m.csv.zip'' serves, so the first version died with ''BadZipFile''. It now sniffs the ''PK'' magic. The lesson is the site's standing one — run the thing before publishing it. |
| | Mistakes caught in review of my own work | (a) the "impossible Alexa version" finding, retracted after hand-checking — §6; (b) an unverifiable ''prevalence'' figure from Zeber et al., dropped — §6; (c) the published sampler crashed on Tranco's real response format — §8; (d) a first draft asserted "coverage error dominates sampling error" flatly, now a footnoted judgement with its reasoning shown. | | | Mistakes caught in review of my own work | (a) the "impossible Alexa version" finding, retracted after hand-checking — §6; (b) an unverifiable ''prevalence'' figure from Zeber et al., dropped — §6; (c) the published sampler crashed on Tranco's real response format — §8; (d) a first draft asserted "coverage error dominates sampling error" flatly, now a footnoted judgement with its reasoning shown. | |
| | Discussion block | None on this page, following the convention set by the other ''provenance:'' pages — comments belong on the content page. Checked against the published ''provenance:design:crawling_location'', which likewise carries no ''<bibtex bibliography>'' block, so the citekeys here render as markers without a reference list. That is the existing convention, not an omission. | | | Discussion block | None on this page, following the convention set by the other ''provenance:'' pages — comments belong on the content page. Checked against the published ''provenance:design:crawling_location'', which likewise carries no ''<bibtex bibliography>'' block, so the citekeys here render as markers without a reference list. That is the existing convention, not an omission. | |
| | Publication order | ''literature:bibliography'' (6 entries) → ''design:sampling'' → ''provenance:design:sampling'', all 2026-08-12. Rendering verified afterwards: 32 inline ''bibtex_citekey'' markers and a ''bibtex_references'' list on the content page, 0 unresolved keys, 3 red links (''statistics:biases'', ''statistics:hypothesis_testing'', ''artifacts'' — all pages ''start'' already promises), and the published ''<file>'' round-trips byte-identical to ''pages/stratified_sample.py'' apart from a stripped trailing newline. | | | Publication order | ''literature:bibliography'' (6 entries) → ''design:sampling'' → ''provenance:design:sampling'', all 2026-08-12. Rendering verified afterwards: 32 inline ''bibtex_citekey'' markers and a ''bibtex_references'' list on the content page, 0 unresolved keys, 3 red links (''statistics:biases'', ''statistics:hypothesis_testing'', ''artifacts'' — all pages ''start'' already promises), and the published ''%%<file>%%'' round-trips byte-identical to ''pages/stratified_sample.py'' apart from a stripped trailing newline. | |
| | Published before Pass D returned | **Yes, deliberately.** Three focused passes had returned and their findings were applied; the generic pass typically returns prose and structure findings, which are a second revision rather than a blocker. Given the run had already been interrupted several times, shipping a reviewed page and revising it was judged better than risking an unpublished one. Whatever Pass D finds is applied as a follow-up edit and recorded above. A reader comparing revisions should know the first published revision predates one of the four reviews. | | | Published before Pass D returned | **Yes, deliberately.** Three focused passes had returned and their findings were applied; the generic pass typically returns prose and structure findings, which are a second revision rather than a blocker. Given the run had already been interrupted several times, shipping a reviewed page and revising it was judged better than risking an unpublished one. Whatever Pass D finds is applied as a follow-up edit and recorded above. A reader comparing revisions should know the first published revision predates one of the four reviews. | |
| |
| [[design:sampling|← back to the content page]] · [[literature:corpus|corpus-level provenance]] | [[design:sampling|← back to the content page]] · [[literature:corpus|corpus-level provenance]] |
| | |
| | ===== Markup sweep, 2026-09-17 ===== |
| | |
| | Mechanical rendering repair only: a fresh live raw/XHTML export of 188 pages was checked with ''check_wrap.mjs'' and ''check_typography.mjs''. Affected plugin tags, CLI flags and heading markup were repaired; no figures or substantive prose were changed. The resulting source and rendered DOM were re-checked after saving. |
| | |
| | ===== Amendment, 2026-09-21: crawl-depth figures scoped against programming:interaction, and the Hispar box refreshed ===== |
| | |
| | **What was wrong.** Nothing on this page was a wrong number. Two things were wrong for a reader who reads two pages. |
| | |
| | * **The depth figures had no scope note.** This page publishes **22.8% of 680** landing-page-only and **15.0%** stating a subpage count; [[programming:interaction]] publishes **37.2% of 417** on the site-depth axis and **12.1% of 857** giving a subpage number. Four figures, two pages, no sentence anywhere telling a reader why they differ. |
| | * **The Hispar box was stale.** Checked 2026-08-12, it said "no maintained equivalent list was found", and the //Open Questions// bullet said the 2020 landing-page result "stands unaddressed, with no tooling to address it". Both had been overtaken by [[programming:interaction]], which documents HTTP Archive's one-secondary-page-per-site crawl (April 2022 onwards, ''is_root_page'' / ''root_page'' columns, ''MAX_DEPTH = 1'' / ''MAX_BREADTH = 1'', the first same-origin link) and measures what that rule misses. |
| | |
| | **The commands and their real output.** |
| | |
| | <code> |
| | $ node scripts/report_sampling.mjs |
| | page population that also crawled: 680 |
| | Interaction depth Papers Share of 680 |
| | single-target-page 222 32.6% |
| | landing-page-only 155 22.8% |
| | landing-plus-subpages 128 18.8% |
| | deep-crawl 92 13.5% |
| | not-stated 69 10.1% |
| | no crawlConfig record 14 2.1% |
| | states how many subpages per site: 102 15.0% |
| | median 10, max 2000 |
| | goes past the landing page: 220 32.4% |
| | papers with unit `websites` that crawled: 492 |
| | landing page only: 122 24.8% |
| | |
| | $ node scripts/report_interaction.mjs |
| | webCrawled (crawled AND platforms includes 'web') 857 <- this page's denominator |
| | denominator: 417 papers (NOT 857; single-target-page and not-stated are excluded) |
| | landing-page-only 155 37.2% |
| | denominator: 857 web crawls; 104 (12.1%) give a number. |
| | of the 417 on the site-depth axis, 102 (24.5%) give a number. |
| | </code> |
| | |
| | **The reconciliation, and why it is a scope note and not a merge.** The two pages report the **same raw count of landing-page-only papers, 155**, and divide it by different denominators: 680 here (every paper with a web-unit population that also crawled, keeping ''single-target-page'', ''not-stated'' and the no-record papers in the denominator) against 417 there (only the web crawls that put a value on the site-depth axis). That is the whole of the 22.8%-against-37.2% gap, and it is now stated on this page in those terms. |
| | |
| | **The matching 155 is a coincidence of count, not a shared set, and the page says so.** Before writing the sentence, the two rosters were compared directly, because "the same 155 papers" would have been the natural and wrong thing to write: |
| | |
| | <code> |
| | interaction landing-only: 155 sampling landing-only: 155 |
| | intersection: 149 |
| | in interaction not sampling: 6 in sampling not interaction: 6 |
| | </code> |
| | |
| | The six each way are the two population definitions, not an error: this page needs a ''population[].unit'' in {websites, domains, web-pages}, that page needs ''platforms'' to include ''web''. Traffic-analysis and censorship papers sit in one and not the other. **No overlap figure was published** — it is not produced by either page's report script, and publishing it would put a number on a page that no script regenerates. It is recorded here instead, which is what this page is for, so that a later run does not "reconcile" two counts that only look identical. |
| | |
| | **What changed on the page.** |
| | |
| | - The //One page per site is the norm// summary gained a paragraph stating the 680-against-417 scope difference explicitly, naming the 222 + 69 + 14 rows that this page keeps and that page excludes, and warning that the 155-paper sets are not identical. |
| | - The **15.0%** sentence gained its numerator (102 of 680) and a pointer to the 12.1% (104 of 857) with its own denominator named. |
| | - The Hispar ''%%<WRAP todo>%%'' box was rewritten: Hispar is still gone (NXDOMAIN, 2026-08-12, unchanged), but the box now records HTTP Archive as the one maintained source of internal pages, states the selection rule that makes it a by-product rather than a sample, and points at the measured demonstration on [[programming:interaction]]. The claim's footnotes and their check dates live on that page and were **not** copied here — one number, one page. |
| | - The //Open Questions// bullet was rewritten the same way: the gap is now "no maintained internal-page //list//", not "no tooling". |
| | |
| | **Rendered-DOM check.** Page body after the table of contents: list items 40, headings 33, tables 13, ''%%<pre>%%'' blocks 2 — **all four identical before and after**, so nothing was swallowed by the rewritten ''%%<WRAP>%%'' box or by the wrapped //Open Questions// bullet. **0** ''wikilink2'' red links. All three cross-page anchors resolve against real heading ids on ''programming:interaction'': ''depth_has_more_than_three_positions'', ''how_deep_the_field_actually_goes'', ''the_one_page_list_that_included_internal_pages_is_gone''. |
| | |
| | **Finding rejected.** The first draft of the scope note compared this page's **24.8% of 492** against the 37.2%, because that is the pairing the task named. It was dropped: 24.8% is the ''unit == websites'' cut (122 of 492) and shares neither numerator nor denominator with anything on the other page, so pairing them would have invented a comparison. The comparable pair is 22.8% of 680 against 37.2% of 417, which share the numerator 155. The 24.8% sentence is left exactly as it was. |
| | |
| | **Reviewers.** One ''sonnet'' figures-against-script pass over both edits and the sections either side of them. No citations pass: no ''{[citekey]}'' and no quoted claim was touched — ''{[aqeel2020_landing]}'' and the Aqeel quote in the same subsection are unchanged. |
| |