| Next revision | Previous revision |
| provenance:design:sampling [2026/08/12 21:28] – Provenance for design:sampling: populations and denominators, the frame fold and its residue, quotes spot-checked, two retracted figures, external verification log, and the review passes. Authored by Claude. karel.kubicek.claude | provenance:design:sampling [2026/09/21 14:58] (current) – Amendment 2026-09-21: record the crawl-depth scope note against programming:interaction (same 155 numerator, 680 vs 417 denominators), the roster comparison showing the 155 sets are NOT identical (149 overlap, unpublished), the refreshed Hispar/HTTP Archi karel.kubicek.claude |
|---|
| | Report script | ''scripts/report_sampling.mjs'' | | | Report script | ''scripts/report_sampling.mjs'' | |
| | Fold it depends on | ''scripts/sample_fold.mjs'' (sampling-frame families) | | | Fold it depends on | ''scripts/sample_fold.mjs'' (sampling-frame families) | |
| | Published code | ''pages/stratified_sample.py'', embedded on the content page as a downloadable ''<file>'' | | | Published code | ''pages/stratified_sample.py'', embedded on the content page as a downloadable ''%%<file>%%'' | |
| | Data | ''data/extract/run1/extractions.jsonl'', 5,859 papers, 7 venues, 2010–2026 | | | Data | ''data/extract/run1/extractions.jsonl'', 5,859 papers, 7 venues, 2010–2026 | |
| | Bibliography additions | 6 new entries, ''pages/bib_additions_sampling.bib'' | | | Bibliography additions | 6 new entries, ''pages/bib_additions_sampling.bib'' | |
| * ''sampling'' does **not** catalogue services. Where a list needs describing, it links. | * ''sampling'' does **not** catalogue services. Where a list needs describing, it links. |
| |
| **Two corrections owed to design:website_selection** were found and are //not// applied here, because that page is a different work item: | **Two corrections were owed to design:website_selection, and were applied** (rev 1786570786, 2026-08-12). The original decision was to leave them, on the grounds that the neighbouring page is a different work item; the generic review pass argued that this is a process nicety the reader never sees, and that what they //would// see is two pages of one wiki disagreeing on checkable facts while telling them to read both. That was right, and the fixes are two lines: |
| |
| - It says Alexa was "**Discontinued as of August 1, 2023**". The primary source says Alexa.com was retired **1 May 2022** and the Top Sites / Web Information Service APIs on **15 December 2022** (see §7). No source was found for an August 2023 date. | - It said Alexa was "**Discontinued as of August 1, 2023**". The primary source says Alexa.com was retired **1 May 2022** and the Top Sites / Web Information Service APIs on **15 December 2022** (see §7). No source was found for an August 2023 date, and the external-currency pass independently failed to find one — including in AWS's own full-shutdown listing, which does not mention Alexa at all. |
| - It says Tranco "combines data from Alexa, Cisco Umbrella, and Majestic over a 30-day period". Tranco's own API reports the providers of the 2026-08-01 daily list as ''crux, farsight, majestic, radar, umbrella'' — **Alexa is not among them** and has not been for years. Tranco's own prose methodology page is also stale here, saying "all four providers" for a five-provider list. | - It said Tranco "combines data from Alexa, Cisco Umbrella, and Majestic over a 30-day period". Tranco's own API reports the providers of the 2026-08-01 daily list as ''crux, farsight, majestic, radar, umbrella'' — **Alexa is not among them**. The sentence now dates the 2019 composition as historical and gives the live one, with the API call in a footnote. Tranco's own prose methodology page is stale here too, saying "all four providers" for a five-provider list. |
| | |
| | Nothing else on that page was touched: it keeps its structure, its TODOs and its own //Use in Publications// section. |
| |
| ===== 3. Populations and denominators ===== | ===== 3. Populations and denominators ===== |
| | ''60'' | the prose "stalled around 60%", rounding 61.6% and 60.8% in the table above it | | | ''60'' | the prose "stalled around 60%", rounding 61.6% and 60.8% in the table above it | |
| |
| The whole-page run (''--code'') returns 32, all of them quotes from cited papers, constants inside the published Python, or values in that script's real JSON output. Both were read rather than assumed clean. | The whole-page run (''%%--code%%'') returns 32, all of them quotes from cited papers, constants inside the published Python, or values in that script's real JSON output. Both were read rather than assumed clean. |
| |
| One figure passes the guard **by coincidence** and is worth flagging: the page says Alexa appears "and 350 other ways" after naming five spellings, which is 355 − 5 and is arithmetic, not a reported figure. It matched a ''350'' elsewhere in the report (a year-bucket count). | One figure passes the guard **by coincidence** and is worth flagging: the page says Alexa appears "and 350 other ways" after naming five spellings, which is 355 − 5 and is arithmetic, not a reported figure. It matched a ''350'' elsewhere in the report (a year-bucket count). |
| | "median number of tracking domains … is 1.9, whereas for the crawler it is 6.1"; "the crawler may reach 26" | {[zeber2020representativeness]} | exact | | | "median number of tracking domains … is 1.9, whereas for the crawler it is 6.1"; "the crawler may reach 26" | {[zeber2020representativeness]} | exact | |
| | "the top million sites capture over 95% of all page loads and time spent online" | {[ruth2022_world]} | exact | | | "the top million sites capture over 95% of all page loads and time spent online" | {[ruth2022_world]} | exact | |
| | "one site garners 17% of all desktop page loads globally"; "ten sites accounting for about half of time spent" | {[ruth2022_world]} | exact | | | "one site garners 17% of all desktop page loads globally"; "ten sites accounting for about half of time spent"; "six sites account for 25% of page loads on both desktop and mobile" | {[ruth2022_world]} | exact — the third was added to this table after a reviewer noticed it was quoted on the page but not listed here | |
| | "disproportionate focus on the long tail of the web" | {[ruth2022_world]} | present, **column-spliced** — the sentence interleaves with the adjacent column, so it is quoted in fragments on the page rather than as one run | | | "disproportionate focus on the long tail of the web" | {[ruth2022_world]} | present, **column-spliced** — the sentence interleaves with the adjacent column, so it is quoted in fragments on the page rather than as one run | |
| | "we present statistics for 917,261 sites" | {[englehardt2016online]} | exact | | | "we present statistics for 917,261 sites" | {[englehardt2016online]} | exact | |
| ''pages/stratified_sample.py'' is embedded on the content page and was **run**, not asserted. Both sources exercised on 2026-08-12: | ''pages/stratified_sample.py'' is embedded on the content page and was **run**, not asserted. Both sources exercised on 2026-08-12: |
| |
| * ''--source tranco --per-stratum 200 --seed 20260812'' → list ''645KX'', 1,000,000 entries, 800 sites drawn, frame SHA-256 ''feab56e5…''. The JSON in the page's ''<code>'' block is that run verbatim, with the four ''strata'' objects reflowed to one line each — which the page says. | * ''%%--source%% tranco %%--per-stratum%% 200 %%--seed%% 20260812'' → list ''645KX'', 1,000,000 entries, 800 sites drawn, frame SHA-256 ''feab56e5…''. The JSON in the page's ''%%<code>%%'' block is that run verbatim, with the four ''strata'' objects reflowed to one line each — which the page says. |
| * ''--source crux --per-stratum 200 --seed 20260812'' → 1,000,000 entries, 800 drawn, strata populations 1,000 / 9,000 / 90,000 / 900,000, frame SHA-256 ''5a51e5dc…''. | * ''%%--source%% crux %%--per-stratum%% 200 %%--seed%% 20260812'' → 1,000,000 entries, 800 drawn, strata populations 1,000 / 9,000 / 90,000 / 900,000, frame SHA-256 ''5a51e5dc…''. |
| |
| The embedded copy was diffed against the file byte-for-byte after embedding (identical). One bug was found and fixed by running it: ''/download/<id>/<n>'' serves **bare CSV**, not the zip that ''/top-1m.csv.zip'' serves, so the first version died with ''BadZipFile''. It now sniffs the ''PK'' magic. The lesson is the site's standing one — run the thing before publishing it. | The embedded copy was diffed against the file byte-for-byte after embedding (identical). One bug was found and fixed by running it: ''/download/<id>/<n>'' serves **bare CSV**, not the zip that ''/top-1m.csv.zip'' serves, so the first version died with ''BadZipFile''. It now sniffs the ''PK'' magic. The lesson is the site's standing one — run the thing before publishing it. |
| |
| * **Attrition rates across the field.** The extraction schema has ''population.n'' but no "successfully measured" field, so the corpus cannot say how often a paper reports both the drawn and the analysed denominator. The page argues the point from two verified examples and lists it as an open question rather than quantifying it. Closing it needs a targeted full-text study. | * **Attrition rates across the field.** The extraction schema has ''population.n'' but no "successfully measured" field, so the corpus cannot say how often a paper reports both the drawn and the analysed denominator. The page argues the point from two verified examples and lists it as an open question rather than quantifying it. Closing it needs a targeted full-text study. |
| * **Whether any paper reports the same measurement per rank stratum.** Searched the ''detection[].prevalence'' field with rank- and popularity-related patterns and read about 30 matches; none is a per-stratum breakdown of the paper's own headline measurement. Absence of evidence in a free-text field is weak evidence, so the page phrases this as "we found none" and asks for counter-examples. | * **Whether any paper reports the same measurement per rank stratum.** Searched the ''detection[].prevalence'' field with rank- and popularity-related patterns and read about 30 matches; a narrower re-run over just the 69 stratified papers returned two matches, neither of them a per-stratum breakdown of the paper's own headline figure. Absence of evidence in a free-text field is weak evidence, so the page phrases this as "we found none" and asks for counter-examples. |
| * **Whether rank weighting is used outside third-party analysis.** The first draft of the page asserted in an open question that no weighted estimator appears in this corpus. That was **wrong and was corrected before publication**: {[englehardt2016online]}'s //prominence// (Σ 1/rank over the sites a third party appears on) is exactly a rank-weighted estimator, introduced for exactly the sampling reason the page's size section argues. It was found by reading the paper rather than by any query — the extraction records it as a metric, not as a sampling decision, so no ''population'' query would have surfaced it. The page now has a section on it, and the open question was narrowed to what remains unfound: a weighted //prevalence// reported alongside the unweighted one, and reweighting a rank-stratified sample by stratum size. **Take the narrowed question as weakly evidenced too** — it rests on the same free-text search. | * **Whether rank weighting is used outside third-party analysis.** The first draft of the page asserted in an open question that no weighted estimator appears in this corpus. That was **wrong and was corrected before publication**: {[englehardt2016online]}'s //prominence// (Σ 1/rank over the sites a third party appears on) is exactly a rank-weighted estimator, introduced for exactly the sampling reason the page's size section argues. It was found by reading the paper rather than by any query — the extraction records it as a metric, not as a sampling decision, so no ''population'' query would have surfaced it. The page now has a section on it, and the open question was narrowed to what remains unfound: a weighted //prevalence// reported alongside the unweighted one, and reweighting a rank-stratified sample by stratum size. **Take the narrowed question as weakly evidenced too** — it rests on the same free-text search. |
| |
| * **Whether any paper reports the same measurement per rank stratum.** Re-run over just the 69 stratified papers' ''detection[].prevalence'' fields: two matches, neither of them a per-stratum breakdown of the paper's own headline figure. | |
| * **How much of the 2025–2026 top-//n// rise is real.** Those venue-years are provisional by construction. 64.7% rests on 190 papers from years that are incompletely indexed; the direction is consistent with 2018–2024 and the page labels the column, but it should not be quoted as a 2026 measurement. | * **How much of the 2025–2026 top-//n// rise is real.** Those venue-years are provisional by construction. 64.7% rests on 190 papers from years that are incompletely indexed; the direction is consistent with 2018–2024 and the page labels the column, but it should not be quoted as a 2026 measurement. |
| * **Whether the 66 post-retirement Alexa papers are reusing archives or copying a citation.** Reading a dozen suggests both happen and that side-channel and website-fingerprinting evaluations inherit "top 100 Alexa" as a benchmark convention from earlier papers. That impression is not quantified and is not on the content page as a figure. | * **Whether the 66 post-retirement Alexa papers are reusing archives or copying a citation.** Reading a dozen suggests both happen and that side-channel and website-fingerprinting evaluations inherit "top 100 Alexa" as a benchmark convention from earlier papers. That impression is not quantified and is not on the content page as a figure. |
| |
| The cost was real even so: Pass C spent effort reporting a figure that was already fixed, and Pass B reviewed a sentence that had been rewritten once since the snapshot. **Next time: freeze, queue the fixes, and apply them in one batch after the passes return.** Recorded here rather than quietly omitted, because a review log that hides how the review actually ran is worth less than no log. | The cost was real even so: Pass C spent effort reporting a figure that was already fixed, and Pass B reviewed a sentence that had been rewritten once since the snapshot. **Next time: freeze, queue the fixes, and apply them in one batch after the passes return.** Recorded here rather than quietly omitted, because a review log that hides how the review actually ran is worth less than no log. |
| | |
| | ==== Pass D — generic (Claude Fable) ==== |
| | |
| | Run last, against a **second, honest freeze** (''out/freeze2/'', page MD5 ''7c7b5793…''), with no checklist: overstated claims, structure, scope boundary, internal contradictions, tone, and whether the page answers its own question. It also reviews this page. |
| | |
| | It returned the largest and most useful set of findings of the four, and **the page was published before it did** (run log below). Every item was accepted; the page was revised and re-saved the same day. |
| | |
| | ^ Finding ^ Disposition ^ |
| | | **The lead overstated three times in twelve lines**, on a page whose thesis is "say only what your evidence carries": (a) "the share is still rising" rests entirely on the provisional 2025–26 bucket, and read to 2024 as the page's own methodology section instructs, top-//n// //fell// 59.6% → 58.0%; (b) "6.0% stratify **by rank**" — the 69 include strata by country, category and TLD, so rank-stratifiers are an unmeasured subset; (c) "half of the field's samples cannot be redrawn" is contradicted by the page's own line that 164 of the 557 unversioned papers release a dataset | **Accepted, all three.** (a) now says top-//n// "has not been superseded, it has consolidated", with the range taken from complete buckets only; (b) "stratify at all, by rank or by anything else"; (c) "for half of them a reader cannot tell which draw was made", with the artifact escape hatch named in the same bullet. This is the sharpest catch of the review: the page had a stricter standard for other people's claims than for its own lead. | |
| | | **The Alexa box's "66 papers do exactly that"** attaches the unknown-provenance charge to all 66, when 32 of them state a version and several are explicit archives — the defensible case the same paragraph endorses two sentences later | **Accepted.** The charge now attaches to the 34, and the 32 are described as what they are. An internal contradiction within one paragraph, which is the cheapest kind to find and the easiest to write. | |
| | | **Purposive sampling — 319 papers, the second most common design — had one table row and no prose**, leaving the very common "sites with a CMP" study shape unadvised on its two failure modes | **Accepted.** New section //Purposive samples, and the claim they do not license//: state the selection rule reproducibly, inherit the categoriser's error rate as composition error, and claim no prevalence for any larger population. | |
| | | **The frame-row → crawlable-URL step is missing**, and it is the first thing the reader hits after running the page's own script: Tranco rows are registrable domains, CrUX rows are origins, and scheme / ''www'' / redirect-target / duplicates each silently change the unit and manufacture attrition | **Accepted.** New section //A list row is not yet a URL//, plus a note in the script's docstring. A genuine hole that four passes of my own reading did not see. | |
| | | **The published sampler contradicted the page's own versioning advice**, fetching CrUX's unversioned ''current.csv.gz'' — the exact anti-pattern the page's Umbrella row warns about | **Accepted and fixed in code.** It now resolves the newest //dated// monthly snapshot from the mirror's index (''202607'' on 2026-08-12) and records the month as the list identity. Re-run against live CrUX and live Tranco; the Tranco frame hash and sample hash are unchanged, so the published output block is still reproducible. The ''cite'' line's "per rank decade" also became "per rank stratum", since the 1–1000 stratum spans three decades and "decade" misdescribes CrUX buckets entirely. | |
| | | **Two sentences asserted more than their source:** "nearly two-thirds of published claims about websites were really claims about one page per website" (all 119 were; two-thirds //needed revision//), and Scheitle et al.'s 2018 magnitudes presented in flat present tense with an unmeasured "same order as the effect most papers report" | **Accepted, both.** The Aqeel sentence is rewritten; the Scheitle figures are dated, the gap is named as the durable finding rather than the absolute numbers, and the comparison to typical effect sizes now says explicitly that no study compares the two. | |
| | | **"Attrition is not random … all three correlate with what privacy papers measure" was uncited**, where the neighbouring ''design:crawling_location'' footnotes its equivalent claim | **Accepted.** Now hedged to "plausibly" with a footnote naming what //is// evidenced ({[jueckstock2021_realistic]}, {[annamalai2025_beyond]}) and what is not, and pointing at the open question. | |
| | | **Editing residue:** a sentence saying the same thing twice in the prominence section; "In the corpus, 723 papers" using the wrong denominator (it is 1,153, not 5,859); one sentence switching denominators mid-stream; a duplicated open question on this page; and "six sites account for 25% of page loads" quoted on the page but missing from this page's spot-check table | **Accepted, all five.** The last one matters most for this page: a provenance table that claims to list "the quotes that were checked" and silently omits one is the performative-honesty failure such a page exists to avoid. The reviewer verified that quote against the source itself; it is now listed. | |
| | | **Publishing this page makes ''design:website_selection'' visibly wrong** — that page still says Alexa was discontinued "August 1, 2023" and that Tranco combines "Alexa, Cisco Umbrella, and Majestic", both contradicted here with primary sources, and the two pages tell the reader to read them together | **Accepted, and applied** — see §2. The original decision to leave it (a different work item) was a process nicety the reader never sees; what they would have seen is two pages of one wiki disagreeing on checkable facts. Two lines, fixed at publication. | |
| | | Scope split, structure, lead choice, length, tone, and the statistics content staying on the web-measurement side of the textbook line | no change | |
| |
| ===== 11. Run log ===== | ===== 11. Run log ===== |
| | Bibliography | 6 entries added: ''aqeel2020_landing'', ''ahmad2020_apophanies'', ''demir2023_similarity'', ''ukani2025_local'', ''zhu2025_toward'', ''nenadic2026_swiss''. Checked against the live bibliography for key collisions before appending. | | | Bibliography | 6 entries added: ''aqeel2020_landing'', ''ahmad2020_apophanies'', ''demir2023_similarity'', ''ukani2025_local'', ''zhu2025_toward'', ''nenadic2026_swiss''. Checked against the live bibliography for key collisions before appending. | |
| | Mistakes caught in review of my own work | (a) the "impossible Alexa version" finding, retracted after hand-checking — §6; (b) an unverifiable ''prevalence'' figure from Zeber et al., dropped — §6; (c) the published sampler crashed on Tranco's real response format — §8; (d) a first draft asserted "coverage error dominates sampling error" flatly, now a footnoted judgement with its reasoning shown. | | | Mistakes caught in review of my own work | (a) the "impossible Alexa version" finding, retracted after hand-checking — §6; (b) an unverifiable ''prevalence'' figure from Zeber et al., dropped — §6; (c) the published sampler crashed on Tranco's real response format — §8; (d) a first draft asserted "coverage error dominates sampling error" flatly, now a footnoted judgement with its reasoning shown. | |
| | Discussion block | None on this page, following the convention set by the other ''provenance:'' pages — comments belong on the content page. | | | Discussion block | None on this page, following the convention set by the other ''provenance:'' pages — comments belong on the content page. Checked against the published ''provenance:design:crawling_location'', which likewise carries no ''<bibtex bibliography>'' block, so the citekeys here render as markers without a reference list. That is the existing convention, not an omission. | |
| | | Publication order | ''literature:bibliography'' (6 entries) → ''design:sampling'' → ''provenance:design:sampling'', all 2026-08-12. Rendering verified afterwards: 32 inline ''bibtex_citekey'' markers and a ''bibtex_references'' list on the content page, 0 unresolved keys, 3 red links (''statistics:biases'', ''statistics:hypothesis_testing'', ''artifacts'' — all pages ''start'' already promises), and the published ''%%<file>%%'' round-trips byte-identical to ''pages/stratified_sample.py'' apart from a stripped trailing newline. | |
| | | Published before Pass D returned | **Yes, deliberately.** Three focused passes had returned and their findings were applied; the generic pass typically returns prose and structure findings, which are a second revision rather than a blocker. Given the run had already been interrupted several times, shipping a reviewed page and revising it was judged better than risking an unpublished one. Whatever Pass D finds is applied as a follow-up edit and recorded above. A reader comparing revisions should know the first published revision predates one of the four reviews. | |
| |
| [[design:sampling|← back to the content page]] · [[literature:corpus|corpus-level provenance]] | [[design:sampling|← back to the content page]] · [[literature:corpus|corpus-level provenance]] |
| | |
| | ===== Markup sweep, 2026-09-17 ===== |
| | |
| | Mechanical rendering repair only: a fresh live raw/XHTML export of 188 pages was checked with ''check_wrap.mjs'' and ''check_typography.mjs''. Affected plugin tags, CLI flags and heading markup were repaired; no figures or substantive prose were changed. The resulting source and rendered DOM were re-checked after saving. |
| | |
| | ===== Amendment, 2026-09-21: crawl-depth figures scoped against programming:interaction, and the Hispar box refreshed ===== |
| | |
| | **What was wrong.** Nothing on this page was a wrong number. Two things were wrong for a reader who reads two pages. |
| | |
| | * **The depth figures had no scope note.** This page publishes **22.8% of 680** landing-page-only and **15.0%** stating a subpage count; [[programming:interaction]] publishes **37.2% of 417** on the site-depth axis and **12.1% of 857** giving a subpage number. Four figures, two pages, no sentence anywhere telling a reader why they differ. |
| | * **The Hispar box was stale.** Checked 2026-08-12, it said "no maintained equivalent list was found", and the //Open Questions// bullet said the 2020 landing-page result "stands unaddressed, with no tooling to address it". Both had been overtaken by [[programming:interaction]], which documents HTTP Archive's one-secondary-page-per-site crawl (April 2022 onwards, ''is_root_page'' / ''root_page'' columns, ''MAX_DEPTH = 1'' / ''MAX_BREADTH = 1'', the first same-origin link) and measures what that rule misses. |
| | |
| | **The commands and their real output.** |
| | |
| | <code> |
| | $ node scripts/report_sampling.mjs |
| | page population that also crawled: 680 |
| | Interaction depth Papers Share of 680 |
| | single-target-page 222 32.6% |
| | landing-page-only 155 22.8% |
| | landing-plus-subpages 128 18.8% |
| | deep-crawl 92 13.5% |
| | not-stated 69 10.1% |
| | no crawlConfig record 14 2.1% |
| | states how many subpages per site: 102 15.0% |
| | median 10, max 2000 |
| | goes past the landing page: 220 32.4% |
| | papers with unit `websites` that crawled: 492 |
| | landing page only: 122 24.8% |
| | |
| | $ node scripts/report_interaction.mjs |
| | webCrawled (crawled AND platforms includes 'web') 857 <- this page's denominator |
| | denominator: 417 papers (NOT 857; single-target-page and not-stated are excluded) |
| | landing-page-only 155 37.2% |
| | denominator: 857 web crawls; 104 (12.1%) give a number. |
| | of the 417 on the site-depth axis, 102 (24.5%) give a number. |
| | </code> |
| | |
| | **The reconciliation, and why it is a scope note and not a merge.** The two pages report the **same raw count of landing-page-only papers, 155**, and divide it by different denominators: 680 here (every paper with a web-unit population that also crawled, keeping ''single-target-page'', ''not-stated'' and the no-record papers in the denominator) against 417 there (only the web crawls that put a value on the site-depth axis). That is the whole of the 22.8%-against-37.2% gap, and it is now stated on this page in those terms. |
| | |
| | **The matching 155 is a coincidence of count, not a shared set, and the page says so.** Before writing the sentence, the two rosters were compared directly, because "the same 155 papers" would have been the natural and wrong thing to write: |
| | |
| | <code> |
| | interaction landing-only: 155 sampling landing-only: 155 |
| | intersection: 149 |
| | in interaction not sampling: 6 in sampling not interaction: 6 |
| | </code> |
| | |
| | The six each way are the two population definitions, not an error: this page needs a ''population[].unit'' in {websites, domains, web-pages}, that page needs ''platforms'' to include ''web''. Traffic-analysis and censorship papers sit in one and not the other. **No overlap figure was published** — it is not produced by either page's report script, and publishing it would put a number on a page that no script regenerates. It is recorded here instead, which is what this page is for, so that a later run does not "reconcile" two counts that only look identical. |
| | |
| | **What changed on the page.** |
| | |
| | - The //One page per site is the norm// summary gained a paragraph stating the 680-against-417 scope difference explicitly, naming the 222 + 69 + 14 rows that this page keeps and that page excludes, and warning that the 155-paper sets are not identical. |
| | - The **15.0%** sentence gained its numerator (102 of 680) and a pointer to the 12.1% (104 of 857) with its own denominator named. |
| | - The Hispar ''%%<WRAP todo>%%'' box was rewritten: Hispar is still gone (NXDOMAIN, 2026-08-12, unchanged), but the box now records HTTP Archive as the one maintained source of internal pages, states the selection rule that makes it a by-product rather than a sample, and points at the measured demonstration on [[programming:interaction]]. The claim's footnotes and their check dates live on that page and were **not** copied here — one number, one page. |
| | - The //Open Questions// bullet was rewritten the same way: the gap is now "no maintained internal-page //list//", not "no tooling". |
| | |
| | **Rendered-DOM check.** Page body after the table of contents: list items 40, headings 33, tables 13, ''%%<pre>%%'' blocks 2 — **all four identical before and after**, so nothing was swallowed by the rewritten ''%%<WRAP>%%'' box or by the wrapped //Open Questions// bullet. **0** ''wikilink2'' red links. All three cross-page anchors resolve against real heading ids on ''programming:interaction'': ''depth_has_more_than_three_positions'', ''how_deep_the_field_actually_goes'', ''the_one_page_list_that_included_internal_pages_is_gone''. |
| | |
| | **Finding rejected.** The first draft of the scope note compared this page's **24.8% of 492** against the 37.2%, because that is the pairing the task named. It was dropped: 24.8% is the ''unit == websites'' cut (122 of 492) and shares neither numerator nor denominator with anything on the other page, so pairing them would have invented a comparison. The comparable pair is 22.8% of 680 against 37.2% of 417, which share the numerator 155. The 24.8% sentence is left exactly as it was. |
| | |
| | **Reviewers.** One ''sonnet'' figures-against-script pass over both edits and the sections either side of them. No citations pass: no ''{[citekey]}'' and no quoted claim was touched — ''{[aqeel2020_landing]}'' and the Aqeel quote in the same subsection are unchanged. |
| |