| Next revision | Previous revision |
| provenance:design:sampling [2026/08/12 21:28] – Provenance for design:sampling: populations and denominators, the frame fold and its residue, quotes spot-checked, two retracted figures, external verification log, and the review passes. Authored by Claude. karel.kubicek.claude | provenance:design:sampling [2026/08/12 21:40] (current) – Record the generic review pass (10 findings, all accepted), the two corrections applied to design:website_selection, and the missing spot-check row. Authored by Claude. karel.kubicek.claude |
|---|
| * ''sampling'' does **not** catalogue services. Where a list needs describing, it links. | * ''sampling'' does **not** catalogue services. Where a list needs describing, it links. |
| |
| **Two corrections owed to design:website_selection** were found and are //not// applied here, because that page is a different work item: | **Two corrections were owed to design:website_selection, and were applied** (rev 1786570786, 2026-08-12). The original decision was to leave them, on the grounds that the neighbouring page is a different work item; the generic review pass argued that this is a process nicety the reader never sees, and that what they //would// see is two pages of one wiki disagreeing on checkable facts while telling them to read both. That was right, and the fixes are two lines: |
| |
| - It says Alexa was "**Discontinued as of August 1, 2023**". The primary source says Alexa.com was retired **1 May 2022** and the Top Sites / Web Information Service APIs on **15 December 2022** (see §7). No source was found for an August 2023 date. | - It said Alexa was "**Discontinued as of August 1, 2023**". The primary source says Alexa.com was retired **1 May 2022** and the Top Sites / Web Information Service APIs on **15 December 2022** (see §7). No source was found for an August 2023 date, and the external-currency pass independently failed to find one — including in AWS's own full-shutdown listing, which does not mention Alexa at all. |
| - It says Tranco "combines data from Alexa, Cisco Umbrella, and Majestic over a 30-day period". Tranco's own API reports the providers of the 2026-08-01 daily list as ''crux, farsight, majestic, radar, umbrella'' — **Alexa is not among them** and has not been for years. Tranco's own prose methodology page is also stale here, saying "all four providers" for a five-provider list. | - It said Tranco "combines data from Alexa, Cisco Umbrella, and Majestic over a 30-day period". Tranco's own API reports the providers of the 2026-08-01 daily list as ''crux, farsight, majestic, radar, umbrella'' — **Alexa is not among them**. The sentence now dates the 2019 composition as historical and gives the live one, with the API call in a footnote. Tranco's own prose methodology page is stale here too, saying "all four providers" for a five-provider list. |
| | |
| | Nothing else on that page was touched: it keeps its structure, its TODOs and its own //Use in Publications// section. |
| |
| ===== 3. Populations and denominators ===== | ===== 3. Populations and denominators ===== |
| | "median number of tracking domains … is 1.9, whereas for the crawler it is 6.1"; "the crawler may reach 26" | {[zeber2020representativeness]} | exact | | | "median number of tracking domains … is 1.9, whereas for the crawler it is 6.1"; "the crawler may reach 26" | {[zeber2020representativeness]} | exact | |
| | "the top million sites capture over 95% of all page loads and time spent online" | {[ruth2022_world]} | exact | | | "the top million sites capture over 95% of all page loads and time spent online" | {[ruth2022_world]} | exact | |
| | "one site garners 17% of all desktop page loads globally"; "ten sites accounting for about half of time spent" | {[ruth2022_world]} | exact | | | "one site garners 17% of all desktop page loads globally"; "ten sites accounting for about half of time spent"; "six sites account for 25% of page loads on both desktop and mobile" | {[ruth2022_world]} | exact — the third was added to this table after a reviewer noticed it was quoted on the page but not listed here | |
| | "disproportionate focus on the long tail of the web" | {[ruth2022_world]} | present, **column-spliced** — the sentence interleaves with the adjacent column, so it is quoted in fragments on the page rather than as one run | | | "disproportionate focus on the long tail of the web" | {[ruth2022_world]} | present, **column-spliced** — the sentence interleaves with the adjacent column, so it is quoted in fragments on the page rather than as one run | |
| | "we present statistics for 917,261 sites" | {[englehardt2016online]} | exact | | | "we present statistics for 917,261 sites" | {[englehardt2016online]} | exact | |
| |
| * **Attrition rates across the field.** The extraction schema has ''population.n'' but no "successfully measured" field, so the corpus cannot say how often a paper reports both the drawn and the analysed denominator. The page argues the point from two verified examples and lists it as an open question rather than quantifying it. Closing it needs a targeted full-text study. | * **Attrition rates across the field.** The extraction schema has ''population.n'' but no "successfully measured" field, so the corpus cannot say how often a paper reports both the drawn and the analysed denominator. The page argues the point from two verified examples and lists it as an open question rather than quantifying it. Closing it needs a targeted full-text study. |
| * **Whether any paper reports the same measurement per rank stratum.** Searched the ''detection[].prevalence'' field with rank- and popularity-related patterns and read about 30 matches; none is a per-stratum breakdown of the paper's own headline measurement. Absence of evidence in a free-text field is weak evidence, so the page phrases this as "we found none" and asks for counter-examples. | * **Whether any paper reports the same measurement per rank stratum.** Searched the ''detection[].prevalence'' field with rank- and popularity-related patterns and read about 30 matches; a narrower re-run over just the 69 stratified papers returned two matches, neither of them a per-stratum breakdown of the paper's own headline figure. Absence of evidence in a free-text field is weak evidence, so the page phrases this as "we found none" and asks for counter-examples. |
| * **Whether rank weighting is used outside third-party analysis.** The first draft of the page asserted in an open question that no weighted estimator appears in this corpus. That was **wrong and was corrected before publication**: {[englehardt2016online]}'s //prominence// (Σ 1/rank over the sites a third party appears on) is exactly a rank-weighted estimator, introduced for exactly the sampling reason the page's size section argues. It was found by reading the paper rather than by any query — the extraction records it as a metric, not as a sampling decision, so no ''population'' query would have surfaced it. The page now has a section on it, and the open question was narrowed to what remains unfound: a weighted //prevalence// reported alongside the unweighted one, and reweighting a rank-stratified sample by stratum size. **Take the narrowed question as weakly evidenced too** — it rests on the same free-text search. | * **Whether rank weighting is used outside third-party analysis.** The first draft of the page asserted in an open question that no weighted estimator appears in this corpus. That was **wrong and was corrected before publication**: {[englehardt2016online]}'s //prominence// (Σ 1/rank over the sites a third party appears on) is exactly a rank-weighted estimator, introduced for exactly the sampling reason the page's size section argues. It was found by reading the paper rather than by any query — the extraction records it as a metric, not as a sampling decision, so no ''population'' query would have surfaced it. The page now has a section on it, and the open question was narrowed to what remains unfound: a weighted //prevalence// reported alongside the unweighted one, and reweighting a rank-stratified sample by stratum size. **Take the narrowed question as weakly evidenced too** — it rests on the same free-text search. |
| |
| * **Whether any paper reports the same measurement per rank stratum.** Re-run over just the 69 stratified papers' ''detection[].prevalence'' fields: two matches, neither of them a per-stratum breakdown of the paper's own headline figure. | |
| * **How much of the 2025–2026 top-//n// rise is real.** Those venue-years are provisional by construction. 64.7% rests on 190 papers from years that are incompletely indexed; the direction is consistent with 2018–2024 and the page labels the column, but it should not be quoted as a 2026 measurement. | * **How much of the 2025–2026 top-//n// rise is real.** Those venue-years are provisional by construction. 64.7% rests on 190 papers from years that are incompletely indexed; the direction is consistent with 2018–2024 and the page labels the column, but it should not be quoted as a 2026 measurement. |
| * **Whether the 66 post-retirement Alexa papers are reusing archives or copying a citation.** Reading a dozen suggests both happen and that side-channel and website-fingerprinting evaluations inherit "top 100 Alexa" as a benchmark convention from earlier papers. That impression is not quantified and is not on the content page as a figure. | * **Whether the 66 post-retirement Alexa papers are reusing archives or copying a citation.** Reading a dozen suggests both happen and that side-channel and website-fingerprinting evaluations inherit "top 100 Alexa" as a benchmark convention from earlier papers. That impression is not quantified and is not on the content page as a figure. |
| |
| The cost was real even so: Pass C spent effort reporting a figure that was already fixed, and Pass B reviewed a sentence that had been rewritten once since the snapshot. **Next time: freeze, queue the fixes, and apply them in one batch after the passes return.** Recorded here rather than quietly omitted, because a review log that hides how the review actually ran is worth less than no log. | The cost was real even so: Pass C spent effort reporting a figure that was already fixed, and Pass B reviewed a sentence that had been rewritten once since the snapshot. **Next time: freeze, queue the fixes, and apply them in one batch after the passes return.** Recorded here rather than quietly omitted, because a review log that hides how the review actually ran is worth less than no log. |
| | |
| | ==== Pass D — generic (Claude Fable) ==== |
| | |
| | Run last, against a **second, honest freeze** (''out/freeze2/'', page MD5 ''7c7b5793…''), with no checklist: overstated claims, structure, scope boundary, internal contradictions, tone, and whether the page answers its own question. It also reviews this page. |
| | |
| | It returned the largest and most useful set of findings of the four, and **the page was published before it did** (run log below). Every item was accepted; the page was revised and re-saved the same day. |
| | |
| | ^ Finding ^ Disposition ^ |
| | | **The lead overstated three times in twelve lines**, on a page whose thesis is "say only what your evidence carries": (a) "the share is still rising" rests entirely on the provisional 2025–26 bucket, and read to 2024 as the page's own methodology section instructs, top-//n// //fell// 59.6% → 58.0%; (b) "6.0% stratify **by rank**" — the 69 include strata by country, category and TLD, so rank-stratifiers are an unmeasured subset; (c) "half of the field's samples cannot be redrawn" is contradicted by the page's own line that 164 of the 557 unversioned papers release a dataset | **Accepted, all three.** (a) now says top-//n// "has not been superseded, it has consolidated", with the range taken from complete buckets only; (b) "stratify at all, by rank or by anything else"; (c) "for half of them a reader cannot tell which draw was made", with the artifact escape hatch named in the same bullet. This is the sharpest catch of the review: the page had a stricter standard for other people's claims than for its own lead. | |
| | | **The Alexa box's "66 papers do exactly that"** attaches the unknown-provenance charge to all 66, when 32 of them state a version and several are explicit archives — the defensible case the same paragraph endorses two sentences later | **Accepted.** The charge now attaches to the 34, and the 32 are described as what they are. An internal contradiction within one paragraph, which is the cheapest kind to find and the easiest to write. | |
| | | **Purposive sampling — 319 papers, the second most common design — had one table row and no prose**, leaving the very common "sites with a CMP" study shape unadvised on its two failure modes | **Accepted.** New section //Purposive samples, and the claim they do not license//: state the selection rule reproducibly, inherit the categoriser's error rate as composition error, and claim no prevalence for any larger population. | |
| | | **The frame-row → crawlable-URL step is missing**, and it is the first thing the reader hits after running the page's own script: Tranco rows are registrable domains, CrUX rows are origins, and scheme / ''www'' / redirect-target / duplicates each silently change the unit and manufacture attrition | **Accepted.** New section //A list row is not yet a URL//, plus a note in the script's docstring. A genuine hole that four passes of my own reading did not see. | |
| | | **The published sampler contradicted the page's own versioning advice**, fetching CrUX's unversioned ''current.csv.gz'' — the exact anti-pattern the page's Umbrella row warns about | **Accepted and fixed in code.** It now resolves the newest //dated// monthly snapshot from the mirror's index (''202607'' on 2026-08-12) and records the month as the list identity. Re-run against live CrUX and live Tranco; the Tranco frame hash and sample hash are unchanged, so the published output block is still reproducible. The ''cite'' line's "per rank decade" also became "per rank stratum", since the 1–1000 stratum spans three decades and "decade" misdescribes CrUX buckets entirely. | |
| | | **Two sentences asserted more than their source:** "nearly two-thirds of published claims about websites were really claims about one page per website" (all 119 were; two-thirds //needed revision//), and Scheitle et al.'s 2018 magnitudes presented in flat present tense with an unmeasured "same order as the effect most papers report" | **Accepted, both.** The Aqeel sentence is rewritten; the Scheitle figures are dated, the gap is named as the durable finding rather than the absolute numbers, and the comparison to typical effect sizes now says explicitly that no study compares the two. | |
| | | **"Attrition is not random … all three correlate with what privacy papers measure" was uncited**, where the neighbouring ''design:crawling_location'' footnotes its equivalent claim | **Accepted.** Now hedged to "plausibly" with a footnote naming what //is// evidenced ({[jueckstock2021_realistic]}, {[annamalai2025_beyond]}) and what is not, and pointing at the open question. | |
| | | **Editing residue:** a sentence saying the same thing twice in the prominence section; "In the corpus, 723 papers" using the wrong denominator (it is 1,153, not 5,859); one sentence switching denominators mid-stream; a duplicated open question on this page; and "six sites account for 25% of page loads" quoted on the page but missing from this page's spot-check table | **Accepted, all five.** The last one matters most for this page: a provenance table that claims to list "the quotes that were checked" and silently omits one is the performative-honesty failure such a page exists to avoid. The reviewer verified that quote against the source itself; it is now listed. | |
| | | **Publishing this page makes ''design:website_selection'' visibly wrong** — that page still says Alexa was discontinued "August 1, 2023" and that Tranco combines "Alexa, Cisco Umbrella, and Majestic", both contradicted here with primary sources, and the two pages tell the reader to read them together | **Accepted, and applied** — see §2. The original decision to leave it (a different work item) was a process nicety the reader never sees; what they would have seen is two pages of one wiki disagreeing on checkable facts. Two lines, fixed at publication. | |
| | | Scope split, structure, lead choice, length, tone, and the statistics content staying on the web-measurement side of the textbook line | no change | |
| |
| ===== 11. Run log ===== | ===== 11. Run log ===== |
| | Bibliography | 6 entries added: ''aqeel2020_landing'', ''ahmad2020_apophanies'', ''demir2023_similarity'', ''ukani2025_local'', ''zhu2025_toward'', ''nenadic2026_swiss''. Checked against the live bibliography for key collisions before appending. | | | Bibliography | 6 entries added: ''aqeel2020_landing'', ''ahmad2020_apophanies'', ''demir2023_similarity'', ''ukani2025_local'', ''zhu2025_toward'', ''nenadic2026_swiss''. Checked against the live bibliography for key collisions before appending. | |
| | Mistakes caught in review of my own work | (a) the "impossible Alexa version" finding, retracted after hand-checking — §6; (b) an unverifiable ''prevalence'' figure from Zeber et al., dropped — §6; (c) the published sampler crashed on Tranco's real response format — §8; (d) a first draft asserted "coverage error dominates sampling error" flatly, now a footnoted judgement with its reasoning shown. | | | Mistakes caught in review of my own work | (a) the "impossible Alexa version" finding, retracted after hand-checking — §6; (b) an unverifiable ''prevalence'' figure from Zeber et al., dropped — §6; (c) the published sampler crashed on Tranco's real response format — §8; (d) a first draft asserted "coverage error dominates sampling error" flatly, now a footnoted judgement with its reasoning shown. | |
| | Discussion block | None on this page, following the convention set by the other ''provenance:'' pages — comments belong on the content page. | | | Discussion block | None on this page, following the convention set by the other ''provenance:'' pages — comments belong on the content page. Checked against the published ''provenance:design:crawling_location'', which likewise carries no ''<bibtex bibliography>'' block, so the citekeys here render as markers without a reference list. That is the existing convention, not an omission. | |
| | | Publication order | ''literature:bibliography'' (6 entries) → ''design:sampling'' → ''provenance:design:sampling'', all 2026-08-12. Rendering verified afterwards: 32 inline ''bibtex_citekey'' markers and a ''bibtex_references'' list on the content page, 0 unresolved keys, 3 red links (''statistics:biases'', ''statistics:hypothesis_testing'', ''artifacts'' — all pages ''start'' already promises), and the published ''<file>'' round-trips byte-identical to ''pages/stratified_sample.py'' apart from a stripped trailing newline. | |
| | | Published before Pass D returned | **Yes, deliberately.** Three focused passes had returned and their findings were applied; the generic pass typically returns prose and structure findings, which are a second revision rather than a blocker. Given the run had already been interrupted several times, shipping a reviewed page and revising it was judged better than risking an unpublished one. Whatever Pass D finds is applied as a follow-up edit and recorded above. A reader comparing revisions should know the first published revision predates one of the four reviews. | |
| |
| [[design:sampling|← back to the content page]] · [[literature:corpus|corpus-level provenance]] | [[design:sampling|← back to the content page]] · [[literature:corpus|corpus-level provenance]] |
| |