User Tools

Site Tools


provenance:design:sampling

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
provenance:design:sampling [2026/08/12 21:29] – Record publication order, rendering verification, the no-bibliography-block convention check, and that the page shipped before the generic review pass returned. Authored by Claude. karel.kubicek.claudeprovenance:design:sampling [2026/08/12 21:40] (current) – Record the generic review pass (10 findings, all accepted), the two corrections applied to design:website_selection, and the missing spot-check row. Authored by Claude. karel.kubicek.claude
Line 32: Line 32:
   * ''sampling'' does **not** catalogue services. Where a list needs describing, it links.   * ''sampling'' does **not** catalogue services. Where a list needs describing, it links.
  
-**Two corrections owed to design:website_selection** were found and are //not// applied herebecause that page is a different work item:+**Two corrections were owed to design:website_selectionand were applied** (rev 1786570786, 2026-08-12). The original decision was to leave themon the grounds that the neighbouring page is a different work item; the generic review pass argued that this is a process nicety the reader never sees, and that what they //would// see is two pages of one wiki disagreeing on checkable facts while telling them to read both. That was right, and the fixes are two lines:
  
-  - It says Alexa was "**Discontinued as of August 1, 2023**". The primary source says Alexa.com was retired **1 May 2022** and the Top Sites / Web Information Service APIs on **15 December 2022** (see §7). No source was found for an August 2023 date. +  - It said Alexa was "**Discontinued as of August 1, 2023**". The primary source says Alexa.com was retired **1 May 2022** and the Top Sites / Web Information Service APIs on **15 December 2022** (see §7). No source was found for an August 2023 date, and the external-currency pass independently failed to find one — including in AWS's own full-shutdown listing, which does not mention Alexa at all
-  - It says Tranco "combines data from Alexa, Cisco Umbrella, and Majestic over a 30-day period". Tranco's own API reports the providers of the 2026-08-01 daily list as ''crux, farsight, majestic, radar, umbrella'' — **Alexa is not among them** and has not been for years. Tranco's own prose methodology page is also stale here, saying "all four providers" for a five-provider list.+  - It said Tranco "combines data from Alexa, Cisco Umbrella, and Majestic over a 30-day period". Tranco's own API reports the providers of the 2026-08-01 daily list as ''crux, farsight, majestic, radar, umbrella'' — **Alexa is not among them**. The sentence now dates the 2019 composition as historical and gives the live one, with the API call in a footnote. Tranco's own prose methodology page is stale here too, saying "all four providers" for a five-provider list
 + 
 +Nothing else on that page was touched: it keeps its structure, its TODOs and its own //Use in Publications// section.
  
 ===== 3. Populations and denominators ===== ===== 3. Populations and denominators =====
Line 163: Line 165:
 | "median number of tracking domains … is 1.9, whereas for the crawler it is 6.1"; "the crawler may reach 26" | {[zeber2020representativeness]} | exact | | "median number of tracking domains … is 1.9, whereas for the crawler it is 6.1"; "the crawler may reach 26" | {[zeber2020representativeness]} | exact |
 | "the top million sites capture over 95% of all page loads and time spent online" | {[ruth2022_world]} | exact | | "the top million sites capture over 95% of all page loads and time spent online" | {[ruth2022_world]} | exact |
-| "one site garners 17% of all desktop page loads globally"; "ten sites accounting for about half of time spent" | {[ruth2022_world]} | exact |+| "one site garners 17% of all desktop page loads globally"; "ten sites accounting for about half of time spent"; "six sites account for 25% of page loads on both desktop and mobile" | {[ruth2022_world]} | exact — the third was added to this table after a reviewer noticed it was quoted on the page but not listed here |
 | "disproportionate focus on the long tail of the web" | {[ruth2022_world]} | present, **column-spliced** — the sentence interleaves with the adjacent column, so it is quoted in fragments on the page rather than as one run | | "disproportionate focus on the long tail of the web" | {[ruth2022_world]} | present, **column-spliced** — the sentence interleaves with the adjacent column, so it is quoted in fragments on the page rather than as one run |
 | "we present statistics for 917,261 sites" | {[englehardt2016online]} | exact | | "we present statistics for 917,261 sites" | {[englehardt2016online]} | exact |
Line 215: Line 217:
  
   * **Attrition rates across the field.** The extraction schema has ''population.n'' but no "successfully measured" field, so the corpus cannot say how often a paper reports both the drawn and the analysed denominator. The page argues the point from two verified examples and lists it as an open question rather than quantifying it. Closing it needs a targeted full-text study.   * **Attrition rates across the field.** The extraction schema has ''population.n'' but no "successfully measured" field, so the corpus cannot say how often a paper reports both the drawn and the analysed denominator. The page argues the point from two verified examples and lists it as an open question rather than quantifying it. Closing it needs a targeted full-text study.
-  * **Whether any paper reports the same measurement per rank stratum.** Searched the ''detection[].prevalence'' field with rank- and popularity-related patterns and read about 30 matches; none is a per-stratum breakdown of the paper's own headline measurement. Absence of evidence in a free-text field is weak evidence, so the page phrases this as "we found none" and asks for counter-examples.+  * **Whether any paper reports the same measurement per rank stratum.** Searched the ''detection[].prevalence'' field with rank- and popularity-related patterns and read about 30 matches; a narrower re-run over just the 69 stratified papers returned two matches, neither of them a per-stratum breakdown of the paper's own headline figure. Absence of evidence in a free-text field is weak evidence, so the page phrases this as "we found none" and asks for counter-examples.
   * **Whether rank weighting is used outside third-party analysis.** The first draft of the page asserted in an open question that no weighted estimator appears in this corpus. That was **wrong and was corrected before publication**: {[englehardt2016online]}'s //prominence// (Σ 1/rank over the sites a third party appears on) is exactly a rank-weighted estimator, introduced for exactly the sampling reason the page's size section argues. It was found by reading the paper rather than by any query — the extraction records it as a metric, not as a sampling decision, so no ''population'' query would have surfaced it. The page now has a section on it, and the open question was narrowed to what remains unfound: a weighted //prevalence// reported alongside the unweighted one, and reweighting a rank-stratified sample by stratum size. **Take the narrowed question as weakly evidenced too** — it rests on the same free-text search.   * **Whether rank weighting is used outside third-party analysis.** The first draft of the page asserted in an open question that no weighted estimator appears in this corpus. That was **wrong and was corrected before publication**: {[englehardt2016online]}'s //prominence// (Σ 1/rank over the sites a third party appears on) is exactly a rank-weighted estimator, introduced for exactly the sampling reason the page's size section argues. It was found by reading the paper rather than by any query — the extraction records it as a metric, not as a sampling decision, so no ''population'' query would have surfaced it. The page now has a section on it, and the open question was narrowed to what remains unfound: a weighted //prevalence// reported alongside the unweighted one, and reweighting a rank-stratified sample by stratum size. **Take the narrowed question as weakly evidenced too** — it rests on the same free-text search.
  
-  * **Whether any paper reports the same measurement per rank stratum.** Re-run over just the 69 stratified papers' ''detection[].prevalence'' fields: two matches, neither of them a per-stratum breakdown of the paper's own headline figure. 
   * **How much of the 2025–2026 top-//n// rise is real.** Those venue-years are provisional by construction. 64.7% rests on 190 papers from years that are incompletely indexed; the direction is consistent with 2018–2024 and the page labels the column, but it should not be quoted as a 2026 measurement.   * **How much of the 2025–2026 top-//n// rise is real.** Those venue-years are provisional by construction. 64.7% rests on 190 papers from years that are incompletely indexed; the direction is consistent with 2018–2024 and the page labels the column, but it should not be quoted as a 2026 measurement.
   * **Whether the 66 post-retirement Alexa papers are reusing archives or copying a citation.** Reading a dozen suggests both happen and that side-channel and website-fingerprinting evaluations inherit "top 100 Alexa" as a benchmark convention from earlier papers. That impression is not quantified and is not on the content page as a figure.   * **Whether the 66 post-retirement Alexa papers are reusing archives or copying a citation.** Reading a dozen suggests both happen and that side-channel and website-fingerprinting evaluations inherit "top 100 Alexa" as a benchmark convention from earlier papers. That impression is not quantified and is not on the content page as a figure.
Line 266: Line 267:
 Run last, against a **second, honest freeze** (''out/freeze2/'', page MD5 ''7c7b5793…''), with no checklist: overstated claims, structure, scope boundary, internal contradictions, tone, and whether the page answers its own question. It also reviews this page. Run last, against a **second, honest freeze** (''out/freeze2/'', page MD5 ''7c7b5793…''), with no checklist: overstated claims, structure, scope boundary, internal contradictions, tone, and whether the page answers its own question. It also reviews this page.
  
-<wrap todo>Findings and disposition to be appended when the pass returns. **The page was published before it did** — see the run log for why, and for what changed afterwards.</wrap>+It returned the largest and most useful set of findings of the four, and **the page was published before it did** (run log below). Every item was accepted; the page was revised and re-saved the same day. 
 + 
 +^ Finding ^ Disposition ^ 
 +| **The lead overstated three times in twelve lines**, on a page whose thesis is "say only what your evidence carries": (a) "the share is still rising" rests entirely on the provisional 2025–26 bucket, and read to 2024 as the page's own methodology section instructs, top-//n// //fell// 59.6% → 58.0%; (b) "6.0% stratify **by rank**" — the 69 include strata by country, category and TLD, so rank-stratifiers are an unmeasured subset; (c) "half of the field's samples cannot be redrawn" is contradicted by the page's own line that 164 of the 557 unversioned papers release a dataset | **Accepted, all three.** (a) now says top-//n// "has not been superseded, it has consolidated", with the range taken from complete buckets only; (b) "stratify at all, by rank or by anything else"; (c) "for half of them a reader cannot tell which draw was made", with the artifact escape hatch named in the same bullet. This is the sharpest catch of the review: the page had a stricter standard for other people's claims than for its own lead. | 
 +| **The Alexa box's "66 papers do exactly that"** attaches the unknown-provenance charge to all 66, when 32 of them state a version and several are explicit archives — the defensible case the same paragraph endorses two sentences later | **Accepted.** The charge now attaches to the 34, and the 32 are described as what they are. An internal contradiction within one paragraph, which is the cheapest kind to find and the easiest to write. | 
 +| **Purposive sampling — 319 papers, the second most common design — had one table row and no prose**, leaving the very common "sites with a CMP" study shape unadvised on its two failure modes | **Accepted.** New section //Purposive samples, and the claim they do not license//: state the selection rule reproducibly, inherit the categoriser's error rate as composition error, and claim no prevalence for any larger population. | 
 +| **The frame-row → crawlable-URL step is missing**, and it is the first thing the reader hits after running the page's own script: Tranco rows are registrable domains, CrUX rows are origins, and scheme / ''www'' / redirect-target / duplicates each silently change the unit and manufacture attrition | **Accepted.** New section //A list row is not yet a URL//, plus a note in the script's docstring. A genuine hole that four passes of my own reading did not see. | 
 +| **The published sampler contradicted the page's own versioning advice**, fetching CrUX's unversioned ''current.csv.gz'' — the exact anti-pattern the page's Umbrella row warns about | **Accepted and fixed in code.** It now resolves the newest //dated// monthly snapshot from the mirror's index (''202607'' on 2026-08-12) and records the month as the list identity. Re-run against live CrUX and live Tranco; the Tranco frame hash and sample hash are unchanged, so the published output block is still reproducible. The ''cite'' line's "per rank decade" also became "per rank stratum", since the 1–1000 stratum spans three decades and "decade" misdescribes CrUX buckets entirely. | 
 +| **Two sentences asserted more than their source:** "nearly two-thirds of published claims about websites were really claims about one page per website" (all 119 were; two-thirds //needed revision//), and Scheitle et al.'s 2018 magnitudes presented in flat present tense with an unmeasured "same order as the effect most papers report" | **Accepted, both.** The Aqeel sentence is rewritten; the Scheitle figures are dated, the gap is named as the durable finding rather than the absolute numbers, and the comparison to typical effect sizes now says explicitly that no study compares the two. | 
 +| **"Attrition is not random … all three correlate with what privacy papers measure" was uncited**, where the neighbouring ''design:crawling_location'' footnotes its equivalent claim | **Accepted.** Now hedged to "plausibly" with a footnote naming what //is// evidenced ({[jueckstock2021_realistic]}, {[annamalai2025_beyond]}) and what is not, and pointing at the open question. | 
 +| **Editing residue:** a sentence saying the same thing twice in the prominence section; "In the corpus, 723 papers" using the wrong denominator (it is 1,153, not 5,859); one sentence switching denominators mid-stream; a duplicated open question on this page; and "six sites account for 25% of page loads" quoted on the page but missing from this page's spot-check table | **Accepted, all five.** The last one matters most for this page: a provenance table that claims to list "the quotes that were checked" and silently omits one is the performative-honesty failure such a page exists to avoid. The reviewer verified that quote against the source itself; it is now listed. | 
 +| **Publishing this page makes ''design:website_selection'' visibly wrong** — that page still says Alexa was discontinued "August 1, 2023" and that Tranco combines "Alexa, Cisco Umbrella, and Majestic", both contradicted here with primary sources, and the two pages tell the reader to read them together | **Accepted, and applied** — see §2. The original decision to leave it (a different work item) was a process nicety the reader never sees; what they would have seen is two pages of one wiki disagreeing on checkable facts. Two lines, fixed at publication. | 
 +| Scope split, structure, lead choice, length, tone, and the statistics content staying on the web-measurement side of the textbook line | no change |
  
 ===== 11. Run log ===== ===== 11. Run log =====
provenance/design/sampling.1786570170.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki