Table of Contents
Provenance: design:website_selection
Working notes behind website_selection — every query with its population and denominator, the report script and its unedited output, the vendor fold and its residue, the Farsight hand map, the quotes that were checked, the external sources that were verified or rejected, and what could not be established. Corpus-level caveats are on corpus and are not restated here.
Contemporaneous. Written during the run that produced the refresh, 2026-08-27, not reconstructed afterwards. Cursor drain sitting (item 184, run 64), not Claude Code.
1. What this page is backing
| Item | Value |
|---|---|
| Content page | website_selection — extending, not creating. Live rev at start of sitting: 1787832045, 10,124 bytes. One remaining <WRAP todo> (Farsight). Radar TODOs had already been filled by the Cloudflare Radar sitting. |
| Report script | scripts/report_website_selection.mjs |
| Fold it depends on | scripts/rank_fold.mjs (multi-label ranking vendors; not sample_fold.mjs, which folds frame kind) |
| External re-fetch | scripts/external_checks_website_selection.sh (exit 0, 2026-08-27T13:58:02Z) |
| Data | data/extract/run1/extractions.jsonl, 5,859 papers, 7 venues, 2010–2026 |
| Bibliography additions | none — every citekey was already in bibliography |
| Neighbour correction | website_classification Alexa footnote no longer says this page gives 1 August 2023 as Alexa's death. CrUX licence one-liners on that page and tranco now match Google CC BY 4.0 vs Tranco's still-stated CC BY-SA 4.0. |
2. Scope: extending, not a new page
The page already existed as a catalogue of ranking services with two borrowed survey figures. The item asked to close remaining TODOs and replace those figures with a corpus “Use in Publications”. sampling already covers how to draw, and already publishes the Alexa → Tranco year table. This sitting:
- filled Farsight (the last TODO)
- added the vendor-fold Use in Publications, keeping Scheitle/Xie as dated historical surveys
- fixed the leftover “discontinued in 2023” sentence (the list bullet already said 1 May 2022; the survey paragraph did not)
- added a methodology section and this provenance page
- did not absorb sampling's method/size/versioning tables
Judgement call: do not create programming:farsight. The ranking is not independently downloadable; the API detail is the Tranco default-list-only trap, already on tranco.
3. Populations and denominators
| Tag | Definition | N |
|---|---|---|
| corpus | all extraction records | 5,859 |
ALL / sampled | population.length > 0 | 5,712 |
| WEB (page population) | ≥1 population[] tuple with unit ∈ {websites, domains, web-pages} | 1,153 |
| names a source | WEB, ≥1 non-sentinel sourceList | 1,143 |
| vendor-users | WEB, ≥1 web-unit sourceList matching a vendor family | 763 |
sample_fold popularity-ranking | WEB, frame-kind fold | 764 |
sample_fold custom-seed | WEB, frame-kind fold | 257 |
| Farsight/DNSDB as pdns dataset | any unit, hand map | 13 papers / 12 strings |
| Farsight as a ranking | vendor family farsight-ranking | 0 |
The item brief said “population.sourceList over 4,207 sampling papers (custom seed list 377, Google Play 183, Tranco 119, Alexa 117)”. Those counts are exact strings on the old 4,322-paper corpus. On this run the same method over ALL yields custom seed list 516, Tranco 180, Google Play 143, Alexa 51. Google Play is 281 after folding and 0 on WEB. The page publishes that trap rather than those four numbers as a ranking ranking.
Sentinels are silence. Papers, not tuples. A string may hit several vendors.
4. Running it
cd /workspace/artifacts/wiki node scripts/rank_fold.mjs # self-test; prints Farsight pdns strings and unnamed residue; exit 1 on unknown Farsight spelling node scripts/report_website_selection.mjs # every figure node scripts/report_website_selection.mjs --wiki node scripts/report_website_selection.mjs --list node scripts/report_website_selection.mjs --quotes 'tranco' bash scripts/external_checks_website_selection.sh python3 pages/compare_ranks.py --wiki # live Tranco/Umbrella/Majestic ranks; dated snapshot node scripts/check_page_numbers.mjs pages/design_website_selection.txt out/report_website_selection.txt \ '===== Use in Publications =====' '===== What to report =====' node scripts/check_page_numbers.mjs pages/design_website_selection.txt out/report_website_selection.txt node scripts/check_tables.mjs pages/design_website_selection.txt node scripts/check_wrap.mjs pages/design_website_selection.txt
Windowed and whole-page number checks passed after the report grew a Y block (sample_fold 764) and Z lines for CC BY 4.0 / Tranco's still-stated CC BY-SA 4.0, rank 5,000, and 1.1.1.1 — those are on the catalogue half of the page, outside the corpus section. The page prints Google's CC BY 4.0 as the licence to use when citing CrUX.
5. The vendor fold, Farsight hand map, residue
Multi-label on purpose: “Tranco and Cloudflare Radar” is both. sample_fold.mjs is exclusive first-match and answers a different question; do not reuse it as a vendor ranking.
Exact-string undercount on WEB (from the report):
| Vendor | Papers | Exact name | Spellings | Undercount |
|---|---|---|---|---|
| Alexa | 463 | 50 | 355 | 89.2% |
| Tranco | 262 | 178 | 91 | 32.1% |
| CrUX | 36 | 9 | 23 | 75.0% |
| Majestic | 26 | 4 | 17 | 84.6% |
| Cisco Umbrella | 24 | 2 | 23 | 91.7% |
FARSIGHT_PDNS (exact strings, classified as a passive-DNS dataset, not a ranking). 12 strings, 13 papers. The report fails if a new /farsight|dnsdb/i spelling is not in this set:
DNSDB Farsight passive DNS database DNSDB and Project Sonar Data Repository Farsight Security passive DNS DNSDB and 360 PassiveDNS Farsight passive DNS Farsight Security DNSDB Farsight's DNSDB Farsight Security Information Exchange (SIE) / Farsight passive DNS database passive DNS dataset similar to Farsight DNSDB, provided by QiAnXin Company Farsight PDNS Farsight DNSDB
Farsight-as-ranking papers: 0. The ranking exists (Tranco methodology, 1M PLD, since 2022-05-01, cache-misses) and is a default Tranco input; nobody in this corpus names it as a sourceList.
Unnamed top-N residue after tightening isUnnamedTop to \btop\b (a first draft matched “desktop website”):
1 Google AdWords top websites
1 Google's Top 1,000 Most-Visited Websites
1 remaining top websites
1 top 100K websites used in Phase I
Four papers. “Google's Top 1,000 Most-Visited Websites” is a ranking with no vendor family; leaving it in residue is correct. The 764 vs 763 gap is this frame-kind-only paper.
Radar regex requires cloudflare radar or radar (domain )?rank so Tracker Radar does not match. Four WEB papers name Cloudflare Radar, all 2025–2026.
6. Quotes and figures spot-checked
Whitespace-normalised against paper.cols.txt.
| Claim | Source | Result |
|---|---|---|
| HTTP/2 26.6% Alexa Top 1M vs 7.84% com/net/org | [1Scheitle, Quirin; Hohlfeld, Oliver; Gamba, Julien; Jelten, Jonas; Zimmermann, Torsten; Strowes, Stephen D.; Vallina-Rodriguez, Narseo (2018): "A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists", in: Proceedings of the Internet Measurement Conference 2018, pp. 478–493. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] IMC 2018 | exact, prose and table |
| 1,790 Alexa top-10K; 70% lower rank-magnitude bucket; 27.2% two or more orders lower; CrUX most accurate across all metrics | [2Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] IMC 2022 | exact (Cloudflare is split / ligatured in .cols) |
| “top 10,000 sites from the Tranco list of October 29, 2019” | USENIX 2020 a-tale-of-two-headers-… | exact in .cols; the extraction quote prefixes “we decided to analyze” which sits across the column break |
| Alexa.com retired 1 May 2022; APIs 15 December 2022 | Wayback of support.alexa.com | exact |
| Tranco latest 46W9X, five providers, dowdall | live API 2026-08-27 | exact |
Rejected as a publishable finding: counting Alexa listVersion strings dated 2023 or later as “impossible draws”. Sampling already tried this; four hand-checks failed because listVersion is not scoped to the tuple. The published figure is the undated share (34 of 66).
7. External sources
All re-fetched 2026-08-27 by external_checks_website_selection.sh (exit 0).
| Claim | How verified |
|---|---|
| Alexa.com 1 May 2022; APIs 15 Dec 2022 | Wayback 20221126115049 of the support article |
| No Alexa retirement on 1 August 2023 | That date is Tranco's provider swap (CrUX+Radar in, Alexa out). AWS shutdown listing does not mention Alexa (recorded on the sampling provenance). |
| Tranco 46W9X; providers crux, farsight, majestic, radar, umbrella | GET /api/lists/date/latest |
| Farsight ranking since 1 May 2022; cache-misses; 1M PLD | Tranco methodology HTML |
| DomainTools “Mirror, Mirror” names 1 May 2022 and Farsight DNSDB | live blog 200 |
| Umbrella top-1m.csv.zip still published | index 200, zip HEAD 200; example rows include TLD com/net |
| Majestic Million page live | 200 |
| SecRank secrank.cn live | 200 |
| CrUX docs live | 200 |
| Radar /domains 403 to curl | 403; documented on cloudflare_radar |
| Quantcast /measure/ → /publisher/measure | 200 after redirect; not treated as a public ranking |
| SimilarWeb homepage | 200 |
| Farsight not in custom-list providers enum | already on tranco; not re-litigated |
Rejected:
- SEO listicles ranking “best website traffic tools 2026”.
- Wikipedia as a primary source for Alexa retirement (the page still links it as a name disambiguator; the date comes from Alexa Support).
- DomainTools acquisition URL
/resources/blog/domaintools-acquires-farsight-security/→ 404. The Mirror, Mirror post itself says they acquired Farsight; that sentence is used, the dead URL is not.
8. What could not be established
- A public download URL for the Farsight ranking (as opposed to DNSDB). Tranco's methodology describes it; DomainTools does not publish a CSV equivalent of Umbrella's.
- Whether Quantcast Measure still sells a top-sites file behind a login. The public URL is publisher analytics.
- A re-run of Ruth et al. against the five-provider Tranco list. The 2022 comparison predates Radar-in-Tranco.
- Full-text (as opposed to
sourceList) mentions of Farsight-as-ranking. Out of scope; the schema question is what people sampled from.
9. The run
- Date: 2026-08-27. Agent: Cursor (not
claude -p/ drain-sandbox). - Item:
design:website_selection: close the remaining TODOs(claimed as id 184 so a night drain could not takedesign:automated_measurementsinstead). - Reviewers: three focused GPT 5.6 Luna medium passes in parallel, then one generic pass with no checklist, after a freeze. Findings in §10.
- No credentials were echoed. Wiki writes go through
scripts/dw.mjs(JSON-RPC). - Deleted the one-off probe
scripts/_probe_ws.mjsafter folding its questions into the report.
10. Review log
Freeze directory: out/freeze_ws/. Reviewers were handed that snapshot. The content page was not edited while the three focused passes ran. Generic ran on the post-focused-review page.
Model for all four: GPT 5.6 Luna medium (not sonnet/fable as the task spec's default split).
Pass 1 — figures vs script
No findings. Report rerun, fold self-test, windowed and whole-page check_page_numbers, tables and wrap: all passed.
Pass 2 — citations and quotes
| Finding | Verdict |
|---|---|
| Blocking. Majestic limitations attributed a CrUX-relative manipulation cost to [3Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)] (2019). CrUX public rankings post-date that paper. | Accepted. Now: link-graph lists are cheap relative to a toolbar list; that paper did not evaluate CrUX. The same anachronism was in the manipulation paragraph (Le Pochat “expensive relative to toolbar” applied to CrUX) and was fixed in the same pass. |
| “Ruth et al. measured every public list” | Accepted. Now names Alexa, Majestic, Umbrella, Tranco and CrUX. |
| Galloway “bucket” put today's Radar API vocabulary in the 2024 paper's mouth | Accepted. Quote is now “consistently achieved a ranking in the top 100,000”; the bucket reading is attributed to today's API. |
| Alexa “user panel and tracking scripts” uncited | Accepted. Now cited to [1Scheitle, Quirin; Hohlfeld, Oliver; Gamba, Julien; Jelten, Jonas; Zimmermann, Torsten; Strowes, Stephen D.; Vallina-Rodriguez, Narseo (2018): "A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists", in: Proceedings of the Internet Measurement Conference 2018, pp. 478–493. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] as toolbar/panel. |
Le Pochat “single HTTP request” unverifiable here (NDSS 2019 Tranco PDF absent from data/fulltext) | Noted, no change. The claim is on tranco and the 2019 abstract; this sitting did not fetch the NDSS PDF. |
Six live citekeys each occur once in bibliography. Scheitle 26.6%/7.84%, Ruth 1,790/70%/27.2%, Galloway $10/top 100,000: verbatim in .cols.
Pass 3 — external currency
| Finding | Verdict |
|---|---|
| CrUX “CC BY-SA 4.0” and “Help improve Chrome” eligibility | Accepted. Google's methodology (fetched 2026-08-27) licences datasets CC BY 4.0 and lists four user-eligibility criteria. Tranco's methodology page still says CC BY-SA 4.0; the page now says so and prefers Google. Neighbour one-liners on tranco and website_classification updated so the three pages do not disagree. |
| “inputs have already changed twice” — Quantcast drop, Farsight add, Alexa→CrUX+Radar is three | Accepted. |
Alexa Wayback, Tranco 46W9X, Umbrella zip, Majestic, SecRank, CrUX docs, Radar 403, Quantcast redirect, SimilarWeb, DomainTools canonical /blog/ | Confirmed, no change. |
Pass 4 — generic, no checklist
| Finding | Verdict |
|---|---|
| Blocking. Lead “Tranco is current practice; CrUX is current evidence” treats a Feb 2022 comparison as a 2026 recommendation, while the page admits it has not been re-run on five-provider Tranco. | Accepted. Lead is now “Tranco is what 2025 papers use; CrUX is what Ruth et al. found most accurate in 2022” plus the not-re-run sentence. |
| “Tranco-plus-CrUX is the combination the accuracy evidence actually supports” | Accepted. Ruth supports them being different instruments, not a composite design. |
| “CrUX is more stable still” / prefer buckets as if that were Le Pochat's CrUX finding | Accepted. 30-day aggregation stays on Le Pochat; CrUX buckets described without an unmeasured stability rate. |
| “Any adoption number run on a top list is biased upward” | Accepted. Scoped to Scheitle's HTTP/2 (and same-direction IPv6/CDN) comparison. |
| “DNS rankings did not become expensive to fake when Alexa died” | Accepted. Scoped to Radar and Tranco-via-Radar. |
| Alexa lag “is a submission cycle” as a causal finding | Accepted. Now “compatible with a submission cycle; we did not read why”. |
| Undated 34 “are a copy of unknown provenance — a mirror, a cached CSV, or inherited” as if those were established | Accepted. Those are listed as possibilities; the measured fact is that they cannot be pinned. |
| Farsight “sees default-resolver and server traffic” | Accepted. Now cache-miss DNS from participating resolvers, including names only infrastructure looks up. |
| Sampling page still says top-n is “weighted towards sites that matter to users” | Rejected for this sitting. True about a neighbour; out of scope. Recorded so the next sampling refresh sees it. |
| Provenance §10 empty while promising findings | Accepted by writing this table. |
11. Same-day rewrite: catalogue → decision page
After item 184 published, Karel asked to question the page against later LLM-written pages (the model was ip_classification) and decide what of the human-era catalogue to keep. Not a new drain item.
Kept from the human page: the vendor catalogue (CrUX / Tranco / Radar / Umbrella / Majestic / Alexa / Farsight / SecRank / SimilarWeb / Quantcast), the DNS-list caveats, the Le Pochat manipulation figure, the Scheitle and Xie survey figures as dated history, Farsight default-list-only, the Alexa 1 May 2022 vs Tranco 1 August 2023 distinction, Best Practices folded into What to report, the split with sampling.
Dropped or rewritten: numbered 1–6 Advantages/Limitations that restated the table; “the provider's own example file begins 1,com / 2,net” (the 2026-08-27 top-1m.csv does not; that is top-1m-TLD.csv.zip); ec2.internal as a live Umbrella example (absent from that day's top-1m); drain-internal voice about the work item that commissioned the corpus section; treating CrUX-first catalogue order as if it were current practice rather than Ruth's 2022 accuracy ranking.
Added, matching IP classification: a decision table (claim → list); a live rank-disagreement table and published compare_ranks.py (the analogue of classify_ips.py); blockquotes from Ruth, Scheitle, Galloway; a 2026 status table with paper counts; Papers to read first; Open questions; custom-seed 257 vs ranking 764 pointed at sampling rather than re-derived as a vendor ranking; gstatic.com at Tranco rank 3 as the page's “not one of them is the BBC” moment.
compare_ranks.py run 2026-08-27T14:26:07Z against Tranco 46W9X, Umbrella top-1m (1,000,000 hostnames) and Majestic Million. Umbrella TLD file: 10,905 rows, com=1, net=2, org=6.
12. Rewrite review log
Focused freeze: out/freeze_ws_rewrite/ (page MD5 b45a6832…). Content page not edited while the three focused passes ran. Generic freeze: out/freeze_ws_generic/ after the currency wording fix. Model for all four: GPT 5.6 Luna medium.
Pass 1 — figures vs script
No findings. Report, fold, windowed-equivalent whole-page check_page_numbers, tables, wrap: passed. Embedded 180 / 1.05 / 1,000 present in Z.
Pass 2 — citations and quotes
| Finding | Verdict |
|---|---|
| Le Pochat “single HTTP request” unverifiable (NDSS 2019 full text absent) | Noted, no change. Same gap as the first sitting. |
| Citekeys, Ruth/Scheitle/Galloway quotes, Galloway not given today's Radar bucket vocabulary, Xie labelled PAM | Confirmed, no change. |
Pass 3 — external currency
| Finding | Verdict |
|---|---|
| Page said “Tranco's methodology still says CC BY-SA 4.0”. Live methodology HTML has no CC BY/CC BY-SA string. | Accepted. The licence line is on Tranco's homepage, not the methodology page. Content, tranco and website_classification now say homepage vs methodology vs Google CC BY 4.0. |
Tranco 46W9X, Farsight paragraph, Alexa Wayback, Umbrella hostnames + TLD file 10,905, Majestic, CrUX CC BY 4.0, Radar 403, Quantcast redirect, SecRank, Farsight not in providers enum, gstatic.com rank 3, compare_ranks.py –wiki | Confirmed 2026-08-27. |
Pass 4 — generic, no checklist
| Finding | Verdict |
|---|---|
| Blocking. Decision table “Pages people actually load → CrUX (or Tranco)” | Accepted. CrUX is the direct choice; Tranco is named as not a substitute. |
| Blocking. “A ranking is a proxy for popularity” while Majestic is backlinks | Accepted. Now “each list ranks an observable”. |
| 2022 Ruth finding reads as a 2026 recommendation | Accepted. Lede is “pick the list that matches the claim”; Ruth labelled 2022 evidence. |
| “Prefer a month of ranks” underspecified | Accepted. Pin one dated frame; Tranco id already is a 30-day Dowdall; CrUX is a month snapshot. |
| “today's Tranco” vs list dated 2026-08-26, fetched 2026-08-27 | Accepted. Prose now uses those two dates. Script function names unchanged. |
| SimilarWeb “paywalled” does not say whether you can pin an export | Accepted. Status row now requires a keepable contract export; no public id. |
| Repeated sampling cross-references | Rejected. The opening split and one What-to-report pointer are the boundary with sampling; cutting them would bury it. |
13. 2026-09-03: site-wide Alexa date audit
Contemporaneous. Written during the sitting, Claude Opus 5, 2026-09-03. No corpus query was run — this item is a vendor date, and extractions.jsonl cannot adjudicate when Amazon switched a service off. The inputs were the wiki's own page sources and four external primary sources. No corpus figure on any page changed.
13.1 The work item, and why its premise was already gone
The item read: “design:website_selection says Alexa was discontinued 1 August 2023 … Fix the date on design:website_selection (and check whether the same date appears elsewhere).” That sentence was already gone. It was created with the page in January 2025 and survived untouched until the sampling sitting replaced it:
| Revision | What it did |
|---|---|
| 1735841558 | Page created (human author). Carries “Discontinued as of August 1, 2023”. |
| 1785947402 | Unrelated link fix; the sentence is still there, verbatim. |
| 1786570786 | Removed it. “Alexa was retired 1 May 2022 (APIs 15 Dec 2022), not August 2023 …” |
| 1787840075 | Did not remove anything — added Tranco's input dates and the explicit “Alexa was not 'discontinued in 2023'” disclaimer. |
So the content page needed no edit, and did not get one. What the item still bought was the second half — check whether the same date appears elsewhere — which nobody had done site-wide, plus an independent re-verification of the date that replaced it. Both are below.
13.2 The date, re-verified from primary sources today
Not carried over from the pages. Each fetched fresh on 2026-09-03 and matched verbatim against the response body, not against a summary of it.
| Claim | Source, fetched 2026-09-03 | Verbatim |
|---|---|---|
alexa.com retired 1 May 2022 | Alexa Support article 4410503838999 via Wayback (support.alexa.com no longer resolves), HTTP 200 | “we made the difficult decision to retire Alexa.com on May 1, 2022” |
| Top Sites / Web Information Service APIs retired 15 December 2022 | same article, HTTP 200 | “The APIs will be retired on December 15, 2022.” |
The end-of-service notice on the alexa.com login page | https://web.archive.org/web/20220315000000/https://www.alexa.com/, HTTP 200 | “we will be retiring Alexa.com on May 1, 2022” |
| 1 August 2023 is Tranco's provider swap, not Amazon's shutdown | https://tranco-list.eu/ front page, HTTP 200 | “The Chrome User Experience Report and Cloudflare Radar rankings have been integrated into the default Tranco list, starting from the daily updated list of August 1, 2023.” and “The Alexa ranking has been removed from the default Tranco list, as it is no longer available.” |
Every Wayback URL cited for these dates across the wiki was fetched and returned HTTP 200 with the quoted text present: the dated capture web/20221126115049/… used by website_selection, the year-wildcard web/2023/… used by sampling, the bare-article web/2022/…4410503838999 used by biases, and the login-page capture used by website_classification. No primary source for an Alexa retirement in 2023 was found, by this sitting or by the sampling sitting searching from the other direction, or by the external-currency reviewer searching Wikipedia's cited sources, contemporary news coverage and AWS's own full-shutdown listing (which contains no occurrence of “Alexa” at all). That is three independent negatives, not a proof — see §13.9.
13.3 The site-wide sweep
| Item | Value |
|---|---|
| Script | scripts/alexa_date_audit.mjs (new this sitting) |
| Mutation test | scripts/alexa_date_audit_mutations.sh (new this sitting) |
| Population | 160 wiki pages — the whole wiki (87 content, 73 provenance:), as the union of ?do=sitemap (148), core.listPages at depth 0 (160) and core.searchPages for Alexa / alexa.com / discontinued / retired / retiring (83 distinct) |
| Unit of analysis | not the page and not the line but the unit: a line's body with ((footnote)) bodies removed, plus each footnote body separately |
| Hits, before this page was written | 18 units pair a death verb with a 2023 token inside one unit on a line mentioning Alexa. 0 SUSPECT. This is the run that cleared the wiki |
| Hits, after this page was published | 37 units. 0 SUSPECT on content pages, 11 provenance-note, 1 corpus-slug, 9 beyond-window — see §13.3.1 |
| Read by hand | all of them, both runs, not taken from the verdict column |
| The excused, honestly | ok-tranco-swap is 14 on the current run, but only about five actually describe the swap: the rest are publication-year sentences (“66 papers from 2023–2026 still name it”, “Tranco overtook Alexa in 2023”) or correction records that the TRANCO rule excused because “Tranco” happened to sit in the ±90 window. The verdict label is a triage bucket, not a claim about each line |
| Unexercised rules | ok-literature-year (0) and corpus-slug (0). Both stay in the script and are printed with their zeros. ok-literature-year reads as 0 partly because TRANCO is tested first and gets there before it — an ordering artefact, not evidence the shape is absent |
| Result | Nothing on the wiki dates Alexa's retirement to 2023. No content page was edited for this defect. |
Why the union of three page lists — and how the first draft still missed two pages. ?do=sitemap is cached and lagged: on 2026-09-03 it listed 148 ids and omitted 10 live pages, including privacy:ads_txt, privacy:browser_protection and statistics:how_many_sites. A search finds only pages containing the search term. core.listPages is complete — but only at depth 0, and its default is depth 1, which returns 88 of 160 pages and just 8 of the 73 provenance: pages. scripts/dw.mjs called it with no arguments, so the first draft of this audit swept 158 pages and reported that as “every page id”.
The generic reviewer found it. The two missed pages were provenance:privacy:browser_extensions and provenance:privacy:privacy_sandbox; neither mentions Alexa, so the PASS survived on outcome — but the section had been about to claim a site-wide sweep while omitting 1.25% of the wiki. listPages() in scripts/dw.mjs now defaults to depth 0 for every caller, which is the real fix: the next script to ask this wiki “what pages exist” will get all of them.
Why footnotes are separate units. The first version of this script analysed whole lines and passed a mutation it should have caught: replacing the body text of website_selection's Alexa bullet with “Discontinued as of August 1, 2023” still exited 0, because the footnote attached to the same line contains the correct “May 1, 2022” and the rule accepted that as vouching for the sentence. On this wiki the claim is in the body and the source is in the footnote, so a footnote's correctness must never excuse a body claim. Splitting the units fixed it; the mutation output below is the proof.
Why a proximity rule rather than a keyword rule. “Alexa” and “2023” co-occur on 102 lines of the 71 pages that mention Alexa at all (grep -i alexa *.txt | grep -c 2023 over the exported sources), nearly all of them legitimate (66 papers published 2023–2026 still name Alexa; IMC/2023/tracking-profiling-…-alexa-echo-… is a corpus slug; Tranco's swap is a 2023 event). Flagging co-occurrence produces a list nobody reads. The defect is specifically a death verb next to a 2023 date, so that is what is matched, with a ±90-character window around the year checked for Tranco/CrUX/Radar/provider/default-list wording that would make it correct.
Both halves of that rule were wrong in the first draft, and reviewers broke them. The window was 60 characters and the death vocabulary was six verbs, so “Alexa was decommissioned in 2023” passed (the verb was not in the list), an ordinary verbose sentence putting “discontinued” ~76 characters from the year passed, and “switched the service off” — a split phrasal verb that website_classification's own footnote uses — was invisible. Now: 29 death-verb alternatives including split phrasal forms, and PROX = 200. All three phrasings are in the mutation suite (M6, M7, M10) so a future edit to those regexes cannot silently undo the fix.
PROX is measured, not asserted, and its residue is printed. On 2026-09-03 the widest verb–year pair the wiki actually contains inside one unit is 188 characters and the narrowest excluded pair is 215, so 200 sits in a real gap rather than at a guessed bound. Pairs beyond it are not dropped: they are reported as beyond-window with their gap, five on this run, and read by hand. Four are the reworded longitudinal footnote and a design:sampling sentence about papers; one is deployment line 1731, where “Nomad moved to BUSL in 2023” shares a line with an unrelated mention of Alexa. That last one is why the window exists at all: with no distance rule the audit reports one false positive and exits non-zero on a clean wiki.
The general lesson outlives this script: a keyword list and a magic constant are the weakest parts of any audit, and the only way to size either is to have someone attack it.
13.3.1 Publishing this audit changed what it measures
The first run of the audit swept 160 pages and found 18 units, 0 SUSPECT. Then §13 was published — and §13 is, by construction, a page densely covered in sentences that pair “Alexa” with 2023 and a death verb: the work item quoted, the old error quoted to correct it, ten mutation strings including a literal “Alexa was shut down on 1 August 2023”, and a review log full of what reviewers found. Re-run against the published wiki, the audit reported 16 SUSPECT — every one of them its own documentation.
That is not a nuisance to suppress. It is the shape of the problem, and it forced two changes that make the script better rather than merely quieter:
- Fenced blocks are stripped.
<code>and<file>contents are not prose claims; the audit's own published listing and its mutation fixtures are the proof.sitemap.mjsstrips the same constructs before hunting red links, for the same reason. Blocks are blanked line-for-line so reported line numbers still point at the right source line. provenance:hits are listed, not failed. A provenance page's job includes recording the error verbatim, so on those pages the pattern is expected. They go to aprovenance-notebucket which is printed in full and does not set the exit code.SUSPECTis reserved for content pages, which is where a wrong date actually misleads a reader — and where the original bug was.
What this costs, stated plainly: the audit no longer polices provenance prose. If a future run writes a wrong Alexa date into a provenance: page, this script will print it in the provenance-note list and exit 0. The mitigation is that the list is short (11 units) and the script's own summary line says “read them, do not trust the label”. A stronger version would keep failing on provenance pages and maintain an explicit allow-list of the sentences that are meant to quote the error — which is a maintenance burden this item did not justify, and is the honest reason it was not built.
13.4 The script
- alexa_date_audit.mjs
// Site-wide audit: does any page on measuretheweb.org still date Alexa's // retirement to 2023? // // node scripts/alexa_date_audit.mjs # exit 0 = clean, 1 = review needed // // Background. design:website_selection once said Alexa was "Discontinued as of // August 1, 2023". The primary source — Alexa Support's own notice — says // alexa.com was retired 1 May 2022 and the Top Sites / Web Information Service // APIs on 15 December 2022. 1 August 2023 is a real date but it belongs to // Tranco: the day CrUX and Cloudflare Radar replaced Alexa in Tranco's default // list. Two facts, one year and one organisation apart, and ~30 pages of this // wiki mention one or the other, so the separation has to be checked site-wide // rather than on the page that got it wrong. // // Method. The page list is the union of three sources because none is complete // on its own: ?do=sitemap is cached and lagged, and a search only finds pages // containing the term. core.listPages IS complete, but only at depth 0 — its // default depth of 1 returned 88 of 160 pages, and the first draft of this // script trusted that default and swept 158 pages instead of 160. The two it // missed happened not to mention Alexa. Each page is split into *units* - // the body with ((footnote)) bodies removed, plus each footnote body on its own — // because a footnote's correct 2022 date must not vouch for a body claim that // says 2023 (that mistake let a mutation test through on the first version of // this script). URLs are stripped before year detection, so an archive // timestamp such as web/2023/ is not read as a claim. // // The defect is a *proximity* pattern: a death verb within PROX characters of a // 2023 token. A hit is excused only if the 2023 token itself sits next to // Tranco/CrUX/Radar/provider/default-list wording, or the sentence is this // wiki's own record of the old error. Every hit is printed with its verdict, so // the part the rules cannot judge stays visible. import fs from 'node:fs'; import path from 'node:path'; import zlib from 'node:zlib'; import { fileURLToPath } from 'node:url'; import { listPages, searchPages } from './dw.mjs'; const BASE = 'https://measuretheweb.org'; const UA = 'measuretheweb-wiki-bot (Claude; contact karel.kubicek@vaultjs.com)'; const CACHE = path.join(path.dirname(fileURLToPath(import.meta.url)), '..', 'out', 'alexafix', 'raw'); const SEARCH_TERMS = ['Alexa', 'alexa.com', 'discontinued', 'retired', 'retiring']; // PROX was 60 in the first draft. A reviewer broke that with an ordinary // verbose sentence putting the verb and the year ~76 chars apart. It is now // 200: measured on 2026-09-03, the widest verb-year pair the wiki actually // contains inside one unit is 188 chars and the narrowest one excluded is 215. // Pairs beyond PROX are NOT dropped - they are reported as 'beyond-window' so // the window is auditable instead of load-bearing. const PROX = 200; // chars between a death verb and a 2023 token const EXCUSE = 90; // chars around the 2023 token searched for Tranco context const toId = (s) => s.trim().toLowerCase().replace(/^:/, ''); async function allPageIds() { const ids = new Set(); const res = await fetch(`${BASE}/doku.php?do=sitemap`, { headers: { 'User-Agent': UA } }); const xml = zlib.gunzipSync(Buffer.from(await res.arrayBuffer())).toString('utf8'); const fromSitemap = [...xml.matchAll(/<loc>(.*?)<\/loc>/g)] .map((m) => toId(m[1].replace(`${BASE}/`, '').replace(/\//g, ':'))); const fromRpc = (await listPages()).map((p) => toId(p.id)); const fromSearch = []; for (const t of SEARCH_TERMS) for (const h of await searchPages(t)) fromSearch.push(toId(h.id)); [...fromSitemap, ...fromRpc, ...fromSearch].forEach((id) => ids.add(id)); console.log(`page ids: sitemap ${fromSitemap.length}, core.listPages ${fromRpc.length}, ` + `search(${SEARCH_TERMS.join('/')}) ${new Set(fromSearch).size} distinct -> union ${ids.size}`); return [...ids].sort(); } async function source(id) { fs.mkdirSync(CACHE, { recursive: true }); const file = path.join(CACHE, `${id.replace(/:/g, '__')}.txt`); if (fs.existsSync(file)) return fs.readFileSync(file, 'utf8'); const r = await fetch(`${BASE}/doku.php?id=${encodeURIComponent(id)}&do=export_raw`, { headers: { 'User-Agent': UA } }); if (!r.ok) throw new Error(`${id}: HTTP ${r.status}`); const text = await r.text(); fs.writeFileSync(file, text); return text; } // Death vocabulary. The first published version had only retire/discontinue/ // dead/shut down/sunset/end-of-service, and a reviewer walked "Alexa was // decommissioned in 2023" straight past it. Anything that ends a service in // English prose belongs here: a false positive costs one line of reading, a // false negative is the entire defect this script exists to find. const DEATH = /(retir\w*|discontinu\w*|decommission\w*|deprecat\w*|dead\b|died|death|shut ?down|shut off|switched off|turned off|sunset\w*|end of service|end[- ]of[- ]life|\bEOL\b|out of service|cease\w*|terminat\w*|wound down|wind[- ]down|went dark|taken offline|pulled offline|killed off|defunct|closed down|no longer (?:available|published|exist\w*|operat\w*|maintain\w*|updat\w*|running|live|provided|a frame)|(?:switch|shut|turn|clos|wind|wound|kill|phas)\w*(?:\s+\w+){1,3}\s+(?:off|down|out)|(?:pull|took|take|taken)\w*(?:\s+\w+){0,3}\s+(?:offline|plug))/gi; const Y2023 = /(1 august 2023|august 1st?,? 2023|2023-08-01|20230801|aug\.? 2023|\b2023\b)/gi; const TRANCO = /tranco|crux|chrome user experience|cloudflare radar|\bradar\b|default (?:daily )?list|provider|integrated into|removed from/i; // The 2023 token is a year of *publications*, not a vendor date: a range that // starts in 2023, or a year sitting next to paper/venue wording. This is what // keeps the wide window from drowning the report in the wiki's many true // sentences about papers published 2023-2026 that still name Alexa. const LITERATURE = /2023\s*[-–—]\s*20\d\d|20\d\d\s*[-–—]\s*2023|2023 or later|\bpapers?\b|\bpublish\w*|\bpublication\w*|\bvenue\b|submission cycle|CCS|IMC|NDSS|PETS|USENIX|WWW|TheWebConf|S&P/i; const HISTORY = /earlier revision|earlier version|previous version|it said|once said|still says|used to say|wrong|incorrect|corrected|correction|not when|belongs to|no source|confus\w+|reject\w+|contradict\w+|this page/i; // Corpus bookkeeping: a venue/year/slug line from a report-script dump. const CORPUS_SLUG = /(?:CCS|IMC|NDSS|PETS|USENIX|WWW|TheWebConf|S&P)\s*\/\s*20\d\d|20\d\d\s*\/\s*(?:CCS|IMC|NDSS|PETS|USENIX|WWW)|alexa-echo|alexaechos/i; // Split a line into units: body without footnote bodies, plus each footnote body. function units(line) { const out = []; const fns = [...line.matchAll(/\(\(([\s\S]*?)\)\)/g)].map((m) => m[1]); out.push({ kind: 'body', text: line.replace(/\(\([\s\S]*?\)\)/g, ' ') }); fns.forEach((f, i) => out.push({ kind: `footnote${i + 1}`, text: f })); return out; } // Remove link targets and bare URLs: an archive timestamp is not a date claim. const stripUrls = (s) => s .replace(/\[\[([^\]|]*)\|([^\]]*)\]\]/g, ' $2 ') .replace(/\[\[([^\]]*)\]\]/g, ' $1 ') .replace(/https?:\/\/\S+/g, ' URL '); // Fenced blocks are not prose claims. This script's own <file> listing and its // mutation fixtures ("Alexa was shut down on 1 August 2023") are published on // provenance:design:website_selection, and without this the audit reads its own // test data as 16 wiki defects. sitemap.mjs strips the same constructs before // looking for links, for the same reason. Blanked line-for-line so reported // line numbers still point at the right line of the source. const stripBlocks = (text) => text.replace( /<(code|file)\b[^>]*>[\s\S]*?<\/\1>/gi, (m) => m.replace(/[^\n]/g, ' ')); const rows = []; const ids = await allPageIds(); for (const id of ids) { const lines = stripBlocks(await source(id)).split('\n'); for (let i = 0; i < lines.length; i++) { if (!/alexa/i.test(lines[i])) continue; for (const u of units(lines[i])) { const t = stripUrls(u.text); if (!/alexa/i.test(t)) continue; const deaths = [...t.matchAll(DEATH)]; const years = [...t.matchAll(Y2023)]; if (!deaths.length || !years.length) continue; for (const y of years) { const gaps = deaths.map((d) => Math.abs(d.index - y.index)); const near = deaths.find((d, k) => gaps[k] <= PROX + Math.max(d[0].length, y[0].length)); if (!near) { // Residue, not silence: a verb and a year in the same unit but further // apart than PROX. Printed so nobody has to trust the window, and so // raising PROX later is an informed decision rather than a guess. rows.push({ id, line: i + 1, unit: u.kind, verdict: 'beyond-window', snippet: `gap ${Math.min(...gaps)} chars: ` + t.slice(Math.max(0, y.index - EXCUSE), y.index + y[0].length + EXCUSE) .replace(/\s+/g, ' ').trim() }); continue; } const window = t.slice(Math.max(0, y.index - EXCUSE), y.index + y[0].length + EXCUSE); let verdict; if (CORPUS_SLUG.test(window)) verdict = 'corpus-slug'; else if (TRANCO.test(window)) verdict = 'ok-tranco-swap'; else if (HISTORY.test(window)) verdict = 'ok-history'; else if (LITERATURE.test(window)) verdict = 'ok-literature-year'; // A provenance: page's job includes recording the error, quoting the work // item, and logging what a reviewer found - all of which look exactly // like the defect. Those are listed for reading, not failed, and they do // not set the exit code. SUSPECT is reserved for content pages, which is // where a wrong date actually misleads a reader. else verdict = id.startsWith('provenance:') ? 'provenance-note' : 'SUSPECT'; rows.push({ id, line: i + 1, unit: u.kind, verdict, snippet: window.replace(/\s+/g, ' ').trim() }); break; // one row per unit is enough to send a human to the line } } } } const order = ['SUSPECT', 'ok-tranco-swap', 'ok-history', 'ok-literature-year', 'corpus-slug', 'provenance-note', 'beyond-window']; console.log(`\npages scanned: ${ids.length}; death verb within ${PROX} chars of a 2023 token, ` + `on a line mentioning Alexa: ${rows.length} unit(s)\n`); for (const v of order) { const g = rows.filter((r) => r.verdict === v); console.log(`== ${v}: ${g.length}`); for (const r of g) console.log(` ${r.id}:${r.line} [${r.unit}] ${r.snippet.slice(0, 300)}`); console.log(''); } const bad = rows.filter((r) => r.verdict === 'SUSPECT').length; const notes = rows.filter((r) => r.verdict === 'provenance-note').length; console.log(bad === 0 ? `PASS: no CONTENT page puts a death verb next to a 2023 date without Tranco ` + `context. ${notes} provenance-note unit(s) listed above are not failed - ` + `read them, do not trust the label.` : `FAIL: ${bad} content unit(s) date Alexa's end to 2023 with no Tranco attribution.`); process.exit(bad === 0 ? 0 : 1);
13.5 Its unedited output
Read the unit count as “as of a revision”, and the SUSPECT count as the claim. This is the real, unedited run against the wiki as it stood immediately before this section was last saved. Re-run it now and the unit count will be higher, because this section is itself part of the input and every save adds sentences that pair Alexa with 2023 (§13.3.1). That is a fixed point nobody can publish their way out of: the output block can never be a run that includes the page carrying it.
So the published claim is the invariant, not the tally: SUSPECT is 0 on content pages in every run of this sitting — before the section existed (18 units), after it was first published (37), and after the revision that added §13.3.1 (39). The committed scripts/alexa_date_audit-output.txt is regenerated after each save and will disagree with the block below in its counts, by design.
page ids: sitemap 148, core.listPages 160, search(Alexa/alexa.com/discontinued/retired/retiring) 83 distinct -> union 160
pages scanned: 160; death verb within 200 chars of a 2023 token, on a line mentioning Alexa: 37 unit(s)
== SUSPECT: 0
== ok-tranco-swap: 14
design:longitudinal:73 [footnote1] ntegrated into the default Tranco list, starting from the daily updated list of August 1, 2023"// (fetched 2026-09-03). It is **Tranco's** date, not Amazon's — ''alexa.com'' was retire
design:website_classification:228 [footnote1] age, URL — retrieved 2026-08-07. An earlier revision of Design:Website Selection gave 1 August 2023; that date is when Tranco dropped Alexa from the default list, not when Amazon switched t
design:website_selection:18 [body] ad now. Alexa.com was retired on **1 May 2022** (APIs 15 December 2022); **66 papers from 2023–2026 still name it**, 34 of them with no version at all. Do not write "the Tranco top 1M"
design:website_selection:263 [footnote1] Support, "We retired Alexa.com on May 1, 2022", archived at web.archive.org . The date **1 August 2023** is when Tranco dropped Alexa from the default list and folded in CrUX and Cloudflare Ra
design:website_selection:338 [body] Tranco overtook Alexa in 2023, the first full year after Alexa.com died. The lag is compatible with a submission cycle;
design:website_selection:346 [body] f crawling papers ends around 2022 (PAM 2024; outside the seven venues). Neither includes 2023–2026, which is the window in which Tranco became the default and Alexa became residue. **
programming:tranco:64 [body] l 2020; Farsight joined on 1 May 2022; Alexa was replaced by CrUX and Cloudflare Radar on 1 August 2023, which is also when the default prefix stopped being ''full'' and became one million. Yo
provenance:design:sampling:280 [body] site_selection'' visibly wrong** — that page still says Alexa was discontinued "August 1, 2023" and that Tranco combines "Alexa, Cisco Umbrella, and Majestic", both contradicted here w
provenance:design:website_selection:16 [body] orrection | design:website_classification Alexa footnote no longer says this page gives 1 August 2023 as Alexa's death. CrUX licence one-liners on that page and programming:tranco now match
provenance:design:website_selection:131 [body] | No Alexa retirement on 1 August 2023 | That date is Tranco's provider swap (CrUX+Radar in, Alexa out). AWS shutdown listing do
provenance:design:website_selection:264 [body] anything — added Tranco's input dates and the explicit //"Alexa was not 'discontinued in 2023'"// disclaimer. |
provenance:design:website_selection:276 [body] | **1 August 2023** is Tranco's provider swap, not Amazon's shutdown | ''%% URL front page, HTTP 200 | //"
provenance:design:website_selection:593 [body] r-set table samples 1 January of each year, and the sentence under it read //"and then in 2023–24 Alexa was replaced by CrUX and Cloudflare Radar"// — the interval between two sampled
statistics:biases:62 [body] ntegrated into the default Tranco list, starting from the daily updated list of August 1, 2023."//
== ok-history: 2
provenance:design:sampling:36 [body] - It said Alexa was "**Discontinued as of August 1, 2023**". The primary source says Alexa.com was retired **1 May 2022** and the Top Sites / Web
provenance:design:sampling:233 [body] | Alexa's retirement dates on this page are correct; the wrong "August 1, 2023" is on ''design:website_selection'', and no source supports an August 2023 date for any A
== ok-literature-year: 0
== corpus-slug: 1
provenance:design:website_selection:299 [body] es), nearly all of them legitimate (66 papers published 2023–2026 still name Alexa; ''IMC/2023/tracking-profiling-…-alexa-echo-…'' is a corpus slug; Tranco's swap is a 2023 event). Fla
== provenance-note: 11
provenance:design:longitudinal:1485 [body] "2023–24" is the gap between two sampled rows, not the date of the change. The footnote then ga
provenance:design:website_selection:258 [body] The item read: //"design:website_selection says Alexa was discontinued 1 August 2023 … Fix the date on design:website_selection (and check whether the same date appears elsew
provenance:design:website_selection:263 [body] 86570786** | **Removed it.** "Alexa was retired 1 May 2022 (APIs 15 Dec 2022), not August 2023 …" |
provenance:design:website_selection:278 [body] e used by design:website_classification . **No primary source for an Alexa retirement in 2023 was found**, by this sitting or by the sampling sitting searching from the other direct
provenance:design:website_selection:287 [body] | Hits | **18 units** pair a death verb with a 2023 token inside one unit on a line mentioning Alexa |
provenance:design:website_selection:291 [body] | Result | Nothing on the wiki dates Alexa's retirement to 2023. **No content page was edited for this defect.** |
provenance:design:website_selection:297 [body] body text of design:website_selection 's Alexa bullet with "Discontinued as of August 1, 2023" still exited 0, because the //footnote// attached to the same line contains the correct
provenance:design:website_selection:301 [body] s 60 characters and the death vocabulary was six verbs, so //"Alexa was decommissioned in 2023"// passed (the verb was not in the list), an ordinary verbose sentence putting "discontin
provenance:design:website_selection:627 [body] * **Whether an Alexa-adjacent Amazon service was in fact retired in 2023.** Not exhaustively excluded — the sampling sitting checked AWS's own full-shutdown listi
provenance:design:website_selection:639 [body] | **Mutation the script missed:** ''**Alexa was decommissioned in 2023.**'' → exit 0, PASS. The ''DEATH'' vocabulary had no entry for decommission / cease / go
provenance:design:website_selection:657 [body] | **No primary source for a 2023 Alexa retirement exists.** Searched Wikipedia's article and its cited sources (which give
== beyond-window: 9
design:longitudinal:73 [body] gap 508 chars: Farsight came in on 1 May 2022, and Alexa was replaced by CrUX and Cloudflare Radar on **1 August 2023**. That last change is the sharpest form of the problem. ''W9ZN9'' and ''25299'' are **co
design:longitudinal:73 [footnote1] gap 571 chars: day's list id and its ''configuration'': the other five rows on 2026-08-27, the 31 July / 1 August 2023 pair on 2026-09-03. ''date/20230731'' gives ''W9ZN9'', ''providers: [alexa, umbrella, maj
design:longitudinal:73 [footnote1] gap 419 chars: ZN9'', ''providers: [alexa, umbrella, majestic, farsight]'', ''listPrefix: full''; ''date/20230801'' gives ''25299'', ''providers: [crux, farsight, majestic, radar, umbrella]'', ''listPref
design:sampling:188 [body] gap 327 chars: a mirror, a cached CSV, or a list inherited from an earlier paper. **66 papers published 2023–2026 in this corpus still draw from Alexa, and 34 of them (51.5%) give no version or date
provenance:design:website_selection:278 [body] gap 211 chars: ture ''web/20221126115049/…'' used by design:website_selection , the year-wildcard ''web/2023/…'' used by design:sampling , the bare-article ''web/2022/…4410503838999'' used by stat
provenance:design:website_selection:299 [body] gap 409 chars: **Why a proximity rule rather than a keyword rule.** "Alexa" and "2023" co-occur on **45 lines** of the **71 pages** that mention Alexa at all (''%%grep -i alex
provenance:design:website_selection:299 [body] gap 297 chars: lines** of the **71 pages** that mention Alexa at all (''%%grep -i alexa *.txt | grep -c 2023%%'' over the exported sources), nearly all of them legitimate (66 papers published 2023–2
provenance:design:website_selection:299 [body] gap 209 chars: c 2023%%'' over the exported sources), nearly all of them legitimate (66 papers published 2023–2026 still name Alexa; ''IMC/2023/tracking-profiling-…-alexa-echo-…'' is a corpus slug; T
provenance:programming:deployment:1731 [body] gap 312 chars: on and does not claim: Redis relicensed in 2024 and again in 2025, Nomad moved to BUSL in 2023.
PASS: no CONTENT page puts a death verb next to a 2023 date without Tranco context. 11 provenance-note unit(s) listed above are not failed - read them, do not trust the label.
13.6 Mutation test: the PASS is not vacuous
An audit that reports “0 SUSPECT” is worth exactly what its ability to find a real one is worth, and the first draft of this one could not (§13.3). So the script is mutation-tested: ten mutations, injected one at a time into the cached copy of website_selection. Eight must make the audit exit 1 (M1–M4, M6–M8, M10) and two are controls that must not be flagged (M5, M9). The interesting ones: M3 puts the true 2022 date in the same sentence after the wrong 2023 one; M4 hides the wrong date in a footnote; M6, M8 and M10 use verb forms absent from the first draft's vocabulary; M7 is a verbose sentence that outran its 60-character window; M9 is a true sentence about papers published 2023–2026 that must not be flagged.
- alexa_date_audit_mutations.sh
#!/bin/sh # Mutation test for alexa_date_audit.mjs: prove the PASS is not vacuous. # # Eight mutations must be caught (exit 1): M1, M2, M3, M4, M6, M7, M8, M10. # Two are controls that must NOT be flagged (exit 0): M5 and M9. # Each is applied to the cached copy of design:website_selection, the audit is # re-run, its exit code is checked, and the cache is restored. # # M1-M5 are the original set: M3 puts the true 2022 date in the same sentence # after the wrong 2023 one, M4 hides the wrong date in a footnote. # M6 and M7 are reviewer mutations that broke the first draft of the audit -- # M6 uses a death verb the vocabulary did not have ("decommissioned"), M7 puts # the verb and the year ~76 characters apart in ordinary prose. M8 is a second # unlisted-verb phrasing ("ceased operations"), M9 a control for a true sentence # about papers published 2023-2026, and M10 a split phrasal verb ("switched the # service off") -- the form a reviewer found in design:website_classification's # own footnote, which the contiguous verb list could not see. They are kept here # so a future edit to the DEATH vocabulary or to PROX cannot silently undo it. set -e cd "$(dirname "$0")/.." P=out/alexafix/raw/design__website_selection.txt cp "$P" /tmp/ws.orig TRUE_LINE='**Discontinued — alexa.com was retired on 1 May 2022 and the Alexa Top Sites and Web Information Service APIs on 15 December 2022.**' run() { name="$1"; want="$2"; repl="$3" python3 - "$P" "$TRUE_LINE" "$repl" <<'PY' import sys p,old,new=sys.argv[1],sys.argv[2],sys.argv[3] s=open(p).read() assert old in s, "anchor sentence not found; update the mutation script" open(p,'w').write(s.replace(old,new,1)) PY set +e; node scripts/alexa_date_audit.mjs >/tmp/mut.out 2>&1; got=$?; set -e cp /tmp/ws.orig "$P" if [ "$got" = "$want" ]; then echo " OK $name: exit $got (wanted $want)" else echo " BAD $name: exit $got (wanted $want)"; grep -E '^(== SUSPECT|PASS|FAIL)' /tmp/mut.out; fi } echo "mutation test of scripts/alexa_date_audit.mjs" run "M1 bare wrong date in body " 1 '**Discontinued as of August 1, 2023.**' run "M2 wrong year, no day " 1 '**Alexa was discontinued in 2023.**' run "M3 wrong date + true date after" 1 '**Discontinued as of August 1, 2023**; alexa.com was retired on 1 May 2022.' run "M4 wrong date in a footnote " 1 '**Discontinued.**((Alexa was shut down on 1 August 2023.))' run "M5 correct Tranco swap sentence" 0 '**Discontinued 1 May 2022.** Tranco removed the Alexa provider from its default list on 1 August 2023.' run "M6 unlisted death verb " 1 '**Alexa was decommissioned in 2023.**' run "M7 verb and year far apart " 1 '**Alexa, the ranking list that this whole page is about, was discontinued by Amazon after a long and well-publicised wind-down, in the month of August 2023.**' run "M8 ceased/went dark phrasing " 1 '**Alexa ceased operations on 1 August 2023.**' run "M9 correct paper-year sentence " 0 '**Discontinued 1 May 2022.** 66 papers published 2023-2026 still name it.' run "M10 split phrasal verb " 1 '**Amazon switched the service off in 2023.**' echo "unmutated:"; node scripts/alexa_date_audit.mjs >/dev/null 2>&1 && echo " OK clean corpus: exit 0"
mutation test of scripts/alexa_date_audit.mjs OK M1 bare wrong date in body : exit 1 (wanted 1) OK M2 wrong year, no day : exit 1 (wanted 1) OK M3 wrong date + true date after: exit 1 (wanted 1) OK M4 wrong date in a footnote : exit 1 (wanted 1) OK M5 correct Tranco swap sentence: exit 0 (wanted 0) OK M6 unlisted death verb : exit 1 (wanted 1) OK M7 verb and year far apart : exit 1 (wanted 1) OK M8 ceased/went dark phrasing : exit 1 (wanted 1) OK M9 correct paper-year sentence : exit 0 (wanted 0) OK M10 split phrasal verb : exit 1 (wanted 1) unmutated: OK clean corpus: exit 0
13.7 The one thing that did change: design:longitudinal
The sweep found no wrong date but one imprecise one, on a page that tells the reader the Tranco default list is not a fixed frame. longitudinal's provider-set table samples 1 January of each year, and the sentence under it read “and then in 2023–24 Alexa was replaced by CrUX and Cloudflare Radar” — the interval between two sampled rows, presented as the date of the change. Its footnote gave Amazon's API retirement (15 December 2022) as the cause but never gave the swap date, so the page had the two events adjacent and neither of them dated to the day. That is the shape the original error grew out of.
Fixed in three saves, the second and third prompted by reviewers:
| Rev | –if-rev | What changed |
|---|---|---|
| 1788465991 | 1787798449 | Sentence names 1 August 2023. Footnote states that the table's 1-January rows bracket the swap rather than dating it, quotes Tranco's front page for the date, and separates Tranco's date from Amazon's — alexa.com 1 May 2022, APIs 15 December 2022, eight months earlier. |
| 1788467004 | 1788465991 | The external-currency reviewer queried an endpoint nobody had: api/lists/date/20230731 returns list W9ZN9, providers: [alexa, umbrella, majestic, farsight], listPrefix: full. The next day's 25299 is [crux, farsight, majestic, radar, umbrella] at listPrefix: 1000000. Both rows added to the table, so the swap is now two consecutive list ids one day apart rather than a bracket — a sharper version of the point the section exists to make. Re-verified before saving, with 30 July (99P42, pre-swap) and 2 August (8289V, post-swap) as controls. |
| 1788467915 | 1788467004 | Generic-reviewer wording fixes: all three provider transitions dated in the sentence (Quantcast after 1 April 2020, Farsight 1 May 2022, Alexa 1 August 2023) instead of only the Alexa one; the 430-character sentence split in three; the 31 July / 1 August pair told once in the footnote instead of twice; “the five year-boundary rows” corrected to “the other five rows” (5 Nov 2019 and 26 Aug 2026 are not year boundaries); and the footnote no longer sends the reader to tranco “for the dated list ids” as though W9ZN9 were there — it says the pre-swap id is only on this page. |
Notes: longitudinal §17.
Structure checked after the second save, counted from the wiki source (unambiguous) and confirmed in the rendered DOM:
| Source | Rendered | |
|---|---|---|
| Headings | 32 (^\s*={2,6}[^=] markers) | — |
| Footnotes | 11 (((…)) pairs) | 11 distinct fn__N ids |
| Tables | — | 14 <table> |
| Red links | — | 0 wikilink2 |
| Citations | 45 markers, 32 distinct keys | 90 bibtex_citekey spans (two per marker), 32 numbered references |
Two of those numbers were wrong in this section's draft — “34 headings, 33 footnotes” — and two independent reviewers caught them. 34 came from grep -c '<h[2-5]' over the whole HTML, which includes the sidebar and page-tools chrome; 33 came from counting class=“fn” / fn_top / fn_bot occurrences, and DokuWiki emits three per footnote, so 11 read as 33. Same trap as the bibtex plugin's two spans per citation marker. Rule for the next run: a rendered-DOM count is only a check if you know how many elements the renderer emits per source construct — count from the source, and use distinct fn__N ids to confirm it rendered.
13.8 Decisions a reasonable person could have made differently
| Decision | Why |
|---|---|
| Did not edit website_selection | Its Alexa bullet, its status table and its lede already carry 1 May 2022 / 15 December 2022, and its footnote already explains what 1 August 2023 actually is. Editing a correct page to close a work item would be theatre. |
| Did not remove the wiki's record of the old error | Eight passages still quote the wrong date in order to say it was wrong — three on content pages (website_selection's Alexa footnote and its §Use in Publications paragraph, and website_classification's Alexa footnote) and five on provenance pages (three on sampling; and on this page, the “Neighbour correction” row in §1 and the rejected-date row in §7 — not §13.7, which uses 1 August 2023 as a correct fact rather than quoting the error). They read as corrections, not as claims, and they are what stops the next run re-deriving the error. Kept, and the audit classifies them as ok-history rather than suppressing them. |
| Edited longitudinal rather than filing it | One line, same defect family, and the page is one of the three that discuss the provider swap. Filing it would have left the wiki with a precise date on two pages and a vague one on the third. |
| Left the corpus figures alone | The corpus figures on the affected pages (463 papers naming Alexa, the 66 post-retirement papers) were neither touched nor re-derived; they belong to §3–§6 above and to sampling. A date audit is not a reason to refresh them. |
Left the W9ZN9 row off tranco | That page's id table would be the natural second home for the pre-swap id, and a reviewer said so. It is a different content page with its own provenance page, and adding a row there is a separate edit with its own review obligation, so it is filed as work rather than smuggled in. The longitudinal footnote says plainly that the id is only there. |
13.9 What could not be established
- Where the “1 August 2023” figure entered the wiki. The oldest revision of website_selection that carries it is the January 2025 human version (rev 1735841558 created the page; the sentence survives untouched until 1786570786). Whether the original author read it off Tranco's front page and mis-attributed it, or got it from a third-party listicle, is not recoverable from the revision history. The Tranco-swap explanation is the only 1 August 2023 event anyone has found, so it remains the most likely origin, but it is inference.
- Whether an Alexa-adjacent Amazon service was in fact retired in 2023. Not exhaustively excluded — the sampling sitting checked AWS's own full-shutdown listing and found no Alexa entry, and this sitting found no primary source either, but “no source exists” is a claim a search cannot close. What can be said is what the pages now say: the Alexa ranking service and its APIs ended in 2022.
- Closed during review. This section originally said no page gave the last pre-swap Tranco list id, only the 2023-08-01 one. The external-currency reviewer simply queried
api/lists/date/20230731and got it —W9ZN9, still[alexa, umbrella, majestic, farsight]. It is now on longitudinal (§13.7). Recorded here because “what could not be established” was, in this case, only what nobody had tried: the endpoint was already documented on tranco and cost onecurl.
13.10 Review log
Four reviewers, all told the context might not be exhaustive and all handed the same bundle: the edited longitudinal footnote, alexa_date_audit.mjs with its output, the mutation script with its output, and this section in draft.
Pass 1 — figures vs script (Claude Sonnet)
| Finding | Verdict |
|---|---|
“34 headings” and “33 footnotes” in §13.7 are wrong. Source has 32 headings and 11 ((…)) pairs; the rendered DOM has 11 distinct fn__N ids. 33 = 11 × the three CSS-class occurrences DokuWiki emits per footnote. | Accepted. §13.7 now counts from the source, states both numbers, and explains where the 3× came from. Re-counted after the second save. |
scripts/alexa_date_audit-output.txt was, transiently, the output of a different script — the pre-rewrite whole-line version (“98 sentences”, a different verdict vocabulary) sitting next to the unit-based script that replaced it. Resolved before publication, but a real instance of exactly the drift this pass exists to catch. | Accepted, and recorded rather than quietly fixed. The output file is regenerated by the same command every time and was re-diffed against a live run immediately before publishing. See §13.11. |
Mutation the script missed: Alexa was decommissioned in 2023. → exit 0, PASS. The DEATH vocabulary had no entry for decommission / cease / go dark / take offline / kill off. | Accepted. Vocabulary widened from 6 verbs to 25; mutation added as M6, and a second phrasing as M8. Both fail correctly now. |
Mutation the script missed: a verbose but ordinary sentence placing “discontinued” ~76 characters from “2023” → exit 0, PASS. PROX = 60 was too narrow for real prose. | Accepted. PROX raised to 200, which is wider than any sentence on this wiki and deliberately shifts the work onto the excuse rules; added as M7. |
CORPUS_SLUG is effectively dead code — 0 hits under the current population. | Accepted in part. Kept, and so is the new ok-literature-year rule, which is also at 0. Both are now listed in §13.3 with their zeros: a rule at zero is visible residue, a deleted rule is a future SUSPECT nobody expected. Removing them would make the report shorter and the next run's diff harder to read. |
| “Eight passages still quote the wrong date” — §13.7 is not one of them, so the total may be 7. | Rejected on the count, accepted on the wording. The five provenance passages are three on sampling plus §1 and §7 of this page; the reviewer read “§7” as “§13.7”. The cell now names both rows explicitly so the ambiguity is gone. |
stripUrls offsets, nested-footnote truncation, population separation of the 463/66-paper figures | Confirmed sound, no change. |
Pass 2 — citations and quotes (Claude Sonnet)
| Finding | Verdict |
|---|---|
| All five external sources fetched, HTTP 200, every quoted string present verbatim in the response body; the extended ellipsis quote on biases elides correctly. | Confirmed, no change. |
| Attribution direction correct in both places and on all six cross-referenced pages: 1 May 2022 / 15 December 2022 → Amazon, 1 August 2023 → Tranco. No reversed attribution anywhere. | Confirmed, no change. |
No new citekeys and no bibliography entries in either the draft or the page edit; literature:bibliography reachable and untouched. | Confirmed, no change. |
| Live website_selection is byte-identical to the cached copy the audit ran against, so the sweep is not stale relative to the wiki. | Confirmed — a check worth more than it looks, since every mutation test runs against that cache. |
| Out of brief: the audit script throws on a non-200 export rather than handling it. | Rejected as a defect, accepted as intentional. Crashing on a failed export is the house rule — a sweep that silently skips a page it could not fetch would report a clean wiki it never read. |
Pass 3 — external currency (Claude Sonnet)
| Finding | Verdict |
|---|---|
No primary source for a 2023 Alexa retirement exists. Searched Wikipedia's article and its cited sources (which give “Discontinued (as of May 1, 2022)” and the category “Products and services discontinued in 2022”), contemporary news coverage, and AWS's own full_shutdown_services.html — which contains no occurrence of “Alexa” at all. | Confirmed independently of the sampling sitting, which reached the same conclusion from the other direction. This is the load-bearing negative for the whole item. |
api/lists/date/20230731 → W9ZN9, providers [alexa, umbrella, majestic, farsight] — dates the swap to the day and closes an open item this page had just declared unclosable. | Accepted and published. Re-verified by me with the neighbouring days, and added to longitudinal as rev 1788467004. The best finding of the review layer, and it cost one curl. |
| “33 footnotes” wrong; actual 11 (found independently of Pass 1). | Accepted — same fix. |
| “34 headings” off by one or two depending on where the content region is cut. | Accepted, and answered differently than asked. Rather than pick a rendered-region convention, §13.7 now publishes the source count (32) as the primary number, because it is reproducible without agreeing where DokuWiki's chrome ends. |
Tranco currency: five-provider default list live (api/lists/date/latest), per-provider licences on the methodology page unchanged, “averaging all four rankings” still stale on the front page, GVWK and XVWN still resolve with the provider sets longitudinal claims, alexa.com → alexa.amazon.com/about in 2 hops, support.alexa.com NXDOMAIN. | Confirmed 2026-09-03, no change. The stale “four rankings” phrase is already footnoted on tranco and was left alone. |
Post-publication self-check (no reviewer; me)
| What | Outcome |
|---|---|
| Re-ran the audit against the published wiki, including this page. It reported 16 SUSPECT, all of them this section's own mutation fixtures, quoted work item and review log. | Real defect in the script's scope, not in the wiki. Fixed by stripping <code>/<file> blocks and routing provenance: hits to a printed-but-not-failed bucket. Both changes and what they cost are in §13.3.1. Found by re-running after publishing rather than before — which is the only order in which it can be found. |
Re-read all 37 units of the post-publication run by hand, including the 11 provenance-note, 9 beyond-window and 1 corpus-slug. | All correct; no content page dates Alexa's retirement to 2023. |
| Re-ran the 10-mutation suite after the scope change. | All 10 behave; the fixtures live in a shell script, and the mutation is written into the page body, so block-stripping does not blunt the test. |
Re-ran the audit once more after saving §13.3.1: 39 units, 0 SUSPECT, 12 provenance-note, 1 corpus-slug, 9 beyond-window. Every unit added since the previous run is this section's own prose. | Confirmed. The count moves with each save and the invariant does not — which is why §13.5 publishes the invariant and labels the tally. |
| Verified the four published listings against the committed files by extracting them from the rendered page. | alexa_date_audit_mutations.sh, both *-output.txt byte-identical; alexa_date_audit.mjs identical after stripping trailing whitespace (DokuWiki pads blank lines inside a <file> block with one space). 186 lines, 0 non-whitespace differences. |
Pass 4 — generic, no checklist (Claude Fable)
The hardest of the four, and the only one that found a defect in the audit's own population.
| Finding | Verdict |
|---|---|
Blocking. The “site-wide” sweep was not site-wide. core.listPages at its default depth 1 returns 88 of 160 pages; scripts/dw.mjs called it with no arguments, so the sweep covered 158 pages and the section was about to claim “every page id”. The two missed pages were provenance:privacy:browser_extensions and provenance:privacy:privacy_sandbox. The script's shipped header also mis-diagnosed the cause as “core.listPages returns a subset”. | Accepted. Verified: core.listPages('', 0) → 160. dw.mjs now defaults to depth 0 for every caller; the audit re-run covers 160 (87 content, 73 provenance) and still finds 0 SUSPECT; the header comment and every count in §13.3 and §13.11 are corrected. Neither missed page mentions Alexa, so the conclusion held — but the claim would have been false. |
Blocking. Heading levels. 13.11 was ===== so it rendered as a sibling of §13, and the review passes were level-equal with §13.10. | Accepted. 13.x are ====, passes are ===, matching the neighbouring provenance pages. |
Blocking. The provenance:design:longitudinal amendment was stale — one rev instead of three, “one-line change”, the wrong 34/33 DOM counts, and a whole paragraph about an open item that had since been closed. | Accepted. Rewritten from scratch against the final page, and its publication order stated. |
Blocking. (pass 4 pending) placeholder would have shipped. | Accepted — this table is what replaced it. A previous sitting shipped an unfilled placeholder and the reviewer was right to assume nothing. |
“9 correctly attribute 2023 to Tranco's provider swap” is false for about 4 of the 9 — several are publication-year sentences excused only because “Tranco” sat in the ±90 window, and ok-literature-year reads 0 because TRANCO is tested first. | Accepted. §13.3 now says the verdict is a triage bucket, that only ~5 of the 11 describe the swap, and that the zero is partly an ordering artefact. This is the “probe hits are not a shared claim” failure and the reviewer named it correctly. |
“PROX = 200 is wider than any sentence on this wiki” is false — the edited sentence alone was ~430 characters. | Accepted, and answered with a measurement. The widest verb–year pair the wiki contains inside one unit is 188 chars, the narrowest excluded is 215. More importantly the script now reports a beyond-window category with each gap, so the constant is auditable instead of load-bearing; 5 residue units on this run, all read. |
| §13.6 arithmetic wrong (“six must exit 1 and three must exit 0”); the mutation script's shipped header still described only M1–M5; “M8 ceased/went dark” tests only “ceased”. | Accepted. Eight must exit 1, two are controls. Header rewritten; M8's label kept but the header now names the phrasing it actually uses. |
| A false-negative shape the suite lacked: split phrasal verbs. website_classification's own footnote says “switched the service off”, which the contiguous verb list could not see. | Accepted, and it changed the result. DEATH now matches split phrasal forms; that footnote is now a hit (correctly excused), the sweep went from 11 units to 18, and M10 locks it in. The best kind of finding: a real gap, demonstrated on real text. |
| “No primary source exists” (§13.2) contradicts “a search cannot close that claim” (§13.9). | Accepted. §13.2 now says “was not found”, names the three independent negatives, and points at §13.9. |
| “That sentence had already been removed twice over” is wrong — 1786570786 removed it; 1787840075 added a disclaimer and removed nothing. | Accepted. §13.1 is now a four-row revision table that says which revision did what. |
| Small stated numbers: “25 death verbs” (the regex has 29 alternatives); the script header said “~30 pages mention one or the other” against §13.3's 71. | Accepted, both corrected. |
design:longitudinal wording: “the five year-boundary rows” (5 Nov 2019 and 26 Aug 2026 are not); the 31 Jul / 1 Aug pair told twice in one footnote; a ~430-character body sentence; the other two provider transitions still bracketed by 1-January rows while only Alexa's was dated; and the footnote pointing at tranco “for the dated list ids” when W9ZN9 is not there. | All accepted, in rev 1788467915 — see the third row of §13.7. |
Add the W9ZN9 row to tranco so the pre-swap id is not on one page only. | Rejected as part of this sitting, filed as work. A different content page, with its own provenance page and its own review obligation; slipping a row in unreviewed is how pages drift. The footnote now states the id is only on longitudinal, so the reader is not misled in the meantime. |
| Proportion: ~38 KB appended for a one-line fix; cut to ~25 KB. Named specific duplications — “no corpus query” ×3, the 34/33 error ×3, the PROX fix ×4, method rationale in both the script header and §13.3. | Accepted in part. The duplications are cut: the lede carries the no-corpus-query note once, §13.2's stale-“four rankings” paragraph is gone, the 34/33 story is one paragraph with a rule, §13.11's bullets are a table, and two §13.8 rows that restated the lede are gone. Rejected: the target. The section grew rather than shrank, because the review layer found four defects in it and their record is the most useful thing here. The code and outputs stay by house rule. |
| “Two things went wrong … both worth more than the item itself” is self-regarding; each confession told twice, each ending in a bolded maxim, where the neighbouring page gives one rule per finding. | Accepted. §13.11 is a two-column table: what happened, and the rule it earns. One each. |
| “published” used to mean “draft handed to reviewers” — on this wiki “published” means live, so a reader would conclude a broken audit went live. It did not. | Accepted. “first draft” throughout. |
| “Claude Code (Opus 5)” vs the neighbours' “Claude Opus 5”. | Accepted. |
| Not a reviewer finding, but it belongs next to them: re-running the audit after publishing §13 turned up 16 SUSPECT, all self-inflicted. | Recorded in §13.3.1 and in the post-publication self-check below. Worth noting that none of the four reviewers could have caught this, because the page they reviewed was not yet part of the audit's input. |
13.11 The run, and three mistakes inside it
| Item | Value |
|---|---|
| Date | 2026-09-03 |
| Wiki at the time | 160 pages: 87 content, 73 provenance: |
| Models | Claude Opus 5 for the audit, the fix and these notes; three sonnet reviewers (figures / citations / external currency) and one fable reviewer (generic) |
| Pages edited | longitudinal (revs 1788465991, 1788467004, 1788467915), longitudinal §17, and this page. website_selection not edited — it was already correct |
| Scripts added | scripts/alexa_date_audit.mjs, scripts/alexa_date_audit_mutations.sh, scripts/mkprov_alexa_date.py; scripts/dw.mjs fixed (listPages depth) |
| Bibliography | untouched, no citekeys added |
| Filed as separate work | the W9ZN9 row on tranco |
Three things went wrong in the making of this, and the rules they earn are worth more than the date was:
| What happened | Rule for the next run |
|---|---|
A committed *-output.txt was, for a few minutes, the output of a different script — the audit had been rewritten from line matching to unit matching and the output file beside it still held the old verdict vocabulary. A reviewer caught it in that window. Nothing went live wrong; the file was regenerated and re-diffed against a live run before saving. | Rewriting a script and re-running it are two actions, and only the second one keeps the audit trail honest. mkprov_alexa_date.py reads the output file rather than a transcript, so a stale file shows up as a diff. |
A patch to the DEATH vocabulary silently did nothing: it used a Python str.replace() whose search string did not match the file, and replace returns the string unchanged with no error. The script was widened in the draft's prose while still narrow in fact, and the mutation suite printed BAD on exactly the mutations the patch was meant to fix. | A silent no-op is indistinguishable from success, so an in-place edit needs assert old in s before it writes. Every later edit in this sitting did; one of them fired and saved a second wrong header. |
The sweep called core.listPages with no arguments and got 88 of 160 pages, then described itself as covering “every page id”. The default depth is 1. | A page list you did not count is not a page list. Print the size of every enumeration next to the union, and check the wrapper's defaults before trusting a remote API's idea of “all”. |
← back to the content page · the longitudinal amendment this prompted · corpus-level provenance
