Table of Contents
Provenance: design:website_selection
Working notes behind website_selection — every query with its population and denominator, the report script and its unedited output, the vendor fold and its residue, the Farsight hand map, the quotes that were checked, the external sources that were verified or rejected, and what could not be established. Corpus-level caveats are on corpus and are not restated here.
Contemporaneous. Written during the run that produced the refresh, 2026-08-27, not reconstructed afterwards. Cursor drain sitting (item 184, run 64), not Claude Code.
1. What this page is backing
| Item | Value |
|---|---|
| Content page | website_selection — extending, not creating. Live rev at start of sitting: 1787832045, 10,124 bytes. One remaining <WRAP todo> (Farsight). Radar TODOs had already been filled by the Cloudflare Radar sitting. |
| Report script | scripts/report_website_selection.mjs |
| Fold it depends on | scripts/rank_fold.mjs (multi-label ranking vendors; not sample_fold.mjs, which folds frame kind) |
| External re-fetch | scripts/external_checks_website_selection.sh (exit 0, 2026-08-27T13:58:02Z) |
| Data | data/extract/run1/extractions.jsonl, 5,859 papers, 7 venues, 2010–2026 |
| Bibliography additions | none — every citekey was already in bibliography |
| Neighbour correction | website_classification Alexa footnote no longer says this page gives 1 August 2023 as Alexa's death. CrUX licence one-liners on that page and tranco now match Google CC BY 4.0 vs Tranco's still-stated CC BY-SA 4.0. |
2. Scope: extending, not a new page
The page already existed as a catalogue of ranking services with two borrowed survey figures. The item asked to close remaining TODOs and replace those figures with a corpus “Use in Publications”. sampling already covers how to draw, and already publishes the Alexa → Tranco year table. This sitting:
- filled Farsight (the last TODO)
- added the vendor-fold Use in Publications, keeping Scheitle/Xie as dated historical surveys
- fixed the leftover “discontinued in 2023” sentence (the list bullet already said 1 May 2022; the survey paragraph did not)
- added a methodology section and this provenance page
- did not absorb sampling's method/size/versioning tables
Judgement call: do not create programming:farsight. The ranking is not independently downloadable; the API detail is the Tranco default-list-only trap, already on tranco.
3. Populations and denominators
| Tag | Definition | N |
|---|---|---|
| corpus | all extraction records | 5,859 |
ALL / sampled | population.length > 0 | 5,712 |
| WEB (page population) | ≥1 population[] tuple with unit ∈ {websites, domains, web-pages} | 1,153 |
| names a source | WEB, ≥1 non-sentinel sourceList | 1,143 |
| vendor-users | WEB, ≥1 web-unit sourceList matching a vendor family | 763 |
sample_fold popularity-ranking | WEB, frame-kind fold | 764 |
sample_fold custom-seed | WEB, frame-kind fold | 257 |
| Farsight/DNSDB as pdns dataset | any unit, hand map | 13 papers / 12 strings |
| Farsight as a ranking | vendor family farsight-ranking | 0 |
The item brief said “population.sourceList over 4,207 sampling papers (custom seed list 377, Google Play 183, Tranco 119, Alexa 117)”. Those counts are exact strings on the old 4,322-paper corpus. On this run the same method over ALL yields custom seed list 516, Tranco 180, Google Play 143, Alexa 51. Google Play is 281 after folding and 0 on WEB. The page publishes that trap rather than those four numbers as a ranking ranking.
Sentinels are silence. Papers, not tuples. A string may hit several vendors.
4. Running it
cd /workspace/artifacts/wiki node scripts/rank_fold.mjs # self-test; prints Farsight pdns strings and unnamed residue; exit 1 on unknown Farsight spelling node scripts/report_website_selection.mjs # every figure node scripts/report_website_selection.mjs --wiki node scripts/report_website_selection.mjs --list node scripts/report_website_selection.mjs --quotes 'tranco' bash scripts/external_checks_website_selection.sh python3 pages/compare_ranks.py --wiki # live Tranco/Umbrella/Majestic ranks; dated snapshot node scripts/check_page_numbers.mjs pages/design_website_selection.txt out/report_website_selection.txt \ '===== Use in Publications =====' '===== What to report =====' node scripts/check_page_numbers.mjs pages/design_website_selection.txt out/report_website_selection.txt node scripts/check_tables.mjs pages/design_website_selection.txt node scripts/check_wrap.mjs pages/design_website_selection.txt
Windowed and whole-page number checks passed after the report grew a Y block (sample_fold 764) and Z lines for CC BY 4.0 / Tranco's still-stated CC BY-SA 4.0, rank 5,000, and 1.1.1.1 — those are on the catalogue half of the page, outside the corpus section. The page prints Google's CC BY 4.0 as the licence to use when citing CrUX.
5. The vendor fold, Farsight hand map, residue
Multi-label on purpose: “Tranco and Cloudflare Radar” is both. sample_fold.mjs is exclusive first-match and answers a different question; do not reuse it as a vendor ranking.
Exact-string undercount on WEB (from the report):
| Vendor | Papers | Exact name | Spellings | Undercount |
|---|---|---|---|---|
| Alexa | 463 | 50 | 355 | 89.2% |
| Tranco | 262 | 178 | 91 | 32.1% |
| CrUX | 36 | 9 | 23 | 75.0% |
| Majestic | 26 | 4 | 17 | 84.6% |
| Cisco Umbrella | 24 | 2 | 23 | 91.7% |
FARSIGHT_PDNS (exact strings, classified as a passive-DNS dataset, not a ranking). 12 strings, 13 papers. The report fails if a new /farsight|dnsdb/i spelling is not in this set:
DNSDB Farsight passive DNS database DNSDB and Project Sonar Data Repository Farsight Security passive DNS DNSDB and 360 PassiveDNS Farsight passive DNS Farsight Security DNSDB Farsight's DNSDB Farsight Security Information Exchange (SIE) / Farsight passive DNS database passive DNS dataset similar to Farsight DNSDB, provided by QiAnXin Company Farsight PDNS Farsight DNSDB
Farsight-as-ranking papers: 0. The ranking exists (Tranco methodology, 1M PLD, since 2022-05-01, cache-misses) and is a default Tranco input; nobody in this corpus names it as a sourceList.
Unnamed top-N residue after tightening isUnnamedTop to \btop\b (a first draft matched “desktop website”):
1 Google AdWords top websites
1 Google's Top 1,000 Most-Visited Websites
1 remaining top websites
1 top 100K websites used in Phase I
Four papers. “Google's Top 1,000 Most-Visited Websites” is a ranking with no vendor family; leaving it in residue is correct. The 764 vs 763 gap is this frame-kind-only paper.
Radar regex requires cloudflare radar or radar (domain )?rank so Tracker Radar does not match. Four WEB papers name Cloudflare Radar, all 2025–2026.
6. Quotes and figures spot-checked
Whitespace-normalised against paper.cols.txt.
| Claim | Source | Result |
|---|---|---|
| HTTP/2 26.6% Alexa Top 1M vs 7.84% com/net/org | [1Scheitle, Quirin; Hohlfeld, Oliver; Gamba, Julien; Jelten, Jonas; Zimmermann, Torsten; Strowes, Stephen D.; Vallina-Rodriguez, Narseo (2018): "A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists", in: Proceedings of the Internet Measurement Conference 2018, pp. 478–493. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] IMC 2018 | exact, prose and table |
| 1,790 Alexa top-10K; 70% lower rank-magnitude bucket; 27.2% two or more orders lower; CrUX most accurate across all metrics | [2Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] IMC 2022 | exact (Cloudflare is split / ligatured in .cols) |
| “top 10,000 sites from the Tranco list of October 29, 2019” | USENIX 2020 a-tale-of-two-headers-… | exact in .cols; the extraction quote prefixes “we decided to analyze” which sits across the column break |
| Alexa.com retired 1 May 2022; APIs 15 December 2022 | Wayback of support.alexa.com | exact |
| Tranco latest 46W9X, five providers, dowdall | live API 2026-08-27 | exact |
Rejected as a publishable finding: counting Alexa listVersion strings dated 2023 or later as “impossible draws”. Sampling already tried this; four hand-checks failed because listVersion is not scoped to the tuple. The published figure is the undated share (34 of 66).
7. External sources
All re-fetched 2026-08-27 by external_checks_website_selection.sh (exit 0).
| Claim | How verified |
|---|---|
| Alexa.com 1 May 2022; APIs 15 Dec 2022 | Wayback 20221126115049 of the support article |
| No Alexa retirement on 1 August 2023 | That date is Tranco's provider swap (CrUX+Radar in, Alexa out). AWS shutdown listing does not mention Alexa (recorded on the sampling provenance). |
| Tranco 46W9X; providers crux, farsight, majestic, radar, umbrella | GET /api/lists/date/latest |
| Farsight ranking since 1 May 2022; cache-misses; 1M PLD | Tranco methodology HTML |
| DomainTools “Mirror, Mirror” names 1 May 2022 and Farsight DNSDB | live blog 200 |
| Umbrella top-1m.csv.zip still published | index 200, zip HEAD 200; example rows include TLD com/net |
| Majestic Million page live | 200 |
| SecRank secrank.cn live | 200 |
| CrUX docs live | 200 |
| Radar /domains 403 to curl | 403; documented on cloudflare_radar |
| Quantcast /measure/ → /publisher/measure | 200 after redirect; not treated as a public ranking |
| SimilarWeb homepage | 200 |
| Farsight not in custom-list providers enum | already on tranco; not re-litigated |
Rejected:
- SEO listicles ranking “best website traffic tools 2026”.
- Wikipedia as a primary source for Alexa retirement (the page still links it as a name disambiguator; the date comes from Alexa Support).
- DomainTools acquisition URL
/resources/blog/domaintools-acquires-farsight-security/→ 404. The Mirror, Mirror post itself says they acquired Farsight; that sentence is used, the dead URL is not.
8. What could not be established
- A public download URL for the Farsight ranking (as opposed to DNSDB). Tranco's methodology describes it; DomainTools does not publish a CSV equivalent of Umbrella's.
- Whether Quantcast Measure still sells a top-sites file behind a login. The public URL is publisher analytics.
- A re-run of Ruth et al. against the five-provider Tranco list. The 2022 comparison predates Radar-in-Tranco.
- Full-text (as opposed to
sourceList) mentions of Farsight-as-ranking. Out of scope; the schema question is what people sampled from.
9. The run
- Date: 2026-08-27. Agent: Cursor (not
claude -p/ drain-sandbox). - Item:
design:website_selection: close the remaining TODOs(claimed as id 184 so a night drain could not takedesign:automated_measurementsinstead). - Reviewers: three focused GPT 5.6 Luna medium passes in parallel, then one generic pass with no checklist, after a freeze. Findings in §10.
- No credentials were echoed. Wiki writes go through
scripts/dw.mjs(JSON-RPC). - Deleted the one-off probe
scripts/_probe_ws.mjsafter folding its questions into the report.
10. Review log
Freeze directory: out/freeze_ws/. Reviewers were handed that snapshot. The content page was not edited while the three focused passes ran. Generic ran on the post-focused-review page.
Model for all four: GPT 5.6 Luna medium (not sonnet/fable as the task spec's default split).
Pass 1 — figures vs script
No findings. Report rerun, fold self-test, windowed and whole-page check_page_numbers, tables and wrap: all passed.
Pass 2 — citations and quotes
| Finding | Verdict |
|---|---|
| Blocking. Majestic limitations attributed a CrUX-relative manipulation cost to [3Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)] (2019). CrUX public rankings post-date that paper. | Accepted. Now: link-graph lists are cheap relative to a toolbar list; that paper did not evaluate CrUX. The same anachronism was in the manipulation paragraph (Le Pochat “expensive relative to toolbar” applied to CrUX) and was fixed in the same pass. |
| “Ruth et al. measured every public list” | Accepted. Now names Alexa, Majestic, Umbrella, Tranco and CrUX. |
| Galloway “bucket” put today's Radar API vocabulary in the 2024 paper's mouth | Accepted. Quote is now “consistently achieved a ranking in the top 100,000”; the bucket reading is attributed to today's API. |
| Alexa “user panel and tracking scripts” uncited | Accepted. Now cited to [1Scheitle, Quirin; Hohlfeld, Oliver; Gamba, Julien; Jelten, Jonas; Zimmermann, Torsten; Strowes, Stephen D.; Vallina-Rodriguez, Narseo (2018): "A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists", in: Proceedings of the Internet Measurement Conference 2018, pp. 478–493. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] as toolbar/panel. |
Le Pochat “single HTTP request” unverifiable here (NDSS 2019 Tranco PDF absent from data/fulltext) | Noted, no change. The claim is on tranco and the 2019 abstract; this sitting did not fetch the NDSS PDF. |
Six live citekeys each occur once in bibliography. Scheitle 26.6%/7.84%, Ruth 1,790/70%/27.2%, Galloway $10/top 100,000: verbatim in .cols.
Pass 3 — external currency
| Finding | Verdict |
|---|---|
| CrUX “CC BY-SA 4.0” and “Help improve Chrome” eligibility | Accepted. Google's methodology (fetched 2026-08-27) licences datasets CC BY 4.0 and lists four user-eligibility criteria. Tranco's methodology page still says CC BY-SA 4.0; the page now says so and prefers Google. Neighbour one-liners on tranco and website_classification updated so the three pages do not disagree. |
| “inputs have already changed twice” — Quantcast drop, Farsight add, Alexa→CrUX+Radar is three | Accepted. |
Alexa Wayback, Tranco 46W9X, Umbrella zip, Majestic, SecRank, CrUX docs, Radar 403, Quantcast redirect, SimilarWeb, DomainTools canonical /blog/ | Confirmed, no change. |
Pass 4 — generic, no checklist
| Finding | Verdict |
|---|---|
| Blocking. Lead “Tranco is current practice; CrUX is current evidence” treats a Feb 2022 comparison as a 2026 recommendation, while the page admits it has not been re-run on five-provider Tranco. | Accepted. Lead is now “Tranco is what 2025 papers use; CrUX is what Ruth et al. found most accurate in 2022” plus the not-re-run sentence. |
| “Tranco-plus-CrUX is the combination the accuracy evidence actually supports” | Accepted. Ruth supports them being different instruments, not a composite design. |
| “CrUX is more stable still” / prefer buckets as if that were Le Pochat's CrUX finding | Accepted. 30-day aggregation stays on Le Pochat; CrUX buckets described without an unmeasured stability rate. |
| “Any adoption number run on a top list is biased upward” | Accepted. Scoped to Scheitle's HTTP/2 (and same-direction IPv6/CDN) comparison. |
| “DNS rankings did not become expensive to fake when Alexa died” | Accepted. Scoped to Radar and Tranco-via-Radar. |
| Alexa lag “is a submission cycle” as a causal finding | Accepted. Now “compatible with a submission cycle; we did not read why”. |
| Undated 34 “are a copy of unknown provenance — a mirror, a cached CSV, or inherited” as if those were established | Accepted. Those are listed as possibilities; the measured fact is that they cannot be pinned. |
| Farsight “sees default-resolver and server traffic” | Accepted. Now cache-miss DNS from participating resolvers, including names only infrastructure looks up. |
| Sampling page still says top-n is “weighted towards sites that matter to users” | Rejected for this sitting. True about a neighbour; out of scope. Recorded so the next sampling refresh sees it. |
| Provenance §10 empty while promising findings | Accepted by writing this table. |
11. Same-day rewrite: catalogue → decision page
After item 184 published, Karel asked to question the page against later LLM-written pages (the model was ip_classification) and decide what of the human-era catalogue to keep. Not a new drain item.
Kept from the human page: the vendor catalogue (CrUX / Tranco / Radar / Umbrella / Majestic / Alexa / Farsight / SecRank / SimilarWeb / Quantcast), the DNS-list caveats, the Le Pochat manipulation figure, the Scheitle and Xie survey figures as dated history, Farsight default-list-only, the Alexa 1 May 2022 vs Tranco 1 August 2023 distinction, Best Practices folded into What to report, the split with sampling.
Dropped or rewritten: numbered 1–6 Advantages/Limitations that restated the table; “the provider's own example file begins 1,com / 2,net” (the 2026-08-27 top-1m.csv does not; that is top-1m-TLD.csv.zip); ec2.internal as a live Umbrella example (absent from that day's top-1m); drain-internal voice about the work item that commissioned the corpus section; treating CrUX-first catalogue order as if it were current practice rather than Ruth's 2022 accuracy ranking.
Added, matching IP classification: a decision table (claim → list); a live rank-disagreement table and published compare_ranks.py (the analogue of classify_ips.py); blockquotes from Ruth, Scheitle, Galloway; a 2026 status table with paper counts; Papers to read first; Open questions; custom-seed 257 vs ranking 764 pointed at sampling rather than re-derived as a vendor ranking; gstatic.com at Tranco rank 3 as the page's “not one of them is the BBC” moment.
compare_ranks.py run 2026-08-27T14:26:07Z against Tranco 46W9X, Umbrella top-1m (1,000,000 hostnames) and Majestic Million. Umbrella TLD file: 10,905 rows, com=1, net=2, org=6.
12. Rewrite review log
Focused freeze: out/freeze_ws_rewrite/ (page MD5 b45a6832…). Content page not edited while the three focused passes ran. Generic freeze: out/freeze_ws_generic/ after the currency wording fix. Model for all four: GPT 5.6 Luna medium.
Pass 1 — figures vs script
No findings. Report, fold, windowed-equivalent whole-page check_page_numbers, tables, wrap: passed. Embedded 180 / 1.05 / 1,000 present in Z.
Pass 2 — citations and quotes
| Finding | Verdict |
|---|---|
| Le Pochat “single HTTP request” unverifiable (NDSS 2019 full text absent) | Noted, no change. Same gap as the first sitting. |
| Citekeys, Ruth/Scheitle/Galloway quotes, Galloway not given today's Radar bucket vocabulary, Xie labelled PAM | Confirmed, no change. |
Pass 3 — external currency
| Finding | Verdict |
|---|---|
| Page said “Tranco's methodology still says CC BY-SA 4.0”. Live methodology HTML has no CC BY/CC BY-SA string. | Accepted. The licence line is on Tranco's homepage, not the methodology page. Content, tranco and website_classification now say homepage vs methodology vs Google CC BY 4.0. |
Tranco 46W9X, Farsight paragraph, Alexa Wayback, Umbrella hostnames + TLD file 10,905, Majestic, CrUX CC BY 4.0, Radar 403, Quantcast redirect, SecRank, Farsight not in providers enum, gstatic.com rank 3, compare_ranks.py –wiki | Confirmed 2026-08-27. |
Pass 4 — generic, no checklist
| Finding | Verdict |
|---|---|
| Blocking. Decision table “Pages people actually load → CrUX (or Tranco)” | Accepted. CrUX is the direct choice; Tranco is named as not a substitute. |
| Blocking. “A ranking is a proxy for popularity” while Majestic is backlinks | Accepted. Now “each list ranks an observable”. |
| 2022 Ruth finding reads as a 2026 recommendation | Accepted. Lede is “pick the list that matches the claim”; Ruth labelled 2022 evidence. |
| “Prefer a month of ranks” underspecified | Accepted. Pin one dated frame; Tranco id already is a 30-day Dowdall; CrUX is a month snapshot. |
| “today's Tranco” vs list dated 2026-08-26, fetched 2026-08-27 | Accepted. Prose now uses those two dates. Script function names unchanged. |
| SimilarWeb “paywalled” does not say whether you can pin an export | Accepted. Status row now requires a keepable contract export; no public id. |
| Repeated sampling cross-references | Rejected. The opening split and one What-to-report pointer are the boundary with sampling; cutting them would bury it. |
