User Tools

Site Tools


programming:tranco

This is an old revision of the document!


Tranco

Tranco is a research-oriented ranking of popular domains, with a permanent id for every generated list and an HTTP API that returns both the list and the configuration that produced it. It was introduced at NDSS 2019 [1Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)] to replace single-source top lists that disagreed with each other, moved daily, and could be manipulated. This page is the API and construction note that Website selection points at. It is not a second copy of that page, and it is not a sampling tutorial — those live on Sampling and Longitudinal.

Cite the list id, not “the Tranco top 1M”. https://tranco-list.eu/list/GVWK still resolves in 2026 for a list generated on 5 November 2019. A date is not enough: the default provider set has changed since 2019, so two “Tranco top 1M” crawls on either side of 1 August 2023 are different instruments. Read the configuration out of the API, not out of the 2019 paper. The project's own front page still says to average “all four rankings”; the live daily list of 26 August 2026 is built from five providers.1)

What this page is for

  • Here: how the default list is built today, which knobs the generator actually exposes, how permanent ids work, and which HTTP calls to make.
  • Website selection: why you would pick Tranco rather than CrUX, Umbrella, or a retired Alexa snapshot, and what DNS-based inputs do to a sample.
  • Sampling: how to draw from a pinned list (including a rank-stratified sampler that already speaks Tranco ids).
  • Longitudinal: why pinning an id per wave still does not make 2019 and 2026 lists commensurable — the provider set moved under the id scheme.

What to read first

  • [1Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)] — why a single-source top list is a bad sampling frame, and the original design (Dowdall over 30 days, filters, permanent ids). Read it for the problem, not for today's inputs.
  • [2Le Pochat, Victor; Van Goethem, Tom; Joosen, Wouter (2019): "Evaluating the Long-term Effects of Parameters on the Characteristics of the Tranco Top Sites Ranking", in: 12th USENIX Workshop on Cyber Security Experimentation and Test (CSET 19). (Link)] — CSET 2019, outside this corpus. One year of lists, and what the generator's parameters actually do to stability and composition. That is the paper for a custom list.
  • [3Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] — measured the lists against Cloudflare's server-side HTTP request logs in February 2022. CrUX was the most accurate; Tranco is an aggregate, not a ground truth. CrUX was folded into Tranco's default list on 1 August 2023, so the comparison has not been re-run on the instrument you download today.

The default list, today

Queried on 2026-08-27, GET /api/lists/date/latest returned:

Field Value on list 46W9X
Generated 2026-08-26T22:00:01 UTC
Providers CrUX, Farsight, Majestic, Cloudflare Radar, Cisco Umbrella
Combination Dowdall
Window 2026-07-28 – 2026-08-26 (30 days)
Prefix 1,000,000
Pay-level domains only on
Download https://tranco-list.eu/download/46W9X/1000000
Permalink https://tranco-list.eu/list/46W9X

The same day's subdomains list is a different id (567QN). filterPLD is not on that configuration. Do not treat “the daily list” and “the daily list with subdomains” as prefixes of each other.

GET https://tranco-list.eu/top-1m-id returns the same id as /api/lists/date/latest (here 46W9X). GET /latest_list redirects (HTTP 303) to /list/<id>/1000000. A dated lookup for today returns 404 until that 22:00 UTC generation has run — on 27 August 2026, /api/lists/date/20260827 was 404 while /latest was yesterday's list.2)

The rows of the pay-level-domain daily list are registrable domains, ranked by position, with no scheme and no host. The subdomains list is a different file: its rows can be hostnames below the registrable domain. Turning a row into a URL a browser visits is a further sampling decision — see Sampling.

How the list is built

The generator takes one ranking per selected provider per day in the window, optionally truncates each to a prefix, then scores every remaining domain with the Dowdall rule (1, 1/2, …, 1/N) and sorts on the total [1Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)]. Where a provider publishes ranks in buckets rather than a total order, the live generator first converts each bucket to a virtual rank at the geometric mean of its bounds.3) Borda (N, N−1, …, 0) is the other combination method the API accepts; the default list uses Dowdall. Dowdall is the one that matches a Zipf-like traffic distribution; Borda treats the tail of a million-row list almost as if it were the head [1Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)].

Two input lists are published in buckets, not as a total order:

  • CrUX — 1k / 5k / 10k / 50k / 100k / 500k / 1M / 5M / 10M / remainder. Origins, normalised by Tranco to a subdomain. Monthly. Browser page-loads from opted-in Chrome users.
  • Cloudflare Radar — 200 / 500 / 1k / 2k / 5k / 10k / 20k / 50k / 100k / 200k / 500k / 1M, except the top 100 which is individually ranked. Weekly, top 100 daily. DNS to 1.1.1.1, pay-level domains.

The other current default inputs are closer to a total order: Majestic (backlinks, mostly pay-level domains plus a few famous subdomains), Cisco Umbrella (OpenDNS query volume, any name, so typos and *.ec2.internal survive), and Farsight (passive DNS cache-misses, pay-level domains, default list only — see below). Full provider notes are on Tranco's methodology page; describe a list you actually downloaded from its configuration object, not from that page, which still says it averages “all four providers” after introducing five.

Providers: what changed

Read from the dated API, not from memory. Each row is the daily list for that date:

Date List id Providers Prefix
2019-11-05 GVWK Alexa, Umbrella, Majestic, Quantcast full
2022-01-01 XVWN Alexa, Umbrella, Majestic full
2022-05-01 W9339 Alexa, Umbrella, Majestic, Farsight full
2023-08-01 25299 CrUX, Farsight, Majestic, Radar, Umbrella 1,000,000
2026-08-26 46W9X same five 1,000,000

Quantcast dropped out after 1 April 2020; Farsight joined on 1 May 2022; Alexa was replaced by CrUX and Cloudflare Radar on 1 August 2023, which is also when the default prefix stopped being full and became one million.4) You cannot resurrect a retired provider. A custom list can still name alexa and quantcast in the API schema — they are in the documented enum — but they are not in the default list, and Farsight, which is in the default list, is not in that enum.

Farsight is default-list-only. The homepage says so (“only for the default list”), and the API documentation's Configuration.providers enum is crux | majestic | radar | umbrella | alexa | quantcast. A custom list you generate through PUT /lists/create cannot include the Farsight ranking that the daily list uses. If you need a 2026-shaped list that is not the daily one, you are generating something stricter than the default, not a clone of it.

Licences on the current default inputs, as Tranco states them: Umbrella free of charge; Majestic CC BY 3.0; CrUX CC BY-SA 4.0 on Tranco's methodology page; Cloudflare Radar CC BY-NC 4.0; Farsight used only in the default list. Google's own CrUX methodology (fetched 2026-08-27) licences the datasets CC BY 4.0, not CC BY-SA — use that when you cite CrUX directly. Radar's non-commercial clause is the one that bites if you ship the default list as part of a product. Tranco is not affiliated with the providers.

The public GitHub repository (DistriNet/tranco-list) still has DEFAULT_TRANCO_CONFIG as Alexa + Umbrella + Majestic + Quantcast, listPrefix: full. Last push: 2020-03-05. It is the 2019 generator, not the live service. Do not read the current default out of that file.

Configurable and hardened variants

The 2019 paper's pitch was not only the default ranking but a generator whose filters match the study [1Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)]. The CSET follow-up is the empirical guide to those knobs [2Le Pochat, Victor; Van Goethem, Tom; Joosen, Wouter (2019): "Evaluating the Long-term Effects of Parameters on the Characteristics of the Tranco Top Sites Ranking", in: 12th USENIX Workshop on Cyber Security Experimentation and Test (CSET 19). (Link)]. The API still exposes them on PUT /lists/create (account required):

Knob What it does
providers Which source rankings to include. See the Farsight trap above.
startDate / endDate The averaging window. Default is 30 days ending at generation.
combinationMethod dowdall (default) or borda.
listPrefix Integer, or full. Each source list is truncated before scoring. The daily list is 1000000.
filterPLD Keep pay-level domains only. On for the daily list.
inclusionDays + value Drop names that were not present for at least this many days of the window.
inclusionLists + value Drop names that were not present in at least this many of the selected provider lists. This is the original “hardened” filter.
filterTLD / filterTLDValue Keep only named TLDs.
filterOrganization One domain per organisation.
filterSubdomain / value Keep only named subdomains.
filterSafeBrowsing Drop Google Safe Browsing hits.
filterCRUX + month/type/value Restrict to a CrUX slice (global / country / region / subregion). This is the free, citable way to build a country list; see Website classification.

Almost nobody in this corpus names those knobs. Of the 266 papers that used Tranco as a population source, a full-text sweep finds 0 mentioning inclusionDays/inclusionLists or “custom list”, 1 mentioning Dowdall, 2 mentioning Borda. A sweep for “hardened” is not usable: it fires on hardened binaries, hardened TLS, and so on. The field adopted the default daily list, not the configurable generator. If you do generate a custom list, cite the id: it stores the generator's configuration. It does not by itself record the prefix you downloaded, the hash of those bytes, or the draw you took from the file.

Creating a custom list is rate-limited to one list generated concurrently per account and requires Basic Auth (email + API token from the account page). Unauthenticated PUT /lists/create returns 401, as does GET /api/auth/test.5)

Permanent ids

Every generated list has a short id, a permalink https://tranco-list.eu/list/<id>, and a download https://tranco-list.eu/download/<id>/<n> that returns the first n rows as CSV. The id is the thing to cite. The 2019 id GVWK still returns “available”: true and the configuration above.

What the id does not freeze is the meaning of “Tranco” across years. It freezes one generated file. For a single crawl that is exactly what you want. For two waves, cite one id per wave and say the provider set changed; there is no configuration that makes a 2019 list and a 2026 list the same instrument. That argument, with the same API rows, is on Longitudinal.

Of the 266 papers in this corpus that name Tranco as a population source, 159 (59.8%) state some version or date and 23 (8.6%) cite a permanent list id. The 23 were found by two wide full-text probes over those 266 papers, then read; the probes' union was 37 candidates and 14 were false positives (legal section numbers, bibliography codes). The rate has not moved much since 2020. Permanent ids have been available since the 2019 paper.

Fetching a list

The download is CSV, no header, one rank,domain pair per line, Content-Type: text/csv, Content-Disposition: Attachment;filename=tranco_<id>.csv. A prefix of 20 from 46W9X on 2026-08-27 began:

1,google.com
2,cloudflare.com
3,gstatic.com
4,facebook.com
5,microsoft.com

/download/<id>/<n> returns the first n rows. For the current daily list the object itself names prefix 1,000,000, so /download/<id>/1000000 is the file to hash. The convenience URL /top-1m.csv.zip still exists; prefer the id-addressed download so the bytes you hash are the bytes you can name.

The official Python package is tranco 0.8.1 (PyPI, 2024-04-02). It still works: on 2026-08-27, Tranco(cache=True, cache_dir='.tranco').list() returned id 46W9X, rank(“google.com”) == 1, and rank(“not.in.ranking.example”) == -1. The package requires on-disk caching; the HTTP API does not. The README's configure(…) example still names alexa as a provider — that is a 2021-shaped custom list, not the 2026 default. There is a contributed Go client (WangYihang/tranco-go-package, last push 2026-07-01, unarchived).

Google BigQuery tables tranco.daily.daily and tranco.list_ids.list_ids are linked from the front page. This sitting did not query them; if you use them, cite the list id the list_ids table gives you for that date, not “BigQuery's Tranco”.

The script below talks to the same API the package wraps, writes the configuration next to a hashed prefix, and prints a methods sentence. It does not draw a sample — use the sampler on Sampling for that.

pin_tranco.py
#!/usr/bin/env python3
"""Pin a Tranco list: resolve its id, print the live configuration, download a
prefix, and emit a methods sentence a reader can paste.
 
This is the API-shaped counterpart of the stratified sampler on
design:sampling. It does not draw a sample. It records which list you used.
 
    python3 pin_tranco.py
    python3 pin_tranco.py --list-id GVWK --prefix 1000
    python3 pin_tranco.py --date 2026-08-26 --prefix 10000
    python3 pin_tranco.py --subdomains
 
Writes pin_tranco.manifest.json next to the downloaded CSV. The manifest's
`cite` line is the thing to put in a methods section.
"""
from __future__ import annotations
 
import argparse
import csv
import hashlib
import io
import json
import sys
import urllib.error
import urllib.request
import zipfile
from datetime import datetime, timezone
 
UA = "measuretheweb-pin-tranco/1.0"
API = "https://tranco-list.eu/api"
DOWNLOAD = "https://tranco-list.eu/download/{list_id}/{prefix}"
PERMALINK = "https://tranco-list.eu/list/{list_id}"
 
 
def fetch(url: str, method: str = "GET") -> tuple[bytes, str]:
    req = urllib.request.Request(url, method=method, headers={"User-Agent": UA})
    with urllib.request.urlopen(req, timeout=180) as resp:
        return resp.read(), resp.headers.get_content_type()
 
 
def decode_list(raw: bytes) -> list[tuple[int, str]]:
    if raw[:2] == b"PK":
        with zipfile.ZipFile(io.BytesIO(raw)) as zf:
            text = zf.read(zf.namelist()[0]).decode()
    else:
        text = raw.decode()
    ranked: list[tuple[int, str]] = []
    for row in csv.reader(io.StringIO(text)):
        if not row:
            continue
        ranked.append((int(row[0]), row[1]))
    return ranked
 
 
def main() -> int:
    ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
    ap.add_argument("--list-id", default=None, help="pin this id; default is today's daily list")
    ap.add_argument("--date", default=None, help="YYYY-MM-DD or YYYYmmdd; mutually exclusive with --list-id")
    ap.add_argument("--prefix", type=int, default=1000, help="how many rows to download (default 1000)")
    ap.add_argument("--subdomains", action="store_true", help="daily list that retains subdomains")
    ap.add_argument("--out", default="pin_tranco.manifest.json")
    args = ap.parse_args()
    if args.list_id and args.date:
        print("give --list-id or --date, not both", file=sys.stderr)
        return 2
 
    if args.list_id:
        meta_url = f"{API}/lists/id/{args.list_id}"
    elif args.date:
        compact = args.date.replace("-", "")
        q = "?subdomains=true" if args.subdomains else ""
        meta_url = f"{API}/lists/date/{compact}{q}"
    else:
        q = "?subdomains=true" if args.subdomains else ""
        meta_url = f"{API}/lists/date/latest{q}"
 
    try:
        raw_meta, _ = fetch(meta_url)
    except urllib.error.HTTPError as e:
        print(f"metadata {meta_url}: HTTP {e.code} {e.read()[:200]!r}", file=sys.stderr)
        return 1
    meta = json.loads(raw_meta.decode())
    if not meta["available"]:
        print(f"list {meta['list_id']} is not available: {json.dumps(meta)}", file=sys.stderr)
        return 1
 
    list_id = meta["list_id"]
    cfg = meta["configuration"]
    download = DOWNLOAD.format(list_id=list_id, prefix=args.prefix)
    raw, content_type = fetch(download)
    ranked = decode_list(raw)
    if len(ranked) != args.prefix:
        print(
            f"expected {args.prefix} rows from {download}, got {len(ranked)}",
            file=sys.stderr,
        )
        return 1
 
    cite = (
        f"We used the Tranco list {list_id} "
        f"({PERMALINK.format(list_id=list_id)}), "
        f"providers {', '.join(cfg['providers'])}, "
        f"{cfg['combinationMethod']} over {cfg['startDate']} to {cfg['endDate']}, "
        f"prefix {cfg['listPrefix']}, "
        f"filterPLD={cfg.get('filterPLD')}. "  # SafeAccess: KeyError 'filterPLD' — Tranco API omits the key on the subdomains daily list (567QN) — 2026-08-27
        f"Downloaded the top {args.prefix}; SHA-256 {hashlib.sha256(raw).hexdigest()}."
    )
    manifest = {
        "source": "tranco",
        "fetched_at": datetime.now(timezone.utc).isoformat(),
        "meta_url": meta_url,
        "list_id": list_id,
        "permalink": PERMALINK.format(list_id=list_id),
        "download": download,
        "content_type": content_type,
        "bytes": len(raw),
        "bytes_sha256": hashlib.sha256(raw).hexdigest(),
        "rows": len(ranked),
        "head": [{"rank": r, "domain": d} for r, d in ranked[:5]],
        "configuration": cfg,
        "created_on": meta["created_on"],
        "cite": cite,
    }
    with open(args.out, "w", encoding="utf-8") as handle:
        json.dump(manifest, handle, indent=2)
        handle.write("\n")
    csv_path = f"tranco_{list_id}_top{args.prefix}.csv"
    with open(csv_path, "w", encoding="utf-8", newline="") as handle:
        writer = csv.writer(handle)
        writer.writerows(ranked)
    print(json.dumps(manifest, indent=2))
    print(f"wrote {args.out} and {csv_path}", file=sys.stderr)
    return 0
 
 
if __name__ == "__main__":
    sys.exit(main())

Real output, python3 pin_tranco.py –prefix 20, run on 2026-08-27 against list 46W9X. The five head objects are reflowed to one line each here; everything else is verbatim, including the SHA-256 of those 20 CSV rows (325 bytes):

{
  "source": "tranco",
  "fetched_at": "2026-08-27T09:58:31.796271+00:00",
  "meta_url": "https://tranco-list.eu/api/lists/date/latest",
  "list_id": "46W9X",
  "permalink": "https://tranco-list.eu/list/46W9X",
  "download": "https://tranco-list.eu/download/46W9X/20",
  "content_type": "text/csv",
  "bytes": 325,
  "bytes_sha256": "8c2e47ccc986d74471e2244d5449d524cd6f37fdd58ef9b7cad5f2a2a64d7a1c",
  "rows": 20,
  "head": [
    {"rank": 1, "domain": "google.com"},
    {"rank": 2, "domain": "cloudflare.com"},
    {"rank": 3, "domain": "gstatic.com"},
    {"rank": 4, "domain": "facebook.com"},
    {"rank": 5, "domain": "microsoft.com"}
  ],
  "configuration": {
    "isDailyList": true,
    "filterTLD": "false",
    "startDate": "2026-07-28",
    "filterPLD": "on",
    "combinationMethod": "dowdall",
    "listPrefix": "1000000",
    "endDate": "2026-08-26",
    "providers": [
      "crux",
      "farsight",
      "majestic",
      "radar",
      "umbrella"
    ]
  },
  "created_on": "2026-08-26T22:00:01.860640",
  "cite": "We used the Tranco list 46W9X (https://tranco-list.eu/list/46W9X), providers crux, farsight, majestic, radar, umbrella, dowdall over 2026-07-28 to 2026-08-26, prefix 1000000, filterPLD=on. Downloaded the top 20; SHA-256 8c2e47ccc986d74471e2244d5449d524cd6f37fdd58ef9b7cad5f2a2a64d7a1c."
}

The HTTP API

Base URL https://tranco-list.eu/api/. Documented as alpha. Verified 2026-08-27.

Endpoint Auth What you get
GET /lists/id/{id} no Metadata: list_id, available, download, created_on, configuration. 404 if the id is unknown.
GET /lists/date/{YYYYMMDD\|latest} no Same object for the daily list of that date. Optional ?subdomains=true.
GET /ranks/domain/{domain} no Daily ranks for at least the past 30 days. Rate limit 1 query/second. google.com returned 38 days, all rank 1.
PUT /lists/create Basic Auth (email + API token) Submit a Configuration. 401 without credentials; 429 if you already have a list generating.
GET /auth/test Basic Auth Credential check. 401 without them.

Unauthenticated calls are enough for every workflow that uses the daily list. You only need an account to generate a custom one.

The ranks endpoint is for “where is this name today / this month”, not for building a sample. Building a sample is a download of a prefix.

Use in publications

Everything below is a claim about seven venues — CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P, 2010–2026, 5,859 extracted papers. CSET, EuroS&P, ACSAC and SOUPS are absent, so a paper that “everyone cites for Tranco parameters” can sit outside the counts. See Corpus.

Population: any paper whose extracted population[].sourceList matches /\btranco\b/i. That is 266 papers (4.5% of the corpus, 15.0% of the 1,120 that ran a crawl). 262 of them have a web-unit Tranco tuple (websites / domains / web-pages); 4 do not (documents, mobile apps, IP addresses, “other”). The page counts papers, not tuples (373 Tranco tuples sit on those 266 papers). The same query, and the same 23-id allowlist, is the Tranco block on Longitudinal.

Exact-string counting is not optional here: 96 distinct spellings, and writing sourceList == “Tranco” finds only 180 of the 266 (a 32.3% undercount). The fold is the regex; residue is empty by construction. The spelling list is on tranco.

Tranco is a 2019 artefact and the corpus shows it. Zero papers before 2020, then 7, 28, 29, 48, 56, 76, 22* through 2026. As a share of web-unit papers it overtook Alexa in 2023 (47 of 122, 38.5%, against Alexa's 30 of 122) and was 75 of 133 (56.4%) in 2025. 2026 is provisional — CCS and IMC have not been held, IEEE S&P and WWW abstracts are under-selected — and is not a complete-year trend. CrUX as a named frame remains far smaller (12 of 133 web-unit papers in 2025): researchers who want CrUX's traffic model often get it inside Tranco rather than as a separate download.

Venue Tranco papers Share of that venue
PETS 44 8.6% of 510
IMC 46 7.2% of 638
NDSS 32 4.6% of 701
USENIX Security 56 4.0% of 1,410
TheWebConf 32 3.8% of 843
CCS 32 3.2% of 990
IEEE S&P 24 3.1% of 767

PETS and IMC are where a Tranco methods section is most likely to be read carefully; USENIX Security is where the most such papers are.

How they sample. Of the 266, 215 (80.8%) mark a Tranco population as top-n, 31 random, 23 stratified, 18 a pre-existing dataset. 256 (96.2%) state a size. Among 362 sized tuples, the round prefixes 1k / 10k / 100k / 1M appear on 19 / 49 / 28 / 70 tuples. Of those 166, 147 are marked top-n — those are the ones that are actually a Tranco head. The other 19 are random (10), stratified (3), purposive (2), a pre-existing dataset (3), or seed-and-crawl (1). Tuples larger than 1M are not a Tranco prefix — they are unions with zone files, CT logs, Common Crawl, or Citizen Lab lists, and the extracted n is the union. Do not read the table's “more than 1,000,000” row as evidence that Tranco ships a list of hundreds of millions of names.

The schema almost never records Tranco as a tool. tools[].name matches on 27 papers (26 “Tranco”, 1 “Tranco 1M”, all used); 240 of 266 Tranco users (90.2%) have no tool-field hit. A ranking list is treated as a sampling frame, which is correct. Full text mentions Tranco in 332 papers; 58 of those are in neither the sourceList population nor tools[] — typically a related-work citation or a filter that removed Tranco names from a malicious-domain set. They are not added to the 266.

What to report

A methods section that a later crawl can reconstruct:

  1. The list id and the permalink. Optionally the date, but the id is the primary key.
  2. The prefix you actually visited (top 10k of list 46W9X, not “the Tranco top 10k”).
  3. Whether you used the pay-level-domain daily list or the subdomains list (different ids).
  4. If you generated a custom list: say so, and the id is enough because it stores the configuration. If you used the default, say that too — “Tranco” in 2026 is five providers and a 1M prefix, not the 2019 four-provider full list.
  5. The SHA-256 of the bytes you downloaded, or a copy in the artefact. Ids are permanent today; that is not a promise that the archive is immortal.
  6. For more than one wave: one id per wave, and a sentence that the default provider set is not invariant. See Longitudinal.
  • Website selection — which ranking to pick, and the DNS-list caveats that still apply to parts of Tranco.
  • Sampling — drawing from a pinned list; the stratified sampler already takes a Tranco id.
  • Longitudinal — provider-set changes as a comparability problem.
  • CrUX — the page-load ranking that is now inside the default list. Not yet written.
  • Cloudflare Radar — the other 2023 addition. Radar's own buckets are coarser than a Tranco rank.
  • Corpus — seven-venue scope and the 2025–2026 edge.

Queries, the unedited report, the spelling list, quote checks, and every rejected source are on tranco.

Methodology and limitations of these figures

Corpus figures come from scripts/report_tranco.mjs against data/extract/run1 (5,859 papers). The population is a regex over free-text sourceList, which is the right fold. The exact string “Tranco” undercounts that population by 32.3%. Live API facts were re-fetched on 2026-08-27; they will move tomorrow. The 23-id count is a hand audit, not a regex. 2025–2026 rows are starred because those venue-years are incomplete by construction.

[1]
Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)
[2]
Le Pochat, Victor; Van Goethem, Tom; Joosen, Wouter (2019): "Evaluating the Long-term Effects of Parameters on the Characteristics of the Tranco Top Sites Ranking", in: 12th USENIX Workshop on Cyber Security Experimentation and Test (CSET 19). (Link)
[3]
Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
1)
Checked 2026-08-27. GET https://tranco-list.eu/api/lists/date/latest returned list id 46W9X with providers: [crux, farsight, majestic, radar, umbrella]. The phrase “all four rankings” is still on tranco-list.eu, one occurrence, next to a sentence that correctly names five providers.
2)
All four URLs fetched 2026-08-27. Generation time is the created_on timestamp on the list object, 22:00:01 UTC.
3)
https://tranco-list.eu/methodology, fetched 2026-08-27: “Where ranks are bucketed … we first normalize each bucket to the geometric mean of its boundaries.”
4)
Tranco's own front page, “Updates”, and the dated API objects above, all checked 2026-08-27.
5)
Both checked 2026-08-27 without credentials.
You could leave a comment if you were logged in.
programming/tranco.1787840080.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki