| Next revision | Previous revision |
| literature:corpus [2026/08/12 15:44] – New page: dataset-wide provenance for the publication corpus behind the site's figures — venue scope, the selection funnel with a number on every row, provisional years, how the extraction was made, quote groundedness re-measured on the current corpus, an karel.kubicek.claude | literature:corpus [2026/08/12 16:06] (current) – Second review round: name the 138 poster records and the 3 title-level duplicates as structural properties of the funnel; add a boxed statement that an aggregate count cannot be independently recounted; widen the hand-checked quote sample to 2025-2026 and karel.kubicek.claude |
|---|
| ====== The Publication Corpus Behind These Pages ====== | ====== The Publication Corpus Behind These Pages ====== |
| |
| Many pages on this site carry a section headed //Use in Publications// whose figures come from one source: a structured extraction over the full text of **5,859 papers** from seven security, privacy and measurement venues, 2010–2026. This page is that source's methodology. It exists so that every figure on this site can be discounted correctly — you cannot judge "31.4% of papers that measured something say where they measured from" without knowing which papers were eligible to be counted, which were never eligible, and how the 31.4% was produced. | Many pages on this site carry a section headed //Use in Publications// whose figures come from one source: a structured extraction over the full text of **5,859 papers** from seven security, privacy and measurement venues, 2010–2026. This page is that source's methodology. It exists so that every figure on this site can be discounted correctly — you cannot judge "1,228 of the 3,908 papers that took a measurement from a vantage point — 31.4% — say where they measured from" without knowing which papers were eligible to be counted, which were never eligible, and how the 31.4% was produced. |
| |
| It is not a description of a public dataset. **The corpus is not published.** See [[#How to check a figure yourself]] for what that means for you. | It is not a description of a public dataset. **The corpus is not published.** See [[#How to check a figure yourself]] for what that means for you. |
| |
| ^ Stage ^ Papers ^ Lost here ^ Is the loss random or systematic? ^ | ^ Stage ^ Papers ^ Lost here ^ Is the loss random or systematic? ^ |
| | Venue-year metadata records (7 venues × 117 venue-years) | 16,864 | — | — | | | Venue-year metadata records (117 venue-years, across 7 venues) | 16,864 | — | — | |
| | …carrying an abstract | 15,800 | −1,064 | **systematic**, in three unrelated blocks — below | | | …carrying an abstract | 15,800 | −1,064 | **systematic**, in three unrelated blocks — below | |
| | Screened on title + abstract | 15,800 | 0 | none: screening covers every abstract | | | Screened on title + abstract | 15,800 | 0 | none: screening covers every abstract | |
| |
| ^ Era ^ Records ^ Why ^ Will waiting fix it? ^ | ^ Era ^ Records ^ Why ^ Will waiting fix it? ^ |
| | 2010–2019 | 443 | **There is no abstract to fetch.** NDSS archive pages for 2010–2018 link straight to the PDF and publish no abstract, and NDSS has no DBLP DOIs, so neither retrieval path applies (260 papers). Pre-PoPETs PETS workshop schedule pages likewise publish none (126). CCS 2011's DOI-less poster entries sit under a DOI-less parent proceedings (40). | **No** | | | 2010–2019 | 443 | **For 426 of them there is no abstract to fetch.** NDSS archive pages for 2010–2018 link straight to the PDF and publish no abstract, and NDSS has no DBLP DOIs, so neither retrieval path applies (260 papers). Pre-PoPETs PETS workshop schedule pages likewise publish none (126). CCS 2011's DOI-less poster entries sit under a DOI-less parent proceedings (40). The remaining 17 are scattered singles across PETS, USENIX and NDSS, mostly front matter. | **No** for the 426; the 17 are recoverable | |
| | 2020–2024 | 58 | Scattered — mostly one pass where the OpenAlex daily quota ran out mid-run. | Yes, with a re-run | | | 2020–2024 | 58 | Scattered — mostly one pass where the OpenAlex daily quota ran out mid-run. | Yes, with a re-run | |
| | 2025–2026 | 563 | OpenAlex has the DOI before it has the abstract: WWW 2026 (324) and IEEE S&P 2026 (194) dominate. | Yes, in a few months | | | 2025–2026 | 563 | OpenAlex has the DOI before it has the abstract: WWW 2026 (324) and IEEE S&P 2026 (194) dominate. | Yes, in a few months | |
| USENIX Security 2026 alone is 45.2% of the loss, for the simple reason that the conference has not happened yet and its papers are not online. Per venue, the share of selected papers with no PDF is 0% for IEEE S&P, NDSS and PETS, 2.5% for WWW, 4.4% for CCS, 7.2% for USENIX and 7.7% for IMC. | USENIX Security 2026 alone is 45.2% of the loss, for the simple reason that the conference has not happened yet and its papers are not online. Per venue, the share of selected papers with no PDF is 0% for IEEE S&P, NDSS and PETS, 2.5% for WWW, 4.4% for CCS, 7.2% for USENIX and 7.7% for IMC. |
| |
| **The 14 papers with prepared text but no extraction record.** One is a USENIX 2017 slide deck excluded deliberately; the other 13 are IEEE S&P papers, twelve of them from 2024, whose extraction requests ended on ''520 status code (no body)'' and were never retried. Across the whole run 559 papers failed at least once and 546 were recovered by a retry; these 13 are the tail nobody went back for. The effect is a mild under-representation of IEEE S&P 2024 in every figure on this site, worth roughly 1.7% of that venue. | **The 14 papers with prepared text but no extraction record.** One is a USENIX 2017 slide deck excluded deliberately; the other 13 are IEEE S&P papers — eleven from 2024, one from 2023, one from 2021 — whose extraction requests ended on ''520 status code (no body)'' and were never retried. Across the whole run 559 papers failed at least once and 546 were recovered by a retry; these 13 are the tail nobody went back for. The effect is a mild under-representation of IEEE S&P 2024 in every figure on this site, worth roughly 1.7% of that venue. |
| | |
| | ==== What "5,859 papers" actually contains ==== |
| | |
| | Two structural properties of the funnel that nothing above shows, both of which change how a silence figure should be read. |
| | |
| | * **138 of the 5,859 records are posters**, 96 from CCS and 42 from IMC, identifiable by their own slug and title. The venue index does not filter front matter and posters out, and screening did not reject them. 89 of them sit inside the 3,908-paper ''measuredFrom'' population and 15 inside the 1,120-paper ''crawled'' one. A poster is one or two pages with its methodology compressed out, so it is systematically more likely to be scored as "did not state", and it biases exactly the reporting-gap figures this site is built on — upward. 251 records are four pages or fewer. If a silence rate matters to you, exclude short records and see whether it moves. |
| | * **Three papers are in the corpus twice.** The consistency check on duplicates is keyed on ''(venue, year, slug)'' and returns zero, but that key cannot see the duplication the venue index warns about: //Detecting and Characterizing Social Spam Campaigns// is indexed under both CCS 2010 and IMC 2010, //UIScope// under both NDSS 2020 and NDSS 2021, and //DRAWN APART// twice within NDSS 2022, once under an ''auto-draft-242'' slug. Keyed on the normalised title instead, the corpus has 3 extra records. Three in 5,859 changes nothing arithmetically; it matters because "zero duplicates" was being reported by a check that could not have found these. |
| |
| ==== The selection rule, verbatim ==== | ==== The selection rule, verbatim ==== |
| |
| Screening produces one label record per abstract with seven fields. A paper is selected if **any** of them is set: | Screening is itself an LLM pass — ''gpt-5.6-luna'', the same model as the extraction, last run on 2026-08-07. It produces one label record per abstract with seven fields, and a paper is selected if **any** of them is set: |
| |
| <code> | <code> |
| </code> | </code> |
| |
| The instruction to the screening model was explicitly over-inclusive: //"Include a paper if it plausibly matches the criteria — be deliberately over-inclusive, since a later pass reads the full text and a wrongly excluded paper is never seen again."// The rule was re-applied for this page and checked both ways: **every one of the 5,873 retrieved papers falls inside it, and no directory exists outside it.** | The instruction to the screening model was explicitly over-inclusive — verbatim from ''scripts/build_selection_batch.mjs'': //"Include a paper if it plausibly matches the criteria — be deliberately over-inclusive, since a later pass reads the full text and a wrongly excluded paper is never seen again."// The rule was re-applied for this page and checked both ways: **every one of the 5,873 retrieved papers falls inside it, and no directory exists outside it.** |
| |
| What has **never** been measured is whether the screening decisions are //correct//. There is no human-screened control set to score them against, and the labels are model output. A relevant paper wrongly dropped at screening leaves no trace anywhere in this pipeline. | What has **never** been measured is whether the screening decisions are //correct//. There is no human-screened control set to score them against, and the labels are model output. A relevant paper wrongly dropped at screening leaves no trace anywhere in this pipeline. |
| ^ ^ ^ | ^ ^ ^ |
| | What | One structured record per paper, describing what the authors did | | | What | One structured record per paper, describing what the authors did | |
| | Input | The paper's whole prepared text | | | Input | The paper's whole text, as ''paper.cols.txt'' (5,852 records) or, where ''pdftotext'' produced a garbled encoding, OCR'd ''paper.mistral.md'' (7). ''paper.norm.txt'' only gates eligibility. | |
| | Method | **An LLM extraction against a fixed schema — a single pass, one request per paper** | | | Method | **An LLM extraction against a fixed schema — a single pass, one request per paper** | |
| | Model | ''gpt-5.6-luna'', structured outputs, all 5,859 records | | | Model | ''gpt-5.6-luna'', structured outputs, all 5,859 records | |
| Three schema decisions shape every query you can run against it: | Three schema decisions shape every query you can run against it: |
| |
| * **''not-stated'' and ''none-mentioned'' are first-class values.** The schema asks the model to record silence rather than omit it. This is what lets a page report "of the 1,120 papers that ran a crawl, 53 say anything about robots.txt" instead of quietly dropping the other 1,067. | * **Silence is a first-class value.** The schema's sentinels are ''not-stated'', ''none-mentioned'', ''not-applicable'', ''unclear'' and ''unknown'', and the model is asked to record them rather than omit the field. This is what lets a page report "of the 1,120 papers that ran a crawl, 53 say anything about robots.txt" instead of quietly dropping the other 1,067. |
| * **''usedOrMentioned'' on anything citable** — ''used'' (84.3% of tool tuples), ''produced'' (9.4%), ''compared'' (5.0%), ''mentioned'' (1.2%), ''unclear'' (0.1%). A baseline the authors compared against reads as adoption if you skip this filter, and the tuples that are not ''used'' concentrate in exactly the names a page is about. | * **''usedOrMentioned'' on anything citable** — ''used'' (84.3% of tool tuples), ''produced'' (9.4%), ''compared'' (5.0%), ''mentioned'' (1.2%), ''unclear'' (0.1%). A baseline the authors compared against reads as adoption if you skip this filter, and the tuples that are not ''used'' concentrate in exactly the names a page is about. |
| * **The measured results live in ''detection[].prevalence''** — 26,316 of the 27,241 detection tuples (96.6%), across 5,572 papers, carry the paper's own figure. But ''prevalence'' is the model's summary of a result, not a quote from it. Grep the full text for any prevalence figure before publishing it, not just the attached quote. | * **The measured results live in ''detection[].prevalence''** — 26,316 of the 27,241 detection tuples (96.6%), across 5,572 papers, carry the paper's own figure. But ''prevalence'' is the model's summary of a result, not a quote from it. Grep the full text for any prevalence figure before publishing it, not just the attached quote. |
| | ''USENIX/2024/you-can-obfuscate-but-you-cannot-hide…'' | "For each destination, bots can gather their available bandwidth through tools like Pathneck [17]." | **A bug in the checker, not the extraction.** This paper's ''pdftotext'' output is a garbled font encoding; its text was OCR-repaired and the extractor read the repaired file. The first version of the script checked the garbled one | | | ''USENIX/2024/you-can-obfuscate-but-you-cannot-hide…'' | "For each destination, bots can gather their available bandwidth through tools like Pathneck [17]." | **A bug in the checker, not the extraction.** This paper's ''pdftotext'' output is a garbled font encoding; its text was OCR-repaired and the extractor read the repaired file. The first version of the script checked the garbled one | |
| |
| A further 15.9% of the not-found quotes contain an ellipsis, meaning the extraction spliced two parts of a source sentence and marked the join — not verbatim, but not invented either. | Those five were the first entries in the script's not-found list, which is filled in sorted-key order and therefore starts in CCS 2010–2012 — the era where column-splice damage is most likely, so it is a favourable sample and not a random one. **Five more were drawn from 2025 and 2026**, where the not-found counts are largest, and read the same way: |
| | |
| | ^ Paper ^ What actually happened ^ |
| | | ''USENIX/2025/a-framework-for-abusability-analysis…'' | Present, spliced through the bibliography: //"rg/10. [other column] 5281/zenodo.14745290."// | |
| | | ''CCS/2025/a-sea-of-cyber-threats…'' | Present, with a participant table interleaved mid-sentence: //"open coding methodology [95], refin-P10 35 Chief Mate … ing the codebook"// | |
| | | ''USENIX/2025/addressing-the-address-books…'' | Present twice; the extraction dropped the "(90.5%)" parenthesis's surrounding clause | |
| | | ''NDSS/2026/a-causal-perspective-for-enhancing-jailbreak…'' | Present; the extraction reconstructed a URL broken across a line, and ''causal-learn'' and ''py-why'' are both in the text | |
| | | ''NDSS/2026/aliens-among-us…'' | **Not a quote at all.** The extraction wrote //"The paper does not report ethics review, affected-party notification, regulator contact…"// — a sentence about the paper's silence rather than a sentence from the paper | |
| | |
| | That last one is a different failure mode from layout damage, and it is the only one worth worrying about. It is also rare: **12 of the 135,025 quotes** describe a paper's silence instead of quoting it, concentrated in ''ethics'' and ''classification''. A further 15.9% of the not-found quotes contain an ellipsis, meaning the extraction spliced two parts of a source sentence and marked the join — not verbatim, but not invented either. |
| |
| A practical consequence for anyone checking a quote by hand: **normalise whitespace before concluding a quote is absent.** ''grep -F'' on a multi-word quote fails about half the time against this text because it still wraps mid-sentence; ''tr -s '[:space:]' ' ' | grep -iF'' finds them. | A practical consequence for anyone checking a quote by hand: **normalise whitespace before concluding a quote is absent.** ''grep -F'' on a multi-word quote fails about half the time against this text because it still wraps mid-sentence; ''tr -s '[:space:]' ' ' | grep -iF'' finds them. |
| |
| The failures are also spread thin rather than concentrated: 1,143 of 5,859 papers (19.5%) have at least one, and the ten worst papers hold only 82 of the 1,708. The rate is flat across sixteen years, between 0.8% and 1.5% in every single year, which is what you would expect of a PDF-layout artefact and not of a model that got worse or better over time. | The failures are also spread thin rather than concentrated: 1,143 of 5,859 papers (19.5%) have at least one, and the ten worst papers hold only 82 of the 1,708. The rate is flat across all seventeen years, between 0.8% and 1.5% in every single one, which is what you would expect of a PDF-layout artefact and not of a model that got worse or better over time. |
| |
| **Two honest caveats about this measurement.** It reproduces the dataset's own audit of the earlier corpus almost exactly — that audit reported 57.9% / 37.0% / 4.2% / 0.9% on the same four buckets — but the two implementations are independent and the "fragmented" test in particular is an approximation of the original, so the 4.0%/1.3% split between the last two rows should be read as one combined 5.3% that is mostly layout damage. And the first measurement of this ran the check against the wrong file for seven papers whose text had been OCR-repaired, which put one paper at the top of the not-found ranking with 29 phantom failures; the script now reads whichever rendering the extractor read. | **Two caveats about this measurement.** It reproduces the dataset's own audit of the earlier corpus almost exactly — that audit reported 57.9% / 37.0% / 4.2% / 0.9% on the same four buckets — but the two implementations are independent and the "fragmented" test in particular is an approximation of the original, so the 4.0%/1.3% split between the last two rows should be read as one combined 5.3% that is mostly layout damage. And the first measurement of this ran the check against the wrong file for seven papers whose text had been OCR-repaired, which put one paper at the top of the not-found ranking with 29 phantom failures; the script now reads whichever rendering the extractor read. |
| |
| ==== Field stability: measured once, on the earlier corpus, never since ==== | ==== Field stability: measured once, on the earlier corpus, never since ==== |
| |
| ^ Run-to-run agreement ^ Fields ^ How to use them ^ | ^ Run-to-run agreement ^ Fields ^ How to use them ^ |
| | **90–100%** | has-an-artifact-link (100), ''crawlConfig'' fired (99), ''.statefulness'' (98), ''.interactionDepth'' (97), ''participants'' fired (96), ''isEmpirical'' (95), ''legal.law'' (95), ''humanAnnotation'' fired (94), ''.consentAction'' (93), ''artifacts.availability'' (90) | Publish the percentage. | | | **90–100%** | has-authors'-own-link (100), ''crawlConfig'' fired (99), ''.statefulness'' (98), ''.interactionDepth'' (97), ''participants'' fired (96), ''isEmpirical'' (95), ''legal.law'' (95), ''humanAnnotation'' fired (94), ''.consentAction'' (93), ''artifacts.availability'' (90) | Publish the percentage. | |
| | **65–80%** | ''platforms'' (77), ''ethics.robotsTxt'' (76), ''temporal.mode'' (69), ''vantage.infrastructure'' (69), ''ethics.reviewOutcome'' (68), ''statistics.kind'' (68), ''ethics.notifiedAffectedParties'' (67) | Publish, with the caveat that a repeat run moves it a few points. | | | **65–80%** | ''platforms'' (77), ''ethics.robotsTxt'' (76), ''temporal.mode'' (69), ''vantage.infrastructure'' (69), ''ethics.reviewOutcome'' (68), ''statistics.kind'' (68), ''ethics.notifiedAffectedParties'' (67) | Publish, with the caveat that a repeat run moves it a few points. | |
| | **under 60%** | ''classification.method'' (58), **''studyTypes'' (57) — the least stable field in the schema** | Report as a ranking or a rough share. Never a precise figure. | | | **under 60%** | ''classification.method'' (58), **''studyTypes'' (57) — the least stable field in the schema** | Report as a ranking or a rough share. Never a precise figure. | |
| Four rules, each with the failure it prevents. If a page on this site breaks one, the figure is wrong. | Four rules, each with the failure it prevents. If a page on this site breaks one, the figure is wrong. |
| |
| **1. Name the denominator before the numerator.** Each query defines its own population: | **1. Name the denominator before the numerator.** Each query defines its own population. The table below is the one legitimate use of 5,859 as a denominator — it exists to show how rarely 5,859 is the right one: |
| |
| ^ Population ^ Papers ^ Share of 5,859 ^ | ^ Population ^ Papers ^ Share of 5,859 ^ |
| | ''legal'' — assessed compliance with a law | 402 | 6.9% | | | ''legal'' — assessed compliance with a law | 402 | 6.9% | |
| |
| **2. A sentinel is never an answer.** ''ethics.robotsTxt'' over the 1,120 crawling papers: 53 papers (4.7%) state a real value; 992 (88.6%) hold //some// value, the rest of which is ''not-stated''. Counting the sentinel turns "one crawling paper in twenty says anything about robots.txt" — which is the finding — into "seven in eight do". | **2. A sentinel is never an answer.** The five sentinels are ''not-stated'', ''none-mentioned'', ''not-applicable'', ''unclear'' and ''unknown''. Worked example — ''ethics.robotsTxt'' over the 1,120 crawling papers: 53 papers (4.7%) state a real value; 992 (88.6%) hold //some// value, the rest of which is ''not-stated''. Counting the sentinel turns "one crawling paper in twenty says anything about robots.txt" — which is the finding — into "seven in eight do". |
| |
| **3. Count papers, never tuples.** A paper naming EasyList three times is one paper: 145 ''EasyList'' tuples across 94 papers, 1.54 per paper. | **3. Count papers, never tuples.** A paper naming EasyList three times is one paper: 145 ''EasyList'' tuples across 94 papers, 1.54 per paper. |
| **A 42.6% undercount, on the exact quantity such a page is about**, and the middle row is what the dataset's own convenience query reports. 308 distinct raw strings fold to the United States: ''US'', ''USA'', ''U.S.'', ''California'', ''New York'', ''Oregon'', ''Los Angeles'', ''US East Coast'', ''Silicon Valley'', ''US-East'' and 298 more. Case-and-punctuation folding does not touch any of them. | **A 42.6% undercount, on the exact quantity such a page is about**, and the middle row is what the dataset's own convenience query reports. 308 distinct raw strings fold to the United States: ''US'', ''USA'', ''U.S.'', ''California'', ''New York'', ''Oregon'', ''Los Angeles'', ''US East Coast'', ''Silicon Valley'', ''US-East'' and 298 more. Case-and-punctuation folding does not touch any of them. |
| |
| Every fold written for a page on this site therefore returns its **unmapped residue**, and the residue is printed in full on that page's ''provenance:'' page. A residue that exists only inside a local script's output is a residue nobody will ever look at. Folds also age: one fold on this site went from 5 unmapped strings to 59 with no code change at all, purely because the 2025–2026 papers introduced new subjects. | Every fold written for a page on this site therefore returns its **unmapped residue**, and the residue is printed in full on that page's ''provenance:'' page. A residue that exists only inside a local script's output is a residue nobody will ever look at. Folds also age. When the corpus was extended to 2025–2026, [[privacy:fingerprinting]]'s subject fold went from 7 unmapped tuples to 59 **with no code change at all** (''scripts/report_fingerprinting.mjs'', not this page's script) — the newer papers fingerprint new things (DPI boxes, LLMs, AR/VR apps). It was extended and is back to 7. A fold that does not print its residue would have absorbed the difference silently. |
| |
| ===== 8. How to check a figure yourself ===== | ===== 8. How to check a figure yourself ===== |
| **What you can check.** Every figure on a corpus-backed page traces to a sentence in a published paper, and the ''provenance:'' page for that content page gives you the query, the population, the denominator, the fold, and the report script's unedited output. Two things follow that you can act on without any access to the corpus: | **What you can check.** Every figure on a corpus-backed page traces to a sentence in a published paper, and the ''provenance:'' page for that content page gives you the query, the population, the denominator, the fold, and the report script's unedited output. Two things follow that you can act on without any access to the corpus: |
| |
| - **Read the paper.** The corpus is public literature. If a page says //n// papers do something and names them, the papers are the evidence, not the dataset. | - **Check the arithmetic and the population.** If a percentage's denominator is not named on the page or on its ''provenance:'' page, that is a defect in the page and worth reporting. So is a share of a population the page never defines. |
| - **Check the arithmetic and the population.** If a percentage's denominator is not named on the page or on its ''provenance:'' page, that is a defect in the page and worth reporting. | - **Read the papers a page names.** Where a page names its papers — the worked examples, the tool tables, the "read these first" lists — the papers are public and are the evidence. The dataset is not. |
| | |
| | **What you cannot check.** The corpus itself — the venue index, the screening labels, the retrieved PDFs and ''extractions.jsonl'' — is a local dataset on the maintainer's machine. It is not published, not downloadable, and there is no API. The largest consequence is the plainest one, and it is not softened anywhere else on this page: |
| | |
| | <WRAP important> |
| | **An aggregate count on this site cannot be independently verified.** When a page says 182 papers crawled from the United States, the list of those 182 exists only in the unpublished extraction. You can check that the denominator is named, that the fold is described, that the arithmetic holds and that the handful of quotes the ''provenance:'' page spot-checked are real. You cannot recount the 182. Weigh these figures accordingly: they are a documented and internally audited measurement, not a reproducible one. |
| | </WRAP> |
| |
| **What you cannot check.** The corpus itself — the venue index, the screening labels, the retrieved PDFs and ''extractions.jsonl'' — is a local dataset on the maintainer's machine. It is not published, not downloadable, and there is no API. Two further consequences are worth being blunt about: | Two further consequences follow: |
| |
| * **The extraction is not reproducible even in principle.** It is generative model output; two runs over identical text disagree, which is the whole point of the stability table in [[#Field stability: measured once, on the earlier corpus, never since]]. Re-running it would produce a similar but not identical dataset. | * **The extraction cannot be reproduced to the same values.** It is generative model output; two runs over identical text disagree, which is the whole point of the stability table in [[#Field stability: measured once, on the earlier corpus, never since]]. Re-running it would produce a similar but not identical dataset. |
| * **The extraction currently exists on one disk with no archive.** That is a known, unresolved risk recorded in the dataset's own runbook, and it is stated here rather than left out because a reader deciding how much to lean on these figures should know it. | * **The extraction currently exists on one disk with no archive.** That is a known, unresolved risk recorded in the dataset's own runbook, and it is stated here rather than left out because a reader deciding how much to lean on these figures should know it. |
| |
| ===== 9. The report script and its output ===== | ===== 9. The report script and its output ===== |
| |
| Every figure on this page comes from one script, which prints all of them with their denominators, plus the checks that the funnel is internally consistent. The output below is what it printed on 2026-08-12, unedited. | Every figure on this page comes from one script, which prints all of them with their denominators, plus the checks that the funnel is internally consistent. Its unedited output, as printed on 2026-08-12, is in the [[#Appendix: the report script's full output|appendix]] at the foot of this page and is downloadable from there. |
| |
| <code bash> | <code bash> |
| </code> | </code> |
| |
| On 2026-08-12 the guard reported exactly **two** figures on this page that the report cannot produce, both from named sources outside the corpus: the **2,870** papers whose token ledger was lost, which comes from the extraction's own ''README.md'', and the **29** phantom quote failures caused by a bug in an early version of this page's own script, described in [[#11. Run log]]. Every other figure on the page, including all four quote-groundedness percentages and the whole funnel, is produced by the script. | Two figures on this page are not produced by the script, and both name their source: the **2,870** papers whose token ledger was lost, which comes from the extraction's own ''README.md'', and the **29** phantom quote failures caused by a bug in an early version of this page's own script, described in [[#11. Run log]]. Every other figure on the page, including all four quote-groundedness percentages and the whole funnel, is produced by the script. |
| |
| The script prints all four checks that the funnel is internally consistent — retrieved papers outside the selection rule, extraction records with no text on disk, duplicate records, abstracts never screened — and all four return zero. A nonzero value in any of them means the funnel above is wrong, not that the corpus is. | What the guard actually reported on 2026-08-12 is worth stating exactly, because it says as much about the guard as about the page. It flagged six figures: the **2,870**, and five digit fragments of quoted paper text (a DOI, a percentage and a participant-table row inside the quotes in [[#Quote groundedness, re-measured on this corpus|section 6]]). It did **not** flag the **29** — because 29 also happens to appear in the report as the count of crawling papers that discuss robots.txt. |
| | |
| | That is the shape of this kind of guard. It asks whether a number //appears// in the report, so it gives a false pass whenever a stale figure collides with an unrelated live one, and it is noisy wherever a page quotes someone else's numbers. It is a floor on staleness checking, not a proof of freshness, and its output has to be read rather than expected to come back clean. |
| | |
| | The script also prints seven consistency checks on the funnel — retrieved papers outside the selection rule, fulltext directories outside it, extraction records with no prepared text on disk, duplicate records, labelled papers with no abstract, abstracts never screened, and truncated extractor inputs. All seven return zero. A nonzero value in any of them means the funnel above is wrong, not that the corpus is. |
| | |
| | |
| | ===== 10. What this page could not establish ===== |
| | |
| | Stated plainly, because a provenance page that overstates its own rigour is worse than none. |
| | |
| | * **Whether the screening is accurate.** No human-screened control set exists. A relevant paper wrongly dropped at abstract screening leaves no trace in any artefact, so the false-negative rate of selection is unknown and unknowable from what is on disk. This is the largest unquantified risk in the funnel and it sits at the widest step of it. |
| | * **Whether field stability has moved.** The agreement figures in §6 are from the 4,322-paper corpus. Re-measuring means a second full extraction pass; it has not been done, and every page quoting those figures should say so. |
| | * **Whether the 14 unextracted papers matter.** Thirteen are IEEE S&P, eleven of them from 2024. Nobody has read them to see whether they would have changed anything. |
| | * **Whether the 309 papers in the nine empty venue-years matter.** The //cause// is established and documented — those venue pages publish no abstracts — but nobody has read the 309 titles to see how much relevant methodology is sitting outside the corpus. That is a cheap check nobody has done. |
| | * **The per-run token and cost ledger.** The extraction's own accounting was lost for roughly the first 2,870 papers of the original run because the ledger was written only on clean exit and every restart killed the process first. The total is a measured tail plus an extrapolation, and it is not re-derivable. |
| | * **Anything about a venue that is not one of the seven.** This is not a limitation that better tooling fixes; it is the scope. |
| | |
| | ===== 11. Run log ===== |
| | |
| | ^ ^ ^ |
| | | Page written | 2026-08-12 | |
| | | Corpus at the time | 5,859 papers, 7 venues, 117 venue-years, 2010–2026 | |
| | | Script | ''scripts/report_corpus.mjs'' (new for this page) | |
| | | Model | Claude Opus 5 | |
| | | Figures carried over from earlier notes | **None.** Every number was re-derived from the artefacts on disk; the dataset's own ''README.md'' still quotes the 4,322-paper figures and was not used as a source. | |
| | | Verified independently | The 333-of-780 IEEE S&P figure, from ''missing_ieee_all.jsonl'' holding exactly 447 records (780 − 447 = 333). The selection rule, by checking that all 5,873 retrieved papers fall inside it and none outside. The ''name_fold'' row of the folding table, against the dataset's own ''site_queries.mjs --page vantage'', which reports the same 367. | |
| | | Not verified | The screening labels' accuracy; the field-stability figures, which are reproduced from the earlier run and labelled as such. | |
| | | Convention settled here | ''provenance:'' pages carry no ''~~DISCUSSION~~'' block and no bibliography — comments belong on the content page. This page keeps a discussion block because it is reader-facing and linked from [[:start]], and cites no papers, so it has no bibliography either. | |
| | | Why there is no ''provenance:literature:corpus'' | This page //is// a provenance page: sections 9 to 11 are its own working log. A provenance page for the provenance page would recurse without adding anything. | |
| | | Wired into | [[:start]]; the six ''provenance:'' pages, which already link here; and the //Methodology and limitations of these figures// section of each corpus-backed content page, where the generic "these venues are absent, 2025–2026 are provisional" text was replaced by a pointer here plus the page-specific consequence. | |
| | | Mistakes caught before review | The quote check initially read ''paper.cols.txt'' for all papers, including the 7 whose extractor input was OCR text; that scored a broken font encoding as 29 fabricated quotes in a single paper and put it top of the not-found ranking. Fixed by keying on each record's own ''textSource''. | |
| | |
| | ==== Review, 2026-08-12 ==== |
| | |
| | Three reviewers (Claude Sonnet 5), each given the page, the script and its output, and each told explicitly that the summary they were given might not be exhaustive. What they found, and what was done about it — the rejections matter as much as the fixes. |
| | |
| | ^ Reviewer ^ Finding ^ Verdict ^ |
| | | Figures vs script | **''report_corpus.mjs'' §9c printed hardcoded literals** — ''13 / 780'', ''104 / 230'' and ''4.0% + 1.3% = 5.3%'' — in the very block whose job is to make derived figures auditable. All three happened to be correct against the current data, and all three would have gone stale silently on the next corpus growth //while the staleness guard kept passing//, because the guard only asks whether a number appears in the report. | **Accepted.** All three now computed. | |
| | | Figures vs script | "twelve of them from 2024" — the real split of the 13 unextracted IEEE S&P papers is 11 from 2024, 1 from 2023, 1 from 2021. Contradicted by the page's own embedded output. Invisible to the guard because the number was spelled out as a word. | **Accepted.** | |
| | | Figures vs script | "flat across sixteen years" — 2010–2026 is seventeen. | **Accepted.** | |
| | | Claims vs sources | The stability table read ''has-an-artifact-link'', where the source says ''has-authors'-own-link''. Those are different claims: any artifact link at all, versus the authors releasing their own. | **Accepted.** | |
| | | Claims vs sources | The three named causes of the 443 pre-2020 missing abstracts sum to 426, not 443. The remaining 17 are a fourth cause of a different kind — scattered singles, mostly front matter, recoverable in principle — so a flat "waiting will not fix it" was wrong for those. | **Accepted.** The report now prints the residue and the table names it. | |
| | | Claims vs sources | The screening instruction is verbatim from ''build_selection_batch.mjs'', not ''label_papers.mjs'', which carries a similar but distinct rubric. Not an error, but easy for a later editor to mis-attribute. | **Accepted** as a clarification; the script is now named on the page. | |
| | | Reader fit | The page's opening example — "31.4% of papers that measured something" — did not name its denominator, which is the exact failure the page exists to prevent. | **Accepted.** It now names the 3,908. | |
| | | Reader fit | The sentinel rule named two sentinels where the schema has five. A reader trusting this page as the authoritative statement would not know ''unclear'' and ''unknown'' also count as silence. | **Accepted.** | |
| | | Reader fit | The raw output block, which is most of the page by line count, sat between "how to check a figure yourself" and the two sections a sceptical reader most needs. | **Accepted.** It is now an appendix after the run log. | |
| | | Reader fit | The population table's header, "Share of 5,859", reads as the page breaking its own rule two sections after stating it. | **Accepted**, reworded to say why that denominator is the point of that particular table. | |
| | | Reader fit | ''[[#How to check a figure yourself]]'' was reported as a broken anchor, because the heading is numbered "8." and the link is not. | **Rejected.** Checked against the published HTML: DokuWiki's ''cleanID'' strips the leading ''8. '' when it builds the section id, so the heading's id //is// ''how_to_check_a_figure_yourself'' and the link resolves. Both the numbered and unnumbered forms work. The rendered page has no broken anchors and no red links. | |
| | |
| | ==== Second review, 2026-08-12 ==== |
| | |
| | A fourth reviewer (Claude Fable 5) was given no checklist and asked for whatever the first three were not looking for. It found more than they did, and two of its findings are properties of the corpus that nobody had written down. |
| | |
| | ^ Finding ^ Verdict ^ |
| | | **138 of the 5,859 records are posters**, unfiltered by the venue index and not rejected by screening, 89 of them inside the ''measuredFrom'' population. Posters compress their methodology out, so they push every //silence// figure on this site upward — which is the class of figure the site is mostly about. | **Accepted.** Measured, and now a named property of the funnel in [[#What "5,859 papers" actually contains]]. | |
| | | **The duplicate check was blind to the duplication mode the venue index warns about.** Keyed on ''(venue, year, slug)'' it returns zero; keyed on the title, three papers are in the corpus twice. "Zero duplicates" was being reported by a check that could not have found them. | **Accepted.** The script now checks both keys and names the three. | |
| | | The five hand-checked quotes were all from CCS 2010–2012, because the script's not-found list is filled in sorted-key order — a favourable sample presented as if it generalised, while the years with the most failures had none checked. | **Accepted.** Five more were drawn from 2025–2026 and read by hand; one of them turned out to be a different failure mode (a sentence about the paper's silence rather than a quote), which was then measured corpus-wide at 12 of 135,025. | |
| | | §8's "read the paper" was technically true and practically misleading: corpus-backed tables do not name their papers, so **an aggregate count cannot be independently recounted at all** — the biggest item, and it was missing from "what you cannot check". | **Accepted.** It is now the boxed statement in [[#8. How to check a figure yourself]]. | |
| | | The script printed a literal ''input paper.norm.txt'' in §7 of its own output, contradicting the ''textSource'' counts three screens above it. Same defect class as the hardcoded arithmetic the first reviewer caught, but hiding in a string. | **Accepted.** Computed now. | |
| | | "twelve of them from 2024" survived in §10 after being fixed in §2 — the first review log said "Accepted", which read as "fixed everywhere". | **Accepted.** Fixed, and worth recording that a review log can flatter a partial fix. | |
| | | The funnel's first row read "(7 venues × 117 venue-years)", which parses as a product. The screening model was never named. "Two honest caveats" and "worth being blunt about" label the page's own candour. | **All accepted.** | |
| | | The dataset's own ''extract/README.md'' and runbook still carry the dead "IEEE S&P is 43% retrieved" caveat and a stale permanent-gaps table. | **Accepted as a note**: if you follow a citation from this page into those files, their IEEE S&P figures predate the 2026-08-10 repair described in [[#3. IEEE S&P: the worked example of a systematic loss]]. | |
| | | Sibling page [[design:crawling_location]] carries the same "sixteen years" slip that was fixed here. | **Accepted**, fixed on that page. | |
| | |
| | Working notes for individual pages are under ''provenance:''; see [[#The provenance namespace]]. |
| | |
| | ===== Appendix: the report script's full output ===== |
| | |
| | Unedited, as printed on 2026-08-12. It is at the end of the page rather than beside the commands because a reader checking a number wants it and a reader learning the caveats does not. |
| |
| <file text report_corpus-output.txt> | <file text report_corpus-output.txt> |
| NDSS 2010-2018 block: 260 (venue pages publish no abstract, and NDSS has no DBLP DOIs) | NDSS 2010-2018 block: 260 (venue pages publish no abstract, and NDSS has no DBLP DOIs) |
| PETS 2010-2014 block: 126 (pre-PoPETs workshop schedule pages publish no abstracts) | PETS 2010-2014 block: 126 (pre-PoPETs workshop schedule pages publish no abstracts) |
| | CCS 2011: 40 (DOI-less poster entries under a DOI-less parent proceedings) |
| | those three account for 426 of the 443 pre-2020 records; |
| | the remaining 17 are scattered singles across PETS, USENIX and NDSS, mostly front matter, |
| | and are recoverable in principle, unlike the three structural blocks. |
| Causes per block are documented in data/corpus2/README.md §5; the counts here are re-derived. | Causes per block are documented in data/corpus2/README.md §5; the counts here are re-derived. |
| |
| CHECK fulltext directories outside the selection rule: 0 | CHECK fulltext directories outside the selection rule: 0 |
| CHECK extraction records with no paper.norm.txt on disk: 0 | CHECK extraction records with no paper.norm.txt on disk: 0 |
| CHECK duplicate extraction records: 0 | CHECK duplicate extraction records, keyed on (venue, year, slug): 0 |
| | CHECK duplicate extraction records, keyed on normalised TITLE: 3 extra records across 3 titles |
| | CCS/2010/detecting-and-characterizing-social-spam-campaigns | IMC/2010/detecting-and-characterizing-social-spam-campaigns |
| | NDSS/2020/uiscope-accurate-instrumentation-free-and-visible-attack-investigation-for-gui-applications | NDSS/2021/uiscope-accurate-instrumentation-free-and-visible-attack-investigation-for-gui-applications |
| | NDSS/2022/auto-draft-242 | NDSS/2022/drawn-apart-a-device-identification-technique-based-on-remote-gpu-fingerprinting |
| | CHECK extraction records that are posters: 138 {"CCS":96,"IMC":42} |
| | of those, in measuredFrom: 89; in crawled: 15 |
| | records of 4 pages or fewer: 251 |
| CHECK labelled papers with no abstract in metadata: 0 | CHECK labelled papers with no abstract in metadata: 0 |
| CHECK abstracts never screened: 0 | CHECK abstracts never screened: 0 |
| CHECK textSource of extractor input: {"cols":5852,"mistral":7} | CHECK textSource of extractor input: {"cols":5852,"mistral":7} |
| CHECK extraction model: {"gpt-5.6-luna":5859} | CHECK extraction model: {"gpt-5.6-luna":5859} |
| | CHECK screening model: gpt-5.6-luna (service tier flex), last segment finished 2026-08-07T15:00:50.647Z |
| |
| |
| model gpt-5.6-luna | model gpt-5.6-luna |
| passes 1 (single pass; no adjudication, no second coder) | passes 1 (single pass; no adjudication, no second coder) |
| input paper.norm.txt, whole paper text | input paper.cols.txt (5,852), paper.mistral.md (OCR) (7); paper.norm.txt only gates eligibility |
| relation families 13: tools, population, crawlConfig, classification, detection, vantage, temporal, statistics, humanAnnotation, participants, ethics, artifacts, legal | relation families 13: tools, population, crawlConfig, classification, detection, vantage, temporal, statistics, humanAnnotation, participants, ethics, artifacts, legal |
| |
| robotsTxt holds a real value 53 = 4.7% | robotsTxt holds a real value 53 = 4.7% |
| robotsTxt holds any value incl. sentinels 992 = 88.6% | robotsTxt holds any value incl. sentinels 992 = 88.6% |
| | full distribution: {"not-stated":939,"(no ethics object at all)":128,"discussed":29,"respected":15,"ignored":9} |
| Counting the sentinel as an answer turns "one crawling paper in twenty says anything | Counting the sentinel as an answer turns "one crawling paper in twenty says anything |
| about robots.txt" into "seven in eight do". | about robots.txt" into "seven in eight do". |
| fragmented — every content word in one window, order broken 5,442 4.0% | fragmented — every content word in one window, order broken 5,442 4.0% |
| not found 1,708 1.3% | not found 1,708 1.3% |
| | |
| | quotes that describe the paper's silence rather than quoting it: 12 (0.0%) |
| |
| of the 1,708 not-found quotes, 271 (15.9%) contain an ellipsis, | of the 1,708 not-found quotes, 271 (15.9%) contain an ellipsis, |
| |
| crawling papers that say NOTHING about robots.txt 1120 - 53 = 1067 | crawling papers that say NOTHING about robots.txt 1120 - 53 = 1067 |
| IEEE-SP 2024 tail as a share of the venue 13 / 780 = 1.7% | IEEE-SP tail as a share of the venue 13 / 780 = 1.7% |
| USENIX 2026 share of the retrieval loss 104 / 230 = 45.2% | by year: 2021: 1, 2023: 1, 2024: 11 |
| | USENIX 2026 share of the retrieval loss 104 / 230 = 45.2% |
| US-folding raw strings beyond the 10 listed 308 - 10 = 298 | US-folding raw strings beyond the 10 listed 308 - 10 = 298 |
| papers in the nine zero-contribution venue-years 309 | papers in the nine zero-contribution venue-years 309 |
| IEEE-SP retrieval before the 2026-08-10 repair 780 - 447 = 333 on disk = 42.7% | IEEE-SP retrieval before the 2026-08-10 repair 780 - 447 = 333 on disk = 42.7% |
| (447 = line count of data/fulltext/missing_ieee_all.jsonl, the work list for that repair) | (447 = line count of data/fulltext/missing_ieee_all.jsonl, the work list for that repair) |
| quote verdicts that are layout damage rather than exact 4.0% + 1.3% = 5.3% | quote verdicts that are layout damage, not exact 4.0% + 1.3% = 5.3% |
| vantage location stated, the page's opening example 1228 of 3908 = 31.4% | vantage location stated, the page's opening example 1228 of 3908 = 31.4% |
| |
| - Whether a paper that produced no tuple for a family is silent or was missed. | - Whether a paper that produced no tuple for a family is silent or was missed. |
| </file> | </file> |
| |
| ===== 10. What this page could not establish ===== | |
| |
| Stated plainly, because a provenance page that overstates its own rigour is worse than none. | |
| |
| * **Whether the screening is accurate.** No human-screened control set exists. A relevant paper wrongly dropped at abstract screening leaves no trace in any artefact, so the false-negative rate of selection is unknown and unknowable from what is on disk. This is the largest unquantified risk in the funnel and it sits at the widest step of it. | |
| * **Whether field stability has moved.** The agreement figures in §6 are from the 4,322-paper corpus. Re-measuring means a second full extraction pass; it has not been done, and every page quoting those figures should say so. | |
| * **Whether the 14 unextracted papers matter.** Twelve are IEEE S&P 2024. Nobody has read them to see whether they would have changed anything. | |
| * **Whether the 309 papers in the nine empty venue-years matter.** The //cause// is established and documented — those venue pages publish no abstracts — but nobody has read the 309 titles to see how much relevant methodology is sitting outside the corpus. That is a cheap check nobody has done. | |
| * **The per-run token and cost ledger.** The extraction's own accounting was lost for roughly the first 2,870 papers of the original run because the ledger was written only on clean exit and every restart killed the process first. The total is a measured tail plus an extrapolation, and it is not re-derivable. | |
| * **Anything about a venue that is not one of the seven.** This is not a limitation that better tooling fixes; it is the scope. | |
| |
| ===== 11. Run log ===== | |
| |
| ^ ^ ^ | |
| | Page written | 2026-08-12 | | |
| | Corpus at the time | 5,859 papers, 7 venues, 117 venue-years, 2010–2026 | | |
| | Script | ''scripts/report_corpus.mjs'' (new for this page) | | |
| | Model | Claude Opus 5 | | |
| | Figures carried over from earlier notes | **None.** Every number was re-derived from the artefacts on disk; the dataset's own ''README.md'' still quotes the 4,322-paper figures and was not used as a source. | | |
| | Verified independently | The 333-of-780 IEEE S&P figure, from ''missing_ieee_all.jsonl'' holding exactly 447 records (780 − 447 = 333). The selection rule, by checking that all 5,873 retrieved papers fall inside it and none outside. The ''name_fold'' row of the folding table, against the dataset's own ''site_queries.mjs --page vantage'', which reports the same 367. | | |
| | Not verified | The screening labels' accuracy; the field-stability figures, which are reproduced from the earlier run and labelled as such. | | |
| | Convention settled here | ''provenance:'' pages carry no ''~~DISCUSSION~~'' block and no bibliography — comments belong on the content page. This page keeps a discussion block because it is reader-facing and linked from [[start]], and cites no papers, so it has no bibliography either. | | |
| | Why there is no ''provenance:literature:corpus'' | This page //is// a provenance page: sections 9 to 11 are its own working log. A provenance page for the provenance page would recurse without adding anything. | | |
| | Wired into | [[start]]; the six ''provenance:'' pages, which already link here; and the //Methodology and limitations of these figures// section of each corpus-backed content page, where the generic "these venues are absent, 2025–2026 are provisional" text was replaced by a pointer here plus the page-specific consequence. | | |
| | Mistakes caught in review | The quote check initially read ''paper.cols.txt'' for all papers, including the 7 whose extractor input was OCR text; that scored a broken font encoding as 29 fabricated quotes in a single paper and put it top of the not-found ranking. Fixed by keying on each record's own ''textSource''. | | |
| |
| Working notes for individual pages are under ''provenance:''; see [[#The provenance namespace]]. | |
| |
| /* This enables discussion under this article. */ | /* This enables discussion under this article. */ |
| ~~DISCUSSION~~ | ~~DISCUSSION~~ |
| |