User Tools

Site Tools


literature:corpus

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
literature:corpus [2026/08/12 15:57] – Report now prints the full ethics.robotsTxt distribution so the 'the rest is not-stated' claim is auditable; section 9 rewritten to say the staleness guard gives false passes on number collisions rather than implying it is a proof of freshness. Authored b karel.kubicek.claudeliterature:corpus [2026/08/12 16:06] (current) – Second review round: name the 138 poster records and the 3 title-level duplicates as structural properties of the funnel; add a boxed statement that an aggregate count cannot be independently recounted; widen the hand-checked quote sample to 2025-2026 and karel.kubicek.claude
Line 51: Line 51:
  
 ^ Stage ^ Papers ^ Lost here ^ Is the loss random or systematic? ^ ^ Stage ^ Papers ^ Lost here ^ Is the loss random or systematic? ^
-| Venue-year metadata records (7 venues × 117 venue-years) | 16,864 | — | — |+| Venue-year metadata records (117 venue-years, across 7 venues) | 16,864 | — | — |
 | …carrying an abstract | 15,800 | −1,064 | **systematic**, in three unrelated blocks — below | | …carrying an abstract | 15,800 | −1,064 | **systematic**, in three unrelated blocks — below |
 | Screened on title + abstract | 15,800 | 0 | none: screening covers every abstract | | Screened on title + abstract | 15,800 | 0 | none: screening covers every abstract |
Line 86: Line 86:
  
 **The 14 papers with prepared text but no extraction record.** One is a USENIX 2017 slide deck excluded deliberately; the other 13 are IEEE S&P papers — eleven from 2024, one from 2023, one from 2021 — whose extraction requests ended on ''520 status code (no body)'' and were never retried. Across the whole run 559 papers failed at least once and 546 were recovered by a retry; these 13 are the tail nobody went back for. The effect is a mild under-representation of IEEE S&P 2024 in every figure on this site, worth roughly 1.7% of that venue. **The 14 papers with prepared text but no extraction record.** One is a USENIX 2017 slide deck excluded deliberately; the other 13 are IEEE S&P papers — eleven from 2024, one from 2023, one from 2021 — whose extraction requests ended on ''520 status code (no body)'' and were never retried. Across the whole run 559 papers failed at least once and 546 were recovered by a retry; these 13 are the tail nobody went back for. The effect is a mild under-representation of IEEE S&P 2024 in every figure on this site, worth roughly 1.7% of that venue.
 +
 +==== What "5,859 papers" actually contains ====
 +
 +Two structural properties of the funnel that nothing above shows, both of which change how a silence figure should be read.
 +
 +  * **138 of the 5,859 records are posters**, 96 from CCS and 42 from IMC, identifiable by their own slug and title. The venue index does not filter front matter and posters out, and screening did not reject them. 89 of them sit inside the 3,908-paper ''measuredFrom'' population and 15 inside the 1,120-paper ''crawled'' one. A poster is one or two pages with its methodology compressed out, so it is systematically more likely to be scored as "did not state", and it biases exactly the reporting-gap figures this site is built on — upward. 251 records are four pages or fewer. If a silence rate matters to you, exclude short records and see whether it moves.
 +  * **Three papers are in the corpus twice.** The consistency check on duplicates is keyed on ''(venue, year, slug)'' and returns zero, but that key cannot see the duplication the venue index warns about: //Detecting and Characterizing Social Spam Campaigns// is indexed under both CCS 2010 and IMC 2010, //UIScope// under both NDSS 2020 and NDSS 2021, and //DRAWN APART// twice within NDSS 2022, once under an ''auto-draft-242'' slug. Keyed on the normalised title instead, the corpus has 3 extra records. Three in 5,859 changes nothing arithmetically; it matters because "zero duplicates" was being reported by a check that could not have found these.
  
 ==== The selection rule, verbatim ==== ==== The selection rule, verbatim ====
  
-Screening produces one label record per abstract with seven fields. A paper is selected if **any** of them is set:+Screening is itself an LLM pass — ''gpt-5.6-luna'', the same model as the extraction, last run on 2026-08-07. It produces one label record per abstract with seven fields, and a paper is selected if **any** of them is set:
  
 <code> <code>
Line 155: Line 162:
 ^ ^ ^ ^ ^ ^
 | What | One structured record per paper, describing what the authors did | | What | One structured record per paper, describing what the authors did |
-| Input | The paper's whole prepared text |+| Input | The paper's whole text, as ''paper.cols.txt'' (5,852 records) or, where ''pdftotext'' produced a garbled encoding, OCR'd ''paper.mistral.md'' (7). ''paper.norm.txt'' only gates eligibility. |
 | Method | **An LLM extraction against a fixed schema — a single pass, one request per paper** | | Method | **An LLM extraction against a fixed schema — a single pass, one request per paper** |
 | Model | ''gpt-5.6-luna'', structured outputs, all 5,859 records | | Model | ''gpt-5.6-luna'', structured outputs, all 5,859 records |
Line 204: Line 211:
 | ''USENIX/2024/you-can-obfuscate-but-you-cannot-hide…'' | "For each destination, bots can gather their available bandwidth through tools like Pathneck [17]." | **A bug in the checker, not the extraction.** This paper's ''pdftotext'' output is a garbled font encoding; its text was OCR-repaired and the extractor read the repaired file. The first version of the script checked the garbled one | | ''USENIX/2024/you-can-obfuscate-but-you-cannot-hide…'' | "For each destination, bots can gather their available bandwidth through tools like Pathneck [17]." | **A bug in the checker, not the extraction.** This paper's ''pdftotext'' output is a garbled font encoding; its text was OCR-repaired and the extractor read the repaired file. The first version of the script checked the garbled one |
  
-A further 15.9% of the not-found quotes contain an ellipsis, meaning the extraction spliced two parts of a source sentence and marked the join — not verbatim, but not invented either.+Those five were the first entries in the script's not-found list, which is filled in sorted-key order and therefore starts in CCS 2010–2012 — the era where column-splice damage is most likely, so it is a favourable sample and not a random one. **Five more were drawn from 2025 and 2026**, where the not-found counts are largest, and read the same way: 
 + 
 +^ Paper ^ What actually happened ^ 
 +| ''USENIX/2025/a-framework-for-abusability-analysis…'' | Present, spliced through the bibliography: //"rg/10. [other column] 5281/zenodo.14745290."//
 +| ''CCS/2025/a-sea-of-cyber-threats…'' | Present, with a participant table interleaved mid-sentence: //"open coding methodology [95], refin-P10 35 Chief Mate … ing the codebook"//
 +| ''USENIX/2025/addressing-the-address-books…'' | Present twice; the extraction dropped the "(90.5%)" parenthesis's surrounding clause | 
 +| ''NDSS/2026/a-causal-perspective-for-enhancing-jailbreak…'' | Present; the extraction reconstructed a URL broken across a line, and ''causal-learn'' and ''py-why'' are both in the text | 
 +| ''NDSS/2026/aliens-among-us…'' | **Not a quote at all.** The extraction wrote //"The paper does not report ethics review, affected-party notification, regulator contact…"// — a sentence about the paper's silence rather than a sentence from the paper | 
 + 
 +That last one is a different failure mode from layout damage, and it is the only one worth worrying about. It is also rare: **12 of the 135,025 quotes** describe a paper's silence instead of quoting it, concentrated in ''ethics'' and ''classification''A further 15.9% of the not-found quotes contain an ellipsis, meaning the extraction spliced two parts of a source sentence and marked the join — not verbatim, but not invented either.
  
 A practical consequence for anyone checking a quote by hand: **normalise whitespace before concluding a quote is absent.** ''grep -F'' on a multi-word quote fails about half the time against this text because it still wraps mid-sentence; ''tr -s '[:space:]' ' ' | grep -iF'' finds them. A practical consequence for anyone checking a quote by hand: **normalise whitespace before concluding a quote is absent.** ''grep -F'' on a multi-word quote fails about half the time against this text because it still wraps mid-sentence; ''tr -s '[:space:]' ' ' | grep -iF'' finds them.
Line 210: Line 226:
 The failures are also spread thin rather than concentrated: 1,143 of 5,859 papers (19.5%) have at least one, and the ten worst papers hold only 82 of the 1,708. The rate is flat across all seventeen years, between 0.8% and 1.5% in every single one, which is what you would expect of a PDF-layout artefact and not of a model that got worse or better over time. The failures are also spread thin rather than concentrated: 1,143 of 5,859 papers (19.5%) have at least one, and the ten worst papers hold only 82 of the 1,708. The rate is flat across all seventeen years, between 0.8% and 1.5% in every single one, which is what you would expect of a PDF-layout artefact and not of a model that got worse or better over time.
  
-**Two honest caveats about this measurement.** It reproduces the dataset's own audit of the earlier corpus almost exactly — that audit reported 57.9% / 37.0% / 4.2% / 0.9% on the same four buckets — but the two implementations are independent and the "fragmented" test in particular is an approximation of the original, so the 4.0%/1.3% split between the last two rows should be read as one combined 5.3% that is mostly layout damage. And the first measurement of this ran the check against the wrong file for seven papers whose text had been OCR-repaired, which put one paper at the top of the not-found ranking with 29 phantom failures; the script now reads whichever rendering the extractor read.+**Two caveats about this measurement.** It reproduces the dataset's own audit of the earlier corpus almost exactly — that audit reported 57.9% / 37.0% / 4.2% / 0.9% on the same four buckets — but the two implementations are independent and the "fragmented" test in particular is an approximation of the original, so the 4.0%/1.3% split between the last two rows should be read as one combined 5.3% that is mostly layout damage. And the first measurement of this ran the check against the wrong file for seven papers whose text had been OCR-repaired, which put one paper at the top of the not-found ranking with 29 phantom failures; the script now reads whichever rendering the extractor read.
  
 ==== Field stability: measured once, on the earlier corpus, never since ==== ==== Field stability: measured once, on the earlier corpus, never since ====
Line 261: Line 277:
 **What you can check.** Every figure on a corpus-backed page traces to a sentence in a published paper, and the ''provenance:'' page for that content page gives you the query, the population, the denominator, the fold, and the report script's unedited output. Two things follow that you can act on without any access to the corpus: **What you can check.** Every figure on a corpus-backed page traces to a sentence in a published paper, and the ''provenance:'' page for that content page gives you the query, the population, the denominator, the fold, and the report script's unedited output. Two things follow that you can act on without any access to the corpus:
  
-  - **Read the paper.** The corpus is public literature. If a page says //n// papers do something and names them, the papers are the evidence, not the dataset. +  - **Check the arithmetic and the population.** If a percentage's denominator is not named on the page or on its ''provenance:'' page, that is a defect in the page and worth reporting. So is a share of a population the page never defines. 
-  - **Check the arithmetic and the population.** If a percentage's denominator is not named on the page or on its ''provenance:'' page, that is a defect in the page and worth reporting.+  - **Read the papers a page names.** Where a page names its papers — the worked examples, the tool tables, the "read these first" lists — the papers are public and are the evidence. The dataset is not.
  
-**What you cannot check.** The corpus itself — the venue index, the screening labels, the retrieved PDFs and ''extractions.jsonl'' — is a local dataset on the maintainer's machine. It is not published, not downloadable, and there is no API. Two further consequences are worth being blunt about:+**What you cannot check.** The corpus itself — the venue index, the screening labels, the retrieved PDFs and ''extractions.jsonl'' — is a local dataset on the maintainer's machine. It is not published, not downloadable, and there is no API. The largest consequence is the plainest one, and it is not softened anywhere else on this page: 
 + 
 +<WRAP important> 
 +**An aggregate count on this site cannot be independently verified.** When a page says 182 papers crawled from the United States, the list of those 182 exists only in the unpublished extraction. You can check that the denominator is named, that the fold is described, that the arithmetic holds and that the handful of quotes the ''provenance:'' page spot-checked are real. You cannot recount the 182. Weigh these figures accordingly: they are a documented and internally audited measurement, not a reproducible one. 
 +</WRAP> 
 + 
 +Two further consequences follow:
  
   * **The extraction cannot be reproduced to the same values.** It is generative model output; two runs over identical text disagree, which is the whole point of the stability table in [[#Field stability: measured once, on the earlier corpus, never since]]. Re-running it would produce a similar but not identical dataset.   * **The extraction cannot be reproduced to the same values.** It is generative model output; two runs over identical text disagree, which is the whole point of the stability table in [[#Field stability: measured once, on the earlier corpus, never since]]. Re-running it would produce a similar but not identical dataset.
Line 284: Line 306:
 Two figures on this page are not produced by the script, and both name their source: the **2,870** papers whose token ledger was lost, which comes from the extraction's own ''README.md'', and the **29** phantom quote failures caused by a bug in an early version of this page's own script, described in [[#11. Run log]]. Every other figure on the page, including all four quote-groundedness percentages and the whole funnel, is produced by the script. Two figures on this page are not produced by the script, and both name their source: the **2,870** papers whose token ledger was lost, which comes from the extraction's own ''README.md'', and the **29** phantom quote failures caused by a bug in an early version of this page's own script, described in [[#11. Run log]]. Every other figure on the page, including all four quote-groundedness percentages and the whole funnel, is produced by the script.
  
-The guard reports only the first of those two, which is worth knowing about the guard rather than about the page: it asks whether a number //appears// in the report, and 29 happens to appear there as the count of crawling papers that discuss robots.txt. A guard of this kind gives false passes whenever a stale figure collides with an unrelated live one, so it is a floor on staleness checking, not a proof of freshness.+What the guard actually reported on 2026-08-12 is worth stating exactly, because it says as much about the guard as about the page. It flagged six figures: the **2,870**, and five digit fragments of quoted paper text (a DOI, a percentage and a participant-table row inside the quotes in [[#Quote groundedness, re-measured on this corpus|section 6]]). It did **not** flag the **29** — because 29 also happens to appear in the report as the count of crawling papers that discuss robots.txt. 
 + 
 +That is the shape of this kind of guard. It asks whether a number //appears// in the report, so it gives false pass whenever a stale figure collides with an unrelated live one, and it is noisy wherever a page quotes someone else's numbers. It is a floor on staleness checking, not a proof of freshness, and its output has to be read rather than expected to come back clean.
  
 The script also prints seven consistency checks on the funnel — retrieved papers outside the selection rule, fulltext directories outside it, extraction records with no prepared text on disk, duplicate records, labelled papers with no abstract, abstracts never screened, and truncated extractor inputs. All seven return zero. A nonzero value in any of them means the funnel above is wrong, not that the corpus is. The script also prints seven consistency checks on the funnel — retrieved papers outside the selection rule, fulltext directories outside it, extraction records with no prepared text on disk, duplicate records, labelled papers with no abstract, abstracts never screened, and truncated extractor inputs. All seven return zero. A nonzero value in any of them means the funnel above is wrong, not that the corpus is.
Line 295: Line 319:
   * **Whether the screening is accurate.** No human-screened control set exists. A relevant paper wrongly dropped at abstract screening leaves no trace in any artefact, so the false-negative rate of selection is unknown and unknowable from what is on disk. This is the largest unquantified risk in the funnel and it sits at the widest step of it.   * **Whether the screening is accurate.** No human-screened control set exists. A relevant paper wrongly dropped at abstract screening leaves no trace in any artefact, so the false-negative rate of selection is unknown and unknowable from what is on disk. This is the largest unquantified risk in the funnel and it sits at the widest step of it.
   * **Whether field stability has moved.** The agreement figures in §6 are from the 4,322-paper corpus. Re-measuring means a second full extraction pass; it has not been done, and every page quoting those figures should say so.   * **Whether field stability has moved.** The agreement figures in §6 are from the 4,322-paper corpus. Re-measuring means a second full extraction pass; it has not been done, and every page quoting those figures should say so.
-  * **Whether the 14 unextracted papers matter.** Eleven are IEEE S&P 2024. Nobody has read them to see whether they would have changed anything.+  * **Whether the 14 unextracted papers matter.** Thirteen are IEEE S&P, eleven of them from 2024. Nobody has read them to see whether they would have changed anything.
   * **Whether the 309 papers in the nine empty venue-years matter.** The //cause// is established and documented — those venue pages publish no abstracts — but nobody has read the 309 titles to see how much relevant methodology is sitting outside the corpus. That is a cheap check nobody has done.   * **Whether the 309 papers in the nine empty venue-years matter.** The //cause// is established and documented — those venue pages publish no abstracts — but nobody has read the 309 titles to see how much relevant methodology is sitting outside the corpus. That is a cheap check nobody has done.
   * **The per-run token and cost ledger.** The extraction's own accounting was lost for roughly the first 2,870 papers of the original run because the ledger was written only on clean exit and every restart killed the process first. The total is a measured tail plus an extrapolation, and it is not re-derivable.   * **The per-run token and cost ledger.** The extraction's own accounting was lost for roughly the first 2,870 papers of the original run because the ledger was written only on clean exit and every restart killed the process first. The total is a measured tail plus an extrapolation, and it is not re-derivable.
Line 331: Line 355:
 | Reader fit | The population table's header, "Share of 5,859", reads as the page breaking its own rule two sections after stating it. | **Accepted**, reworded to say why that denominator is the point of that particular table. | | Reader fit | The population table's header, "Share of 5,859", reads as the page breaking its own rule two sections after stating it. | **Accepted**, reworded to say why that denominator is the point of that particular table. |
 | Reader fit | ''[[#How to check a figure yourself]]'' was reported as a broken anchor, because the heading is numbered "8." and the link is not. | **Rejected.** Checked against the published HTML: DokuWiki's ''cleanID'' strips the leading ''8. '' when it builds the section id, so the heading's id //is// ''how_to_check_a_figure_yourself'' and the link resolves. Both the numbered and unnumbered forms work. The rendered page has no broken anchors and no red links. | | Reader fit | ''[[#How to check a figure yourself]]'' was reported as a broken anchor, because the heading is numbered "8." and the link is not. | **Rejected.** Checked against the published HTML: DokuWiki's ''cleanID'' strips the leading ''8. '' when it builds the section id, so the heading's id //is// ''how_to_check_a_figure_yourself'' and the link resolves. Both the numbered and unnumbered forms work. The rendered page has no broken anchors and no red links. |
 +
 +==== Second review, 2026-08-12 ====
 +
 +A fourth reviewer (Claude Fable 5) was given no checklist and asked for whatever the first three were not looking for. It found more than they did, and two of its findings are properties of the corpus that nobody had written down.
 +
 +^ Finding ^ Verdict ^
 +| **138 of the 5,859 records are posters**, unfiltered by the venue index and not rejected by screening, 89 of them inside the ''measuredFrom'' population. Posters compress their methodology out, so they push every //silence// figure on this site upward — which is the class of figure the site is mostly about. | **Accepted.** Measured, and now a named property of the funnel in [[#What "5,859 papers" actually contains]]. |
 +| **The duplicate check was blind to the duplication mode the venue index warns about.** Keyed on ''(venue, year, slug)'' it returns zero; keyed on the title, three papers are in the corpus twice. "Zero duplicates" was being reported by a check that could not have found them. | **Accepted.** The script now checks both keys and names the three. |
 +| The five hand-checked quotes were all from CCS 2010–2012, because the script's not-found list is filled in sorted-key order — a favourable sample presented as if it generalised, while the years with the most failures had none checked. | **Accepted.** Five more were drawn from 2025–2026 and read by hand; one of them turned out to be a different failure mode (a sentence about the paper's silence rather than a quote), which was then measured corpus-wide at 12 of 135,025. |
 +| §8's "read the paper" was technically true and practically misleading: corpus-backed tables do not name their papers, so **an aggregate count cannot be independently recounted at all** — the biggest item, and it was missing from "what you cannot check". | **Accepted.** It is now the boxed statement in [[#8. How to check a figure yourself]]. |
 +| The script printed a literal ''input paper.norm.txt'' in §7 of its own output, contradicting the ''textSource'' counts three screens above it. Same defect class as the hardcoded arithmetic the first reviewer caught, but hiding in a string. | **Accepted.** Computed now. |
 +| "twelve of them from 2024" survived in §10 after being fixed in §2 — the first review log said "Accepted", which read as "fixed everywhere". | **Accepted.** Fixed, and worth recording that a review log can flatter a partial fix. |
 +| The funnel's first row read "(7 venues × 117 venue-years)", which parses as a product. The screening model was never named. "Two honest caveats" and "worth being blunt about" label the page's own candour. | **All accepted.** |
 +| The dataset's own ''extract/README.md'' and runbook still carry the dead "IEEE S&P is 43% retrieved" caveat and a stale permanent-gaps table. | **Accepted as a note**: if you follow a citation from this page into those files, their IEEE S&P figures predate the 2026-08-10 repair described in [[#3. IEEE S&P: the worked example of a systematic loss]]. |
 +| Sibling page [[design:crawling_location]] carries the same "sixteen years" slip that was fixed here. | **Accepted**, fixed on that page. |
  
 Working notes for individual pages are under ''provenance:''; see [[#The provenance namespace]]. Working notes for individual pages are under ''provenance:''; see [[#The provenance namespace]].
Line 424: Line 463:
 CHECK  fulltext directories outside the selection rule: 0 CHECK  fulltext directories outside the selection rule: 0
 CHECK  extraction records with no paper.norm.txt on disk: 0 CHECK  extraction records with no paper.norm.txt on disk: 0
-CHECK  duplicate extraction records: 0+CHECK  duplicate extraction records, keyed on (venue, year, slug): 0 
 +CHECK  duplicate extraction records, keyed on normalised TITLE: 3 extra records across 3 titles 
 +         CCS/2010/detecting-and-characterizing-social-spam-campaigns  |  IMC/2010/detecting-and-characterizing-social-spam-campaigns 
 +         NDSS/2020/uiscope-accurate-instrumentation-free-and-visible-attack-investigation-for-gui-applications  |  NDSS/2021/uiscope-accurate-instrumentation-free-and-visible-attack-investigation-for-gui-applications 
 +         NDSS/2022/auto-draft-242  |  NDSS/2022/drawn-apart-a-device-identification-technique-based-on-remote-gpu-fingerprinting 
 +CHECK  extraction records that are posters: 138 {"CCS":96,"IMC":42} 
 +         of those, in measuredFrom: 89; in crawled: 15 
 +         records of 4 pages or fewer: 251
 CHECK  labelled papers with no abstract in metadata: 0 CHECK  labelled papers with no abstract in metadata: 0
 CHECK  abstracts never screened: 0 CHECK  abstracts never screened: 0
Line 430: Line 476:
 CHECK  textSource of extractor input: {"cols":5852,"mistral":7} CHECK  textSource of extractor input: {"cols":5852,"mistral":7}
 CHECK  extraction model: {"gpt-5.6-luna":5859} CHECK  extraction model: {"gpt-5.6-luna":5859}
 +CHECK  screening model: gpt-5.6-luna (service tier flex), last segment finished 2026-08-07T15:00:50.647Z
  
  
Line 587: Line 634:
 model                  gpt-5.6-luna model                  gpt-5.6-luna
 passes                 1 (single pass; no adjudication, no second coder) passes                 1 (single pass; no adjudication, no second coder)
-input                  paper.norm.txt, whole paper text+input                  paper.cols.txt (5,852), paper.mistral.md (OCR) (7); paper.norm.txt only gates eligibility
 relation families      13: tools, population, crawlConfig, classification, detection, vantage, temporal, statistics, humanAnnotation, participants, ethics, artifacts, legal relation families      13: tools, population, crawlConfig, classification, detection, vantage, temporal, statistics, humanAnnotation, participants, ethics, artifacts, legal
  
Line 709: Line 756:
   fragmented — every content word in one window, order broken    5,442   4.0%               fragmented — every content word in one window, order broken    5,442   4.0%            
   not found                                                      1,708   1.3%               not found                                                      1,708   1.3%            
 +
 +quotes that describe the paper's silence rather than quoting it: 12 (0.0%)
  
 of the 1,708 not-found quotes, 271 (15.9%) contain an ellipsis, of the 1,708 not-found quotes, 271 (15.9%) contain an ellipsis,
literature/corpus.1786550231.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki