| Next revision | Previous revision |
| provenance:literature:corpus [2026/09/04 08:12] – New page: working log for the 2026-09-04 corpus-wide pypdf-vs-paper.cols.txt audit — 73.2% of unlocatable quotes are in the PDF, ~24.6% of pages left two-column. Authored by Claude karel.kubicek.claude | provenance:literature:corpus [2026/09/04 08:29] (current) – Review-section headings: DokuWiki does not process '' '' inside a heading, so it rendered the quote marks literally. Authored by Claude karel.kubicek.claude |
|---|
| Working log behind [[literature:corpus]]. This page holds the working notes for that page's **own** figures; the corpus-level caveats every //other// ''provenance:'' page inherits — the venue scope, the funnel, the provisional years — are on [[literature:corpus]] itself and are not restated here. | Working log behind [[literature:corpus]]. This page holds the working notes for that page's **own** figures; the corpus-level caveats every //other// ''provenance:'' page inherits — the venue scope, the funnel, the provisional years — are on [[literature:corpus]] itself and are not restated here. |
| |
| **Its whole content, as of 2026-09-04, is one audit:** the corpus-wide diff of an independent PDF rendering against the stored text renderings, which settles what share of [[literature:corpus#quote_groundedness_re-measured_on_this_corpus|§6]]'s 1.3% unlocatable-quote rate is a text-extraction artefact rather than an extraction defect. [[literature:corpus]]'s other sections were written on 2026-08-12 with their notes inline, in its own §9 and §11; this page does not duplicate them. | **Its whole content, as of 2026-09-04, is one audit:** the corpus-wide diff of an independent PDF rendering against the stored text renderings, which **estimates** what share of [[literature:corpus#quote_groundedness_re-measured_on_this_corpus|§6]]'s 1.3% unlocatable-quote rate is a text-extraction artefact rather than an extraction defect. Estimates, not settles: it rests on one alternative PDF reader whose own failure rate is worse, and on a page-layout detector with no negative control. Both limits are in §7. [[literature:corpus]]'s other sections were written on 2026-08-12 with their notes inline, in its own §9 and §11; this page does not duplicate them. |
| |
| ===== 1. Why this was run, and what it retracted ===== | ===== 1. Why this was run, and what it retracted ===== |
| | Q8 | How many pages did the two-column repair leave untouched? | pages of every ''paper.txt'' | 91,509 pages | 29,347 (32.1%); 0 rejected by the word-multiset check | | | Q8 | How many pages did the two-column repair leave untouched? | pages of every ''paper.txt'' | 91,509 pages | 29,347 (32.1%); 0 rejected by the word-multiset check | |
| | Q9 | How many of those were really two-column? | stratified sample, PDF geometry | 400 untouched pages sampled, 60 reflowed as control | est. **22,475** of 91,509 pages (24.6%, CI 21.5–26.6%) | | | Q9 | How many of those were really two-column? | stratified sample, PDF geometry | 400 untouched pages sampled, 60 reflowed as control | est. **22,475** of 91,509 pages (24.6%, CI 21.5–26.6%) | |
| | Q10 | Does the untouched-page rate predict the quote verdict? | papers with ≥1 quote and both audits | 5,848 papers | yes, monotonically: ''exact'' 70.6% → 36.6% | | | Q10 | Does the untouched-page rate predict the quote verdict? | papers with ≥1 quote and both audits | 5,853 papers | yes, monotonically: ''exact'' 70.6% → 36.6% | |
| | Q11 | Is any text actually lost, or only re-ordered? | papers with both renderings | 5,856 papers | median de-spaced ratio 1.005; 34 papers (0.6%) below 0.95, of which 30 survive classification | | | Q11 | Is any text actually lost, or only re-ordered? | papers with both renderings | 5,856 papers | median de-spaced ratio 1.005; 34 papers (0.6%) below 0.95, of which 30 survive classification | |
| | Q12 | Why do 4 records have no ''paper.cols.txt''? | records with no cols file | 4 papers | by design — no usable text layer, OCR-repaired into ''paper.mistral.md'', which the extractor read | | | Q12 | Why do 4 records have no ''paper.cols.txt''? | records with no cols file | 4 papers | by design — no usable text layer, OCR-repaired into ''paper.mistral.md'', which the extractor read | |
| | Q13 | Which papers could not be rendered at all? | every extraction record | 5,859 papers | **1** after the limit raise: ''NDSS/2025/mtzk-…'', which needs the absent ''cryptography'' package. It holds 21 quotes | | | Q13 | Which papers could not be rendered at all? | every extraction record | 5,859 papers | **1** after the limit raise: ''NDSS/2025/mtzk-…'', which needs the absent ''cryptography'' package. It holds 21 quotes | |
| |
| **Denominators that are easy to mix up here.** 135,025 is every evidence quote; **135,004** is the subset with a verdict on both sides, and the 21-quote difference is exactly the one paper that still cannot be rendered. 97,505 was the page count before the phantom trailing segment was removed; **91,509** is the real one. 5,859 is every record; **5,858** have a cached rendering, **5,856** have a comparable character count, **5,848** have both a quote and a page verdict. The 1,806 / 931 / 541 / 118 / 216 figures quoted from [[provenance:statistics:hypothesis_testing]] are hypothesis-test quotes only, scored by ''ht_quotecheck.mjs'''s much looser 60%-of-5-word-windows rule — **not** by the four-verdict rule everything else here uses, and not comparable to it. | **Denominators that are easy to mix up here.** 135,025 is every evidence quote; **135,004** is the subset with a verdict on both sides, and the 21-quote difference is exactly the one paper that still cannot be rendered. 97,505 was the page count before the phantom trailing segment was removed; **91,509** is the real one. 5,859 is every record; **5,858** have a cached rendering, **5,856** have a comparable character count, **5,853** have both a quote and a page verdict. The 1,806 / 931 / 541 / 118 / 216 figures quoted from [[provenance:statistics:hypothesis_testing]] are hypothesis-test quotes only, scored by ''ht_quotecheck.mjs'''s much looser 60%-of-5-word-windows rule — **not** by the four-verdict rule everything else here uses, and not comparable to it. |
| |
| ===== 4. The reproduction check that makes the rest usable ===== | ===== 4. The reproduction check that makes the rest usable ===== |
| |
| <WRAP todo> | <WRAP todo> |
| * **Whether the 458 quotes neither rendering locates are all present in their papers.** 12 were read; the pattern is ellipsis splices the extraction marked itself, reflowed table rows, small capitals (%%"S ABOT"%%), symbol runs, sentences about a paper's silence, and at least one one-word paraphrase. The other 446 were not read. **0.34% is a ceiling on unlocatable, not a fabrication rate**, and the true fabrication rate is not measured by this audit or by any other on this site. | * **Whether the 458 quotes neither rendering locates are all present in their papers.** The 12 the report prints were **read as text** — the pattern is ellipsis splices the extraction marked itself, reflowed table rows, small capitals (%%"S ABOT"%%), symbol runs, sentences about a paper's silence, and one one-word paraphrase — but only **one** of them (''IMC/2011/pingin-in-the-rain'', §5) was checked against its paper. The other 457 were not opened. **0.34% is a ceiling on unlocatable, not a fabrication rate**, and the true fabrication rate is not measured by this audit or by any other on this site. |
| * **Whether re-rendering would actually fix it.** The obvious repair — lower the 50% vote threshold in ''findGutter'', or replace it with the x-histogram test used here for ground truth — was **not implemented and not tested.** The dataset is mounted read-only in this container, so it could not be tried end to end, and a threshold change has a failure mode in the other direction: reflowing a genuinely single-column page shuffles it. The 60-of-60 control says the current detector has no false positives to lose, which is an argument for trying, not a measurement of trying. Filed as a task against the dataset repository rather than left here. | * **Whether re-rendering would actually fix it.** The obvious repair — lower the 50% vote threshold in ''findGutter'', or replace it with the x-histogram test used here for ground truth — was **not implemented and not tested.** The dataset is mounted read-only in this container, so it could not be tried end to end, and a threshold change has a failure mode in the other direction: reflowing a genuinely single-column page shuffles it. The 60-of-60 control says the current detector has no false positives to lose, which is an argument for trying, not a measurement of trying. Filed as a task against the dataset repository rather than left here. |
| | * **The geometry detector has no negative control.** It is 60 for 60 on pages ''decolumn.mjs'' reflowed, which bounds its false negatives. Nothing bounds its **false positives**: no set of known single-column pages was assembled and put through it. If it over-calls two-column, the 22,475 estimate is too high, and the 0.6–1.0 band's unexplained 47.5% is the shape that would suggest. This is the weakest link in that number. |
| * **Whether a third reader would agree with ''pypdf''.** ''pdftotext'', ''pdfminer'', ''PyMuPDF'', ''pdfplumber'', ''mutool'', ''qpdf'' and ''ghostscript'' are all absent from this container. Every "present in the PDF" verdict rests on one reader whose own not-found rate (2.3%) is worse than the rendering it audits (1.3%). The 73.2% is a lower bound in one sense — a better reader would rescue more — and unreplicated in another. | * **Whether a third reader would agree with ''pypdf''.** ''pdftotext'', ''pdfminer'', ''PyMuPDF'', ''pdfplumber'', ''mutool'', ''qpdf'' and ''ghostscript'' are all absent from this container. Every "present in the PDF" verdict rests on one reader whose own not-found rate (2.3%) is worse than the rendering it audits (1.3%). The 73.2% is a lower bound in one sense — a better reader would rescue more — and unreplicated in another. |
| * **One paper still has no second rendering.** ''NDSS/2025/mtzk-testing-and-exploring-bugs-in-zero-knowledge-zk-compilers'' uses AES encryption, so ''pypdf'' needs the ''cryptography'' package. ''pip install cryptography'' is refused here by PEP 668, and ''--break-system-packages'' was **declined** rather than risk the container for one paper's 21 quotes. Anyone with a normal Python environment can close this in one command. | * **One paper still has no second rendering.** ''NDSS/2025/mtzk-testing-and-exploring-bugs-in-zero-knowledge-zk-compilers'' uses AES encryption, so ''pypdf'' needs the ''cryptography'' package. ''pip install cryptography'' is refused here by PEP 668, and ''--break-system-packages'' was **declined** rather than risk the container for one paper's 21 quotes. Anyone with a normal Python environment can close this in one command. |
| * **The ~60 pages that already publish a per-page quote-check figure.** Each was computed against the stored rendering alone and each overstates its failure rate. They were not re-run. Only [[statistics:hypothesis_testing]] was refreshed, because its own figures prompted the audit. Filed as deferred work. | * **The 57 pages that already publish a per-page quote-check figure.** Each was computed against the stored rendering alone and each overstates its failure rate. They were not re-run; only [[statistics:hypothesis_testing]] was, because its own figures prompted the audit. The list, so the next editor does not re-derive it, is every page matching ''quote_check %%|%% quotecheck %%|%% below threshold %%|%% five-word windows %%|%% 5-word windows'' over a fresh export: ''design:website:classification'', ''privacy:fingerprinting'', ''privacy:tcf:consent:strings'', ''programming:crawler'', ''programming:crawler:detection'', ''programming:crawler:foxhound'', ''programming:crawler:llm:agents'', ''programming:crawler:pagegraph'', ''programming:crawler:panoptichrome'', ''programming:crawler:webxray'', ''provenance:artifacts'', ''provenance:design:archives'', ''provenance:design:existing:datasets'', ''provenance:design:ip:classification'', ''provenance:design:longitudinal'', ''provenance:design:mobile:and:app:measurement'', ''provenance:design:platforms'', ''provenance:design:website:classification'', ''provenance:practices:ethics'', ''provenance:practices:legal:enforcement'', ''provenance:privacy:ads:txt'', ''provenance:privacy:browser:extensions'', ''provenance:privacy:browser:protection'', ''provenance:privacy:browser:storage'', ''provenance:privacy:consent'', ''provenance:privacy:cookie:syncing'', ''provenance:privacy:email:tracking'', ''provenance:privacy:fingerprinting'', ''provenance:privacy:javascript'', ''provenance:privacy:privacy:sandbox'', ''provenance:privacy:server:side:tracking'', ''provenance:privacy:tcf:consent:strings'', ''provenance:programming:crawler'', ''provenance:programming:crawler:detection'', ''provenance:programming:crawler:foxhound'', ''provenance:programming:crawler:llm:agents'', ''provenance:programming:crawler:openwpm'', ''provenance:programming:crawler:pagegraph'', ''provenance:programming:crawler:panoptichrome'', ''provenance:programming:crawler:webxray'', ''provenance:programming:interaction'', ''provenance:programming:interaction:outputs'', ''provenance:programming:stateful:stateless'', ''provenance:programming:traffic:files'', ''provenance:programming:tranco'', ''provenance:security:tls:certificates'', ''provenance:security:web:vulnerabilities'', ''provenance:statistics:biases'', ''provenance:statistics:how:many:sites'', ''provenance:statistics:hypothesis:testing'', ''provenance:statistics:pvalue:corrections'', ''provenance:statistics:regression'', |