User Tools

Site Tools


provenance:literature:corpus

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Next revision
Previous revision
provenance:literature:corpus [2026/09/04 08:12] – New page: working log for the 2026-09-04 corpus-wide pypdf-vs-paper.cols.txt audit — 73.2% of unlocatable quotes are in the PDF, ~24.6% of pages left two-column. Authored by Claude karel.kubicek.claudeprovenance:literature:corpus [2026/09/04 08:29] (current) – Review-section headings: DokuWiki does not process '' '' inside a heading, so it rendered the quote marks literally. Authored by Claude karel.kubicek.claude
Line 3: Line 3:
 Working log behind [[literature:corpus]]. This page holds the working notes for that page's **own** figures; the corpus-level caveats every //other// ''provenance:'' page inherits — the venue scope, the funnel, the provisional years — are on [[literature:corpus]] itself and are not restated here. Working log behind [[literature:corpus]]. This page holds the working notes for that page's **own** figures; the corpus-level caveats every //other// ''provenance:'' page inherits — the venue scope, the funnel, the provisional years — are on [[literature:corpus]] itself and are not restated here.
  
-**Its whole content, as of 2026-09-04, is one audit:** the corpus-wide diff of an independent PDF rendering against the stored text renderings, which settles what share of [[literature:corpus#quote_groundedness_re-measured_on_this_corpus|§6]]'s 1.3% unlocatable-quote rate is a text-extraction artefact rather than an extraction defect. [[literature:corpus]]'s other sections were written on 2026-08-12 with their notes inline, in its own §9 and §11; this page does not duplicate them.+**Its whole content, as of 2026-09-04, is one audit:** the corpus-wide diff of an independent PDF rendering against the stored text renderings, which **estimates** what share of [[literature:corpus#quote_groundedness_re-measured_on_this_corpus|§6]]'s 1.3% unlocatable-quote rate is a text-extraction artefact rather than an extraction defect. Estimates, not settles: it rests on one alternative PDF reader whose own failure rate is worse, and on a page-layout detector with no negative control. Both limits are in §7. [[literature:corpus]]'s other sections were written on 2026-08-12 with their notes inline, in its own §9 and §11; this page does not duplicate them.
  
 ===== 1. Why this was run, and what it retracted ===== ===== 1. Why this was run, and what it retracted =====
Line 123: Line 123:
 | Q8 | How many pages did the two-column repair leave untouched? | pages of every ''paper.txt'' | 91,509 pages | 29,347 (32.1%); 0 rejected by the word-multiset check | | Q8 | How many pages did the two-column repair leave untouched? | pages of every ''paper.txt'' | 91,509 pages | 29,347 (32.1%); 0 rejected by the word-multiset check |
 | Q9 | How many of those were really two-column? | stratified sample, PDF geometry | 400 untouched pages sampled, 60 reflowed as control | est. **22,475** of 91,509 pages (24.6%, CI 21.5–26.6%) | | Q9 | How many of those were really two-column? | stratified sample, PDF geometry | 400 untouched pages sampled, 60 reflowed as control | est. **22,475** of 91,509 pages (24.6%, CI 21.5–26.6%) |
-| Q10 | Does the untouched-page rate predict the quote verdict? | papers with ≥1 quote and both audits | 5,848 papers | yes, monotonically: ''exact'' 70.6% → 36.6% |+| Q10 | Does the untouched-page rate predict the quote verdict? | papers with ≥1 quote and both audits | 5,853 papers | yes, monotonically: ''exact'' 70.6% → 36.6% |
 | Q11 | Is any text actually lost, or only re-ordered? | papers with both renderings | 5,856 papers | median de-spaced ratio 1.005; 34 papers (0.6%) below 0.95, of which 30 survive classification | | Q11 | Is any text actually lost, or only re-ordered? | papers with both renderings | 5,856 papers | median de-spaced ratio 1.005; 34 papers (0.6%) below 0.95, of which 30 survive classification |
 | Q12 | Why do 4 records have no ''paper.cols.txt''? | records with no cols file | 4 papers | by design — no usable text layer, OCR-repaired into ''paper.mistral.md'', which the extractor read | | Q12 | Why do 4 records have no ''paper.cols.txt''? | records with no cols file | 4 papers | by design — no usable text layer, OCR-repaired into ''paper.mistral.md'', which the extractor read |
 | Q13 | Which papers could not be rendered at all? | every extraction record | 5,859 papers | **1** after the limit raise: ''NDSS/2025/mtzk-…'', which needs the absent ''cryptography'' package. It holds 21 quotes | | Q13 | Which papers could not be rendered at all? | every extraction record | 5,859 papers | **1** after the limit raise: ''NDSS/2025/mtzk-…'', which needs the absent ''cryptography'' package. It holds 21 quotes |
  
-**Denominators that are easy to mix up here.** 135,025 is every evidence quote; **135,004** is the subset with a verdict on both sides, and the 21-quote difference is exactly the one paper that still cannot be rendered. 97,505 was the page count before the phantom trailing segment was removed; **91,509** is the real one. 5,859 is every record; **5,858** have a cached rendering, **5,856** have a comparable character count, **5,848** have both a quote and a page verdict. The 1,806 / 931 / 541 / 118 / 216 figures quoted from [[provenance:statistics:hypothesis_testing]] are hypothesis-test quotes only, scored by ''ht_quotecheck.mjs'''s much looser 60%-of-5-word-windows rule — **not** by the four-verdict rule everything else here uses, and not comparable to it.+**Denominators that are easy to mix up here.** 135,025 is every evidence quote; **135,004** is the subset with a verdict on both sides, and the 21-quote difference is exactly the one paper that still cannot be rendered. 97,505 was the page count before the phantom trailing segment was removed; **91,509** is the real one. 5,859 is every record; **5,858** have a cached rendering, **5,856** have a comparable character count, **5,853** have both a quote and a page verdict. The 1,806 / 931 / 541 / 118 / 216 figures quoted from [[provenance:statistics:hypothesis_testing]] are hypothesis-test quotes only, scored by ''ht_quotecheck.mjs'''s much looser 60%-of-5-word-windows rule — **not** by the four-verdict rule everything else here uses, and not comparable to it.
  
 ===== 4. The reproduction check that makes the rest usable ===== ===== 4. The reproduction check that makes the rest usable =====
Line 178: Line 178:
  
 <WRAP todo> <WRAP todo>
-  * **Whether the 458 quotes neither rendering locates are all present in their papers.** 12 were readthe pattern is ellipsis splices the extraction marked itself, reflowed table rows, small capitals (%%"S ABOT"%%), symbol runs, sentences about a paper's silence, and at least one one-word paraphrase. The other 446 were not read. **0.34% is a ceiling on unlocatable, not a fabrication rate**, and the true fabrication rate is not measured by this audit or by any other on this site.+  * **Whether the 458 quotes neither rendering locates are all present in their papers.** The 12 the report prints were **read as text** — the pattern is ellipsis splices the extraction marked itself, reflowed table rows, small capitals (%%"S ABOT"%%), symbol runs, sentences about a paper's silence, and one one-word paraphrase — but only **one** of them (''IMC/2011/pingin-in-the-rain'', §5) was checked against its paper. The other 457 were not opened. **0.34% is a ceiling on unlocatable, not a fabrication rate**, and the true fabrication rate is not measured by this audit or by any other on this site.
   * **Whether re-rendering would actually fix it.** The obvious repair — lower the 50% vote threshold in ''findGutter'', or replace it with the x-histogram test used here for ground truth — was **not implemented and not tested.** The dataset is mounted read-only in this container, so it could not be tried end to end, and a threshold change has a failure mode in the other direction: reflowing a genuinely single-column page shuffles it. The 60-of-60 control says the current detector has no false positives to lose, which is an argument for trying, not a measurement of trying. Filed as a task against the dataset repository rather than left here.   * **Whether re-rendering would actually fix it.** The obvious repair — lower the 50% vote threshold in ''findGutter'', or replace it with the x-histogram test used here for ground truth — was **not implemented and not tested.** The dataset is mounted read-only in this container, so it could not be tried end to end, and a threshold change has a failure mode in the other direction: reflowing a genuinely single-column page shuffles it. The 60-of-60 control says the current detector has no false positives to lose, which is an argument for trying, not a measurement of trying. Filed as a task against the dataset repository rather than left here.
 +  * **The geometry detector has no negative control.** It is 60 for 60 on pages ''decolumn.mjs'' reflowed, which bounds its false negatives. Nothing bounds its **false positives**: no set of known single-column pages was assembled and put through it. If it over-calls two-column, the 22,475 estimate is too high, and the 0.6–1.0 band's unexplained 47.5% is the shape that would suggest. This is the weakest link in that number.
   * **Whether a third reader would agree with ''pypdf''.** ''pdftotext'', ''pdfminer'', ''PyMuPDF'', ''pdfplumber'', ''mutool'', ''qpdf'' and ''ghostscript'' are all absent from this container. Every "present in the PDF" verdict rests on one reader whose own not-found rate (2.3%) is worse than the rendering it audits (1.3%). The 73.2% is a lower bound in one sense — a better reader would rescue more — and unreplicated in another.   * **Whether a third reader would agree with ''pypdf''.** ''pdftotext'', ''pdfminer'', ''PyMuPDF'', ''pdfplumber'', ''mutool'', ''qpdf'' and ''ghostscript'' are all absent from this container. Every "present in the PDF" verdict rests on one reader whose own not-found rate (2.3%) is worse than the rendering it audits (1.3%). The 73.2% is a lower bound in one sense — a better reader would rescue more — and unreplicated in another.
   * **One paper still has no second rendering.** ''NDSS/2025/mtzk-testing-and-exploring-bugs-in-zero-knowledge-zk-compilers'' uses AES encryption, so ''pypdf'' needs the ''cryptography'' package. ''pip install cryptography'' is refused here by PEP 668, and ''--break-system-packages'' was **declined** rather than risk the container for one paper's 21 quotes. Anyone with a normal Python environment can close this in one command.   * **One paper still has no second rendering.** ''NDSS/2025/mtzk-testing-and-exploring-bugs-in-zero-knowledge-zk-compilers'' uses AES encryption, so ''pypdf'' needs the ''cryptography'' package. ''pip install cryptography'' is refused here by PEP 668, and ''--break-system-packages'' was **declined** rather than risk the container for one paper's 21 quotes. Anyone with a normal Python environment can close this in one command.
-  * **The ~60 pages that already publish a per-page quote-check figure.** Each was computed against the stored rendering alone and each overstates its failure rate. They were not re-run. Only [[statistics:hypothesis_testing]] was refreshed, because its own figures prompted the audit. Filed as deferred work.+  * **The 57 pages that already publish a per-page quote-check figure.** Each was computed against the stored rendering alone and each overstates its failure rate. They were not re-run; only [[statistics:hypothesis_testing]] was, because its own figures prompted the audit. The list, so the next editor does not re-derive it, is every page matching ''quote_check %%|%% quotecheck %%|%% below threshold %%|%% five-word windows %%|%% 5-word windows'' over a fresh export: ''design:website:classification'', ''privacy:fingerprinting'', ''privacy:tcf:consent:strings'', ''programming:crawler'', ''programming:crawler:detection'', ''programming:crawler:foxhound'', ''programming:crawler:llm:agents'', ''programming:crawler:pagegraph'', ''programming:crawler:panoptichrome'', ''programming:crawler:webxray'', ''provenance:artifacts'', ''provenance:design:archives'', ''provenance:design:existing:datasets'', ''provenance:design:ip:classification'', ''provenance:design:longitudinal'', ''provenance:design:mobile:and:app:measurement'', ''provenance:design:platforms'', ''provenance:design:website:classification'', ''provenance:practices:ethics'', ''provenance:practices:legal:enforcement'', ''provenance:privacy:ads:txt'', ''provenance:privacy:browser:extensions'', ''provenance:privacy:browser:protection'', ''provenance:privacy:browser:storage'', ''provenance:privacy:consent'', ''provenance:privacy:cookie:syncing'', ''provenance:privacy:email:tracking'', ''provenance:privacy:fingerprinting'', ''provenance:privacy:javascript'', ''provenance:privacy:privacy:sandbox'', ''provenance:privacy:server:side:tracking'', ''provenance:privacy:tcf:consent:strings'', ''provenance:programming:crawler'', ''provenance:programming:crawler:detection'', ''provenance:programming:crawler:foxhound'', ''provenance:programming:crawler:llm:agents'', ''provenance:programming:crawler:openwpm'', ''provenance:programming:crawler:pagegraph'', ''provenance:programming:crawler:panoptichrome'', ''provenance:programming:crawler:webxray'', ''provenance:programming:interaction'', ''provenance:programming:interaction:outputs'', ''provenance:programming:stateful:stateless'', ''provenance:programming:traffic:files'', ''provenance:programming:tranco'', ''provenance:security:tls:certificates'', ''provenance:security:web:vulnerabilities'', ''provenance:statistics:biases'', ''provenance:statistics:how:many:sites'', ''provenance:statistics:hypothesis:testing'', ''provenance:statistics:pvalue:corrections'', ''provenance:statistics:regression'', ''security:tls:certificates'', ''statistics:biases'', ''statistics:hypothesis:testing'', ''statistics:pvalue:corrections'', ''statistics:regression''. Filed as deferred work rather than left as a note here.
   * **Field stability.** Still measured only on the 4,322-paper corpus, still not re-measured, and untouched by this audit.   * **Field stability.** Still measured only on the 4,322-paper corpus, still not re-measured, and untouched by this audit.
   * **Four other pages carry a lowercase %%<wrap>%% box**, found while fixing the one on [[provenance:statistics:hypothesis_testing]]: ''provenance:practices:ethics'', ''provenance:practices:notifying_websites'', ''provenance:practices:public_relations'' and ''provenance:statistics:pvalue_corrections''. Each renders its box as a ''<span>'' full of literal asterisks. Not fixed here — a different job from this audit, and filed as such. ((An earlier draft said //eight//. That count came from ''grep -l "<wrap "'' over a site export, which also matches the five pages that merely //mention// the tag inside ''%%…%%'', where it is correctly escaped and renders fine. ''check_wrap.mjs'' over the same export gives five, one of which is the page being fixed. The figures reviewer caught it. A count from a grep is not a count from the guard that owns the rule.))   * **Four other pages carry a lowercase %%<wrap>%% box**, found while fixing the one on [[provenance:statistics:hypothesis_testing]]: ''provenance:practices:ethics'', ''provenance:practices:notifying_websites'', ''provenance:practices:public_relations'' and ''provenance:statistics:pvalue_corrections''. Each renders its box as a ''<span>'' full of literal asterisks. Not fixed here — a different job from this audit, and filed as such. ((An earlier draft said //eight//. That count came from ''grep -l "<wrap "'' over a site export, which also matches the five pages that merely //mention// the tag inside ''%%…%%'', where it is correctly escaped and renders fine. ''check_wrap.mjs'' over the same export gives five, one of which is the page being fixed. The figures reviewer caught it. A count from a grep is not a count from the guard that owns the rule.))
Line 199: Line 200:
 | **The 0.6–1.0 score bands are reported pooled** | Four rows of //n// ≤ 25 read as precision that is not there. The corpus-wide estimate still uses them separately, so the pooling affects only what is displayed | | **The 0.6–1.0 score bands are reported pooled** | Four rows of //n// ≤ 25 read as precision that is not there. The corpus-wide estimate still uses them separately, so the pooling affects only what is displayed |
 | **The fix is a fallback in the checkers, not a re-render** | A re-render is the real fix and is not mine to make: the mount is read-only. A fallback is strictly additive — it can only move a quote from FAILED to RESCUED — so it cannot make any existing figure worse | | **The fix is a fallback in the checkers, not a re-render** | A re-render is the real fix and is not mine to make: the mount is read-only. A fallback is strictly additive — it can only move a quote from FAILED to RESCUED — so it cannot make any existing figure worse |
-| **''statistics:hypothesis_testing'' was refreshed and ~60 other pages were not** | Refreshing all of them is a bigger job than this audit and each needs its own figures re-derived and re-reviewed. Doing one and declaring the rest stale is worse than doing none only if the declaration is hidden; it is in §7 and on [[literature:corpus]] |+| **''statistics:hypothesis_testing'' was refreshed and the other 56 were not** | Refreshing all of them is a bigger job than this audit and each needs its own figures re-derived and re-reviewed. Doing one and declaring the rest stale is worse than doing none only if the declaration is hidden; it is in §7 and on [[literature:corpus]] |
  
 ===== 9. The run itself ===== ===== 9. The run itself =====
Line 231: Line 232:
 Four passes, all given the page text, the scripts, their output and these notes, and all told explicitly that the author's context might not be exhaustive. Three ran in parallel first; their findings were fixed and the figures re-derived before the fourth. Four passes, all given the page text, the scripts, their output and these notes, and all told explicitly that the author's context might not be exhaustive. Three ran in parallel first; their findings were fixed and the figures re-derived before the fourth.
  
-==== Figures against the scripts (''sonnet''====+==== Figures against the scripts — sonnet ====
  
 ^ # ^ Finding ^ Disposition ^ ^ # ^ Finding ^ Disposition ^
Line 239: Line 240:
 | F4 | ''audit_decolumn_pages.mjs'' counts a word-bag-rejected page in both ''rejected'' and ''untouched'' | **NOTED, not fixed.** The corpus has 0 rejected pages, so it affects no published number. Left as a latent defect with this note | | F4 | ''audit_decolumn_pages.mjs'' counts a word-bag-rejected page in both ''rejected'' and ''untouched'' | **NOTED, not fixed.** The corpus has 0 rejected pages, so it affects no published number. Left as a latent defect with this note |
  
-==== Quotes and attribution (''sonnet''====+==== Quotes and attribution — sonnet ====
  
 ^ # ^ Finding ^ Disposition ^ ^ # ^ Finding ^ Disposition ^
Line 249: Line 250:
 | Q6 | Cross-page figures, citekeys and the ''decolumn.mjs''/''normalize_text.mjs'' source claims all check out; the deleted footnote orphaned no citation; the base files are byte-identical to the live pages | No change | | Q6 | Cross-page figures, citekeys and the ''decolumn.mjs''/''normalize_text.mjs'' source claims all check out; the deleted footnote orphaned no citation; the base files are byte-identical to the live pages | No change |
  
-==== External and tool currency (''sonnet''====+==== External and tool currency — sonnet ====
  
 ^ # ^ Finding ^ Disposition ^ ^ # ^ Finding ^ Disposition ^
Line 260: Line 261:
 | C7 | ''visitor_text'''s signature and ''tm[4]'' are current and not deprecated; ''extraction_mode="layout"'' does not support the visitor callbacks, so it is not an alternative. DokuWiki's %%<WRAP>%%/%%<wrap>%% div-vs-span behaviour, ''%%''…''%%'' being monospace rather than nowiki, and %%<file text name.txt>%% are all current. Every internal link resolves, including the ''#quote_groundedness_re-measured_on_this_corpus'' anchor | No change | | C7 | ''visitor_text'''s signature and ''tm[4]'' are current and not deprecated; ''extraction_mode="layout"'' does not support the visitor callbacks, so it is not an alternative. DokuWiki's %%<WRAP>%%/%%<wrap>%% div-vs-span behaviour, ''%%''…''%%'' being monospace rather than nowiki, and %%<file text name.txt>%% are all current. Every internal link resolves, including the ''#quote_groundedness_re-measured_on_this_corpus'' anchor | No change |
  
-==== Generic (''fable''====+==== Generic — fable ====
  
-//Filled in below when that pass has run.//+Run after the three above were fixed, with no checklist. 
 + 
 +^ # ^ Finding ^ Disposition ^ 
 +| G1 | **[[literature:corpus]] published a stale copy of the dose-response table** — three rows and the "5,848 papers" denominator predated the ''pypdf'' limit raise, and contradicted this page's own Appendix A two hundred lines below | **ACCEPTED and FIXED**, and the cause fixed too: the content page's tables were hand-typed literals. ''edit_corpus.py'' now **slices them out of the report script's own ''--wiki'' output**, so they cannot drift again. The re-run figures pass found the same defect independently, which is the strongest signal in this review log | 
 +| G2 | **The two quote-scoring rules were still conflated** on both hypothesis-testing pages: each described this page as "the corpus-wide version of the same measurement" when the thresholds differ | **ACCEPTED and FIXED** on both, plus a 73.1% that should have been 73.2% | 
 +| G3 | **[[literature:corpus]] now contradicts its own untouched sections**: §11 had a row headed //"Why there is no ''provenance:literature:corpus''"//, §9 said every figure on the page comes from one script, and §11 said "the six ''provenance:'' pages" | **ACCEPTED and FIXED.** This is the //carve-out leaves day-one drift// pattern exactly: adding a section left literal sentences elsewhere on the page false | 
 +| G4 | **"12 read examples" overstated the hand-reading**: §5 records six quotes read, only one of which is in the 458 | **ACCEPTED and FIXED.** §7 now says the 12 were read as text and one was checked against its paper | 
 +| G5 | **Seven overstatements**, in order of load: "1,250 are demonstrably in the paper" (137 are only fragmented); the heading "Nothing is missing"; "settled" for a second automated detector with only a positive control; "very largely a measure of de-columning" for a paper-level association with obvious confounds; this page's intro "settles"; "below-threshold mostly is not the extraction's fault" for 118 of 334; and a garbled "one paper in every 6.8" | **ALL SEVEN ACCEPTED AND FIXED.** The most useful was the third: the reviewer pointed out there is **no negative control** on the geometry detector, which is now stated in §7 as the weakest link in the 22,475 estimate. Nothing on either page had said so | 
 +| G6 | **The grep story was told three times in full**, and the content page's copy fails the //no textbook// rule — its reader has neither the harness shell nor the corpus | **ACCEPTED for the content page**, cut to one caveat sentence plus a link. **PARTLY REJECTED for [[provenance:statistics:hypothesis_testing]]**: that is the page whose claim was withdrawn, and a retraction that does not show the cause is not a retraction. The transcript stays there. The "with the ugrep transcript" wording was fixed — there is no ugrep | 
 +| G7 | **§6 of the content page is out of proportion** — 10.7 KB against the 5.3 KB it extends, five tables, an undefined "loose gutter score" column | **PARTLY ACCEPTED.** The calibration band table and the container-specific detail moved here; the two verdict tables, the dose-response table and the practical rule stayed, because they are what a reader checking a figure needs. The suggestion to move the character-ratio detail was **rejected**: "the order is wrong, almost nothing is missing" is the single most load-bearing sentence for anyone deciding whether to trust a quote, and it needs its number on the page that makes the claim | 
 +| G8 | **The retraction is handled well, with two gaps**: the page's run table still dates the page 2026-08-13 with no pointer to the correction pass, and "it is kept here" overstates what was kept | **ACCEPTED and FIXED**, both | 
 +| G9 | **A live placeholder**: this section said "filled in when that pass has run" | **ACCEPTED.** You are reading the fix | 
 +| G10 | **Deferred work was declared but not addressable** — "~60 pages" with no list | **ACCEPTED and FIXED.** All 57 are named in §7 | 
 +| G11 | What not to change: §3's query/denominator table, §4's reproduction check, §8's judgement calls, the two "not fixed not re-run" bullets on the content page, the RESCUED verdict and ''textSource'' keying, and this log's finding/disposition format | No change. Recorded because a review that lists only defects gives no signal about what to leave alone | 
 + 
 +==== Figures, re-run after the fixes above — sonnet ==== 
 + 
 +^ # ^ Finding ^ Disposition ^ 
 +| R1 | The stale dose-response table and its 5,848, found independently of G1 | **ACCEPTED and FIXED**; see G1 | 
 +| R2 | **Q10 and the "denominators" paragraph on this page also said 5,848** where the correct count is 5,853 | **ACCEPTED and FIXED** | 
 +| R3 | **''report_cols_mechanism.mjs'' printed a header count (5,855) two higher than its own table's sum (5,853)**, because two papers are skipped inside the loop for want of a page record | **ACCEPTED and FIXED in the script**, so the header and the table agree. A pre-existing defect, not introduced by the fixes, and it was embedded in Appendix B | 
 +| R4 | Everything else reproduced byte-for-byte: all four live pages identical to the local drafts, all nine scripts re-run, the reconciliation (15+6+0+0 = 21), every derived share recomputed, the mechanism figures unchanged, the limit raise confirmed in effect in every worker before any parse, the cache confirmed free of pre-raise entries for the five rescued papers, ''ht_quotecheck.mjs'''s three-candidate fallback confirmed never to fire on this corpus, and ''spread()'' confirmed deterministic and count-neutral | No change |
  
 ===== Related ===== ===== Related =====
Line 463: Line 485:
 reason alone. What matters here is the gradient, not the level. reason alone. What matters here is the gradient, not the level.
  
-papers with >=1 evidence quote and both audits complete  5,855+papers with >=1 evidence quote and both audits complete  5,853
  
   untouched≥0.35   papers   quotes   below thr.       rescued   untouched≥0.35   papers   quotes   below thr.       rescued
provenance/literature/corpus.1788509525.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki