User Tools

Site Tools


provenance:literature:bibliography

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Next revision
Previous revision
provenance:literature:bibliography [2026/09/04 03:30] – Provenance for literature:bibliography: 2026-09-04 audit of USENIX author lists (18 of 90 cached lists carried an affiliation as an author; 0 defects reached the page, all 159 live USENIX entries verified against their landing pages) and re-check of PoPET karel.kubicek.claudeprovenance:literature:bibliography [2026/09/04 18:43] (current) – Second-sitting audit: apply the four review passes' findings (autoKey STOP-list rule, 20-page union, 28/26 candidate split, 160-vs-161, first-sitting labelling, render-check off-by-one) and publish the full review log. Authored by Claude karel.kubicek.claude
Line 2: Line 2:
  
 Back to [[:literature:bibliography|the bibliography]]. Corpus-wide selection and Back to [[:literature:bibliography|the bibliography]]. Corpus-wide selection and
-extraction notes belong on [[:literature:corpus]] (not yet written). This page is +extraction notes are on [[:literature:corpus]], the root of this namespace. 
-the query log for the bibliography itself: where each entry's metadata came from,+This page is the query log for the bibliography itself: where each entry's metadata came from,
 what has been checked against a primary source, and what is known to be wrong what has been checked against a primary source, and what is known to be wrong
 with it. with it.
Line 14: Line 14:
 ===== Run record ===== ===== Run record =====
  
-  * Run date: 2026-09-04 (UTC). First revision of this page.+  * Run date: 2026-09-04 (UTC). First sitting; first revision of this page.
   * Authoring agent: Claude Opus 5, executing the drain item ''audit-usenix-author-lists'' non-interactively. No human in the loop during the run.   * Authoring agent: Claude Opus 5, executing the drain item ''audit-usenix-author-lists'' non-interactively. No human in the loop during the run.
   * Corpus at run time: 5,859 extracted papers, 2010–2026, over CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf and IEEE S&P. Read-only inputs under ''/workspace/publications_dataset/data/''.   * Corpus at run time: 5,859 extracted papers, 2010–2026, over CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf and IEEE S&P. Read-only inputs under ''/workspace/publications_dataset/data/''.
-  * Live bibliography read at revision **1788471646**, 382,620 bytes, **837 entries**, exported with ''?do=export_raw'' and kept at ''out/usenix_audit/bib_live.txt''. Every figure below is against that snapshot.+  * Live bibliography read at revision **1788471646**, 382,620 bytes, **837 entries**, exported with ''?do=export_raw'' and kept at ''out/usenix_audit/bib_live.txt''. Every figure in the first-sitting sections below is against that snapshot; the second sitting's figures are against its own, later snapshots, named where they are used.
   * Nothing on ''literature:bibliography'' was edited by this run. The only content change is the one-line pointer to this page.   * Nothing on ''literature:bibliography'' was edited by this run. The only content change is the one-line pointer to this page.
   * ''scripts/fetch_authors.py'' and ''out/authors.json'' were changed. ''out/usenix_audit/authors.json.before'' is the pre-run snapshot and is what the audit diffs against.   * ''scripts/fetch_authors.py'' and ''out/authors.json'' were changed. ''out/usenix_audit/authors.json.before'' is the pre-run snapshot and is what the audit diffs against.
 +  * **Second sitting, 2026-09-04, later the same day:** drain item ''dedup-regional-filter-lists-bibkey'', Claude Fable 5.1, non-interactive, no human in the loop. Live bibliography read at revision **1788526615**, 390,804 bytes, **855 entries** — 18 more than the first sitting, added by other items in between — and saved at revision 1788543026 with **850**. This sitting **did** edit ''literature:bibliography'', and 20 other pages; every change is listed under //Audit 2026-09-04, second sitting// below.
  
 ===== What the audits found, in one table ===== ===== What the audits found, in one table =====
Line 32: Line 33:
 | Is any live USENIX entry's author list wrong in any way? | all 159 live USENIX Security entries, names **and** order, against their own landing pages (pass 3) | **no, 0** | | Is any live USENIX entry's author list wrong in any way? | all 159 live USENIX Security entries, names **and** order, against their own landing pages (pass 3) | **no, 0** |
 | Is any PoPETs DOI on the wrong prefix, or dead? | 84 PoPETs entries with a DOI; the prefix rule covers 83 | **0 wrong prefix, 0 dead** | | Is any PoPETs DOI on the wrong prefix, or dead? | 84 PoPETs entries with a DOI; the prefix rule covers 83 | **0 wrong prefix, 0 dead** |
-| Is the same paper in the bibliography twice? | all 837 entries | **yes, 5 pairs** — latent, no page cites both keys of a pair |+| Is the same paper in the bibliography twice? | all 837 entries (first sitting) | **yes, 5 pairs** — latent, no page cited both keys of a pair 
 +| … and after the second sitting? | all 850 entries on the saved page | **0 pairs**, both scans; 23 markers on 13 pages repointed first, then the 5 entries deleted | 
 +| Did the umlaut slip (''bottger''/''boettger'') recur for any other author? | 855 entries before and 850 after, under all four rules — shared DOI, shared title, near-title, and same year with the first-author surname folded so Böttger = Boettger = Bottger | **no** — 58 candidates printed, all read, none a duplicate |
 | Are the checks above capable of failing? | 13 mutations | all 13 change the reported counts | | Are the checks above capable of failing? | 13 mutations | all 13 change the reported counts |
  
Line 50: Line 53:
 | every other venue | 12,599 | 1 | 1,618 | | every other venue | 12,599 | 1 | 1,618 |
  
-(That table is corpus-level and belongs on [[:literature:corpus]] once it +(The per-venue **record counts** in that table are already published on 
-exists; it is here because the author gap is the reason this audit had to +[[:literature:corpus]] and agree with these exactly — CCS 3,381, IEEE S&P 1,837, 
-happen at all.) So the author gap is PETS and USENIX and essentially nothing +IMC 862, NDSS 1,577, PETS 1,253, USENIX 3,012, WWW 4,942. What is new here is 
-else — 4,265 of+the author and DOI breakdown, which is the reason this audit had to happen at 
 +all.) So the author gap is PETS and USENIX and essentially nothing else — 4,265 of
 16,864 records, plus a single stray NDSS record 16,864 records, plus a single stray NDSS record
 (''NDSS/2018/veil-private-browsing-semantics-without-browser-side-assistance'', (''NDSS/2018/veil-private-browsing-semantics-without-browser-side-assistance'',
Line 570: Line 574:
 established. established.
  
-===== Found and deferred: five papers are in the bibliography twice =====+===== Found at 837 entries: five papers were in the bibliography twice =====
  
 A citekey-collision check passes while the same paper sits in the file under two A citekey-collision check passes while the same paper sits in the file under two
Line 585: Line 589:
  
 **No page currently cites both keys of a pair**, checked by **No page currently cites both keys of a pair**, checked by
-''scripts/bib_dupe_cocitation.py'' over the 160 other pages on the wiki (161 in+''scripts/bib_dupe_cocitation.py'' over the 160 other pages on the wiki at the time (161 in
 ''dw.mjs pages'', less ''literature:bibliography'' itself). 26 pages cite at ''dw.mjs pages'', less ''literature:bibliography'' itself). 26 pages cite at
-least one key of some pair; ''bouhoula2024_automated'' alone is on 11so no reference list renders the paper twice today. This is latent, not+least one key of some pair; ''bouhoula2024_automated'' alone is on 11so no reference list renders the paper twice today. This is latent, not
 visible. Consolidating means choosing one key per pair and rewriting %%{[key]}%% visible. Consolidating means choosing one key per pair and rewriting %%{[key]}%%
 markers across the 26 pages that cite one, which is a different piece of work markers across the 26 pages that cite one, which is a different piece of work
Line 593: Line 597:
 ''dedupe-bibliography-entries'' rather than half-done here. ''dedupe-bibliography-entries'' rather than half-done here.
  
-===== What could not be established =====+**Closed later the same day**, but under a different item name than the one 
 +filed here. ''dedup-regional-filter-lists-bibkey'' had been raised on 2026-08-14 
 +for the Böttger pair alone, with the instruction to "check for other 
 +duplicate-title pairs in the same pass"; it therefore already covered the work 
 +''dedupe-bibliography-entries'' was filed for, and was the one that ran. The 
 +consolidation and its invariants are the section //Audit 2026-09-04, second 
 +sitting// below. ''dedupe-bibliography-entries'' is superseded by it and should 
 +be closed as such; this run could add items to the queue but not close them. 
 +The table and counts in this section describe the 837-entry file and are left 
 +as they were. 
 + 
 +===== Audit 2026-09-04, second sitting: the five duplicates consolidated ===== 
 + 
 +Drain item ''dedup-regional-filter-lists-bibkey'', raised on 2026-08-14 while 
 +[[:programming:crawler:openwpm]] was being written and its author noticed 
 +''bottger2025_regional'' next to ''boettger2025_regional''. The first sitting 
 +above found all five pairs and deferred them; this sitting closed them. Every 
 +figure here is against a fresh ''?do=export_raw'' of the bibliography at 
 +revision 1788526615 (855 entries) and of all 161 other pages, taken at 17:21 
 +UTC and re-checked against ''dw.mjs pages'' immediately before saving: no 
 +revision had moved. **161, where the first sitting says 160**: this provenance 
 +page did not exist when that count was taken, and it is one of the 26 that name 
 +a deleted key in prose. Both numbers are right for their own date. 
 + 
 +==== Which key was kept, and why ==== 
 + 
 +None of the ten keys can have been minted from an author list: PETS and USENIX 
 +index records carry none, so ''autoKey'' would have fallen back to the surname 
 +in the landing URL or required an explicit ''--key''. The tie-break is the 
 +file's own majority convention, **surname + year + ''_'' + first title word, 
 +diacritics dropped** — where "first title word" is ''bibgen.mjs'' ''autoKey'': 
 +the first word longer than three letters that is not in its ''STOP'' list. 
 +''Understanding'' is in ''STOP'', which is why the kept key is 
 +''bottger2025_regional'' and not ''bottger2025_understanding''. ''scripts/bib_key_convention.py'' counts how the 25 entries whose 
 +first author's surname carries a diacritic spell it in the key. It runs on the 
 +**855-entry file, before the deletion**, so the Böttger paper is in it twice, 
 +once in each row it is the example for; after the save 24 entries remain, 19 of 
 +them stripped: 
 + 
 +^ Key spelling of the diacritic ^ Entries (of 25) ^ Examples ^ 
 +| stripped — Böttger → ''bottger'' | **19** | ''sjosten2020_filter'', ''bosch2016_tales'', ''tornberg2024_bestpractices'', ''gross2021_iuipc''
 +| German digraph — Böttger → ''boettger'' | 4 | ''boettger2025_regional'', ''stoever2023_owners'', ''rueth2018_digging'', ''schoeni2024_cookieblock''
 +| letter dropped — Kührer → ''khrer'' | 2 | ''khrer2015_going'', ''som2017_content'': the ''[^a-z]'' filter in ''bibgen.mjs'' ''autoKey'' eating a non-ASCII letter — a bug, not a convention, and not touched here | 
 + 
 +^ Paper ^ Kept ^ Deleted ^ Content pages citing kept / deleted, before ^ What the deleted entry had that the kept one lacked ^ 
 +| Fouad et al., PoPETs 2022 | ''fouad2022_cookie'' | ''fouad2022my'' | 2 / **4** | a ''url'' that repeats the DOI — dropped | 
 +| Böttger et al., PoPETs 2025 | ''bottger2025_regional'' | ''boettger2025_regional'' | 2 / 1 (+1 provenance page) | nothing; the kept entry additionally carries volume, number and pages | 
 +| Ahmad et al., PoPETs 2026 | ''ahmad2026_more'' | ''ahmad2026_ipfp'' | 1 / 1 | nothing; the kept title braces ''{IPv6}''
 +| Bouhoula et al., USENIX Sec 2024 | ''bouhoula2024_automated'' | ''bouhoula2024automated'' | 6 (+5 provenance) / 3 (+1 provenance) | ''pages = {1723--1739}'' — **carried over**, after reading USENIX's own BibTeX block and the ''citation_firstpage''/''citation_lastpage'' meta tags on ''usenixsecurity24/presentation/bouhoula'' on 2026-09-04, both of which say 1723–1739. ''isbn'', ''address'', ''publisher'', ''month'' and an author-homepage ''url'' — dropped; no other entry in the file carries any of them | 
 +| Lerner et al., USENIX Sec 2016 | ''lerner2016_internet'' | ''lerner2016internet'' | 2 / 2 (+1 provenance) | nothing | 
 + 
 +The per-page lists behind the fourth column are in the apply output below. 
 +Three of the five deleted keys were the Google-Scholar style (no underscore), 
 +and one of those, ''fouad2022my'', was on **more** content pages than its 
 +replacement — four against two. Majority-of-pages would have kept it. Convention 
 +won, because the convention is what the next ''bibgen.mjs'' run will mint, and a 
 +file with two live key styles is how these five pairs arose. 
 + 
 +==== What was done, in order ==== 
 + 
 +  - Export every page and the bibliography fresh: 161 + 1 files. ''dw.mjs search'' was not used; its index is stale (first sitting). 
 +  - Run ''scripts/bib_dedup_scan.py'' — the wider scan this item asked for — on the export: 5 definite pairs, 58 candidates. 
 +  - Run ''scripts/bib_dedup_apply.py'' offline. It rewrites the markers and the bibliography into ''out/dedup_apply/'' and asserts the invariants in its docstring. It saves nothing. 
 +  - Save the **20 pages first** — the 13 whose markers changed plus the 10 given an amendment, three pages being in both sets — each with ''--if-rev'' against the revision read at export, then the bibliography: 21 saves. In that order nothing ever renders unresolved: the kept keys existed while the deleted ones were still cited, and the deleted ones were gone only once no page cited them. 
 +  - Purge the bibtex4dw cache on the bibliography and every repointed page (''?purge=true''; see the first sitting for why a rendered check passes on a stale cache without this). 
 +  - Re-fetch the 13 repointed pages and compare rendered reference-list length and marker count against copies fetched before the edit (''scripts/bib_dedup_render_check.py''). 
 +  - Re-export the bibliography and all 21 saved pages and confirm byte identity with the saved files; re-run both duplicate scans on the live bibliography — 0 pairs each. 
 + 
 +==== The invariants, and the real output ==== 
 + 
 +Each figure on this page about the consolidation is a line of this output. 
 + 
 +<code> 
 +bibliography : out/live_bibliography_20260904_1721.txt 
 +entries      : 855 -> 850  (removed 5: fouad2022my, boettger2025_regional, ahmad2026_ipfp, bouhoula2024automated, lerner2016internet) 
 +field added  : bouhoula2024_automated  [('pages', '1723--1739')] 
 +pages read   : 161 (bibliography excluded) 
 + 
 +loser-key marker occurrences before, by key: 
 +  fouad2022my                7 
 +  boettger2025_regional      3 
 +  ahmad2026_ipfp             1 
 +  bouhoula2024automated      7 
 +  lerner2016internet         5 
 + 
 +pages citing each key BEFORE (content pages; provenance pages in brackets): 
 +  kept    fouad2022_cookie           2 [0]  practices:legal_enforcement, statistics:hypothesis_testing 
 +  deleted fouad2022my                4 [0]  privacy:browser_storage, privacy:fingerprinting, programming:crawler:openwpm, programming:stateful_stateless 
 +  kept    bottger2025_regional       2 [0]  programming:crawler:openwpm, programming:filter_lists 
 +  deleted boettger2025_regional      1 [1]  statistics:pvalue_corrections  [provenance:statistics:pvalue_corrections] 
 +  kept    ahmad2026_more             1 [0]  programming:cloudflare_radar 
 +  deleted ahmad2026_ipfp             1 [0]  design:longitudinal 
 +  kept    bouhoula2024_automated     6 [5]  privacy:darkpatterns, privacy:tcf_consent_strings, programming:crawler:openwpm, programming:deployment, programming:multilingual_support, statistics:pvalue_corrections  [provenance:privacy:tcf_consent_strings, provenance:programming:crawler:openwpm, provenance:programming:deployment, provenance:programming:multilingual_support, provenance:statistics:pvalue_corrections] 
 +  deleted bouhoula2024automated      3 [1]  design:crawling_location, privacy:consent, privacy:requests  [provenance:privacy:consent] 
 +  kept    lerner2016_internet        2 [0]  privacy:fingerprinting, statistics:biases 
 +  deleted lerner2016internet         2 [1]  design:archives, privacy:browser_storage  [provenance:design:archives] 
 + 
 +^ page ^ markers repointed ^ distinct keys before ^ after ^ pairs ^ 
 +| design:archives | 3 | 14 | 14 | lerner2016internet→lerner2016_internet | 
 +| design:crawling_location | 2 | 15 | 15 | bouhoula2024automated→bouhoula2024_automated | 
 +| design:longitudinal | 1 | 32 | 32 | ahmad2026_ipfp→ahmad2026_more | 
 +| privacy:browser_storage | 2 | 29 | 29 | fouad2022my→fouad2022_cookie; lerner2016internet→lerner2016_internet | 
 +| privacy:consent | 3 | 28 | 28 | bouhoula2024automated→bouhoula2024_automated | 
 +| privacy:fingerprinting | 2 | 35 | 35 | fouad2022my→fouad2022_cookie | 
 +| privacy:requests | 1 | 39 | 39 | bouhoula2024automated→bouhoula2024_automated | 
 +| programming:crawler:openwpm | 1 | 26 | 26 | fouad2022my→fouad2022_cookie | 
 +| programming:stateful_stateless | 3 | 28 | 28 | fouad2022my→fouad2022_cookie | 
 +| provenance:design:archives | 1 | 7 | 7 | lerner2016internet→lerner2016_internet | 
 +| provenance:privacy:consent | 1 | 3 | 3 | bouhoula2024automated→bouhoula2024_automated | 
 +| provenance:statistics:pvalue_corrections | 2 | 18 | 18 | boettger2025_regional→bottger2025_regional | 
 +| statistics:pvalue_corrections | 1 | 28 | 28 | boettger2025_regional→bottger2025_regional | 
 + 
 +pages changed          : 13 
 +markers repointed      : 23  (loser occurrences before: 23) 
 +provenance amendments  : 10 
 +files written          : 21 under out/dedup_apply/ 
 + 
 +markers naming a key the bibliography does not define, BEFORE: 22 on 4 key(s) 
 +  ...  (1 page(s)) 
 +  cite  (1 page(s)) 
 +  citekey  (2 page(s)) 
 +  key  (18 page(s)) 
 +same, AFTER: 22 on 4 key(s) 
 +  ...  (1 page(s)) 
 +  cite  (1 page(s)) 
 +  citekey  (2 page(s)) 
 +  key  (18 page(s)) 
 +distinct pages carrying any such example marker, AFTER: 20 
 + 
 +deleted keys still named in PROSE (not markers), left as historical record: 26 page(s) 
 +  provenance:design:archives                       boettger2025_regional, bouhoula2024automated, fouad2022my, lerner2016internet 
 +  provenance:design:crawling_location              bouhoula2024automated 
 +  provenance:design:dns                            ahmad2026_ipfp, boettger2025_regional, bouhoula2024automated, fouad2022my, lerner2016internet 
 +  provenance:design:longitudinal                   ahmad2026_ipfp, lerner2016internet 
 +  provenance:design:platforms                      ahmad2026_ipfp, boettger2025_regional, bouhoula2024automated, fouad2022my, lerner2016internet 
 +  provenance:design:website_classification         ahmad2026_ipfp, boettger2025_regional, bouhoula2024automated, fouad2022my, lerner2016internet 
 +  provenance:literature:bibliography               ahmad2026_ipfp, boettger2025_regional, bouhoula2024automated, fouad2022my, lerner2016internet 
 +  provenance:privacy:browser_extensions            lerner2016internet 
 +  provenance:privacy:browser_protection            ahmad2026_ipfp, boettger2025_regional, fouad2022my 
 +  provenance:privacy:browser_storage               fouad2022my, lerner2016internet 
 +  provenance:privacy:consent                       bouhoula2024automated 
 +  provenance:privacy:cookie_syncing                boettger2025_regional, fouad2022my 
 +  provenance:privacy:darkpatterns                  bouhoula2024automated 
 +  provenance:privacy:fingerprinting                fouad2022my 
 +  provenance:privacy:privacy_sandbox               ahmad2026_ipfp, boettger2025_regional, bouhoula2024automated, fouad2022my, lerner2016internet 
 +  provenance:privacy:requests                      bouhoula2024automated 
 +  provenance:privacy:server_side_tracking          fouad2022my 
 +  provenance:privacy:tcf_consent_strings           boettger2025_regional, bouhoula2024automated, fouad2022my, lerner2016internet 
 +  provenance:programming:crawler:openwpm           boettger2025_regional, bouhoula2024automated, fouad2022my 
 +  provenance:programming:deployment                bouhoula2024automated 
 +  provenance:programming:filter_lists              boettger2025_regional 
 +  provenance:programming:stateful_stateless        boettger2025_regional, fouad2022my 
 +  provenance:statistics:biases                     boettger2025_regional, bouhoula2024automated, lerner2016internet 
 +  provenance:statistics:how_many_sites             lerner2016internet 
 +  provenance:statistics:pvalue_corrections         boettger2025_regional 
 +  provenance:statistics:regression                 boettger2025_regional, bouhoula2024automated, fouad2022my, lerner2016internet 
 + 
 +all invariants hold 
 +</code> 
 + 
 +Reading it: 23 deleted-key occurrences before, 23 repointed; every page'
 +distinct-key count is unchanged, which is the check that no page cited both 
 +keys of a pair (a page that had would have lost a reference silently); and the 
 +four "unresolved" strings — ''key'', ''citekey'', ''cite'', ''...'' — are 
 +documentation examples written as literal markers in prose on **20** distinct 
 +pages (''key'' alone is on 18; the per-string counts overlap, which is why the 
 +script prints the union), identical before and after. They are a pre-existing 
 +residue this run did not touch and did not create. 
 + 
 +==== Rendered check ==== 
 + 
 +<code> 
 +^ page ^ references (dt) before → after ^ citekey spans before → after ^ deleted-key strings after ^ verdict ^ 
 +| design:archives | 14 → 14 | 84 → 84 | 0 | unchanged | 
 +| design:crawling_location | 15 → 15 | 40 → 40 | 0 | unchanged | 
 +| design:longitudinal | 32 → 32 | 90 → 90 | 0 | unchanged | 
 +| privacy:browser_storage | 28 → 28 | 96 → 96 | 0 | unchanged | 
 +| privacy:consent | 28 → 28 | 138 → 138 | 0 | unchanged | 
 +| privacy:fingerprinting | 35 → 35 | 118 → 118 | 0 | unchanged | 
 +| privacy:requests | 38 → 38 | 156 → 156 | 0 | unchanged | 
 +| programming:crawler:openwpm | 26 → 26 | 96 → 96 | 0 | unchanged | 
 +| programming:stateful_stateless | 28 → 28 | 150 → 150 | 0 | unchanged | 
 +| provenance:design:archives | -1 → -1 | 46 → 46 | 5 | unchanged | 
 +| provenance:privacy:consent | 3 → 3 | 6 → 6 | 3 | unchanged | 
 +| provenance:statistics:pvalue_corrections | 18 → 18 | 48 → 48 | 1 | unchanged | 
 +| statistics:pvalue_corrections | 27 → 27 | 112 → 112 | 0 | unchanged | 
 + 
 +pages checked: 13   pages whose counts moved: 0 
 +</code> 
 + 
 +Three rows sit one below the apply table's distinct-key count — 
 +''privacy:browser_storage'' 29 vs 28, ''privacy:requests'' 39 vs 38, 
 +''statistics:pvalue_corrections'' 28 vs 27 — because each carries a 
 +**documentation example** written as a real marker: ''%%{[key]}%%'' on the 
 +first, and a contributor instruction naming ''LePochat2019_tranco'' on the 
 +other two, which the plugin does not render at all. The apply script counts 
 +those as citations, so its distinct-key column and its "pages citing each key" 
 +lists are one high on those three pages. It does not touch the consolidation: 
 +none of the five pairs is involved. ''-1'' is [[:provenance:design:archives]], 
 +which cites papers but carries no 
 +''%%<bibtex bibliography>%%'' block, so it has no reference list to count. The 
 +"deleted-key strings" on the three provenance pages are their dated amendment 
 +and their historical notes, not citations — the same rows show their marker 
 +counts unchanged. 
 + 
 +==== The wider scan: the umlaut slip did not recur ==== 
 + 
 +The item asked whether the transliteration mistake existed for other 
 +German-named authors. ''bib_dedup_scan.py'' adds two candidate generators to 
 +the DOI-and-title scan of the first sitting: near-identical squashed titles 
 +(''difflib'' ratio ≥ 0.85, or one title a prefix of the other, for a dropped 
 +subtitle), and **same year plus same folded first-author surname**, where the 
 +fold strips diacritics and collapses oe/ue/ae/ss to o/u/a/s so that Böttger, 
 +''B"ottger'', Boettger and Bottger compare equal. On the 855-entry file it 
 +reports the five definite pairs and **58 candidates**; on the saved 850-entry 
 +file, **0 definite** and the same 58. 
 + 
 +All 58 were read by hand and **none is a duplicate**. Three are DuckDuckGo 
 +artefacts whose "surname" is the vendor name; one is a replication and its 
 +original, caught by the title rule (''bratton2019_replication''
 +''sumner2014_exaggeration''). Of the remaining 54, the script's own 
 +first-author comparison splits them **28 / 26**: 28 are the same first author 
 +with two or three different papers in one year (Durumeric 2013–2015, Vastel 
 +2018, Oest 2020, Kancherla 2025 …), and 26 have first-author strings that 
 +differ: 25 are **different people who share a surname and a year** — Li, Zhang, 
 +Liu, Lin, Wu, Agarwal, Tang — and one is a single person spelled two ways 
 +(''Al Roomi'' / ''Alroomi'', 2023). That 
 +second class is the price of a rule that folds surnames: it is the only way 
 +Böttger and Boettger can be paired at all. No second umlaut pair exists in the file. The list is the residue 
 +of the fold and is printed in full, with both first-author strings on every 
 +row, so the judgement can be checked: 
 + 
 +<code> 
 +bibliography : out/live_bibliography_20260904_1721.txt 
 +entries      : 855   distinct citekeys: 855 
 + 
 +DEFINITE duplicate pairs (rule A or B): 5 
 +  [ABCD] ahmad2026_ipfp  /  ahmad2026_more 
 +       More Space, Less Privacy? Measuring the Effectiveness of IP-based Website Fingerprinting i 
 +       More Space, Less Privacy? Measuring the Effectiveness of IP-based Website Fingerprinting i 
 +  [ABCD] boettger2025_regional  /  bottger2025_regional 
 +       Understanding Regional Filter Lists: Efficacy and Impact 
 +       Understanding Regional Filter Lists: Efficacy and Impact 
 +  [BCD] bouhoula2024_automated  /  bouhoula2024automated 
 +       Automated Large-Scale Analysis of Cookie Notice Compliance 
 +       Automated Large-Scale Analysis of Cookie Notice Compliance 
 +  [ABCD] fouad2022_cookie  /  fouad2022my 
 +       My Cookie is a phoenix: detection, measurement, and lawfulness of cookie respawning with b 
 +       My Cookie is a phoenix: Detection, measurement, and lawfulness of cookie respawning with b 
 +  [BCD] lerner2016_internet  /  lerner2016internet 
 +       Internet Jones and the Raiders of the Lost Trackers: An Archaeological Study of Web Tracki 
 +       Internet Jones and the Raiders of the lost trackers: An archaeological study of web tracki 
 + 
 +CANDIDATE pairs (rule C or D only) — judged by hand, see the provenance page: 58 
 +  [D] LePochat2019_tranco  /  LePochat2019_tranco_eval   first author: SAME ({Le Pochat}, Victor | {Le Pochat}, Victor) 
 +       2019 Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation 
 +       2019 Evaluating the Long-term Effects of Parameters on the Characteristics of the {Tranco} Top Sites Rank 
 +  [D] agarwal2024_peeking  /  agarwal2024_poster   first author: differs (Agarwal, Shubham | Agarwal, Sharad) 
 +       2024 Peeking through the window: Fingerprinting Browser Extensions through Page-Visible Execution Traces  
 +       2024 Poster: A Comprehensive Categorization of SMS Scams 
 +  [D] agarwal2025_dropped  /  agarwal2025_fishing   first author: SAME (Agarwal, Sharad | Agarwal, Sharad) 
 +       2025 'Hey mum, I dropped my phone down the toilet': Investigating Hi Mum and Dad SMS Scams in the United  
 +       2025 Fishing for Smishing: Understanding SMS Phishing Infrastructure and Strategies by Mining Public User 
 +  [D] agarwal2025_dropped  /  agarwal2025_mindsets   first author: differs (Agarwal, Sharad | Agarwal, Shubham) 
 +       2025 'Hey mum, I dropped my phone down the toilet': Investigating Hi Mum and Dad SMS Scams in the United  
 +       2025 "I have no idea how to make it safer": Studying Security and Privacy Mindsets of Browser Extension D 
 +  [D] agarwal2025_fishing  /  agarwal2025_mindsets   first author: differs (Agarwal, Sharad | Agarwal, Shubham) 
 +       2025 Fishing for Smishing: Understanding SMS Phishing Infrastructure and Strategies by Mining Public User 
 +       2025 "I have no idea how to make it safer": Studying Security and Privacy Mindsets of Browser Extension D 
 +  [D] alroomi2023_login  /  alroomi2023_password   first author: differs (Al Roomi, Suood | Alroomi, Suood) 
 +       2023 A Large-Scale Measurement of Website Login Policies 
 +       2023 Measuring Website Password Creation Policies At Scale 
 +  [D] bahrami2025_bytedefender  /  bahrami2025_cookieguard   first author: SAME (Nikkhah Bahrami, Pouneh | Nikkhah Bahrami, Pouneh) 
 +       2025 Byte by Byte: Unmasking Browser Fingerprinting at the Function Level Using V8 Bytecode Transformers 
 +       2025 {CookieGuard}: Characterizing and Isolating the First-Party Cookie Jar 
 +  [D] bashir2019_adstxt  /  bashir2019_quantity   first author: SAME (Bashir, Muhammad Ahmad | Bashir, Muhammad Ahmad) 
 +       2019 A Longitudinal Analysis of the ads.txt Standard 
 +       2019 Quantity vs. Quality: Evaluating User Interest Profiles Using Ad Preference Managers 
 +  [D] bhuiyan2025_digital  /  bhuiyan2025_visitors   first author: SAME (Bhuiyan, Masudul Hasan Masud | Bhuiyan, Masudul Hasan Masud) 
 +       2025 Digital Disparities: A Comparative Web Measurement Study Across Economic Boundaries 
 +       2025 Not All Visitors are Bilingual: A Measurement Study of the Multilingual Web from an Accessibility Pe 
 +  [C] bratton2019_replication  /  sumner2014_exaggeration   first author: differs (Bratton, Luke | Sumner, Petroc) 
 +       2019 The Association Between Exaggeration in Health-Related Science News and Academic Press Releases: A R 
 +       2014 The Association Between Exaggeration in Health Related Science News and Academic Press Releases: Ret 
 +  [D] chen2021_cookieswap  /  chen2021_detecting   first author: SAME (Chen, Quan | Chen, Quan) 
 +       2021 Cookie Swap Party: Abusing First-Party Cookies for Web Tracking 
 +       2021 Detecting Filter List Evasion with Event-Loop-Turn Granularity JavaScript Signatures 
 +  [D] chen2025_parents  /  chen2025_semantics   first author: differs (Chen, Xiaowei | Chen, Baiqi) 
 +       2025 Empowering Parents to Support Children's Online Security and Privacy: Findings from a Randomized Con 
 +       2025 Semantics-Aware Cookie Purpose Compliance 
 +  [CD] duckduckgo_tracker_radar_2026  /  duckduckgo_tracker_radar_detector_2026   first author: SAME ({DuckDuckGo} | {DuckDuckGo}) 
 +       2026 DuckDuckGo Tracker Radar 
 +       2026 DuckDuckGo Tracker Radar Detector 
 +  [D] duckduckgo_tracker_radar_2026  /  duckduckgo_trc_2026   first author: SAME ({DuckDuckGo} | {DuckDuckGo}) 
 +       2026 DuckDuckGo Tracker Radar 
 +       2026 Tracker Radar Collector 
 +  [D] duckduckgo_tracker_radar_detector_2026  /  duckduckgo_trc_2026   first author: SAME ({DuckDuckGo} | {DuckDuckGo}) 
 +       2026 DuckDuckGo Tracker Radar Detector 
 +       2026 Tracker Radar Collector 
 +  [D] durumeric2013_https  /  durumeric2013_zmap   first author: SAME (Durumeric, Zakir | Durumeric, Zakir) 
 +       2013 Analysis of the HTTPS certificate ecosystem 
 +       2013 {ZMap}: Fast Internet-wide Scanning and Its Security Applications 
 +  [D] durumeric2014_heartbleed  /  durumeric2014_view   first author: SAME (Durumeric, Zakir | Durumeric, Zakir) 
 +       2014 The Matter of Heartbleed 
 +       2014 An Internet-Wide View of Internet-Wide Scanning 
 +  [D] durumeric2015_neither  /  durumeric2015_search   first author: SAME (Durumeric, Zakir | Durumeric, Zakir) 
 +       2015 Neither Snow Nor Rain Nor MITM...: An Empirical Analysis of Email Delivery Security 
 +       2015 A Search Engine Backed by Internet-Wide Scanning 
 +  [D] edu2022_alexa  /  edu2022_exploring   first author: SAME (Edu, Jide S. | Edu, Jide S.) 
 +       2022 Measuring Alexa Skill Privacy Practices across Three Years 
 +       2022 Exploring the security and privacy risks of chatbots in messaging services 
 +  [D] iqbal2022_khaleesi  /  iqbal2022_left   first author: differs (Umar Iqbal | Iqbal, Hassan) 
 +       2022 Khaleesi: Breaker of Advertising and Tracking Request Chains 
 +       2022 Left or Right: A Peek into the Political Biases in Email Spam Filtering Algorithms During US Electio 
 +  [D] kancherla2025_johnny  /  kancherla2025_least   first author: SAME (Kancherla, Gayatri Priyadarsini | Kancherla, Gayatri Priyadarsini) 
 +       2025 Johnny Can't Revoke Consent Either: Measuring Compliance of Consent Revocation on the Web 
 +       2025 Least Privilege Access for Persistent Storage Mechanisms in Web Browsers 
 +  [D] kirchner2024_black  /  kirchner2024_dancer   first author: SAME (Kirchner, Robin | Kirchner, Robin) 
 +       2024 A Black-Box Privacy Analysis of Messaging Service Providers' Chat Message Processing 
 +       2024 Dancer in the Dark: Synthesizing and Evaluating Polyglots for Blind Cross-Site Scripting 
 +  [D] lee2023_adtargeting  /  lee2023_track   first author: differs (Lee, Hao-Ping (Hank) | Lee, Dongkeun) 
 +       2023 When and Why Do People Want Ad Targeting Explanations? Evidence from a Four-Week, Mixed-Methods Fiel 
 +       2023 Net-track: Generic Web Tracking Detection Using Packet Metadata 
 +  [D] li2016_remedying  /  li2016_youve   first author: SAME (Li, Frank | Li, Frank) 
 +       2016 Remedying Web Hijacking: Notification Effectiveness and Webmaster Comprehension 
 +       2016 You've Got Vulnerability: Exploring Effective Vulnerability Notifications 
 +  [D] li2017_radar  /  li2017_security   first author: differs (Li, Zhenhua | Li, Frank) 
 +       2017 FBS-Radar: Uncovering Fake Base Stations at Scale in the Wild 
 +       2017 A Large-Scale Empirical Study of Security Patches 
 +  [D] li2017_radar  /  li2017_static   first author: differs (Li, Zhenhua | Li, Li) 
 +       2017 FBS-Radar: Uncovering Fake Base Stations at Scale in the Wild 
 +       2017 Static analysis of Android apps: A systematic literature review 
 +  [D] li2017_security  /  li2017_static   first author: differs (Li, Frank | Li, Li) 
 +       2017 A Large-Scale Empirical Study of Security Patches 
 +       2017 Static analysis of Android apps: A systematic literature review 
 +  [D] li2024_bounce  /  li2024_wellinformed   first author: differs (Li, Ruixuan | Li, Shuai) 
 +       2024 Bounce in the Wild: A Deep Dive into Email Delivery Failures from a Large Email Service Provider 
 +       2024 Are We Getting Well-informed? An In-depth Study of Runtime Privacy Notice Practice in Mobile Apps 
 +  [D] li2024_bounce  /  li2024_worldwide   first author: SAME (Li, Ruixuan | Li, Ruixuan) 
 +       2024 Bounce in the Wild: A Deep Dive into Email Delivery Failures from a Large Email Service Provider 
 +       2024 A Worldwide View on the Reachability of Encrypted DNS Services 
 +  [D] li2024_wellinformed  /  li2024_worldwide   first author: differs (Li, Shuai | Li, Ruixuan) 
 +       2024 Are We Getting Well-informed? An In-depth Study of Runtime Privacy Notice Practice in Mobile Apps 
 +       2024 A Worldwide View on the Reachability of Encrypted DNS Services 
 +  [D] liao2016_characterizing  /  liao2016_seeking   first author: SAME (Liao, Xiaojing | Liao, Xiaojing) 
 +       2016 Characterizing Long-tail SEO Spam on Cloud Web Hosting Services 
 +       2016 Seeking Nonsense, Looking for Trouble: Efficient Promotional-Infection Detection through Semantic In 
 +  [D] lin2021_longitudinal  /  lin2021_phishpedia   first author: differs (Lin, Fuqi | Lin, Yun) 
 +       2021 A Longitudinal Study of Removed Apps in {iOS} App Store 
 +       2021 Phishpedia: A Hybrid Deep Learning Based Approach to Visually Identify Phishing Webpages 
 +  [D] lin2022_investigating  /  lin2022_sheep   first author: differs (Lin, Su-Chin | Lin, Xu) 
 +       2022 Investigating Advertisers' Domain-changing Behaviors and Their Impacts on Ad-blocker Filter Lists 
 +       2022 Phish in Sheep's Clothing: Exploring the Authentication Pitfalls of Browser Fingerprinting 
 +  [D] liu2024_opted  /  liu2024_promises   first author: differs (Liu, Zengrui | Liu, Xiaoyin) 
 +       2024 Opted Out, Yet Tracked: Are Regulations Enough to Protect Your Privacy? 
 +       2024 From Promises to Practice: Evaluating the Private Browsing Modes of Android Browser Apps 
 +  [D] liu2025_domino  /  liu2025_fingerprinting   first author: differs (Liu, Zhengyu | Liu, Zengrui) 
 +       2025 The DOMino Effect: Detecting and Exploiting DOM Clobbering Gadgets via Concolic Execution with Symbo 
 +       2025 The First Early Evidence of the Use of Browser Fingerprinting for Online Tracking 
 +  [D] liu2025_domino  /  liu2025_somesite   first author: differs (Liu, Zhengyu | Liu, Enze) 
 +       2025 The DOMino Effect: Detecting and Exploiting DOM Clobbering Gadgets via Concolic Execution with Symbo 
 +       2025 Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Craw 
 +  [D] liu2025_fingerprinting  /  liu2025_somesite   first author: differs (Liu, Zengrui | Liu, Enze) 
 +       2025 The First Early Evidence of the Use of Browser Fingerprinting for Online Tracking 
 +       2025 Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Craw 
 +  [D] nguyen2025_breaking  /  nguyen2025_please   first author: SAME (Nguyen, Hoang Dai | Nguyen, Hoang Dai) 
 +       2025 Breaking the Shield: Analyzing and Attacking Canvas Fingerprinting Defenses in the Wild 
 +       2025 "Please don't send that bot anything": A Mixed-methods Study of Personal Impersonation Attacks Targe 
 +  [D] nisenoff2023_awareness  /  nisenoff2023_defining   first author: SAME (Nisenoff, Alexandra | Nisenoff, Alexandra) 
 +       2023 User Awareness and Behaviors Concerning Encrypted {DNS} Settings in Web Browsers 
 +       2023 Defining "Broken": User Experiences and Remediation Tactics When Ad-Blocking or Tracking-Protection  
 +  [D] oest2020_phishtime  /  oest2020_sunrise   first author: SAME (Oest, Adam | Oest, Adam) 
 +       2020 PhishTime: Continuous Longitudinal Measurement of the Effectiveness of Anti-phishing Blacklists 
 +       2020 Sunrise to Sunset: Analyzing the End-to-end Life Cycle and Effectiveness of Phishing Attacks at Scal 
 +  [D] papadogiannakis2025_before  /  papadogiannakis2025_darkside   first author: SAME (Papadogiannakis, Emmanouil | Papadogiannakis, Emmanouil) 
 +       2025 Before \& After: The Effect of EU's 2022 Code of Practice on Disinformation 
 +       2025 Welcome to the Dark Side: Analyzing the Revenue Flows of Fraud in the Online Ad Ecosystem 
 +  [D] ruth2022_toppling  /  ruth2022_world   first author: SAME (Ruth, Kimberly | Ruth, Kimberly) 
 +       2022 Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists 
 +       2022 A World Wide View of Browsing the World Wide Web 
 +  [D] scheitle2018_long  /  scheitle2018_rise   first author: SAME (Scheitle, Quirin | Scheitle, Quirin) 
 +       2018 A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists 
 +       2018 The Rise of Certificate Transparency and Its Implications on the Internet Ecosystem 
 +  [D] starov2017_extended  /  starov2017_xhound   first author: SAME (Starov, Oleksii | Starov, Oleksii) 
 +       2017 Extended Tracking Powers: Measuring the Privacy Diffusion Enabled by Browser Extensions 
 +       2017 XHOUND: Quantifying the Fingerprintability of Browser Extensions 
 +  [D] tang2025_misuse  /  tang2025_navigating   first author: differs (Tang, Jenny | Tang, Brian) 
 +       2025 Misuse, Misreporting, Misinterpretation of Statistical Methods in Usable Privacy and Security Papers 
 +       2025 Navigating Cookie Consent Violations Across the Globe 
 +  [D] utz2023_comparing  /  utz2023_rarely   first author: SAME (Utz, Christine | Utz, Christine) 
 +       2023 Comparing Large-Scale Privacy and Security Notifications 
 +       2023 Privacy Rarely Considered: Exploring Considerations in the Adoption of Third-Party Services by Websi 
 +  [D] vastel2018_scanner  /  vastel2018_stalker   first author: SAME (Vastel, Antoine | Vastel, Antoine) 
 +       2018 Fp-Scanner: The Privacy Implications of Browser Fingerprint Inconsistencies 
 +       2018 FP-STALKER: Tracking Browser Fingerprint Evolutions 
 +  [D] vekaria2025_bighelp  /  vekaria2025_soktracking   first author: SAME (Vekaria, Yash | Vekaria, Yash) 
 +       2025 Big Help or Big Brother? Auditing Tracking, Profiling, and Personalization in Generative AI Assistan 
 +       2025 SoK: Advances and Open Problems in Web Tracking 
 +  [D] venkatadri2019_auditing  /  venkatadri2019_investigating   first author: SAME (Venkatadri, Giridhari | Venkatadri, Giridhari) 
 +       2019 Auditing Offline Data Brokers via Facebook's Advertising Platform 
 +       2019 Investigating sources of PII used in Facebook’s targeted advertising 
 +  [D] wang2026_masks  /  wang2026_sipconfusion   first author: differs (Wang, Weihong | Wang, Qi) 
 +       2026 The Masks We (Think We) Wear: Privacy Threats of Browser-Extension Wallets in the Web3 Ecosystem 
 +       2026 SIPConfusion: Exploiting SIP Semantic Ambiguities for Caller ID and SMS Spoofing 
 +  [D] wu2025_appprivacyreport  /  wu2025_revealing   first author: differs (Wu, Xiaoyuan | Wu, Mengying) 
 +       2025 Transparency or Information Overload? Evaluating Users’ Comprehension and Perceptions of the iOS App 
 +       2025 Revealing the Black Box of Device Search Engine: Scanning Assets, Strategies, and Ethical Considerat 
 +  [D] wu2026_email  /  wu2026_tracking   first author: differs (Wu, Mengying | Wu, Wenhao) 
 +       2026 One Email, Many Faces: A Deep Dive into Identity Confusion in Email Aliases 
 +       2026 Tracking the Stray Sheep: Understanding DNS Response Manipulation in the Wild 
 +  [D] xie2024_arcanum  /  xie2024_crawling   first author: SAME (Xie, Qinge | Xie, Qinge) 
 +       2024 Arcanum: Detecting and Evaluating the Privacy Risks of Browser Extensions on Web Pages and Web Conte 
 +       2024 Crawling to the Top: An Empirical Evaluation of Top List Use 
 +  [D] yang2022_extensive  /  yang2022_wtagraph   first author: differs (Yang, Mingshuo | Yang, Zhiju) 
 +       2022 An Extensive Study of Residential Proxies in China 
 +       2022 WTAGRAPH: Web Tracking and Advertising Detection using Graph Neural Networks 
 +  [D] zhang2022_harpo  /  zhang2022_spartacus   first author: differs (Zhang, Jiang | Zhang, Penghui) 
 +       2022 HARPO: Learning to Subvert Online Behavioral Advertising 
 +       2022 I'm SPARTACUS, No, I'm SPARTACUS: Proactively Protecting Users from Phishing by Intentionally Trigge 
 +  [D] zhang2024_inbox  /  zhang2024_quic   first author: differs (Zhang, Jiahe | Zhang, Xumiao) 
 +       2024 Inbox Invasion: Exploiting MIME Ambiguities to Evade Email Attachment Detectors 
 +       2024 QUIC is not Quick Enough over Fast Internet 
 +  [D] zhang2025_abusability  /  zhang2025_qrcode   first author: differs (Zhang, Shirley | Zhang, Xin) 
 +       2025 Abusability of Automation Apps in Intimate Partner Violence 
 +       2025 Demystifying the (In)Security of {QR} Code-based Login in Real-world Deployments 
 +  [D] zhu2020_label  /  zhu2020_vtset   first author: SAME (Zhu, Shuofei | Zhu, Shuofei) 
 +       2020 Measuring and Modeling the Label Dynamics of Online Anti-Malware Engines 
 +       2020 Demo: Benchmarking Label Dynamics of VirusTotal Engines 
 + 
 +candidates whose first-author string is identical on both sides: 31 of 58; different first author (shared surname, or a non-person 'author'): 27 
 +</code> 
 + 
 +==== Prose that still names the deleted keys ==== 
 + 
 +26 provenance pages name a deleted key in a run record or review log — "found 
 +and left alone", "this page cites ''bouhoula2024automated''". Those are 
 +statements about the bibliography as it was when they were written and were 
 +**not rewritten**; the list is the last block of the apply output above. The 
 +ten provenance pages whose content page was repointed each got a dated 
 +//Amendment, 2026-09-04: citekey consolidation// section instead, so a reader 
 +who follows one of those statements finds the correction on the same page. 
 + 
 +==== What could not be established, second sitting ==== 
 + 
 +  * Whether anything **outside the wiki** cites a deleted key — a BibTeX file someone exported from the site, a draft that copied ''fouad2022my''. Nothing on the site does; nothing off it can be checked. 
 +  * Whether ''khrer2015_going'' and ''som2017_content'' should be re-keyed. They are the ''[^a-z]'' bug in ''bibgen.mjs'', not duplicates, and each is cited; re-keying them is the same kind of wiki-wide rewrite as this item and was not done inside it. 
 + 
 +==== Judgement calls, second sitting ==== 
 + 
 +^ Call ^ Alternative a reasonable person would pick ^ Why this one ^ 
 +| Keep the stripped-diacritic, underscore key even where the other key was older and on more pages | Keep whichever key more pages used | The convention is what ''bibgen.mjs'' mints; keeping a minority style alive is how the pairs arose. 19 of 25 diacritic surnames in the file are already spelled this way | 
 +| Carry ''pages'' over to ''bouhoula2024_automated''; drop ''isbn'', ''address'', ''publisher'', ''month'' | Merge every field of the deleted entry | ''pages'' was verified against USENIX's page; the other four appear on no other entry, and adding them to one would make that entry the odd one | 
 +| Leave the 26 provenance pages' historical prose alone; amend only the 10 whose content page was repointed | Rewrite every mention of a deleted key | A review log records what was true when written. Rewriting it is the mistake the first sitting's own review found — a debunked claim surviving in the log — turned inside out | 
 +| Pages first, bibliography last, one ''--if-rev'' per save | Bibliography first, or one combined pass | The other order leaves 23 markers unresolved between two saves; ''--if-rev'' is what makes a 21-save run safe against an edit landing in the middle | 
 +| Publish the 58 candidates in full | Publish the count and the verdict | A hand judgement over a list nobody can read is not checkable | 
 +| Extend this page rather than create ''provenance:literature:bibliography:dedup'' | A page per item | The provenance page mirrors the content page's id; the bibliography has one id. Two sittings, two dated audits, one page | 
 + 
 +===== What could not be established, first sitting =====
  
   * **Whether defects 1 and 2 ever reached the wiki.** Zero missing authors survive in ''out/authors.json'' and zero in the live bibliography, but the cache carries no history, so it cannot be shown whether a dropped author was ever cached and later corrected or was never cached at all. The live page is clean either way, which is the question that matters.   * **Whether defects 1 and 2 ever reached the wiki.** Zero missing authors survive in ''out/authors.json'' and zero in the live bibliography, but the cache carries no history, so it cannot be shown whether a dropped author was ever cached and later corrected or was never cached at all. The live page is clean either way, which is the question that matters.
Line 603: Line 1064:
   * **Whether ''bibgen.mjs'' output was hand-cleaned, or the affected entries were written from another source.** The live entries are correct and the cache was not; the intermediate step left no record. The conclusion "the parser contaminated the cache and never the page" is a statement about the two endpoints, not about what happened between them.   * **Whether ''bibgen.mjs'' output was hand-cleaned, or the affected entries were written from another source.** The live entries are correct and the cache was not; the intermediate step left no record. The conclusion "the parser contaminated the cache and never the page" is a statement about the two endpoints, not about what happened between them.
  
-===== Judgement calls =====+===== Judgement calls, first sitting =====
  
 ^ Call ^ Alternative a reasonable person would pick ^ Why this one ^ ^ Call ^ Alternative a reasonable person would pick ^ Why this one ^
Line 613: Line 1074:
 | Correct ''fetch_authors.py'' and ''bibgen.mjs'' as part of this item | Report the parser bug and stop | The item asked for the cache and the bibliography to be corrected. Leaving the producer broken would mean the next fetch re-introduces exactly the fragments this run removed | | Correct ''fetch_authors.py'' and ''bibgen.mjs'' as part of this item | Report the parser bug and stop | The item asked for the cache and the bibliography to be corrected. Leaving the producer broken would mean the next fetch re-introduces exactly the fragments this run removed |
 | Extend the audit to all 159 live USENIX entries (pass 3), beyond the 90 the item named | Stop at the 90 keys | Pass 2 can only see an affiliation that got **in**, never an author that fell **out**, and 74 live entries are invisible to a cache-based check. 72 extra page fetches closed the gap | | Extend the audit to all 159 live USENIX entries (pass 3), beyond the 90 the item named | Stop at the 90 keys | Pass 2 can only see an affiliation that got **in**, never an author that fell **out**, and 74 live entries are invisible to a cache-based check. 72 extra page fetches closed the gap |
-| Defer the 5 duplicate entries | Consolidate them in the same sitting | ~30 pages would need %%{[key]}%% rewrites; nothing renders wrong today | +| Defer the 5 duplicate entries | Consolidate them in the same sitting | ~30 pages would need %%{[key]}%% rewrites; nothing renders wrong today. Done as its own item later the same day: 13 pages and 23 markers, not ~30 — see the second-sitting audit 
-| Put this page at ''provenance:literature:bibliography'' | Fold it into ''literature:corpus'' The spec mirrors the content id under ''provenance:''. ''literature:corpus'is corpus-level and is its own item; this is bibliography-level |+| Put this page at ''provenance:literature:bibliography'' | Fold it into [[:literature:corpus]] [[:literature:corpus]] states the rule this follows: a provenance page mirrors its content page's id exactly. It is also corpus-level — selection, the funnel, extraction reliability — while this is bibliography-level: where one file's author lists and DOIs came from | 
 +| Keep ''%%<bibtex bibliography>%%'' on a provenance page | Follow [[:literature:corpus]], which says provenance pages carry no bibliography | That convention was settled on a page that cites no papers. This one cites nine, and without the block the markers render as bare numbers with no reference list. The neighbouring provenance pages that cite papers all carry it. No ''%%~~DISCUSSION~~%%'' block, per the same convention |
  
 ===== Scripts ===== ===== Scripts =====
Line 2353: Line 2815:
   This is a smoke test over parse_popets output, not the three-source audit the USENIX keys got.   This is a smoke test over parse_popets output, not the three-source audit the USENIX keys got.
 </code> </code>
 +
 +==== Second sitting — the wider duplicate scan ====
 +
 +Run on the 855-entry export before the change (output above, under the audit)
 +and on the saved 850-entry page:
 +
 +<file python bib_dedup_scan.py>
 +#!/usr/bin/env python3
 +"""Find every pair of literature:bibliography entries that may describe one paper.
 +
 +scripts/bib_doi_duplicates.py matches on exact DOI and exact squashed title.
 +That found five pairs on 2026-09-04, and one of them was an umlaut
 +transliteration in the citekey (bottger / boettger). The same slip can also
 +land in the TITLE (a subtitle dropped, "{IPv6}" braced) or leave no shared DOI,
 +so this scan casts wider and prints its candidates for a human to judge:
 +
 +  A  same DOI                                   (definite)
 +  B  same squashed title                        (definite)
 +  C  near-identical title, difflib ratio >= 0.85 on the squashed form,
 +     or one squashed title a prefix of the other (subtitle dropped)
 +  D  same year AND same folded first-author surname — the fold strips
 +     diacritics and collapses the German digraphs oe/ue/ae/ss to o/u/a/s, so
 +     Böttger, B\"ottger, Boettger and Bottger all compare equal
 +
 +C and D are candidate generators, not verdicts. Every C/D pair that is not
 +already in A/B is printed with both titles so the residue is visible; the
 +verdicts recorded on provenance:literature:bibliography are hand judgements
 +over that printed list, not the script's.
 +
 +    python3 scripts/bib_dedup_scan.py --bib out/live_bibliography_YYYYMMDD.txt
 +"""
 +import difflib
 +import os
 +import re
 +import sys
 +import unicodedata
 +from collections import defaultdict
 +
 +sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
 +from usenix_bib_diff import field, squash
 +
 +LATEX = {r'\"o': 'oe', r'\"u': 'ue', r'\"a': 'ae', r'\"O': 'Oe', r'\"U': 'Ue',
 +         r'\"A': 'Ae', r'\ss': 'ss'}
 +
 +
 +def fold_surname(author_field):
 +    """First author's surname, lower-cased, diacritics stripped, German
 +    digraphs collapsed. Handles 'Last, First' and 'First Last'."""
 +    first = re.split(r"\s+and\s+", author_field)[0]
 +    s = first
 +    for k, v in LATEX.items():
 +        s = s.replace(k, v)
 +    s = re.sub(r"\\[\'`^~vcH=.]\{?(\w)\}?", r"\1", s)        # \'e, \v{c}, {\'e}
 +    s = re.sub(r"[{}]", "", s)
 +    s = unicodedata.normalize("NFKD", s)
 +    s = "".join(c for c in s if not unicodedata.combining(c))
 +    last = s.split(",")[0].strip() if "," in s else s.strip().split()[-1]
 +    last = last.lower()
 +    last = re.sub(r"[^a-z]", "", last)
 +    for dg, one in (("oe", "o"), ("ue", "u"), ("ae", "a"), ("ss", "s")):
 +        last = last.replace(dg, one)
 +    return last
 +
 +
 +def main():
 +    bib = sys.argv[sys.argv.index("--bib") + 1]
 +    text = open(bib, encoding="utf-8").read()
 +    entries = re.findall(r"@\w+\{[^@]*?\n\}", text, re.S)
 +    rec = {}
 +    for e in entries:
 +        k = re.match(r"@\w+\{([^,]+),", e).group(1)
 +        if k in rec:
 +            raise SystemExit(f"colliding citekey {k}")
 +        doi = field(e, "doi") or ""
 +        doi = re.sub(r"^https?://(dx\.)?doi\.org/", "", doi).lower()
 +        title = field(e, "title") or ""
 +        author = field(e, "author") or ""
 +        year = field(e, "year") or ""
 +        first = re.split(r"\s+and\s+", author)[0].strip() if author else ""
 +        rec[k] = dict(doi=doi, title=title, sq=squash(title), year=year,
 +                      surname=fold_surname(author) if author else "", first=first)
 +    print(f"bibliography : {bib}")
 +    print(f"entries      : {len(entries)}   distinct citekeys: {len(rec)}")
 +
 +    pairs = {}  # frozenset(k1,k2) -> set of rules
 +
 +    def add(a, b, rule):
 +        pairs.setdefault(frozenset((a, b)), set()).add(rule)
 +
 +    by = defaultdict(list)
 +    for k, r in rec.items():
 +        if r["doi"]:
 +            by[("A", r["doi"])].append(k)
 +        if r["sq"]:
 +            by[("B", r["sq"])].append(k)
 +        if r["year"] and r["surname"]:
 +            by[("D", r["year"], r["surname"])].append(k)
 +    for key, ks in by.items():
 +        for i in range(len(ks)):
 +            for j in range(i + 1, len(ks)):
 +                add(ks[i], ks[j], key[0])
 +    keys = sorted(rec)
 +    for i in range(len(keys)):
 +        a = rec[keys[i]]["sq"]
 +        if len(a) < 20:
 +            continue
 +        for j in range(i + 1, len(keys)):
 +            b = rec[keys[j]]["sq"]
 +            if len(b) < 20:
 +                continue
 +            if a.startswith(b) or b.startswith(a) or \
 +               difflib.SequenceMatcher(None, a, b).ratio() >= 0.85:
 +                add(keys[i], keys[j], "C")
 +
 +    definite = {p: r for p, r in pairs.items() if r & {"A", "B"}}
 +    candidates = {p: r for p, r in pairs.items() if not r & {"A", "B"}}
 +    print(f"\nDEFINITE duplicate pairs (rule A or B): {len(definite)}")
 +    for p, rules in sorted(definite.items(), key=lambda x: sorted(x[0])):
 +        a, b = sorted(p)
 +        print(f"  [{''.join(sorted(rules))}] {a}  /  {b}")
 +        print(f"       {rec[a]['title'][:90]}")
 +        print(f"       {rec[b]['title'][:90]}")
 +    print(f"\nCANDIDATE pairs (rule C or D only) — judged by hand, see the "
 +          f"provenance page: {len(candidates)}")
 +    # Rule D fires on a folded SURNAME, so it also pairs different people who
 +    # share a common surname (Li, Zhang, Liu). Print the first author's full
 +    # name string for both sides and count how many pairs are the same string,
 +    # so the residue can be described without hand-counting.
 +    same_person = 0
 +    for p, rules in sorted(candidates.items(), key=lambda x: sorted(x[0])):
 +        a, b = sorted(p)
 +        same = rec[a]["first"] == rec[b]["first"]
 +        same_person += same
 +        print(f"  [{''.join(sorted(rules))}] {a}  /  {b}   first author: "
 +              f"{'SAME' if same else 'differs'} ({rec[a]['first']} | {rec[b]['first']})")
 +        print(f"       {rec[a]['year']} {rec[a]['title'][:100]}")
 +        print(f"       {rec[b]['year']} {rec[b]['title'][:100]}")
 +    print(f"\ncandidates whose first-author string is identical on both sides: "
 +          f"{same_person} of {len(candidates)}; different first author (shared "
 +          f"surname, or a non-person 'author'): {len(candidates) - same_person}")
 +    return 1 if definite else 0
 +
 +
 +if __name__ == "__main__":
 +    sys.exit(main())
 +</file>
 +
 +<code>
 +bibliography : out/live_bibliography_20260904_after.txt
 +entries      : 850   distinct citekeys: 850
 +
 +DEFINITE duplicate pairs (rule A or B): 0
 +
 +CANDIDATE pairs (rule C or D only) — judged by hand, see the provenance page: 58
 +  [D] LePochat2019_tranco  /  LePochat2019_tranco_eval   first author: SAME ({Le Pochat}, Victor | {Le Pochat}, Victor)
 +       2019 Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation
 +       2019 Evaluating the Long-term Effects of Parameters on the Characteristics of the {Tranco} Top Sites Rank
 +  [D] agarwal2024_peeking  /  agarwal2024_poster   first author: differs (Agarwal, Shubham | Agarwal, Sharad)
 +       2024 Peeking through the window: Fingerprinting Browser Extensions through Page-Visible Execution Traces 
 +       2024 Poster: A Comprehensive Categorization of SMS Scams
 +  [D] agarwal2025_dropped  /  agarwal2025_fishing   first author: SAME (Agarwal, Sharad | Agarwal, Sharad)
 +       2025 'Hey mum, I dropped my phone down the toilet': Investigating Hi Mum and Dad SMS Scams in the United 
 +       2025 Fishing for Smishing: Understanding SMS Phishing Infrastructure and Strategies by Mining Public User
 +  [D] agarwal2025_dropped  /  agarwal2025_mindsets   first author: differs (Agarwal, Sharad | Agarwal, Shubham)
 +       2025 'Hey mum, I dropped my phone down the toilet': Investigating Hi Mum and Dad SMS Scams in the United 
 +       2025 "I have no idea how to make it safer": Studying Security and Privacy Mindsets of Browser Extension D
 +  [D] agarwal2025_fishing  /  agarwal2025_mindsets   first author: differs (Agarwal, Sharad | Agarwal, Shubham)
 +       2025 Fishing for Smishing: Understanding SMS Phishing Infrastructure and Strategies by Mining Public User
 +       2025 "I have no idea how to make it safer": Studying Security and Privacy Mindsets of Browser Extension D
 +  [D] alroomi2023_login  /  alroomi2023_password   first author: differs (Al Roomi, Suood | Alroomi, Suood)
 +       2023 A Large-Scale Measurement of Website Login Policies
 +       2023 Measuring Website Password Creation Policies At Scale
 +  [D] bahrami2025_bytedefender  /  bahrami2025_cookieguard   first author: SAME (Nikkhah Bahrami, Pouneh | Nikkhah Bahrami, Pouneh)
 +       2025 Byte by Byte: Unmasking Browser Fingerprinting at the Function Level Using V8 Bytecode Transformers
 +       2025 {CookieGuard}: Characterizing and Isolating the First-Party Cookie Jar
 +  [D] bashir2019_adstxt  /  bashir2019_quantity   first author: SAME (Bashir, Muhammad Ahmad | Bashir, Muhammad Ahmad)
 +       2019 A Longitudinal Analysis of the ads.txt Standard
 +       2019 Quantity vs. Quality: Evaluating User Interest Profiles Using Ad Preference Managers
 +  [D] bhuiyan2025_digital  /  bhuiyan2025_visitors   first author: SAME (Bhuiyan, Masudul Hasan Masud | Bhuiyan, Masudul Hasan Masud)
 +       2025 Digital Disparities: A Comparative Web Measurement Study Across Economic Boundaries
 +       2025 Not All Visitors are Bilingual: A Measurement Study of the Multilingual Web from an Accessibility Pe
 +  [C] bratton2019_replication  /  sumner2014_exaggeration   first author: differs (Bratton, Luke | Sumner, Petroc)
 +       2019 The Association Between Exaggeration in Health-Related Science News and Academic Press Releases: A R
 +       2014 The Association Between Exaggeration in Health Related Science News and Academic Press Releases: Ret
 +  [D] chen2021_cookieswap  /  chen2021_detecting   first author: SAME (Chen, Quan | Chen, Quan)
 +       2021 Cookie Swap Party: Abusing First-Party Cookies for Web Tracking
 +       2021 Detecting Filter List Evasion with Event-Loop-Turn Granularity JavaScript Signatures
 +  [D] chen2025_parents  /  chen2025_semantics   first author: differs (Chen, Xiaowei | Chen, Baiqi)
 +       2025 Empowering Parents to Support Children's Online Security and Privacy: Findings from a Randomized Con
 +       2025 Semantics-Aware Cookie Purpose Compliance
 +  [CD] duckduckgo_tracker_radar_2026  /  duckduckgo_tracker_radar_detector_2026   first author: SAME ({DuckDuckGo} | {DuckDuckGo})
 +       2026 DuckDuckGo Tracker Radar
 +       2026 DuckDuckGo Tracker Radar Detector
 +  [D] duckduckgo_tracker_radar_2026  /  duckduckgo_trc_2026   first author: SAME ({DuckDuckGo} | {DuckDuckGo})
 +       2026 DuckDuckGo Tracker Radar
 +       2026 Tracker Radar Collector
 +  [D] duckduckgo_tracker_radar_detector_2026  /  duckduckgo_trc_2026   first author: SAME ({DuckDuckGo} | {DuckDuckGo})
 +       2026 DuckDuckGo Tracker Radar Detector
 +       2026 Tracker Radar Collector
 +  [D] durumeric2013_https  /  durumeric2013_zmap   first author: SAME (Durumeric, Zakir | Durumeric, Zakir)
 +       2013 Analysis of the HTTPS certificate ecosystem
 +       2013 {ZMap}: Fast Internet-wide Scanning and Its Security Applications
 +  [D] durumeric2014_heartbleed  /  durumeric2014_view   first author: SAME (Durumeric, Zakir | Durumeric, Zakir)
 +       2014 The Matter of Heartbleed
 +       2014 An Internet-Wide View of Internet-Wide Scanning
 +  [D] durumeric2015_neither  /  durumeric2015_search   first author: SAME (Durumeric, Zakir | Durumeric, Zakir)
 +       2015 Neither Snow Nor Rain Nor MITM...: An Empirical Analysis of Email Delivery Security
 +       2015 A Search Engine Backed by Internet-Wide Scanning
 +  [D] edu2022_alexa  /  edu2022_exploring   first author: SAME (Edu, Jide S. | Edu, Jide S.)
 +       2022 Measuring Alexa Skill Privacy Practices across Three Years
 +       2022 Exploring the security and privacy risks of chatbots in messaging services
 +  [D] iqbal2022_khaleesi  /  iqbal2022_left   first author: differs (Umar Iqbal | Iqbal, Hassan)
 +       2022 Khaleesi: Breaker of Advertising and Tracking Request Chains
 +       2022 Left or Right: A Peek into the Political Biases in Email Spam Filtering Algorithms During US Electio
 +  [D] kancherla2025_johnny  /  kancherla2025_least   first author: SAME (Kancherla, Gayatri Priyadarsini | Kancherla, Gayatri Priyadarsini)
 +       2025 Johnny Can't Revoke Consent Either: Measuring Compliance of Consent Revocation on the Web
 +       2025 Least Privilege Access for Persistent Storage Mechanisms in Web Browsers
 +  [D] kirchner2024_black  /  kirchner2024_dancer   first author: SAME (Kirchner, Robin | Kirchner, Robin)
 +       2024 A Black-Box Privacy Analysis of Messaging Service Providers' Chat Message Processing
 +       2024 Dancer in the Dark: Synthesizing and Evaluating Polyglots for Blind Cross-Site Scripting
 +  [D] lee2023_adtargeting  /  lee2023_track   first author: differs (Lee, Hao-Ping (Hank) | Lee, Dongkeun)
 +       2023 When and Why Do People Want Ad Targeting Explanations? Evidence from a Four-Week, Mixed-Methods Fiel
 +       2023 Net-track: Generic Web Tracking Detection Using Packet Metadata
 +  [D] li2016_remedying  /  li2016_youve   first author: SAME (Li, Frank | Li, Frank)
 +       2016 Remedying Web Hijacking: Notification Effectiveness and Webmaster Comprehension
 +       2016 You've Got Vulnerability: Exploring Effective Vulnerability Notifications
 +  [D] li2017_radar  /  li2017_security   first author: differs (Li, Zhenhua | Li, Frank)
 +       2017 FBS-Radar: Uncovering Fake Base Stations at Scale in the Wild
 +       2017 A Large-Scale Empirical Study of Security Patches
 +  [D] li2017_radar  /  li2017_static   first author: differs (Li, Zhenhua | Li, Li)
 +       2017 FBS-Radar: Uncovering Fake Base Stations at Scale in the Wild
 +       2017 Static analysis of Android apps: A systematic literature review
 +  [D] li2017_security  /  li2017_static   first author: differs (Li, Frank | Li, Li)
 +       2017 A Large-Scale Empirical Study of Security Patches
 +       2017 Static analysis of Android apps: A systematic literature review
 +  [D] li2024_bounce  /  li2024_wellinformed   first author: differs (Li, Ruixuan | Li, Shuai)
 +       2024 Bounce in the Wild: A Deep Dive into Email Delivery Failures from a Large Email Service Provider
 +       2024 Are We Getting Well-informed? An In-depth Study of Runtime Privacy Notice Practice in Mobile Apps
 +  [D] li2024_bounce  /  li2024_worldwide   first author: SAME (Li, Ruixuan | Li, Ruixuan)
 +       2024 Bounce in the Wild: A Deep Dive into Email Delivery Failures from a Large Email Service Provider
 +       2024 A Worldwide View on the Reachability of Encrypted DNS Services
 +  [D] li2024_wellinformed  /  li2024_worldwide   first author: differs (Li, Shuai | Li, Ruixuan)
 +       2024 Are We Getting Well-informed? An In-depth Study of Runtime Privacy Notice Practice in Mobile Apps
 +       2024 A Worldwide View on the Reachability of Encrypted DNS Services
 +  [D] liao2016_characterizing  /  liao2016_seeking   first author: SAME (Liao, Xiaojing | Liao, Xiaojing)
 +       2016 Characterizing Long-tail SEO Spam on Cloud Web Hosting Services
 +       2016 Seeking Nonsense, Looking for Trouble: Efficient Promotional-Infection Detection through Semantic In
 +  [D] lin2021_longitudinal  /  lin2021_phishpedia   first author: differs (Lin, Fuqi | Lin, Yun)
 +       2021 A Longitudinal Study of Removed Apps in {iOS} App Store
 +       2021 Phishpedia: A Hybrid Deep Learning Based Approach to Visually Identify Phishing Webpages
 +  [D] lin2022_investigating  /  lin2022_sheep   first author: differs (Lin, Su-Chin | Lin, Xu)
 +       2022 Investigating Advertisers' Domain-changing Behaviors and Their Impacts on Ad-blocker Filter Lists
 +       2022 Phish in Sheep's Clothing: Exploring the Authentication Pitfalls of Browser Fingerprinting
 +  [D] liu2024_opted  /  liu2024_promises   first author: differs (Liu, Zengrui | Liu, Xiaoyin)
 +       2024 Opted Out, Yet Tracked: Are Regulations Enough to Protect Your Privacy?
 +       2024 From Promises to Practice: Evaluating the Private Browsing Modes of Android Browser Apps
 +  [D] liu2025_domino  /  liu2025_fingerprinting   first author: differs (Liu, Zhengyu | Liu, Zengrui)
 +       2025 The DOMino Effect: Detecting and Exploiting DOM Clobbering Gadgets via Concolic Execution with Symbo
 +       2025 The First Early Evidence of the Use of Browser Fingerprinting for Online Tracking
 +  [D] liu2025_domino  /  liu2025_somesite   first author: differs (Liu, Zhengyu | Liu, Enze)
 +       2025 The DOMino Effect: Detecting and Exploiting DOM Clobbering Gadgets via Concolic Execution with Symbo
 +       2025 Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Craw
 +  [D] liu2025_fingerprinting  /  liu2025_somesite   first author: differs (Liu, Zengrui | Liu, Enze)
 +       2025 The First Early Evidence of the Use of Browser Fingerprinting for Online Tracking
 +       2025 Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Craw
 +  [D] nguyen2025_breaking  /  nguyen2025_please   first author: SAME (Nguyen, Hoang Dai | Nguyen, Hoang Dai)
 +       2025 Breaking the Shield: Analyzing and Attacking Canvas Fingerprinting Defenses in the Wild
 +       2025 "Please don't send that bot anything": A Mixed-methods Study of Personal Impersonation Attacks Targe
 +  [D] nisenoff2023_awareness  /  nisenoff2023_defining   first author: SAME (Nisenoff, Alexandra | Nisenoff, Alexandra)
 +       2023 User Awareness and Behaviors Concerning Encrypted {DNS} Settings in Web Browsers
 +       2023 Defining "Broken": User Experiences and Remediation Tactics When Ad-Blocking or Tracking-Protection 
 +  [D] oest2020_phishtime  /  oest2020_sunrise   first author: SAME (Oest, Adam | Oest, Adam)
 +       2020 PhishTime: Continuous Longitudinal Measurement of the Effectiveness of Anti-phishing Blacklists
 +       2020 Sunrise to Sunset: Analyzing the End-to-end Life Cycle and Effectiveness of Phishing Attacks at Scal
 +  [D] papadogiannakis2025_before  /  papadogiannakis2025_darkside   first author: SAME (Papadogiannakis, Emmanouil | Papadogiannakis, Emmanouil)
 +       2025 Before \& After: The Effect of EU's 2022 Code of Practice on Disinformation
 +       2025 Welcome to the Dark Side: Analyzing the Revenue Flows of Fraud in the Online Ad Ecosystem
 +  [D] ruth2022_toppling  /  ruth2022_world   first author: SAME (Ruth, Kimberly | Ruth, Kimberly)
 +       2022 Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists
 +       2022 A World Wide View of Browsing the World Wide Web
 +  [D] scheitle2018_long  /  scheitle2018_rise   first author: SAME (Scheitle, Quirin | Scheitle, Quirin)
 +       2018 A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists
 +       2018 The Rise of Certificate Transparency and Its Implications on the Internet Ecosystem
 +  [D] starov2017_extended  /  starov2017_xhound   first author: SAME (Starov, Oleksii | Starov, Oleksii)
 +       2017 Extended Tracking Powers: Measuring the Privacy Diffusion Enabled by Browser Extensions
 +       2017 XHOUND: Quantifying the Fingerprintability of Browser Extensions
 +  [D] tang2025_misuse  /  tang2025_navigating   first author: differs (Tang, Jenny | Tang, Brian)
 +       2025 Misuse, Misreporting, Misinterpretation of Statistical Methods in Usable Privacy and Security Papers
 +       2025 Navigating Cookie Consent Violations Across the Globe
 +  [D] utz2023_comparing  /  utz2023_rarely   first author: SAME (Utz, Christine | Utz, Christine)
 +       2023 Comparing Large-Scale Privacy and Security Notifications
 +       2023 Privacy Rarely Considered: Exploring Considerations in the Adoption of Third-Party Services by Websi
 +  [D] vastel2018_scanner  /  vastel2018_stalker   first author: SAME (Vastel, Antoine | Vastel, Antoine)
 +       2018 Fp-Scanner: The Privacy Implications of Browser Fingerprint Inconsistencies
 +       2018 FP-STALKER: Tracking Browser Fingerprint Evolutions
 +  [D] vekaria2025_bighelp  /  vekaria2025_soktracking   first author: SAME (Vekaria, Yash | Vekaria, Yash)
 +       2025 Big Help or Big Brother? Auditing Tracking, Profiling, and Personalization in Generative AI Assistan
 +       2025 SoK: Advances and Open Problems in Web Tracking
 +  [D] venkatadri2019_auditing  /  venkatadri2019_investigating   first author: SAME (Venkatadri, Giridhari | Venkatadri, Giridhari)
 +       2019 Auditing Offline Data Brokers via Facebook's Advertising Platform
 +       2019 Investigating sources of PII used in Facebook’s targeted advertising
 +  [D] wang2026_masks  /  wang2026_sipconfusion   first author: differs (Wang, Weihong | Wang, Qi)
 +       2026 The Masks We (Think We) Wear: Privacy Threats of Browser-Extension Wallets in the Web3 Ecosystem
 +       2026 SIPConfusion: Exploiting SIP Semantic Ambiguities for Caller ID and SMS Spoofing
 +  [D] wu2025_appprivacyreport  /  wu2025_revealing   first author: differs (Wu, Xiaoyuan | Wu, Mengying)
 +       2025 Transparency or Information Overload? Evaluating Users’ Comprehension and Perceptions of the iOS App
 +       2025 Revealing the Black Box of Device Search Engine: Scanning Assets, Strategies, and Ethical Considerat
 +  [D] wu2026_email  /  wu2026_tracking   first author: differs (Wu, Mengying | Wu, Wenhao)
 +       2026 One Email, Many Faces: A Deep Dive into Identity Confusion in Email Aliases
 +       2026 Tracking the Stray Sheep: Understanding DNS Response Manipulation in the Wild
 +  [D] xie2024_arcanum  /  xie2024_crawling   first author: SAME (Xie, Qinge | Xie, Qinge)
 +       2024 Arcanum: Detecting and Evaluating the Privacy Risks of Browser Extensions on Web Pages and Web Conte
 +       2024 Crawling to the Top: An Empirical Evaluation of Top List Use
 +  [D] yang2022_extensive  /  yang2022_wtagraph   first author: differs (Yang, Mingshuo | Yang, Zhiju)
 +       2022 An Extensive Study of Residential Proxies in China
 +       2022 WTAGRAPH: Web Tracking and Advertising Detection using Graph Neural Networks
 +  [D] zhang2022_harpo  /  zhang2022_spartacus   first author: differs (Zhang, Jiang | Zhang, Penghui)
 +       2022 HARPO: Learning to Subvert Online Behavioral Advertising
 +       2022 I'm SPARTACUS, No, I'm SPARTACUS: Proactively Protecting Users from Phishing by Intentionally Trigge
 +  [D] zhang2024_inbox  /  zhang2024_quic   first author: differs (Zhang, Jiahe | Zhang, Xumiao)
 +       2024 Inbox Invasion: Exploiting MIME Ambiguities to Evade Email Attachment Detectors
 +       2024 QUIC is not Quick Enough over Fast Internet
 +  [D] zhang2025_abusability  /  zhang2025_qrcode   first author: differs (Zhang, Shirley | Zhang, Xin)
 +       2025 Abusability of Automation Apps in Intimate Partner Violence
 +       2025 Demystifying the (In)Security of {QR} Code-based Login in Real-world Deployments
 +  [D] zhu2020_label  /  zhu2020_vtset   first author: SAME (Zhu, Shuofei | Zhu, Shuofei)
 +       2020 Measuring and Modeling the Label Dynamics of Online Anti-Malware Engines
 +       2020 Demo: Benchmarking Label Dynamics of VirusTotal Engines
 +
 +candidates whose first-author string is identical on both sides: 31 of 58; different first author (shared surname, or a non-person 'author'): 27
 +</code>
 +
 +==== Second sitting — citekey convention census ====
 +
 +<file python bib_key_convention.py>
 +#!/usr/bin/env python3
 +"""How does literature:bibliography spell a first author's diacritic in the citekey?
 +
 +Needed to pick between bottger2025_regional and boettger2025_regional on
 +evidence rather than taste. For every entry whose FIRST AUTHOR'S SURNAME
 +carries a non-ASCII letter or a LaTeX accent command, classify how the key's
 +surname part renders it:
 +
 +  stripped   diacritic removed, base letter kept   (Böttger -> bottger, Sjösten -> sjosten)
 +  digraph    German transliteration                (Böttger -> boettger, Rüth -> rueth)
 +  dropped    the accented letter vanished entirely (Kührer -> khrer — the
 +             bibgen.mjs [^a-z] filter, not a convention)
 +  other      none of the above (printed; judge by hand)
 +
 +    python3 scripts/bib_key_convention.py --bib out/live_bibliography_X.txt
 +"""
 +import os
 +import re
 +import sys
 +import unicodedata
 +from collections import Counter
 +
 +sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
 +from usenix_bib_diff import field
 +
 +LATEX_ACCENT = re.compile(r"\\[\"'`^~vcH=.]\s*\{?\\?(\w)\}?|\{\\[\"'`^~vcH=.]\s*(\w)\}")
 +LETTER_CMD = {r"\ss": "ss", r"\o": "o", r"\O": "O", r"\l": "l", r"\L": "L",
 +              r"\ae": "ae", r"\i": "i"}
 +
 +
 +def first_surname(author):
 +    first = re.split(r"\s+and\s+", author)[0]
 +    return first.split(",")[0].strip() if "," in first else first.strip().split()[-1]
 +
 +
 +def de_latex(s):
 +    s = LATEX_ACCENT.sub(lambda m: m.group(1) or m.group(2), s)
 +    for k, v in LETTER_CMD.items():
 +        s = s.replace(k, v)
 +    return re.sub(r"[{}]", "", s)
 +
 +
 +def has_diacritic(s):
 +    return bool(re.search(r"[^\x00-\x7f]", s) or LATEX_ACCENT.search(s)
 +                or any(k in s for k in LETTER_CMD))
 +
 +
 +def stripped(s):
 +    s = unicodedata.normalize("NFKD", de_latex(s))
 +    s = "".join(c for c in s if not unicodedata.combining(c))
 +    return re.sub(r"[^a-z]", "", s.lower().replace("ß", "ss").replace("ø", "o"))
 +
 +
 +def digraph(s):
 +    # LaTeX umlauts first (\"o, {\"o}, \"{o}), then literal ones, then de_latex
 +    # for whatever accents remain.
 +    s = re.sub(r'\{?\\"\{?([aouAOU])\}?\}?', lambda m: m.group(1) + "e", s)
 +    for a, b in (("ö", "oe"), ("ü", "ue"), ("ä", "ae"), ("Ö", "Oe"), ("Ü", "Ue"),
 +                 ("Ä", "Ae"), ("ß", "ss")):
 +        s = s.replace(a, b)
 +    s = de_latex(s)
 +    s = unicodedata.normalize("NFKD", s)
 +    s = "".join(c for c in s if not unicodedata.combining(c))
 +    return re.sub(r"[^a-z]", "", s.lower())
 +
 +
 +def dropped(s):
 +    s = de_latex(s)
 +    return re.sub(r"[^a-z]", "", "".join(c for c in s if ord(c) < 128).lower())
 +
 +
 +def main():
 +    bib = sys.argv[sys.argv.index("--bib") + 1]
 +    text = open(bib, encoding="utf-8").read()
 +    rows, kinds = [], Counter()
 +    for e in re.findall(r"@\w+\{[^@]*?\n\}", text, re.S):
 +        key = re.match(r"@\w+\{([^,]+),", e).group(1)
 +        author = field(e, "author")
 +        if not author:
 +            continue
 +        sur = first_surname(author)
 +        if not has_diacritic(sur):
 +            continue
 +        keysur = re.match(r"[A-Za-z]+", key).group(0).lower()
 +        if keysur == stripped(sur):
 +            kind = "stripped"
 +        elif keysur == digraph(sur) and digraph(sur) != stripped(sur):
 +            kind = "digraph"
 +        elif keysur == dropped(sur):
 +            kind = "dropped"
 +        else:
 +            kind = "other"
 +        kinds[kind] += 1
 +        rows.append((kind, key, sur))
 +    print(f"bibliography : {bib}")
 +    print(f"entries whose first author's SURNAME carries a diacritic: {len(rows)}")
 +    for k in ("stripped", "digraph", "dropped", "other"):
 +        print(f"  {k:9s} {kinds[k]}")
 +    print()
 +    for kind, key, sur in sorted(rows):
 +        print(f"  {kind:9s} {key:30s} {sur}")
 +
 +
 +if __name__ == "__main__":
 +    main()
 +</file>
 +
 +<code>
 +bibliography : out/live_bibliography_20260904_1721.txt
 +entries whose first author's SURNAME carries a diacritic: 25
 +  stripped  19
 +  digraph   4
 +  dropped   2
 +  other     0
 +
 +  digraph   boettger2025_regional          Böttger
 +  digraph   rueth2018_digging              R{\"u}th
 +  digraph   schoeni2024_cookieblock        Sch\"oni
 +  digraph   stoever2023_owners             Stöver
 +  dropped   khrer2015_going                Kührer
 +  dropped   som2017_content                Somé
 +  stripped  bosch2016_tales                B\"osch
 +  stripped  bottger2025_regional           B\"ottger
 +  stripped  gomezboix2018_hiding           Gómez-Boix
 +  stripped  gross2021_iuipc                Groß
 +  stripped  gulyas2016_near                Gulyás
 +  stripped  korczynski2016_zone            Korczy{\'n}ski
 +  stripped  kubicek2022_emails             Kub{\'i}{\v{c}}ek
 +  stripped  lecuyer2014_xray               Lécuyer
 +  stripped  lecuyer2015_sunlight           Lécuyer
 +  stripped  lukic2026_mv3                  Luki\'c
 +  stripped  nenadic2026_swiss              Nenadi{\'c}
 +  stripped  sanchezrola2021_journey        Sánchez-Rola
 +  stripped  sanchezrola2023_rods           Sánchez-Rola
 +  stripped  sjosten2019_latex              Sj\"osten
 +  stripped  sjosten2020_filter             Sj\"osten
 +  stripped  some2019_empoweb               Somé
 +  stripped  sorensen2019_beforeafter       Sørensen
 +  stripped  tornberg2024_bestpractices     T\"{o}rnberg
 +  stripped  tramer2019_adversarial         Tram\`er
 +</code>
 +
 +==== Second sitting — the consolidation itself ====
 +
 +<file python bib_dedup_apply.py>
 +#!/usr/bin/env python3
 +"""Consolidate the five duplicate entries in literature:bibliography.
 +
 +Does NOT touch the wiki. It reads a fresh ?do=export_raw of every page and of
 +the bibliography, writes the rewritten files under --out, and prints a report
 +whose invariants must all hold before anything is saved:
 +
 +  * each loser key is present in the bibliography exactly once and is removed;
 +    each winner is present exactly once and is kept;
 +  * entry count drops by exactly len(MERGES);
 +  * every {[...]} marker naming a loser is rewritten to the winner; the number
 +    of keys replaced equals the number of loser occurrences counted beforehand;
 +  * no page cites both keys of a pair (else the rewrite would silently collapse
 +    two markers into one reference and the page's distinct-key count would
 +    drop) — the distinct-key count of every page is asserted unchanged;
 +  * after the rewrite, no marker anywhere names a key the new bibliography
 +    does not define, other than keys that were ALREADY unresolved before
 +    (printed as residue, never one of the losers).
 +
 +Winners follow the file's own majority convention — surname + year + '_' +
 +first title word, diacritics dropped — which is also what scripts/bibgen.mjs
 +mints. Which key of each pair is the winner is a decision recorded on
 +provenance:literature:bibliography, not something this script infers.
 +
 +    python3 scripts/bib_dedup_apply.py --bib out/live_bibliography_X.txt \
 +        --pages out/dedup_pages --out out/dedup_apply
 +"""
 +import glob
 +import os
 +import re
 +import sys
 +from collections import Counter
 +
 +# loser -> winner
 +MERGES = {
 +    "fouad2022my": "fouad2022_cookie",
 +    "boettger2025_regional": "bottger2025_regional",
 +    "ahmad2026_ipfp": "ahmad2026_more",
 +    "bouhoula2024automated": "bouhoula2024_automated",
 +    "lerner2016internet": "lerner2016_internet",
 +}
 +# A field the loser carried that the winner lacks and that was verified against
 +# the venue page on 2026-09-04 (USENIX's own BibTeX block and citation_firstpage /
 +# citation_lastpage meta tags on usenixsecurity24/presentation/bouhoula).
 +EXTRA_FIELDS = {
 +    "bouhoula2024_automated": [("pages", "1723--1739")],
 +}
 +BIBPAGE = "literature__bibliography"
 +MARKER = re.compile(r"\{\[([A-Za-z0-9_:.\-, ]+)\]\}")
 +TEMPLATE_KEYS = {"CitationKey"  # the i_template's example marker
 +
 +# Provenance pages that get a dated amendment because their content page's
 +# markers were repointed. content page id -> provenance page id.
 +PROVENANCE_OF = {
 +    "privacy:browser_storage": "provenance:privacy:browser_storage",
 +    "privacy:fingerprinting": "provenance:privacy:fingerprinting",
 +    "programming:crawler:openwpm": "provenance:programming:crawler:openwpm",
 +    "programming:stateful_stateless": "provenance:programming:stateful_stateless",
 +    "statistics:pvalue_corrections": "provenance:statistics:pvalue_corrections",
 +    "design:longitudinal": "provenance:design:longitudinal",
 +    "design:crawling_location": "provenance:design:crawling_location",
 +    "privacy:consent": "provenance:privacy:consent",
 +    "privacy:requests": "provenance:privacy:requests",
 +    "design:archives": "provenance:design:archives",
 +}
 +
 +
 +def pid(fname):
 +    return os.path.basename(fname)[:-4].replace("__", ":")
 +
 +
 +def entries_of(text):
 +    return re.findall(r"@\w+\{[^@]*?\n\}", text, re.S)
 +
 +
 +def key_of(entry):
 +    return re.match(r"@\w+\{([^,]+),", entry).group(1)
 +
 +
 +def rewrite_bib(text):
 +    ents = entries_of(text)
 +    keys = Counter(key_of(e) for e in ents)
 +    for lo, wi in MERGES.items():
 +        assert keys[lo] == 1, f"loser {lo} appears {keys[lo]} times"
 +        assert keys[wi] == 1, f"winner {wi} appears {keys[wi]} times"
 +    out = text
 +    for e in ents:
 +        k = key_of(e)
 +        if k in MERGES:
 +            # remove the entry and one preceding blank line
 +            assert out.count(e) == 1
 +            out = out.replace("\n" + e, "", 1) if ("\n" + e) in out else out.replace(e, "", 1)
 +        elif k in EXTRA_FIELDS:
 +            new = e
 +            for fname, val in EXTRA_FIELDS[k]:
 +                assert not re.search(r"\b" + fname + r"\s*=", e), f"{k} already has {fname}"
 +                new = new[:-1].rstrip("\n") + f"\n  {fname:<13} = {{{val}}},\n}}"
 +            assert out.count(e) == 1
 +            out = out.replace(e, new, 1)
 +    after = entries_of(out)
 +    assert len(after) == len(ents) - len(MERGES), (len(ents), len(after))
 +    for lo in MERGES:
 +        assert not re.search(r"\b" + re.escape(lo) + r"\b", out), f"{lo} still in bib"
 +    return out, len(ents), len(after), {key_of(e) for e in after}
 +
 +
 +def rewrite_markers(text):
 +    """Return (new_text, replaced_count, distinct_before, distinct_after)."""
 +    replaced = 0
 +    before, after = set(), set()
 +
 +    def sub(m):
 +        nonlocal replaced
 +        keys = [k.strip() for k in m.group(1).split(",")]
 +        before.update(keys)
 +        new = []
 +        for k in keys:
 +            if k in MERGES:
 +                replaced += 1
 +                k = MERGES[k]
 +            if k not in new:
 +                new.append(k)
 +        after.update(new)
 +        return "{[" + ",".join(new) + "]}"
 +
 +    return MARKER.sub(sub, text), replaced, before, after
 +
 +
 +def amendment(content_id, pairs, n_content, n_prov):
 +    lines = ["", "===== Amendment, 2026-09-04: citekey consolidation =====", ""]
 +    for lo, wi in pairs:
 +        lines.append(
 +            f"  * ''{lo}'' was one of two keys for the same paper in "
 +            f"[[:literature:bibliography]]. The wiki-wide consolidation of "
 +            f"2026-09-04 (drain item ''dedup-regional-filter-lists-bibkey'') kept "
 +            f"''{wi}'' and deleted the other entry.")
 +    def n_markers(n):
 +        return f"{n} citation marker" + ("s" if n != 1 else "")
 +    where = f"{n_markers(n_content)} on [[:{content_id}]]"
 +    if n_prov:
 +        where += f" and {n_markers(n_prov)} on this page"
 +    verb = "were" if n_content + n_prov != 1 else "was"
 +    lines.append(
 +        f"  * {where} {verb} repointed to the kept key. No prose on either page "
 +        f"changed, and no figure moved. Statements above that name the deleted "
 +        f"key describe the state when they were written. Full query log and the "
 +        f"invariants checked before saving: "
 +        f"[[:provenance:literature:bibliography]].")
 +    lines.append("")
 +    return "\n".join(lines)
 +
 +
 +def insert_amendment(text, note):
 +    i = text.find("====== References ======")
 +    if i >= 0:
 +        return text[:i].rstrip("\n") + "\n" + note + "\n" + text[i:]
 +    body = text.rstrip("\n").split("\n")
 +    if body[-1].startswith("[["):            # trailing back-link line
 +        return "\n".join(body[:-1]).rstrip("\n") + "\n" + note + "\n" + body[-1] + "\n"
 +    return text.rstrip("\n") + "\n" + note
 +
 +
 +def main():
 +    a = sys.argv
 +    bib_path = a[a.index("--bib") + 1]
 +    pages_dir = a[a.index("--pages") + 1]
 +    out_dir = a[a.index("--out") + 1]
 +    os.makedirs(out_dir, exist_ok=True)
 +
 +    bib_text = open(bib_path, encoding="utf-8").read()
 +    new_bib, n_before, n_after, new_keys = rewrite_bib(bib_text)
 +    open(os.path.join(out_dir, BIBPAGE + ".txt"), "w", encoding="utf-8").write(new_bib)
 +    old_keys = {key_of(e) for e in entries_of(bib_text)}
 +    print(f"bibliography : {bib_path}")
 +    print(f"entries      : {n_before} -> {n_after}  (removed {n_before - n_after}: "
 +          f"{', '.join(MERGES)})")
 +    for k, fs in EXTRA_FIELDS.items():
 +        print(f"field added  : {k}  {fs}")
 +
 +    files = sorted(f for f in glob.glob(os.path.join(pages_dir, "*.txt"))
 +                   if pid(f) != BIBPAGE.replace("__", ":"))
 +    print(f"pages read   : {len(files)} (bibliography excluded)")
 +
 +    # occurrences of loser keys inside markers, before
 +    loser_occ = Counter()
 +    for f in files:
 +        for m in MARKER.finditer(open(f, encoding="utf-8").read()):
 +            for k in (x.strip() for x in m.group(1).split(",")):
 +                if k in MERGES:
 +                    loser_occ[k] += 1
 +    print("\nloser-key marker occurrences before, by key:")
 +    for lo in MERGES:
 +        print(f"  {lo:26s} {loser_occ[lo]}")
 +
 +    # Which pages cited each key of a pair before the rewrite — content pages and
 +    # provenance pages separately, so the "which key was on more pages" question
 +    # on the provenance page has a printed answer.
 +    def citing(key):
 +        out = []
 +        for f in files:
 +            keys = {k.strip() for m in MARKER.finditer(open(f, encoding="utf-8").read())
 +                    for k in m.group(1).split(",")}
 +            if key in keys:
 +                out.append(pid(f))
 +        return out
 +
 +    print("\npages citing each key BEFORE (content pages; provenance pages in brackets):")
 +    for lo, wi in MERGES.items():
 +        for k in (wi, lo):
 +            ps = citing(k)
 +            c = [p_ for p_ in ps if not p_.startswith("provenance:")]
 +            pr = [p_ for p_ in ps if p_.startswith("provenance:")]
 +            tag = "kept   " if k == wi else "deleted"
 +            print(f"  {tag} {k:26s} {len(c)} [{len(pr)}]  {', '.join(c)}"
 +                  + (f"  [{', '.join(pr)}]" if pr else ""))
 +
 +    changed = {}
 +    total_replaced = 0
 +    unresolved_before, unresolved_after = Counter(), Counter()
 +    all_after_keys = set()
 +    print("\n^ page ^ markers repointed ^ distinct keys before ^ after ^ pairs ^")
 +    for f in files:
 +        text = open(f, encoding="utf-8").read()
 +        new, n, kb, ka = rewrite_markers(text)
 +        kb -= TEMPLATE_KEYS
 +        ka -= TEMPLATE_KEYS
 +        for k in kb - old_keys:
 +            unresolved_before[k] += 1
 +        for k in ka - new_keys:
 +            unresolved_after[k] += 1
 +        all_after_keys |= ka
 +        if n:
 +            assert len(kb) == len(ka), f"{pid(f)}: distinct-key count changed {len(kb)}->{len(ka)}"
 +            pairs = sorted((lo, MERGES[lo]) for lo in kb if lo in MERGES)
 +            changed[pid(f)] = (new, n, pairs)
 +            total_replaced += n
 +            print(f"| {pid(f)} | {n} | {len(kb)} | {len(ka)} | "
 +                  f"{'; '.join(f'{lo}→{wi}' for lo, wi in pairs)} |")
 +    print(f"\npages changed          : {len(changed)}")
 +    print(f"markers repointed      : {total_replaced}  (loser occurrences before: "
 +          f"{sum(loser_occ.values())})")
 +    assert total_replaced == sum(loser_occ.values())
 +    for lo in MERGES:
 +        assert lo not in all_after_keys, f"{lo} survives in a marker"
 +
 +    # provenance amendments
 +    prov_texts = {pid(f): open(f, encoding="utf-8").read() for f in files}
 +    n_prov_notes = 0
 +    for cid, prov in PROVENANCE_OF.items():
 +        assert cid in changed, f"{cid} listed in PROVENANCE_OF but unchanged"
 +        _, n_content, pairs = changed[cid]
 +        n_prov = changed[prov][1] if prov in changed else 0
 +        base = changed[prov][0] if prov in changed else prov_texts[prov]
 +        assert "citekey consolidation" not in base, f"{prov} already amended"
 +        changed[prov] = (insert_amendment(base, amendment(cid, pairs, n_content, n_prov)),
 +                         n_prov, pairs)
 +        n_prov_notes += 1
 +    for cid in changed:
 +        if not cid.startswith("provenance:"):
 +            assert cid in PROVENANCE_OF, f"{cid} changed but has no provenance page listed"
 +    print(f"provenance amendments  : {n_prov_notes}")
 +
 +    for id_, (text, _, _) in changed.items():
 +        open(os.path.join(out_dir, id_.replace(":", "__") + ".txt"), "w",
 +             encoding="utf-8").write(text)
 +    print(f"files written          : {len(changed) + 1} under {out_dir}/")
 +
 +    print(f"\nmarkers naming a key the bibliography does not define, BEFORE: "
 +          f"{sum(unresolved_before.values())} on {len(unresolved_before)} key(s)")
 +    for k, n in sorted(unresolved_before.items()):
 +        print(f"  {k}  ({n} page(s))")
 +    print(f"same, AFTER: {sum(unresolved_after.values())} on {len(unresolved_after)} key(s)")
 +    for k, n in sorted(unresolved_after.items()):
 +        print(f"  {k}  ({n} page(s))")
 +    # The per-key counts above overlap: one page can carry several example
 +    # strings. The number of DISTINCT pages carrying any of them is what the
 +    # provenance page quotes, so print it rather than leave it to be summed.
 +    union = set()
 +    for f in files:
 +        text = changed[pid(f)][0] if pid(f) in changed else open(f, encoding="utf-8").read()
 +        for m in MARKER.finditer(text):
 +            if any(k.strip() in unresolved_after for k in m.group(1).split(",")):
 +                union.add(pid(f))
 +    print(f"distinct pages carrying any such example marker, AFTER: {len(union)}")
 +    assert set(unresolved_after) == set(unresolved_before), "the rewrite created an unresolved key"
 +
 +    # Deleted keys that survive as PROSE on other pages (inside ''...'' or %%...%%
 +    # in review logs and run records). Those are historical statements about the
 +    # bibliography as it was, not citations, and are left as written; they are
 +    # printed so the residue is visible rather than silently ignored.
 +    prose = {}
 +    for f in files:
 +        text = changed[pid(f)][0] if pid(f) in changed else open(f, encoding="utf-8").read()
 +        stripped = MARKER.sub("", text)
 +        hits = sorted(lo for lo in MERGES if re.search(r"\b" + re.escape(lo) + r"\b", stripped))
 +        if hits:
 +            prose[pid(f)] = hits
 +    print(f"\ndeleted keys still named in PROSE (not markers), left as historical record: "
 +          f"{len(prose)} page(s)")
 +    for p_, hits in sorted(prose.items()):
 +        print(f"  {p_:48s} {', '.join(hits)}")
 +    print("\nall invariants hold")
 +
 +
 +if __name__ == "__main__":
 +    main()
 +</file>
 +
 +==== Second sitting — rendered before/after ====
 +
 +<file python bib_dedup_render_check.py>
 +#!/usr/bin/env python3
 +"""Rendered-DOM check for the 2026-09-04 citekey consolidation.
 +
 +For each page whose markers were repointed, compare the rendered page fetched
 +BEFORE the edit with the one fetched AFTER the bibliography was saved and the
 +bibtex4dw cache purged. Both counts must be unchanged: a repointed marker is
 +still one marker (citekey spans), and because no page cited both keys of a pair
 +the reference list keeps its length (<dt> inside dl.bibtex_references). -1
 +means the page has no reference list at all (provenance pages that carry no
 +<bibtex bibliography> block). Counts are scoped to the wikipage start/stop
 +comments. "deleted-key strings" counts the deleted keys as TEXT anywhere in the
 +body — on provenance pages that is the dated amendment and the historical
 +notes, not a citation, so it is printed rather than asserted.
 +
 +    python3 scripts/bib_dedup_render_check.py out/dedup_render_before out/dedup_render_after
 +"""
 +import os
 +import re
 +import sys
 +
 +DELETED = ["fouad2022my", "boettger2025_regional", "ahmad2026_ipfp",
 +           "bouhoula2024automated", "lerner2016internet"]
 +
 +
 +def stats(path):
 +    h = open(path, encoding="utf-8", errors="replace").read()
 +    s, e = h.find("<!-- wikipage start -->"), h.find("<!-- wikipage stop -->")
 +    body = h[s:e] if 0 <= s < e else h
 +    dl = re.search(r'<dl class="bibtex_references">(.*?)</dl>', body, re.S)
 +    dts = len(re.findall(r"<dt", dl.group(1))) if dl else -1
 +    spans = len(re.findall(r"bibtex_citekey", body))
 +    strings = sum(body.count(k) for k in DELETED)
 +    return dts, spans, strings
 +
 +
 +def main():
 +    before, after = sys.argv[1], sys.argv[2]
 +    bad = 0
 +    print("^ page ^ references (dt) before → after ^ citekey spans before → after ^ "
 +          "deleted-key strings after ^ verdict ^")
 +    for f in sorted(os.listdir(after)):
 +        pid = f[:-5].replace("__", ":")
 +        b, a = stats(os.path.join(before, f)), stats(os.path.join(after, f))
 +        ok = (a[0], a[1]) == (b[0], b[1])
 +        bad += not ok
 +        print(f"| {pid} | {b[0]} → {a[0]} | {b[1]} → {a[1]} | {a[2]} | "
 +              f"{'unchanged' if ok else 'CHANGED'} |")
 +    print(f"\npages checked: {len(os.listdir(after))}   pages whose counts moved: {bad}")
 +    return 1 if bad else 0
 +
 +
 +if __name__ == "__main__":
 +    sys.exit(main())
 +</file>
  
 ===== Review log ===== ===== Review log =====
Line 2454: Line 3761:
 | Cut the moralising sentences; the bug narratives read as honest, the commentary on them reads as performance | **Accepted.** Five removed, and the heading "What the reviewers did not find, and this run did" became "One bug no reviewer found" | | Cut the moralising sentences; the bug narratives read as honest, the commentary on them reads as performance | **Accepted.** Five removed, and the heading "What the reviewers did not find, and this run did" became "One bug no reviewer found" |
 | Redundancy: the three adjudications appear four times, the Benoît bug twice | **Partly accepted.** The quote-table repeat was cut. The Benoît account stays in both places: one is the narrative, the other is a decoder comment that has to stand on its own | | Redundancy: the three adjudications appear four times, the Benoît bug twice | **Partly accepted.** The quote-table repeat was cut. The Benoît account stays in both places: one is the narrative, the other is a decoder comment that has to stand on its own |
-| The corpus-wide 16,864-record table is [[:literature:corpus]] material | **Accepted as a pointer, rejected as a removal.** It is the reason this audit had to happen; the page now says where it belongs and links onward |+| The corpus-wide 16,864-record table is [[:literature:corpus]] material | **Accepted as a pointer, rejected as a removal.** Checking this finding turned up something neither the reviewer nor this run had noticed: [[:literature:corpus]] **already exists** (75 KB, rev 1786550805) and this page had been calling it "not yet written" in its own first paragraph. Its published per-venue record counts agree with this page's exactly, which is an independent confirmation of the denominator; the author and DOI breakdown is new here and stays |
  
 Two findings from the earlier passes were **not** accepted as stated: Two findings from the earlier passes were **not** accepted as stated:
Line 2460: Line 3767:
   * "''programming:crawler_detection'' cites no Doupé-authored paper at all." It cites two, ''zhang2021_crawlphish'' and ''zhang2022_spartacus''. The reviewer conclusion was right for a different reason — both spell the name with a **literal** é, so that page verified nothing about LaTeX escapes — and the fix was made on that basis, not the one offered.   * "''programming:crawler_detection'' cites no Doupé-authored paper at all." It cites two, ''zhang2021_crawlphish'' and ''zhang2022_spartacus''. The reviewer conclusion was right for a different reason — both spell the name with a **literal** é, so that page verified nothing about LaTeX escapes — and the fix was made on that basis, not the one offered.
   * "5,859 extracted papers is wrong; there are 5,869." Two different populations: 5,859 records in ''extractions.jsonl'', 5,869 papers with full text on disk. Recorded as a clarification, not a correction; the run record now gives both and says neither is used here.   * "5,859 extracted papers is wrong; there are 5,869." Two different populations: 5,859 records in ''extractions.jsonl'', 5,869 papers with full text on disk. Recorded as a clarification, not a correction; the run record now gives both and says neither is used here.
 +
 +==== Second sitting, 2026-09-04: citekey consolidation ====
 +
 +Three focused passes ran against the frozen draft of this section, each handed
 +the page text, the scripts, their committed outputs, the 21 saved files and the
 +rendered before/after HTML, and told the author's context might not be
 +exhaustive. The section was **first published before their findings arrived**
 +(revision 1788543550, with a line saying so); the findings below were then
 +applied and the page re-saved. A follow-up item
 +''dedup-bibkey-review-followup'' was filed at the first save for exactly this
 +work and is overtaken by it.
 +
 +^ Reviewer ^ Finding ^ Disposition ^
 +| Sonnet — figures vs script | "18 pages" carry a literal example marker: that is the count for the ''key'' string alone; the union over all four strings is **20** (''citekey'' adds ''provenance:statistics:interrater_agreement'', ''...'' adds ''provenance:artifacts'') | **Accepted.** ''bib_dedup_apply.py'' now prints the distinct-page union; prose says 20 and names the mistake |
 +| Sonnet — figures vs script | "Fifty-four are one first author with two or three papers in one year" is wrong for about half: 28 pairs share a first-author string, 26 are different people sharing a surname (Zhenhua Li / Frank Li / Li Li, Jiang Zhang / Penghui Zhang …) | **Accepted; the finding of this review.** ''bib_dedup_scan.py'' now prints both first-author strings on every candidate row and the same/different split (31 of 58 identical, of which 3 are DuckDuckGo; 27 differ, of which 1 is Bratton/Sumner). Prose rewritten to 28 / 26 and to say that this class is what a surname fold necessarily produces |
 +| Sonnet — figures vs script | The first sitting filed the work as ''dedupe-bibliography-entries''; the new text credits ''dedup-regional-filter-lists-bibkey'' and describes it as raised 2026-08-14 for a narrower reason, without reconciling the two | **Accepted.** The "Closed later the same day" paragraph now explains that the 2026-08-14 item already included "check for other duplicate pairs", so it covered the filed work, and says plainly that the older item should be closed as superseded and that this run could not close it |
 +| Sonnet — figures vs script | All five scripts reproduce their committed outputs byte for byte; embedded ''%%<file>%%'' bodies byte-identical to disk; the five-row table, 19/4/2, 23/13/20/21/10/26/58, 855→850 and both revision numbers verified; umlaut fold tested on four spellings; marker regex misses no real marker (9 near-misses are code fragments in ''%%<code>%%'' blocks); two mutations (deleting a MERGES pair; disabling the distinct-key assert and injecting a both-keys marker) both change the outcome; live bibliography has 850 entries and none of the deleted keys; three live repointed pages carry no deleted key in a marker | Verified, no change |
 +| Sonnet — citations and claims | All ten entries of the five pairs read: same paper, author lists identical including order; every "what was dropped" cell correct. Independent marker-resolution check over the 21 saved files and this page: zero unresolved beyond the documented example tokens. 1723–1739 confirmed from cached and fresh USENIX fetch. 2026-08-14 discovery note found verbatim on ''provenance:programming:crawler:openwpm''. "No page cited both keys", 26 prose pages (identical per-page lists), 19/25 (8 fields spot-checked; the two ''dropped'' cases traced to ''bibgen.mjs'' line 110), "three of five without underscore; fouad2022my 4 vs 2" all recomputed. All 21 diffs are marker substitutions only. 10 of 58 candidates read: distinct papers | Verified, no change. The reviewer reported two false alarms of its own (an ''endswith'' filter that excluded this page from a count; a ''grep -A15'' spilling into the next entry) and resolved both itself |
 +| Sonnet — external currency | DOIs 10.56553/popets-2022-0063, -2025-0063, -2026-0109 resolve (200) to pages whose ''citation_title''/''citation_author'' match the kept entries in order; all in the 10.56553 range. Böttger volume 2025 / issue 2 / pages 309–325 match ''petsymposium.org'' byte for byte. Both USENIX pages return 200 with identical ''citation_author'' lists; ''pages = {1723--1739}'' confirmed from USENIX's BibTeX and meta; no ''citation_doi'' on either, so "no DOI" is current | Verified, no change |
 +
 +Then the generic pass, with no checklist, after those fixes were applied.
 +
 +^ Reviewer ^ Finding ^ Disposition ^
 +| Fable — generic | The second sitting's "what could not be established" and "judgement calls" precede the first sitting's, which are unlabelled, so a reader meets the closers out of order | **Accepted**, by labelling the two older sections //first sitting// rather than moving 400 lines |
 +| Fable — generic | "161 other pages" here against "160" in the first sitting — both true, neither explained; the first sitting's own review log had already closed a 160-vs-161 finding | **Accepted.** This page did not exist when the earlier count was taken; that is now said where the 161 appears |
 +| Fable — generic | "First revision of this page" and "every figure below is against that snapshot" (837 entries) are made false by the second-sitting bullet directly beneath them | **Accepted.** Both scoped to the first sitting |
 +| Fable — generic | The tie-break rule is stated as "first title word", but ''autoKey'' takes the first word over three letters **not in its ''STOP'' list** — and ''Understanding'' is in ''STOP''. Under the rule as written, two of the five kept keys would be wrong | **Accepted; the best finding of this pass.** The whole tie-break rests on "the convention is what ''bibgen.mjs'' mints", so stating a rule that does not mint the kept keys undercut it. The real rule is now given, with ''bottger2025_regional'' as the worked example |
 +| Fable — generic | Three rows of the rendered check sit one below the apply table's distinct-key count, unexplained: those pages carry a documentation example written as a real marker, so ''bib_dedup_apply.py'' counts a template marker as a citation | **Accepted as a disclosure, rejected as a code change.** The paragraph now names the three pages and the cause, and says the inflation cannot touch the consolidation. Teaching the marker regex to skip instruction lines would change a number nothing depends on, in a script whose output is already published |
 +| Fable — generic | "26 are different people … including one spelling variant of a single person" contradicts itself | **Accepted.** Now 25 different people and one person spelled two ways |
 +| Fable — generic | The headline row describes only the surname rule, though the 58 candidates include a title match and the scan also ran on the 850-entry file | **Accepted.** The row now names all four rules and both files |
 +| Fable — generic | 13 changed + 10 amended is stated as 20 pages and 21 files without saying three pages are in both sets | **Accepted.** One clause added |
 +| Fable — generic | "all ten were typed by hand" is an inference: ''autoKey'' has a no-authors path that derives a surname from ''landingUrl'', and callers can pass ''--key'' | **Accepted.** Softened to what the evidence supports — none can have been minted from an author list |
 +| Fable — generic | The 19-of-25 census runs on the 855-entry file, so it counts the Böttger paper twice; after the save it is 19 of 24 | **Accepted.** Labelled, with both figures |
 +| Fable — generic | Both reviewer corrections are narrated in the body and again in the review log | **Accepted.** Cut from the body; the log is the right place |
 +| Fable — generic | "''bouhoula2024_automated'' alone is on 11. so no reference list…" — a lowercase sentence start, pre-existing | **Accepted**, fixed in passing |
 +| Fable — generic | DokuWiki syntax checked: no literal closing tag inside the four new ''%%<file>%%'' blocks, WRAP balanced, ''%%'' spans even, no pipe in a table cell; every number in the new prose matches the committed outputs | Verified, no change |
  
 ====== References ====== ====== References ======
provenance/literature/bibliography.1788492653.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki