provenance:literature:bibliography
Differences
This shows you the differences between two versions of the page.
| Both sides previous revisionPrevious revisionNext revision | Previous revision | ||
| provenance:literature:bibliography [2026/09/04 03:33] – Correct three claims about literature:corpus (it exists; its per-venue counts agree with this page's), add the co-citation and PETS smoke-test script, record the second review round. Authored by Claude karel.kubicek.claude | provenance:literature:bibliography [2026/09/04 18:43] (current) – Second-sitting audit: apply the four review passes' findings (autoKey STOP-list rule, 20-page union, 28/26 candidate split, 160-vs-161, first-sitting labelling, render-check off-by-one) and publish the full review log. Authored by Claude karel.kubicek.claude | ||
|---|---|---|---|
| Line 14: | Line 14: | ||
| ===== Run record ===== | ===== Run record ===== | ||
| - | * Run date: 2026-09-04 (UTC). First revision of this page. | + | * Run date: 2026-09-04 (UTC). First sitting; first revision of this page. |
| * Authoring agent: Claude Opus 5, executing the drain item '' | * Authoring agent: Claude Opus 5, executing the drain item '' | ||
| * Corpus at run time: 5,859 extracted papers, 2010–2026, | * Corpus at run time: 5,859 extracted papers, 2010–2026, | ||
| - | * Live bibliography read at revision **1788471646**, | + | * Live bibliography read at revision **1788471646**, |
| * Nothing on '' | * Nothing on '' | ||
| * '' | * '' | ||
| + | * **Second sitting, 2026-09-04, later the same day:** drain item '' | ||
| ===== What the audits found, in one table ===== | ===== What the audits found, in one table ===== | ||
| Line 32: | Line 33: | ||
| | Is any live USENIX entry' | | Is any live USENIX entry' | ||
| | Is any PoPETs DOI on the wrong prefix, or dead? | 84 PoPETs entries with a DOI; the prefix rule covers 83 | **0 wrong prefix, 0 dead** | | | Is any PoPETs DOI on the wrong prefix, or dead? | 84 PoPETs entries with a DOI; the prefix rule covers 83 | **0 wrong prefix, 0 dead** | | ||
| - | | Is the same paper in the bibliography twice? | all 837 entries | **yes, 5 pairs** — latent, no page cites both keys of a pair | | + | | Is the same paper in the bibliography twice? | all 837 entries |
| + | | … and after the second sitting? | all 850 entries on the saved page | **0 pairs**, both scans; 23 markers on 13 pages repointed first, then the 5 entries deleted | | ||
| + | | Did the umlaut slip ('' | ||
| | Are the checks above capable of failing? | 13 mutations | all 13 change the reported counts | | | Are the checks above capable of failing? | 13 mutations | all 13 change the reported counts | | ||
| Line 571: | Line 574: | ||
| established. | established. | ||
| - | ===== Found and deferred: five papers | + | ===== Found at 837 entries: five papers |
| A citekey-collision check passes while the same paper sits in the file under two | A citekey-collision check passes while the same paper sits in the file under two | ||
| Line 586: | Line 589: | ||
| **No page currently cites both keys of a pair**, checked by | **No page currently cites both keys of a pair**, checked by | ||
| - | '' | + | '' |
| '' | '' | ||
| - | least one key of some pair; '' | + | least one key of some pair; '' |
| visible. Consolidating means choosing one key per pair and rewriting %%{[key]}%% | visible. Consolidating means choosing one key per pair and rewriting %%{[key]}%% | ||
| markers across the 26 pages that cite one, which is a different piece of work | markers across the 26 pages that cite one, which is a different piece of work | ||
| Line 594: | Line 597: | ||
| '' | '' | ||
| - | ===== What could not be established ===== | + | **Closed later the same day**, but under a different item name than the one |
| + | filed here. '' | ||
| + | for the Böttger pair alone, with the instruction to "check for other | ||
| + | duplicate-title pairs in the same pass"; it therefore already covered the work | ||
| + | '' | ||
| + | consolidation and its invariants are the section //Audit 2026-09-04, second | ||
| + | sitting// below. '' | ||
| + | be closed as such; this run could add items to the queue but not close them. | ||
| + | The table and counts in this section describe the 837-entry file and are left | ||
| + | as they were. | ||
| + | |||
| + | ===== Audit 2026-09-04, second sitting: the five duplicates consolidated ===== | ||
| + | |||
| + | Drain item '' | ||
| + | [[: | ||
| + | '' | ||
| + | above found all five pairs and deferred them; this sitting closed them. Every | ||
| + | figure here is against a fresh ''? | ||
| + | revision 1788526615 (855 entries) and of all 161 other pages, taken at 17:21 | ||
| + | UTC and re-checked against '' | ||
| + | revision had moved. **161, where the first sitting says 160**: this provenance | ||
| + | page did not exist when that count was taken, and it is one of the 26 that name | ||
| + | a deleted key in prose. Both numbers are right for their own date. | ||
| + | |||
| + | ==== Which key was kept, and why ==== | ||
| + | |||
| + | None of the ten keys can have been minted from an author list: PETS and USENIX | ||
| + | index records carry none, so '' | ||
| + | in the landing URL or required an explicit '' | ||
| + | file's own majority convention, **surname + year + '' | ||
| + | diacritics dropped** — where "first title word" is '' | ||
| + | the first word longer than three letters that is not in its '' | ||
| + | '' | ||
| + | '' | ||
| + | first author' | ||
| + | **855-entry file, before the deletion**, so the Böttger paper is in it twice, | ||
| + | once in each row it is the example for; after the save 24 entries remain, 19 of | ||
| + | them stripped: | ||
| + | |||
| + | ^ Key spelling of the diacritic ^ Entries (of 25) ^ Examples ^ | ||
| + | | stripped — Böttger → '' | ||
| + | | German digraph — Böttger → '' | ||
| + | | letter dropped — Kührer → '' | ||
| + | |||
| + | ^ Paper ^ Kept ^ Deleted ^ Content pages citing kept / deleted, before ^ What the deleted entry had that the kept one lacked ^ | ||
| + | | Fouad et al., PoPETs 2022 | '' | ||
| + | | Böttger et al., PoPETs 2025 | '' | ||
| + | | Ahmad et al., PoPETs 2026 | '' | ||
| + | | Bouhoula et al., USENIX Sec 2024 | '' | ||
| + | | Lerner et al., USENIX Sec 2016 | '' | ||
| + | |||
| + | The per-page lists behind the fourth column are in the apply output below. | ||
| + | Three of the five deleted keys were the Google-Scholar style (no underscore), | ||
| + | and one of those, '' | ||
| + | replacement — four against two. Majority-of-pages would have kept it. Convention | ||
| + | won, because the convention is what the next '' | ||
| + | file with two live key styles is how these five pairs arose. | ||
| + | |||
| + | ==== What was done, in order ==== | ||
| + | |||
| + | - Export every page and the bibliography fresh: 161 + 1 files. '' | ||
| + | - Run '' | ||
| + | - Run '' | ||
| + | - Save the **20 pages first** — the 13 whose markers changed plus the 10 given an amendment, three pages being in both sets — each with '' | ||
| + | - Purge the bibtex4dw cache on the bibliography and every repointed page (''? | ||
| + | - Re-fetch the 13 repointed pages and compare rendered reference-list length and marker count against copies fetched before the edit ('' | ||
| + | - Re-export the bibliography and all 21 saved pages and confirm byte identity with the saved files; re-run both duplicate scans on the live bibliography — 0 pairs each. | ||
| + | |||
| + | ==== The invariants, and the real output ==== | ||
| + | |||
| + | Each figure on this page about the consolidation is a line of this output. | ||
| + | |||
| + | < | ||
| + | bibliography : out/ | ||
| + | entries | ||
| + | field added : bouhoula2024_automated | ||
| + | pages read : 161 (bibliography excluded) | ||
| + | |||
| + | loser-key marker occurrences before, by key: | ||
| + | fouad2022my | ||
| + | boettger2025_regional | ||
| + | ahmad2026_ipfp | ||
| + | bouhoula2024automated | ||
| + | lerner2016internet | ||
| + | |||
| + | pages citing each key BEFORE (content pages; provenance pages in brackets): | ||
| + | kept fouad2022_cookie | ||
| + | deleted fouad2022my | ||
| + | kept bottger2025_regional | ||
| + | deleted boettger2025_regional | ||
| + | kept ahmad2026_more | ||
| + | deleted ahmad2026_ipfp | ||
| + | kept bouhoula2024_automated | ||
| + | deleted bouhoula2024automated | ||
| + | kept lerner2016_internet | ||
| + | deleted lerner2016internet | ||
| + | |||
| + | ^ page ^ markers repointed ^ distinct keys before ^ after ^ pairs ^ | ||
| + | | design: | ||
| + | | design: | ||
| + | | design: | ||
| + | | privacy: | ||
| + | | privacy: | ||
| + | | privacy: | ||
| + | | privacy: | ||
| + | | programming: | ||
| + | | programming: | ||
| + | | provenance: | ||
| + | | provenance: | ||
| + | | provenance: | ||
| + | | statistics: | ||
| + | |||
| + | pages changed | ||
| + | markers repointed | ||
| + | provenance amendments | ||
| + | files written | ||
| + | |||
| + | markers naming a key the bibliography does not define, BEFORE: 22 on 4 key(s) | ||
| + | ... (1 page(s)) | ||
| + | cite (1 page(s)) | ||
| + | citekey | ||
| + | key (18 page(s)) | ||
| + | same, AFTER: 22 on 4 key(s) | ||
| + | ... (1 page(s)) | ||
| + | cite (1 page(s)) | ||
| + | citekey | ||
| + | key (18 page(s)) | ||
| + | distinct pages carrying any such example marker, AFTER: 20 | ||
| + | |||
| + | deleted keys still named in PROSE (not markers), left as historical record: 26 page(s) | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | provenance: | ||
| + | |||
| + | all invariants hold | ||
| + | </ | ||
| + | |||
| + | Reading it: 23 deleted-key occurrences before, 23 repointed; every page' | ||
| + | distinct-key count is unchanged, which is the check that no page cited both | ||
| + | keys of a pair (a page that had would have lost a reference silently); and the | ||
| + | four " | ||
| + | documentation examples written as literal markers in prose on **20** distinct | ||
| + | pages ('' | ||
| + | script prints the union), identical before and after. They are a pre-existing | ||
| + | residue this run did not touch and did not create. | ||
| + | |||
| + | ==== Rendered check ==== | ||
| + | |||
| + | < | ||
| + | ^ page ^ references (dt) before → after ^ citekey spans before → after ^ deleted-key strings after ^ verdict ^ | ||
| + | | design: | ||
| + | | design: | ||
| + | | design: | ||
| + | | privacy: | ||
| + | | privacy: | ||
| + | | privacy: | ||
| + | | privacy: | ||
| + | | programming: | ||
| + | | programming: | ||
| + | | provenance: | ||
| + | | provenance: | ||
| + | | provenance: | ||
| + | | statistics: | ||
| + | |||
| + | pages checked: 13 pages whose counts moved: 0 | ||
| + | </ | ||
| + | |||
| + | Three rows sit one below the apply table' | ||
| + | '' | ||
| + | '' | ||
| + | **documentation example** written as a real marker: '' | ||
| + | first, and a contributor instruction naming '' | ||
| + | other two, which the plugin does not render at all. The apply script counts | ||
| + | those as citations, so its distinct-key column and its "pages citing each key" | ||
| + | lists are one high on those three pages. It does not touch the consolidation: | ||
| + | none of the five pairs is involved. '' | ||
| + | which cites papers but carries no | ||
| + | '' | ||
| + | " | ||
| + | and their historical notes, not citations — the same rows show their marker | ||
| + | counts unchanged. | ||
| + | |||
| + | ==== The wider scan: the umlaut slip did not recur ==== | ||
| + | |||
| + | The item asked whether the transliteration mistake existed for other | ||
| + | German-named authors. '' | ||
| + | the DOI-and-title scan of the first sitting: near-identical squashed titles | ||
| + | ('' | ||
| + | subtitle), and **same year plus same folded first-author surname**, where the | ||
| + | fold strips diacritics and collapses oe/ue/ae/ss to o/u/a/s so that Böttger, | ||
| + | '' | ||
| + | reports the five definite pairs and **58 candidates**; | ||
| + | file, **0 definite** and the same 58. | ||
| + | |||
| + | All 58 were read by hand and **none is a duplicate**. Three are DuckDuckGo | ||
| + | artefacts whose " | ||
| + | original, caught by the title rule ('' | ||
| + | '' | ||
| + | first-author comparison splits them **28 / 26**: 28 are the same first author | ||
| + | with two or three different papers in one year (Durumeric 2013–2015, | ||
| + | 2018, Oest 2020, Kancherla 2025 …), and 26 have first-author strings that | ||
| + | differ: 25 are **different people who share a surname and a year** — Li, Zhang, | ||
| + | Liu, Lin, Wu, Agarwal, Tang — and one is a single person spelled two ways | ||
| + | ('' | ||
| + | second class is the price of a rule that folds surnames: it is the only way | ||
| + | Böttger and Boettger can be paired at all. No second umlaut pair exists in the file. The list is the residue | ||
| + | of the fold and is printed in full, with both first-author strings on every | ||
| + | row, so the judgement can be checked: | ||
| + | |||
| + | < | ||
| + | bibliography : out/ | ||
| + | entries | ||
| + | |||
| + | DEFINITE duplicate pairs (rule A or B): 5 | ||
| + | [ABCD] ahmad2026_ipfp | ||
| + | More Space, Less Privacy? Measuring the Effectiveness of IP-based Website Fingerprinting i | ||
| + | More Space, Less Privacy? Measuring the Effectiveness of IP-based Website Fingerprinting i | ||
| + | [ABCD] boettger2025_regional | ||
| + | | ||
| + | | ||
| + | [BCD] bouhoula2024_automated | ||
| + | | ||
| + | | ||
| + | [ABCD] fouad2022_cookie | ||
| + | My Cookie is a phoenix: detection, measurement, | ||
| + | My Cookie is a phoenix: Detection, measurement, | ||
| + | [BCD] lerner2016_internet | ||
| + | | ||
| + | | ||
| + | |||
| + | CANDIDATE pairs (rule C or D only) — judged by hand, see the provenance page: 58 | ||
| + | [D] LePochat2019_tranco | ||
| + | 2019 Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation | ||
| + | 2019 Evaluating the Long-term Effects of Parameters on the Characteristics of the {Tranco} Top Sites Rank | ||
| + | [D] agarwal2024_peeking | ||
| + | 2024 Peeking through the window: Fingerprinting Browser Extensions through Page-Visible Execution Traces | ||
| + | 2024 Poster: A Comprehensive Categorization of SMS Scams | ||
| + | [D] agarwal2025_dropped | ||
| + | 2025 'Hey mum, I dropped my phone down the toilet': | ||
| + | 2025 Fishing for Smishing: Understanding SMS Phishing Infrastructure and Strategies by Mining Public User | ||
| + | [D] agarwal2025_dropped | ||
| + | 2025 'Hey mum, I dropped my phone down the toilet': | ||
| + | 2025 "I have no idea how to make it safer": | ||
| + | [D] agarwal2025_fishing | ||
| + | 2025 Fishing for Smishing: Understanding SMS Phishing Infrastructure and Strategies by Mining Public User | ||
| + | 2025 "I have no idea how to make it safer": | ||
| + | [D] alroomi2023_login | ||
| + | 2023 A Large-Scale Measurement of Website Login Policies | ||
| + | 2023 Measuring Website Password Creation Policies At Scale | ||
| + | [D] bahrami2025_bytedefender | ||
| + | 2025 Byte by Byte: Unmasking Browser Fingerprinting at the Function Level Using V8 Bytecode Transformers | ||
| + | 2025 {CookieGuard}: | ||
| + | [D] bashir2019_adstxt | ||
| + | 2019 A Longitudinal Analysis of the ads.txt Standard | ||
| + | 2019 Quantity vs. Quality: Evaluating User Interest Profiles Using Ad Preference Managers | ||
| + | [D] bhuiyan2025_digital | ||
| + | 2025 Digital Disparities: | ||
| + | 2025 Not All Visitors are Bilingual: A Measurement Study of the Multilingual Web from an Accessibility Pe | ||
| + | [C] bratton2019_replication | ||
| + | 2019 The Association Between Exaggeration in Health-Related Science News and Academic Press Releases: A R | ||
| + | 2014 The Association Between Exaggeration in Health Related Science News and Academic Press Releases: Ret | ||
| + | [D] chen2021_cookieswap | ||
| + | 2021 Cookie Swap Party: Abusing First-Party Cookies for Web Tracking | ||
| + | 2021 Detecting Filter List Evasion with Event-Loop-Turn Granularity JavaScript Signatures | ||
| + | [D] chen2025_parents | ||
| + | 2025 Empowering Parents to Support Children' | ||
| + | 2025 Semantics-Aware Cookie Purpose Compliance | ||
| + | [CD] duckduckgo_tracker_radar_2026 | ||
| + | 2026 DuckDuckGo Tracker Radar | ||
| + | 2026 DuckDuckGo Tracker Radar Detector | ||
| + | [D] duckduckgo_tracker_radar_2026 | ||
| + | 2026 DuckDuckGo Tracker Radar | ||
| + | 2026 Tracker Radar Collector | ||
| + | [D] duckduckgo_tracker_radar_detector_2026 | ||
| + | 2026 DuckDuckGo Tracker Radar Detector | ||
| + | 2026 Tracker Radar Collector | ||
| + | [D] durumeric2013_https | ||
| + | 2013 Analysis of the HTTPS certificate ecosystem | ||
| + | 2013 {ZMap}: Fast Internet-wide Scanning and Its Security Applications | ||
| + | [D] durumeric2014_heartbleed | ||
| + | 2014 The Matter of Heartbleed | ||
| + | 2014 An Internet-Wide View of Internet-Wide Scanning | ||
| + | [D] durumeric2015_neither | ||
| + | 2015 Neither Snow Nor Rain Nor MITM...: An Empirical Analysis of Email Delivery Security | ||
| + | 2015 A Search Engine Backed by Internet-Wide Scanning | ||
| + | [D] edu2022_alexa | ||
| + | 2022 Measuring Alexa Skill Privacy Practices across Three Years | ||
| + | 2022 Exploring the security and privacy risks of chatbots in messaging services | ||
| + | [D] iqbal2022_khaleesi | ||
| + | 2022 Khaleesi: Breaker of Advertising and Tracking Request Chains | ||
| + | 2022 Left or Right: A Peek into the Political Biases in Email Spam Filtering Algorithms During US Electio | ||
| + | [D] kancherla2025_johnny | ||
| + | 2025 Johnny Can't Revoke Consent Either: Measuring Compliance of Consent Revocation on the Web | ||
| + | 2025 Least Privilege Access for Persistent Storage Mechanisms in Web Browsers | ||
| + | [D] kirchner2024_black | ||
| + | 2024 A Black-Box Privacy Analysis of Messaging Service Providers' | ||
| + | 2024 Dancer in the Dark: Synthesizing and Evaluating Polyglots for Blind Cross-Site Scripting | ||
| + | [D] lee2023_adtargeting | ||
| + | 2023 When and Why Do People Want Ad Targeting Explanations? | ||
| + | 2023 Net-track: Generic Web Tracking Detection Using Packet Metadata | ||
| + | [D] li2016_remedying | ||
| + | 2016 Remedying Web Hijacking: Notification Effectiveness and Webmaster Comprehension | ||
| + | 2016 You've Got Vulnerability: | ||
| + | [D] li2017_radar | ||
| + | 2017 FBS-Radar: Uncovering Fake Base Stations at Scale in the Wild | ||
| + | 2017 A Large-Scale Empirical Study of Security Patches | ||
| + | [D] li2017_radar | ||
| + | 2017 FBS-Radar: Uncovering Fake Base Stations at Scale in the Wild | ||
| + | 2017 Static analysis of Android apps: A systematic literature review | ||
| + | [D] li2017_security | ||
| + | 2017 A Large-Scale Empirical Study of Security Patches | ||
| + | 2017 Static analysis of Android apps: A systematic literature review | ||
| + | [D] li2024_bounce | ||
| + | 2024 Bounce in the Wild: A Deep Dive into Email Delivery Failures from a Large Email Service Provider | ||
| + | 2024 Are We Getting Well-informed? | ||
| + | [D] li2024_bounce | ||
| + | 2024 Bounce in the Wild: A Deep Dive into Email Delivery Failures from a Large Email Service Provider | ||
| + | 2024 A Worldwide View on the Reachability of Encrypted DNS Services | ||
| + | [D] li2024_wellinformed | ||
| + | 2024 Are We Getting Well-informed? | ||
| + | 2024 A Worldwide View on the Reachability of Encrypted DNS Services | ||
| + | [D] liao2016_characterizing | ||
| + | 2016 Characterizing Long-tail SEO Spam on Cloud Web Hosting Services | ||
| + | 2016 Seeking Nonsense, Looking for Trouble: Efficient Promotional-Infection Detection through Semantic In | ||
| + | [D] lin2021_longitudinal | ||
| + | 2021 A Longitudinal Study of Removed Apps in {iOS} App Store | ||
| + | 2021 Phishpedia: A Hybrid Deep Learning Based Approach to Visually Identify Phishing Webpages | ||
| + | [D] lin2022_investigating | ||
| + | 2022 Investigating Advertisers' | ||
| + | 2022 Phish in Sheep' | ||
| + | [D] liu2024_opted | ||
| + | 2024 Opted Out, Yet Tracked: Are Regulations Enough to Protect Your Privacy? | ||
| + | 2024 From Promises to Practice: Evaluating the Private Browsing Modes of Android Browser Apps | ||
| + | [D] liu2025_domino | ||
| + | 2025 The DOMino Effect: Detecting and Exploiting DOM Clobbering Gadgets via Concolic Execution with Symbo | ||
| + | 2025 The First Early Evidence of the Use of Browser Fingerprinting for Online Tracking | ||
| + | [D] liu2025_domino | ||
| + | 2025 The DOMino Effect: Detecting and Exploiting DOM Clobbering Gadgets via Concolic Execution with Symbo | ||
| + | 2025 Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Craw | ||
| + | [D] liu2025_fingerprinting | ||
| + | 2025 The First Early Evidence of the Use of Browser Fingerprinting for Online Tracking | ||
| + | 2025 Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Craw | ||
| + | [D] nguyen2025_breaking | ||
| + | 2025 Breaking the Shield: Analyzing and Attacking Canvas Fingerprinting Defenses in the Wild | ||
| + | 2025 " | ||
| + | [D] nisenoff2023_awareness | ||
| + | 2023 User Awareness and Behaviors Concerning Encrypted {DNS} Settings in Web Browsers | ||
| + | 2023 Defining " | ||
| + | [D] oest2020_phishtime | ||
| + | 2020 PhishTime: Continuous Longitudinal Measurement of the Effectiveness of Anti-phishing Blacklists | ||
| + | 2020 Sunrise to Sunset: Analyzing the End-to-end Life Cycle and Effectiveness of Phishing Attacks at Scal | ||
| + | [D] papadogiannakis2025_before | ||
| + | 2025 Before \& After: The Effect of EU's 2022 Code of Practice on Disinformation | ||
| + | 2025 Welcome to the Dark Side: Analyzing the Revenue Flows of Fraud in the Online Ad Ecosystem | ||
| + | [D] ruth2022_toppling | ||
| + | 2022 Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists | ||
| + | 2022 A World Wide View of Browsing the World Wide Web | ||
| + | [D] scheitle2018_long | ||
| + | 2018 A Long Way to the Top: Significance, | ||
| + | 2018 The Rise of Certificate Transparency and Its Implications on the Internet Ecosystem | ||
| + | [D] starov2017_extended | ||
| + | 2017 Extended Tracking Powers: Measuring the Privacy Diffusion Enabled by Browser Extensions | ||
| + | 2017 XHOUND: Quantifying the Fingerprintability of Browser Extensions | ||
| + | [D] tang2025_misuse | ||
| + | 2025 Misuse, Misreporting, | ||
| + | 2025 Navigating Cookie Consent Violations Across the Globe | ||
| + | [D] utz2023_comparing | ||
| + | 2023 Comparing Large-Scale Privacy and Security Notifications | ||
| + | 2023 Privacy Rarely Considered: Exploring Considerations in the Adoption of Third-Party Services by Websi | ||
| + | [D] vastel2018_scanner | ||
| + | 2018 Fp-Scanner: The Privacy Implications of Browser Fingerprint Inconsistencies | ||
| + | 2018 FP-STALKER: Tracking Browser Fingerprint Evolutions | ||
| + | [D] vekaria2025_bighelp | ||
| + | 2025 Big Help or Big Brother? Auditing Tracking, Profiling, and Personalization in Generative AI Assistan | ||
| + | 2025 SoK: Advances and Open Problems in Web Tracking | ||
| + | [D] venkatadri2019_auditing | ||
| + | 2019 Auditing Offline Data Brokers via Facebook' | ||
| + | 2019 Investigating sources of PII used in Facebook’s targeted advertising | ||
| + | [D] wang2026_masks | ||
| + | 2026 The Masks We (Think We) Wear: Privacy Threats of Browser-Extension Wallets in the Web3 Ecosystem | ||
| + | 2026 SIPConfusion: | ||
| + | [D] wu2025_appprivacyreport | ||
| + | 2025 Transparency or Information Overload? Evaluating Users’ Comprehension and Perceptions of the iOS App | ||
| + | 2025 Revealing the Black Box of Device Search Engine: Scanning Assets, Strategies, and Ethical Considerat | ||
| + | [D] wu2026_email | ||
| + | 2026 One Email, Many Faces: A Deep Dive into Identity Confusion in Email Aliases | ||
| + | 2026 Tracking the Stray Sheep: Understanding DNS Response Manipulation in the Wild | ||
| + | [D] xie2024_arcanum | ||
| + | 2024 Arcanum: Detecting and Evaluating the Privacy Risks of Browser Extensions on Web Pages and Web Conte | ||
| + | 2024 Crawling to the Top: An Empirical Evaluation of Top List Use | ||
| + | [D] yang2022_extensive | ||
| + | 2022 An Extensive Study of Residential Proxies in China | ||
| + | 2022 WTAGRAPH: Web Tracking and Advertising Detection using Graph Neural Networks | ||
| + | [D] zhang2022_harpo | ||
| + | 2022 HARPO: Learning to Subvert Online Behavioral Advertising | ||
| + | 2022 I'm SPARTACUS, No, I'm SPARTACUS: Proactively Protecting Users from Phishing by Intentionally Trigge | ||
| + | [D] zhang2024_inbox | ||
| + | 2024 Inbox Invasion: Exploiting MIME Ambiguities to Evade Email Attachment Detectors | ||
| + | 2024 QUIC is not Quick Enough over Fast Internet | ||
| + | [D] zhang2025_abusability | ||
| + | 2025 Abusability of Automation Apps in Intimate Partner Violence | ||
| + | 2025 Demystifying the (In)Security of {QR} Code-based Login in Real-world Deployments | ||
| + | [D] zhu2020_label | ||
| + | 2020 Measuring and Modeling the Label Dynamics of Online Anti-Malware Engines | ||
| + | 2020 Demo: Benchmarking Label Dynamics of VirusTotal Engines | ||
| + | |||
| + | candidates whose first-author string is identical on both sides: 31 of 58; different first author (shared surname, or a non-person ' | ||
| + | </ | ||
| + | |||
| + | ==== Prose that still names the deleted keys ==== | ||
| + | |||
| + | 26 provenance pages name a deleted key in a run record or review log — " | ||
| + | and left alone", | ||
| + | statements about the bibliography as it was when they were written and were | ||
| + | **not rewritten**; | ||
| + | ten provenance pages whose content page was repointed each got a dated | ||
| + | // | ||
| + | who follows one of those statements finds the correction on the same page. | ||
| + | |||
| + | ==== What could not be established, | ||
| + | |||
| + | * Whether anything **outside the wiki** cites a deleted key — a BibTeX file someone exported from the site, a draft that copied '' | ||
| + | * Whether '' | ||
| + | |||
| + | ==== Judgement calls, second sitting ==== | ||
| + | |||
| + | ^ Call ^ Alternative a reasonable person would pick ^ Why this one ^ | ||
| + | | Keep the stripped-diacritic, | ||
| + | | Carry '' | ||
| + | | Leave the 26 provenance pages' historical prose alone; amend only the 10 whose content page was repointed | Rewrite every mention of a deleted key | A review log records what was true when written. Rewriting it is the mistake the first sitting' | ||
| + | | Pages first, bibliography last, one '' | ||
| + | | Publish the 58 candidates in full | Publish the count and the verdict | A hand judgement over a list nobody can read is not checkable | | ||
| + | | Extend this page rather than create '' | ||
| + | |||
| + | ===== What could not be established, first sitting | ||
| * **Whether defects 1 and 2 ever reached the wiki.** Zero missing authors survive in '' | * **Whether defects 1 and 2 ever reached the wiki.** Zero missing authors survive in '' | ||
| Line 604: | Line 1064: | ||
| * **Whether '' | * **Whether '' | ||
| - | ===== Judgement calls ===== | + | ===== Judgement calls, first sitting |
| ^ Call ^ Alternative a reasonable person would pick ^ Why this one ^ | ^ Call ^ Alternative a reasonable person would pick ^ Why this one ^ | ||
| Line 614: | Line 1074: | ||
| | Correct '' | | Correct '' | ||
| | Extend the audit to all 159 live USENIX entries (pass 3), beyond the 90 the item named | Stop at the 90 keys | Pass 2 can only see an affiliation that got **in**, never an author that fell **out**, and 74 live entries are invisible to a cache-based check. 72 extra page fetches closed the gap | | | Extend the audit to all 159 live USENIX entries (pass 3), beyond the 90 the item named | Stop at the 90 keys | Pass 2 can only see an affiliation that got **in**, never an author that fell **out**, and 74 live entries are invisible to a cache-based check. 72 extra page fetches closed the gap | | ||
| - | | Defer the 5 duplicate entries | Consolidate them in the same sitting | ~30 pages would need %%{[key]}%% rewrites; nothing renders wrong today | | + | | Defer the 5 duplicate entries | Consolidate them in the same sitting | ~30 pages would need %%{[key]}%% rewrites; nothing renders wrong today. Done as its own item later the same day: 13 pages and 23 markers, not ~30 — see the second-sitting audit | |
| | Put this page at '' | | Put this page at '' | ||
| | Keep '' | | Keep '' | ||
| Line 2355: | Line 2815: | ||
| This is a smoke test over parse_popets output, not the three-source audit the USENIX keys got. | This is a smoke test over parse_popets output, not the three-source audit the USENIX keys got. | ||
| </ | </ | ||
| + | |||
| + | ==== Second sitting — the wider duplicate scan ==== | ||
| + | |||
| + | Run on the 855-entry export before the change (output above, under the audit) | ||
| + | and on the saved 850-entry page: | ||
| + | |||
| + | <file python bib_dedup_scan.py> | ||
| + | # | ||
| + | """ | ||
| + | |||
| + | scripts/ | ||
| + | That found five pairs on 2026-09-04, and one of them was an umlaut | ||
| + | transliteration in the citekey (bottger / boettger). The same slip can also | ||
| + | land in the TITLE (a subtitle dropped, " | ||
| + | so this scan casts wider and prints its candidates for a human to judge: | ||
| + | |||
| + | A same DOI | ||
| + | B same squashed title (definite) | ||
| + | C near-identical title, difflib ratio >= 0.85 on the squashed form, | ||
| + | or one squashed title a prefix of the other (subtitle dropped) | ||
| + | D same year AND same folded first-author surname — the fold strips | ||
| + | | ||
| + | | ||
| + | |||
| + | C and D are candidate generators, not verdicts. Every C/D pair that is not | ||
| + | already in A/B is printed with both titles so the residue is visible; the | ||
| + | verdicts recorded on provenance: | ||
| + | over that printed list, not the script' | ||
| + | |||
| + | python3 scripts/ | ||
| + | """ | ||
| + | import difflib | ||
| + | import os | ||
| + | import re | ||
| + | import sys | ||
| + | import unicodedata | ||
| + | from collections import defaultdict | ||
| + | |||
| + | sys.path.insert(0, | ||
| + | from usenix_bib_diff import field, squash | ||
| + | |||
| + | LATEX = {r' | ||
| + | | ||
| + | |||
| + | |||
| + | def fold_surname(author_field): | ||
| + | """ | ||
| + | digraphs collapsed. Handles 'Last, First' and 'First Last' | ||
| + | first = re.split(r" | ||
| + | s = first | ||
| + | for k, v in LATEX.items(): | ||
| + | s = s.replace(k, | ||
| + | s = re.sub(r" | ||
| + | s = re.sub(r" | ||
| + | s = unicodedata.normalize(" | ||
| + | s = "" | ||
| + | last = s.split("," | ||
| + | last = last.lower() | ||
| + | last = re.sub(r" | ||
| + | for dg, one in ((" | ||
| + | last = last.replace(dg, | ||
| + | return last | ||
| + | |||
| + | |||
| + | def main(): | ||
| + | bib = sys.argv[sys.argv.index(" | ||
| + | text = open(bib, encoding=" | ||
| + | entries = re.findall(r" | ||
| + | rec = {} | ||
| + | for e in entries: | ||
| + | k = re.match(r" | ||
| + | if k in rec: | ||
| + | raise SystemExit(f" | ||
| + | doi = field(e, " | ||
| + | doi = re.sub(r" | ||
| + | title = field(e, " | ||
| + | author = field(e, " | ||
| + | year = field(e, " | ||
| + | first = re.split(r" | ||
| + | rec[k] = dict(doi=doi, | ||
| + | surname=fold_surname(author) if author else "", | ||
| + | print(f" | ||
| + | print(f" | ||
| + | |||
| + | pairs = {} # frozenset(k1, | ||
| + | |||
| + | def add(a, b, rule): | ||
| + | pairs.setdefault(frozenset((a, | ||
| + | |||
| + | by = defaultdict(list) | ||
| + | for k, r in rec.items(): | ||
| + | if r[" | ||
| + | by[(" | ||
| + | if r[" | ||
| + | by[(" | ||
| + | if r[" | ||
| + | by[(" | ||
| + | for key, ks in by.items(): | ||
| + | for i in range(len(ks)): | ||
| + | for j in range(i + 1, len(ks)): | ||
| + | add(ks[i], ks[j], key[0]) | ||
| + | keys = sorted(rec) | ||
| + | for i in range(len(keys)): | ||
| + | a = rec[keys[i]][" | ||
| + | if len(a) < 20: | ||
| + | continue | ||
| + | for j in range(i + 1, len(keys)): | ||
| + | b = rec[keys[j]][" | ||
| + | if len(b) < 20: | ||
| + | continue | ||
| + | if a.startswith(b) or b.startswith(a) or \ | ||
| + | | ||
| + | add(keys[i], | ||
| + | |||
| + | definite = {p: r for p, r in pairs.items() if r & {" | ||
| + | candidates = {p: r for p, r in pairs.items() if not r & {" | ||
| + | print(f" | ||
| + | for p, rules in sorted(definite.items(), | ||
| + | a, b = sorted(p) | ||
| + | print(f" | ||
| + | print(f" | ||
| + | print(f" | ||
| + | print(f" | ||
| + | f" | ||
| + | # Rule D fires on a folded SURNAME, so it also pairs different people who | ||
| + | # share a common surname (Li, Zhang, Liu). Print the first author' | ||
| + | # name string for both sides and count how many pairs are the same string, | ||
| + | # so the residue can be described without hand-counting. | ||
| + | same_person = 0 | ||
| + | for p, rules in sorted(candidates.items(), | ||
| + | a, b = sorted(p) | ||
| + | same = rec[a][" | ||
| + | same_person += same | ||
| + | print(f" | ||
| + | f" | ||
| + | print(f" | ||
| + | print(f" | ||
| + | print(f" | ||
| + | f" | ||
| + | f" | ||
| + | return 1 if definite else 0 | ||
| + | |||
| + | |||
| + | if __name__ == " | ||
| + | sys.exit(main()) | ||
| + | </ | ||
| + | |||
| + | < | ||
| + | bibliography : out/ | ||
| + | entries | ||
| + | |||
| + | DEFINITE duplicate pairs (rule A or B): 0 | ||
| + | |||
| + | CANDIDATE pairs (rule C or D only) — judged by hand, see the provenance page: 58 | ||
| + | [D] LePochat2019_tranco | ||
| + | 2019 Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation | ||
| + | 2019 Evaluating the Long-term Effects of Parameters on the Characteristics of the {Tranco} Top Sites Rank | ||
| + | [D] agarwal2024_peeking | ||
| + | 2024 Peeking through the window: Fingerprinting Browser Extensions through Page-Visible Execution Traces | ||
| + | 2024 Poster: A Comprehensive Categorization of SMS Scams | ||
| + | [D] agarwal2025_dropped | ||
| + | 2025 'Hey mum, I dropped my phone down the toilet': | ||
| + | 2025 Fishing for Smishing: Understanding SMS Phishing Infrastructure and Strategies by Mining Public User | ||
| + | [D] agarwal2025_dropped | ||
| + | 2025 'Hey mum, I dropped my phone down the toilet': | ||
| + | 2025 "I have no idea how to make it safer": | ||
| + | [D] agarwal2025_fishing | ||
| + | 2025 Fishing for Smishing: Understanding SMS Phishing Infrastructure and Strategies by Mining Public User | ||
| + | 2025 "I have no idea how to make it safer": | ||
| + | [D] alroomi2023_login | ||
| + | 2023 A Large-Scale Measurement of Website Login Policies | ||
| + | 2023 Measuring Website Password Creation Policies At Scale | ||
| + | [D] bahrami2025_bytedefender | ||
| + | 2025 Byte by Byte: Unmasking Browser Fingerprinting at the Function Level Using V8 Bytecode Transformers | ||
| + | 2025 {CookieGuard}: | ||
| + | [D] bashir2019_adstxt | ||
| + | 2019 A Longitudinal Analysis of the ads.txt Standard | ||
| + | 2019 Quantity vs. Quality: Evaluating User Interest Profiles Using Ad Preference Managers | ||
| + | [D] bhuiyan2025_digital | ||
| + | 2025 Digital Disparities: | ||
| + | 2025 Not All Visitors are Bilingual: A Measurement Study of the Multilingual Web from an Accessibility Pe | ||
| + | [C] bratton2019_replication | ||
| + | 2019 The Association Between Exaggeration in Health-Related Science News and Academic Press Releases: A R | ||
| + | 2014 The Association Between Exaggeration in Health Related Science News and Academic Press Releases: Ret | ||
| + | [D] chen2021_cookieswap | ||
| + | 2021 Cookie Swap Party: Abusing First-Party Cookies for Web Tracking | ||
| + | 2021 Detecting Filter List Evasion with Event-Loop-Turn Granularity JavaScript Signatures | ||
| + | [D] chen2025_parents | ||
| + | 2025 Empowering Parents to Support Children' | ||
| + | 2025 Semantics-Aware Cookie Purpose Compliance | ||
| + | [CD] duckduckgo_tracker_radar_2026 | ||
| + | 2026 DuckDuckGo Tracker Radar | ||
| + | 2026 DuckDuckGo Tracker Radar Detector | ||
| + | [D] duckduckgo_tracker_radar_2026 | ||
| + | 2026 DuckDuckGo Tracker Radar | ||
| + | 2026 Tracker Radar Collector | ||
| + | [D] duckduckgo_tracker_radar_detector_2026 | ||
| + | 2026 DuckDuckGo Tracker Radar Detector | ||
| + | 2026 Tracker Radar Collector | ||
| + | [D] durumeric2013_https | ||
| + | 2013 Analysis of the HTTPS certificate ecosystem | ||
| + | 2013 {ZMap}: Fast Internet-wide Scanning and Its Security Applications | ||
| + | [D] durumeric2014_heartbleed | ||
| + | 2014 The Matter of Heartbleed | ||
| + | 2014 An Internet-Wide View of Internet-Wide Scanning | ||
| + | [D] durumeric2015_neither | ||
| + | 2015 Neither Snow Nor Rain Nor MITM...: An Empirical Analysis of Email Delivery Security | ||
| + | 2015 A Search Engine Backed by Internet-Wide Scanning | ||
| + | [D] edu2022_alexa | ||
| + | 2022 Measuring Alexa Skill Privacy Practices across Three Years | ||
| + | 2022 Exploring the security and privacy risks of chatbots in messaging services | ||
| + | [D] iqbal2022_khaleesi | ||
| + | 2022 Khaleesi: Breaker of Advertising and Tracking Request Chains | ||
| + | 2022 Left or Right: A Peek into the Political Biases in Email Spam Filtering Algorithms During US Electio | ||
| + | [D] kancherla2025_johnny | ||
| + | 2025 Johnny Can't Revoke Consent Either: Measuring Compliance of Consent Revocation on the Web | ||
| + | 2025 Least Privilege Access for Persistent Storage Mechanisms in Web Browsers | ||
| + | [D] kirchner2024_black | ||
| + | 2024 A Black-Box Privacy Analysis of Messaging Service Providers' | ||
| + | 2024 Dancer in the Dark: Synthesizing and Evaluating Polyglots for Blind Cross-Site Scripting | ||
| + | [D] lee2023_adtargeting | ||
| + | 2023 When and Why Do People Want Ad Targeting Explanations? | ||
| + | 2023 Net-track: Generic Web Tracking Detection Using Packet Metadata | ||
| + | [D] li2016_remedying | ||
| + | 2016 Remedying Web Hijacking: Notification Effectiveness and Webmaster Comprehension | ||
| + | 2016 You've Got Vulnerability: | ||
| + | [D] li2017_radar | ||
| + | 2017 FBS-Radar: Uncovering Fake Base Stations at Scale in the Wild | ||
| + | 2017 A Large-Scale Empirical Study of Security Patches | ||
| + | [D] li2017_radar | ||
| + | 2017 FBS-Radar: Uncovering Fake Base Stations at Scale in the Wild | ||
| + | 2017 Static analysis of Android apps: A systematic literature review | ||
| + | [D] li2017_security | ||
| + | 2017 A Large-Scale Empirical Study of Security Patches | ||
| + | 2017 Static analysis of Android apps: A systematic literature review | ||
| + | [D] li2024_bounce | ||
| + | 2024 Bounce in the Wild: A Deep Dive into Email Delivery Failures from a Large Email Service Provider | ||
| + | 2024 Are We Getting Well-informed? | ||
| + | [D] li2024_bounce | ||
| + | 2024 Bounce in the Wild: A Deep Dive into Email Delivery Failures from a Large Email Service Provider | ||
| + | 2024 A Worldwide View on the Reachability of Encrypted DNS Services | ||
| + | [D] li2024_wellinformed | ||
| + | 2024 Are We Getting Well-informed? | ||
| + | 2024 A Worldwide View on the Reachability of Encrypted DNS Services | ||
| + | [D] liao2016_characterizing | ||
| + | 2016 Characterizing Long-tail SEO Spam on Cloud Web Hosting Services | ||
| + | 2016 Seeking Nonsense, Looking for Trouble: Efficient Promotional-Infection Detection through Semantic In | ||
| + | [D] lin2021_longitudinal | ||
| + | 2021 A Longitudinal Study of Removed Apps in {iOS} App Store | ||
| + | 2021 Phishpedia: A Hybrid Deep Learning Based Approach to Visually Identify Phishing Webpages | ||
| + | [D] lin2022_investigating | ||
| + | 2022 Investigating Advertisers' | ||
| + | 2022 Phish in Sheep' | ||
| + | [D] liu2024_opted | ||
| + | 2024 Opted Out, Yet Tracked: Are Regulations Enough to Protect Your Privacy? | ||
| + | 2024 From Promises to Practice: Evaluating the Private Browsing Modes of Android Browser Apps | ||
| + | [D] liu2025_domino | ||
| + | 2025 The DOMino Effect: Detecting and Exploiting DOM Clobbering Gadgets via Concolic Execution with Symbo | ||
| + | 2025 The First Early Evidence of the Use of Browser Fingerprinting for Online Tracking | ||
| + | [D] liu2025_domino | ||
| + | 2025 The DOMino Effect: Detecting and Exploiting DOM Clobbering Gadgets via Concolic Execution with Symbo | ||
| + | 2025 Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Craw | ||
| + | [D] liu2025_fingerprinting | ||
| + | 2025 The First Early Evidence of the Use of Browser Fingerprinting for Online Tracking | ||
| + | 2025 Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Craw | ||
| + | [D] nguyen2025_breaking | ||
| + | 2025 Breaking the Shield: Analyzing and Attacking Canvas Fingerprinting Defenses in the Wild | ||
| + | 2025 " | ||
| + | [D] nisenoff2023_awareness | ||
| + | 2023 User Awareness and Behaviors Concerning Encrypted {DNS} Settings in Web Browsers | ||
| + | 2023 Defining " | ||
| + | [D] oest2020_phishtime | ||
| + | 2020 PhishTime: Continuous Longitudinal Measurement of the Effectiveness of Anti-phishing Blacklists | ||
| + | 2020 Sunrise to Sunset: Analyzing the End-to-end Life Cycle and Effectiveness of Phishing Attacks at Scal | ||
| + | [D] papadogiannakis2025_before | ||
| + | 2025 Before \& After: The Effect of EU's 2022 Code of Practice on Disinformation | ||
| + | 2025 Welcome to the Dark Side: Analyzing the Revenue Flows of Fraud in the Online Ad Ecosystem | ||
| + | [D] ruth2022_toppling | ||
| + | 2022 Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists | ||
| + | 2022 A World Wide View of Browsing the World Wide Web | ||
| + | [D] scheitle2018_long | ||
| + | 2018 A Long Way to the Top: Significance, | ||
| + | 2018 The Rise of Certificate Transparency and Its Implications on the Internet Ecosystem | ||
| + | [D] starov2017_extended | ||
| + | 2017 Extended Tracking Powers: Measuring the Privacy Diffusion Enabled by Browser Extensions | ||
| + | 2017 XHOUND: Quantifying the Fingerprintability of Browser Extensions | ||
| + | [D] tang2025_misuse | ||
| + | 2025 Misuse, Misreporting, | ||
| + | 2025 Navigating Cookie Consent Violations Across the Globe | ||
| + | [D] utz2023_comparing | ||
| + | 2023 Comparing Large-Scale Privacy and Security Notifications | ||
| + | 2023 Privacy Rarely Considered: Exploring Considerations in the Adoption of Third-Party Services by Websi | ||
| + | [D] vastel2018_scanner | ||
| + | 2018 Fp-Scanner: The Privacy Implications of Browser Fingerprint Inconsistencies | ||
| + | 2018 FP-STALKER: Tracking Browser Fingerprint Evolutions | ||
| + | [D] vekaria2025_bighelp | ||
| + | 2025 Big Help or Big Brother? Auditing Tracking, Profiling, and Personalization in Generative AI Assistan | ||
| + | 2025 SoK: Advances and Open Problems in Web Tracking | ||
| + | [D] venkatadri2019_auditing | ||
| + | 2019 Auditing Offline Data Brokers via Facebook' | ||
| + | 2019 Investigating sources of PII used in Facebook’s targeted advertising | ||
| + | [D] wang2026_masks | ||
| + | 2026 The Masks We (Think We) Wear: Privacy Threats of Browser-Extension Wallets in the Web3 Ecosystem | ||
| + | 2026 SIPConfusion: | ||
| + | [D] wu2025_appprivacyreport | ||
| + | 2025 Transparency or Information Overload? Evaluating Users’ Comprehension and Perceptions of the iOS App | ||
| + | 2025 Revealing the Black Box of Device Search Engine: Scanning Assets, Strategies, and Ethical Considerat | ||
| + | [D] wu2026_email | ||
| + | 2026 One Email, Many Faces: A Deep Dive into Identity Confusion in Email Aliases | ||
| + | 2026 Tracking the Stray Sheep: Understanding DNS Response Manipulation in the Wild | ||
| + | [D] xie2024_arcanum | ||
| + | 2024 Arcanum: Detecting and Evaluating the Privacy Risks of Browser Extensions on Web Pages and Web Conte | ||
| + | 2024 Crawling to the Top: An Empirical Evaluation of Top List Use | ||
| + | [D] yang2022_extensive | ||
| + | 2022 An Extensive Study of Residential Proxies in China | ||
| + | 2022 WTAGRAPH: Web Tracking and Advertising Detection using Graph Neural Networks | ||
| + | [D] zhang2022_harpo | ||
| + | 2022 HARPO: Learning to Subvert Online Behavioral Advertising | ||
| + | 2022 I'm SPARTACUS, No, I'm SPARTACUS: Proactively Protecting Users from Phishing by Intentionally Trigge | ||
| + | [D] zhang2024_inbox | ||
| + | 2024 Inbox Invasion: Exploiting MIME Ambiguities to Evade Email Attachment Detectors | ||
| + | 2024 QUIC is not Quick Enough over Fast Internet | ||
| + | [D] zhang2025_abusability | ||
| + | 2025 Abusability of Automation Apps in Intimate Partner Violence | ||
| + | 2025 Demystifying the (In)Security of {QR} Code-based Login in Real-world Deployments | ||
| + | [D] zhu2020_label | ||
| + | 2020 Measuring and Modeling the Label Dynamics of Online Anti-Malware Engines | ||
| + | 2020 Demo: Benchmarking Label Dynamics of VirusTotal Engines | ||
| + | |||
| + | candidates whose first-author string is identical on both sides: 31 of 58; different first author (shared surname, or a non-person ' | ||
| + | </ | ||
| + | |||
| + | ==== Second sitting — citekey convention census ==== | ||
| + | |||
| + | <file python bib_key_convention.py> | ||
| + | # | ||
| + | """ | ||
| + | |||
| + | Needed to pick between bottger2025_regional and boettger2025_regional on | ||
| + | evidence rather than taste. For every entry whose FIRST AUTHOR' | ||
| + | carries a non-ASCII letter or a LaTeX accent command, classify how the key's | ||
| + | surname part renders it: | ||
| + | |||
| + | stripped | ||
| + | digraph | ||
| + | dropped | ||
| + | | ||
| + | other none of the above (printed; judge by hand) | ||
| + | |||
| + | python3 scripts/ | ||
| + | """ | ||
| + | import os | ||
| + | import re | ||
| + | import sys | ||
| + | import unicodedata | ||
| + | from collections import Counter | ||
| + | |||
| + | sys.path.insert(0, | ||
| + | from usenix_bib_diff import field | ||
| + | |||
| + | LATEX_ACCENT = re.compile(r" | ||
| + | LETTER_CMD = {r" | ||
| + | r" | ||
| + | |||
| + | |||
| + | def first_surname(author): | ||
| + | first = re.split(r" | ||
| + | return first.split("," | ||
| + | |||
| + | |||
| + | def de_latex(s): | ||
| + | s = LATEX_ACCENT.sub(lambda m: m.group(1) or m.group(2), s) | ||
| + | for k, v in LETTER_CMD.items(): | ||
| + | s = s.replace(k, | ||
| + | return re.sub(r" | ||
| + | |||
| + | |||
| + | def has_diacritic(s): | ||
| + | return bool(re.search(r" | ||
| + | or any(k in s for k in LETTER_CMD)) | ||
| + | |||
| + | |||
| + | def stripped(s): | ||
| + | s = unicodedata.normalize(" | ||
| + | s = "" | ||
| + | return re.sub(r" | ||
| + | |||
| + | |||
| + | def digraph(s): | ||
| + | # LaTeX umlauts first (\"o, {\"o}, \" | ||
| + | # for whatever accents remain. | ||
| + | s = re.sub(r' | ||
| + | for a, b in ((" | ||
| + | | ||
| + | s = s.replace(a, | ||
| + | s = de_latex(s) | ||
| + | s = unicodedata.normalize(" | ||
| + | s = "" | ||
| + | return re.sub(r" | ||
| + | |||
| + | |||
| + | def dropped(s): | ||
| + | s = de_latex(s) | ||
| + | return re.sub(r" | ||
| + | |||
| + | |||
| + | def main(): | ||
| + | bib = sys.argv[sys.argv.index(" | ||
| + | text = open(bib, encoding=" | ||
| + | rows, kinds = [], Counter() | ||
| + | for e in re.findall(r" | ||
| + | key = re.match(r" | ||
| + | author = field(e, " | ||
| + | if not author: | ||
| + | continue | ||
| + | sur = first_surname(author) | ||
| + | if not has_diacritic(sur): | ||
| + | continue | ||
| + | keysur = re.match(r" | ||
| + | if keysur == stripped(sur): | ||
| + | kind = " | ||
| + | elif keysur == digraph(sur) and digraph(sur) != stripped(sur): | ||
| + | kind = " | ||
| + | elif keysur == dropped(sur): | ||
| + | kind = " | ||
| + | else: | ||
| + | kind = " | ||
| + | kinds[kind] += 1 | ||
| + | rows.append((kind, | ||
| + | print(f" | ||
| + | print(f" | ||
| + | for k in (" | ||
| + | print(f" | ||
| + | print() | ||
| + | for kind, key, sur in sorted(rows): | ||
| + | print(f" | ||
| + | |||
| + | |||
| + | if __name__ == " | ||
| + | main() | ||
| + | </ | ||
| + | |||
| + | < | ||
| + | bibliography : out/ | ||
| + | entries whose first author' | ||
| + | stripped | ||
| + | digraph | ||
| + | dropped | ||
| + | other 0 | ||
| + | |||
| + | digraph | ||
| + | digraph | ||
| + | digraph | ||
| + | digraph | ||
| + | dropped | ||
| + | dropped | ||
| + | stripped | ||
| + | stripped | ||
| + | stripped | ||
| + | stripped | ||
| + | stripped | ||
| + | stripped | ||
| + | stripped | ||
| + | stripped | ||
| + | stripped | ||
| + | stripped | ||
| + | stripped | ||
| + | stripped | ||
| + | stripped | ||
| + | stripped | ||
| + | stripped | ||
| + | stripped | ||
| + | stripped | ||
| + | stripped | ||
| + | stripped | ||
| + | </ | ||
| + | |||
| + | ==== Second sitting — the consolidation itself ==== | ||
| + | |||
| + | <file python bib_dedup_apply.py> | ||
| + | # | ||
| + | """ | ||
| + | |||
| + | Does NOT touch the wiki. It reads a fresh ? | ||
| + | the bibliography, | ||
| + | whose invariants must all hold before anything is saved: | ||
| + | |||
| + | * each loser key is present in the bibliography exactly once and is removed; | ||
| + | each winner is present exactly once and is kept; | ||
| + | * entry count drops by exactly len(MERGES); | ||
| + | * every {[...]} marker naming a loser is rewritten to the winner; the number | ||
| + | of keys replaced equals the number of loser occurrences counted beforehand; | ||
| + | * no page cites both keys of a pair (else the rewrite would silently collapse | ||
| + | two markers into one reference and the page's distinct-key count would | ||
| + | drop) — the distinct-key count of every page is asserted unchanged; | ||
| + | * after the rewrite, no marker anywhere names a key the new bibliography | ||
| + | does not define, other than keys that were ALREADY unresolved before | ||
| + | (printed as residue, never one of the losers). | ||
| + | |||
| + | Winners follow the file's own majority convention — surname + year + ' | ||
| + | first title word, diacritics dropped — which is also what scripts/ | ||
| + | mints. Which key of each pair is the winner is a decision recorded on | ||
| + | provenance: | ||
| + | |||
| + | python3 scripts/ | ||
| + | --pages out/ | ||
| + | """ | ||
| + | import glob | ||
| + | import os | ||
| + | import re | ||
| + | import sys | ||
| + | from collections import Counter | ||
| + | |||
| + | # loser -> winner | ||
| + | MERGES = { | ||
| + | " | ||
| + | " | ||
| + | " | ||
| + | " | ||
| + | " | ||
| + | } | ||
| + | # A field the loser carried that the winner lacks and that was verified against | ||
| + | # the venue page on 2026-09-04 (USENIX' | ||
| + | # citation_lastpage meta tags on usenixsecurity24/ | ||
| + | EXTRA_FIELDS = { | ||
| + | " | ||
| + | } | ||
| + | BIBPAGE = " | ||
| + | MARKER = re.compile(r" | ||
| + | TEMPLATE_KEYS = {" | ||
| + | |||
| + | # Provenance pages that get a dated amendment because their content page's | ||
| + | # markers were repointed. content page id -> provenance page id. | ||
| + | PROVENANCE_OF = { | ||
| + | " | ||
| + | " | ||
| + | " | ||
| + | " | ||
| + | " | ||
| + | " | ||
| + | " | ||
| + | " | ||
| + | " | ||
| + | " | ||
| + | } | ||
| + | |||
| + | |||
| + | def pid(fname): | ||
| + | return os.path.basename(fname)[: | ||
| + | |||
| + | |||
| + | def entries_of(text): | ||
| + | return re.findall(r" | ||
| + | |||
| + | |||
| + | def key_of(entry): | ||
| + | return re.match(r" | ||
| + | |||
| + | |||
| + | def rewrite_bib(text): | ||
| + | ents = entries_of(text) | ||
| + | keys = Counter(key_of(e) for e in ents) | ||
| + | for lo, wi in MERGES.items(): | ||
| + | assert keys[lo] == 1, f" | ||
| + | assert keys[wi] == 1, f" | ||
| + | out = text | ||
| + | for e in ents: | ||
| + | k = key_of(e) | ||
| + | if k in MERGES: | ||
| + | # remove the entry and one preceding blank line | ||
| + | assert out.count(e) == 1 | ||
| + | out = out.replace(" | ||
| + | elif k in EXTRA_FIELDS: | ||
| + | new = e | ||
| + | for fname, val in EXTRA_FIELDS[k]: | ||
| + | assert not re.search(r" | ||
| + | new = new[: | ||
| + | assert out.count(e) == 1 | ||
| + | out = out.replace(e, | ||
| + | after = entries_of(out) | ||
| + | assert len(after) == len(ents) - len(MERGES), | ||
| + | for lo in MERGES: | ||
| + | assert not re.search(r" | ||
| + | return out, len(ents), len(after), {key_of(e) for e in after} | ||
| + | |||
| + | |||
| + | def rewrite_markers(text): | ||
| + | """ | ||
| + | replaced = 0 | ||
| + | before, after = set(), set() | ||
| + | |||
| + | def sub(m): | ||
| + | nonlocal replaced | ||
| + | keys = [k.strip() for k in m.group(1).split("," | ||
| + | before.update(keys) | ||
| + | new = [] | ||
| + | for k in keys: | ||
| + | if k in MERGES: | ||
| + | replaced += 1 | ||
| + | k = MERGES[k] | ||
| + | if k not in new: | ||
| + | new.append(k) | ||
| + | after.update(new) | ||
| + | return " | ||
| + | |||
| + | return MARKER.sub(sub, | ||
| + | |||
| + | |||
| + | def amendment(content_id, | ||
| + | lines = ["", | ||
| + | for lo, wi in pairs: | ||
| + | lines.append( | ||
| + | f" | ||
| + | f" | ||
| + | f" | ||
| + | f"'' | ||
| + | def n_markers(n): | ||
| + | return f"{n} citation marker" | ||
| + | where = f" | ||
| + | if n_prov: | ||
| + | where += f" and {n_markers(n_prov)} on this page" | ||
| + | verb = " | ||
| + | lines.append( | ||
| + | f" | ||
| + | f" | ||
| + | f"key describe the state when they were written. Full query log and the " | ||
| + | f" | ||
| + | f" | ||
| + | lines.append("" | ||
| + | return " | ||
| + | |||
| + | |||
| + | def insert_amendment(text, | ||
| + | i = text.find(" | ||
| + | if i >= 0: | ||
| + | return text[: | ||
| + | body = text.rstrip(" | ||
| + | if body[-1].startswith(" | ||
| + | return " | ||
| + | return text.rstrip(" | ||
| + | |||
| + | |||
| + | def main(): | ||
| + | a = sys.argv | ||
| + | bib_path = a[a.index(" | ||
| + | pages_dir = a[a.index(" | ||
| + | out_dir = a[a.index(" | ||
| + | os.makedirs(out_dir, | ||
| + | |||
| + | bib_text = open(bib_path, | ||
| + | new_bib, n_before, n_after, new_keys = rewrite_bib(bib_text) | ||
| + | open(os.path.join(out_dir, | ||
| + | old_keys = {key_of(e) for e in entries_of(bib_text)} | ||
| + | print(f" | ||
| + | print(f" | ||
| + | f" | ||
| + | for k, fs in EXTRA_FIELDS.items(): | ||
| + | print(f" | ||
| + | |||
| + | files = sorted(f for f in glob.glob(os.path.join(pages_dir, | ||
| + | if pid(f) != BIBPAGE.replace(" | ||
| + | print(f" | ||
| + | |||
| + | # occurrences of loser keys inside markers, before | ||
| + | loser_occ = Counter() | ||
| + | for f in files: | ||
| + | for m in MARKER.finditer(open(f, | ||
| + | for k in (x.strip() for x in m.group(1).split("," | ||
| + | if k in MERGES: | ||
| + | loser_occ[k] += 1 | ||
| + | print(" | ||
| + | for lo in MERGES: | ||
| + | print(f" | ||
| + | |||
| + | # Which pages cited each key of a pair before the rewrite — content pages and | ||
| + | # provenance pages separately, so the "which key was on more pages" question | ||
| + | # on the provenance page has a printed answer. | ||
| + | def citing(key): | ||
| + | out = [] | ||
| + | for f in files: | ||
| + | keys = {k.strip() for m in MARKER.finditer(open(f, | ||
| + | for k in m.group(1).split("," | ||
| + | if key in keys: | ||
| + | out.append(pid(f)) | ||
| + | return out | ||
| + | |||
| + | print(" | ||
| + | for lo, wi in MERGES.items(): | ||
| + | for k in (wi, lo): | ||
| + | ps = citing(k) | ||
| + | c = [p_ for p_ in ps if not p_.startswith(" | ||
| + | pr = [p_ for p_ in ps if p_.startswith(" | ||
| + | tag = " | ||
| + | print(f" | ||
| + | + (f" | ||
| + | |||
| + | changed = {} | ||
| + | total_replaced = 0 | ||
| + | unresolved_before, | ||
| + | all_after_keys = set() | ||
| + | print(" | ||
| + | for f in files: | ||
| + | text = open(f, encoding=" | ||
| + | new, n, kb, ka = rewrite_markers(text) | ||
| + | kb -= TEMPLATE_KEYS | ||
| + | ka -= TEMPLATE_KEYS | ||
| + | for k in kb - old_keys: | ||
| + | unresolved_before[k] += 1 | ||
| + | for k in ka - new_keys: | ||
| + | unresolved_after[k] += 1 | ||
| + | all_after_keys |= ka | ||
| + | if n: | ||
| + | assert len(kb) == len(ka), f" | ||
| + | pairs = sorted((lo, MERGES[lo]) for lo in kb if lo in MERGES) | ||
| + | changed[pid(f)] = (new, n, pairs) | ||
| + | total_replaced += n | ||
| + | print(f" | ||
| + | f" | ||
| + | print(f" | ||
| + | print(f" | ||
| + | f" | ||
| + | assert total_replaced == sum(loser_occ.values()) | ||
| + | for lo in MERGES: | ||
| + | assert lo not in all_after_keys, | ||
| + | |||
| + | # provenance amendments | ||
| + | prov_texts = {pid(f): open(f, encoding=" | ||
| + | n_prov_notes = 0 | ||
| + | for cid, prov in PROVENANCE_OF.items(): | ||
| + | assert cid in changed, f" | ||
| + | _, n_content, pairs = changed[cid] | ||
| + | n_prov = changed[prov][1] if prov in changed else 0 | ||
| + | base = changed[prov][0] if prov in changed else prov_texts[prov] | ||
| + | assert " | ||
| + | changed[prov] = (insert_amendment(base, | ||
| + | | ||
| + | n_prov_notes += 1 | ||
| + | for cid in changed: | ||
| + | if not cid.startswith(" | ||
| + | assert cid in PROVENANCE_OF, | ||
| + | print(f" | ||
| + | |||
| + | for id_, (text, _, _) in changed.items(): | ||
| + | open(os.path.join(out_dir, | ||
| + | | ||
| + | print(f" | ||
| + | |||
| + | print(f" | ||
| + | f" | ||
| + | for k, n in sorted(unresolved_before.items()): | ||
| + | print(f" | ||
| + | print(f" | ||
| + | for k, n in sorted(unresolved_after.items()): | ||
| + | print(f" | ||
| + | # The per-key counts above overlap: one page can carry several example | ||
| + | # strings. The number of DISTINCT pages carrying any of them is what the | ||
| + | # provenance page quotes, so print it rather than leave it to be summed. | ||
| + | union = set() | ||
| + | for f in files: | ||
| + | text = changed[pid(f)][0] if pid(f) in changed else open(f, encoding=" | ||
| + | for m in MARKER.finditer(text): | ||
| + | if any(k.strip() in unresolved_after for k in m.group(1).split("," | ||
| + | union.add(pid(f)) | ||
| + | print(f" | ||
| + | assert set(unresolved_after) == set(unresolved_before), | ||
| + | |||
| + | # Deleted keys that survive as PROSE on other pages (inside '' | ||
| + | # in review logs and run records). Those are historical statements about the | ||
| + | # bibliography as it was, not citations, and are left as written; they are | ||
| + | # printed so the residue is visible rather than silently ignored. | ||
| + | prose = {} | ||
| + | for f in files: | ||
| + | text = changed[pid(f)][0] if pid(f) in changed else open(f, encoding=" | ||
| + | stripped = MARKER.sub("", | ||
| + | hits = sorted(lo for lo in MERGES if re.search(r" | ||
| + | if hits: | ||
| + | prose[pid(f)] = hits | ||
| + | print(f" | ||
| + | f" | ||
| + | for p_, hits in sorted(prose.items()): | ||
| + | print(f" | ||
| + | print(" | ||
| + | |||
| + | |||
| + | if __name__ == " | ||
| + | main() | ||
| + | </ | ||
| + | |||
| + | ==== Second sitting — rendered before/ | ||
| + | |||
| + | <file python bib_dedup_render_check.py> | ||
| + | # | ||
| + | """ | ||
| + | |||
| + | For each page whose markers were repointed, compare the rendered page fetched | ||
| + | BEFORE the edit with the one fetched AFTER the bibliography was saved and the | ||
| + | bibtex4dw cache purged. Both counts must be unchanged: a repointed marker is | ||
| + | still one marker (citekey spans), and because no page cited both keys of a pair | ||
| + | the reference list keeps its length (<dt> inside dl.bibtex_references). -1 | ||
| + | means the page has no reference list at all (provenance pages that carry no | ||
| + | <bibtex bibliography> | ||
| + | comments. " | ||
| + | body — on provenance pages that is the dated amendment and the historical | ||
| + | notes, not a citation, so it is printed rather than asserted. | ||
| + | |||
| + | python3 scripts/ | ||
| + | """ | ||
| + | import os | ||
| + | import re | ||
| + | import sys | ||
| + | |||
| + | DELETED = [" | ||
| + | " | ||
| + | |||
| + | |||
| + | def stats(path): | ||
| + | h = open(path, encoding=" | ||
| + | s, e = h.find("< | ||
| + | body = h[s:e] if 0 <= s < e else h | ||
| + | dl = re.search(r'< | ||
| + | dts = len(re.findall(r"< | ||
| + | spans = len(re.findall(r" | ||
| + | strings = sum(body.count(k) for k in DELETED) | ||
| + | return dts, spans, strings | ||
| + | |||
| + | |||
| + | def main(): | ||
| + | before, after = sys.argv[1], | ||
| + | bad = 0 | ||
| + | print(" | ||
| + | " | ||
| + | for f in sorted(os.listdir(after)): | ||
| + | pid = f[: | ||
| + | b, a = stats(os.path.join(before, | ||
| + | ok = (a[0], a[1]) == (b[0], b[1]) | ||
| + | bad += not ok | ||
| + | print(f" | ||
| + | f" | ||
| + | print(f" | ||
| + | return 1 if bad else 0 | ||
| + | |||
| + | |||
| + | if __name__ == " | ||
| + | sys.exit(main()) | ||
| + | </ | ||
| ===== Review log ===== | ===== Review log ===== | ||
| Line 2462: | Line 3767: | ||
| * "'' | * "'' | ||
| * "5,859 extracted papers is wrong; there are 5, | * "5,859 extracted papers is wrong; there are 5, | ||
| + | |||
| + | ==== Second sitting, 2026-09-04: citekey consolidation ==== | ||
| + | |||
| + | Three focused passes ran against the frozen draft of this section, each handed | ||
| + | the page text, the scripts, their committed outputs, the 21 saved files and the | ||
| + | rendered before/ | ||
| + | exhaustive. The section was **first published before their findings arrived** | ||
| + | (revision 1788543550, with a line saying so); the findings below were then | ||
| + | applied and the page re-saved. A follow-up item | ||
| + | '' | ||
| + | work and is overtaken by it. | ||
| + | |||
| + | ^ Reviewer ^ Finding ^ Disposition ^ | ||
| + | | Sonnet — figures vs script | "18 pages" carry a literal example marker: that is the count for the '' | ||
| + | | Sonnet — figures vs script | " | ||
| + | | Sonnet — figures vs script | The first sitting filed the work as '' | ||
| + | | Sonnet — figures vs script | All five scripts reproduce their committed outputs byte for byte; embedded '' | ||
| + | | Sonnet — citations and claims | All ten entries of the five pairs read: same paper, author lists identical including order; every "what was dropped" | ||
| + | | Sonnet — external currency | DOIs 10.56553/ | ||
| + | |||
| + | Then the generic pass, with no checklist, after those fixes were applied. | ||
| + | |||
| + | ^ Reviewer ^ Finding ^ Disposition ^ | ||
| + | | Fable — generic | The second sitting' | ||
| + | | Fable — generic | "161 other pages" here against " | ||
| + | | Fable — generic | "First revision of this page" and "every figure below is against that snapshot" | ||
| + | | Fable — generic | The tie-break rule is stated as "first title word", but '' | ||
| + | | Fable — generic | Three rows of the rendered check sit one below the apply table' | ||
| + | | Fable — generic | "26 are different people … including one spelling variant of a single person" | ||
| + | | Fable — generic | The headline row describes only the surname rule, though the 58 candidates include a title match and the scan also ran on the 850-entry file | **Accepted.** The row now names all four rules and both files | | ||
| + | | Fable — generic | 13 changed + 10 amended is stated as 20 pages and 21 files without saying three pages are in both sets | **Accepted.** One clause added | | ||
| + | | Fable — generic | "all ten were typed by hand" is an inference: '' | ||
| + | | Fable — generic | The 19-of-25 census runs on the 855-entry file, so it counts the Böttger paper twice; after the save it is 19 of 24 | **Accepted.** Labelled, with both figures | | ||
| + | | Fable — generic | Both reviewer corrections are narrated in the body and again in the review log | **Accepted.** Cut from the body; the log is the right place | | ||
| + | | Fable — generic | "'' | ||
| + | | Fable — generic | DokuWiki syntax checked: no literal closing tag inside the four new '' | ||
| ====== References ====== | ====== References ====== | ||
provenance/literature/bibliography.1788492795.txt.gz · Last modified: by karel.kubicek.claude
