| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| provenance:design:website_classification [2026/09/03 22:17] – Add sections 12.1-12.12: LLM-classification currency, 2026-09-03. Every query with its denominator, the decision to lead the LLM section with corpus evidence rather than the preprint, the retraction of the curated-database 'open sources returning' reading karel.kubicek.claude | provenance:design:website_classification [2026/09/21 14:40] (current) – Review follow-ups: 12.8's 'task item exists' bullet closed by 12.14, and 12.14's page-sweep claim corrected (provenance:statistics:annotation carries the raw string, not a fold row). Authored by Claude karel.kubicek.claude |
|---|
| ^ Item ^ Value ^ | ^ Item ^ Value ^ |
| | Content page | [[design:website_classification]] | | | Content page | [[design:website_classification]] | |
| | Report script | ''scripts/report_website_classification.mjs'' (''--wiki'', ''--list'', ''--quotes <regex>'') | | | Report script | ''scripts/report_website_classification.mjs'' (''%%--wiki%%'', ''%%--list%%'', ''%%--quotes%% <regex>'') | |
| | Folds | ''scripts/webcat_fold.mjs'' — a **task** fold and a **resource** fold | | | Folds | ''scripts/webcat_fold.mjs'' — a **task** fold and a **resource** fold | |
| | Data | ''data/extract/run1/extractions.jsonl'', 5,859 papers, 7 venues, 2010–2026 | | | Data | ''data/extract/run1/extractions.jsonl'', 5,859 papers, 7 venues, 2010–2026 | |
| * The matcher was **substring**, not word-boundary, so ''report.includes('59')'' was satisfied by ''11.59 bits''. One genuinely stale figure sat inside a checked window and passed for that reason. | * The matcher was **substring**, not word-boundary, so ''report.includes('59')'' was satisfied by ''11.59 bits''. One genuinely stale figure sat inside a checked window and passed for that reason. |
| |
| Both are fixed in ''scripts/check_page_numbers.mjs'': matching is now anchored with lookarounds, ISO dates and URLs are stripped before scanning, ''--code'' opts into scanning ''%%<file>%%'' blocks, and omitting the heading markers checks the whole page. **Run it windowed //and// whole-page.** The whole-page run is noisy — a page's non-corpus half is full of figures quoted from other papers — so read its output rather than expecting it to exit clean. | Both are fixed in ''scripts/check_page_numbers.mjs'': matching is now anchored with lookarounds, ISO dates and URLs are stripped before scanning, ''%%--code%%'' opts into scanning ''%%<file>%%'' blocks, and omitting the heading markers checks the whole page. **Run it windowed //and// whole-page.** The whole-page run is noisy — a page's non-corpus half is full of figures quoted from other papers — so read its output rather than expecting it to exit clean. |
| Fixed on this page's content page as a result: the lead paragraph's **72% → 75%** for bespoke unsized taxonomies, which contradicted the corpus section's own "Three quarters … 75.4%" four screens down. | Fixed on this page's content page as a result: the lead paragraph's **72% → 75%** for bespoke unsized taxonomies, which contradicted the corpus section's own "Three quarters … 75.4%" four screens down. |
| |
| | validated labels | LLM papers = 177 | **148 (83.6%)** | | | validated labels | LLM papers = 177 | **148 (83.6%)** | |
| | per target | papers classifying **that target** at all | the table in [[design:website_classification#Which task, though]] | | | per target | papers classifying **that target** at all | the table in [[design:website_classification#Which task, though]] | |
| | model named to an artefact | LLM papers = 177 | A dated hosted snapshot **11 (6.2%)**; B open-weight checkpoint with a size **25 (14.1%)**; **A+B 36 (20.3%)**; C family only **130 (73.4%)**; D no model **11 (6.2%)** | | | model named to an artefact | the **175** papers that used or produced LLM labels | A dated hosted snapshot **10 (5.7%)**; B open-weight checkpoint with a size **24 (13.7%)**; **A+B 34 (19.4%)**; C family only **131 (74.9%)**; D no model **10 (5.7%)** | |
| |
| **Three denominator traps this section had to avoid, and one it fell into.** | **Three denominator traps this section had to avoid, and one it fell into.** |
| |
| - **No cut clears p < 0.05.** 0.16 for 2025 alone, 0.23 for 2026 alone, 0.095 for the full window, 0.079 for the five-venue control. | - **No cut clears p < 0.05.** 0.16 for 2025 alone, 0.23 for 2026 alone, 0.095 for the full window, 0.079 for the five-venue control. |
| - **The row does not contain open directories.** ''report_website_classification.mjs'' now prints every name behind it. The twelve recent-window papers name: Tracker Radar Entity List, Cloudflare Radar, SimilarWeb, Symantec SiteReview, AllSides, MediaBias/FactCheck (×2), Science Feedback, IAB taxonomy, Homepage2Vec, NAICSlite, and "predefined source rules (custom)". Three commercial vendors, one model, one taxonomy, four media-bias raters with closed editorial processes. **Not one is DMOZ, Curlie or anything like them.** By contrast the eight in 2022–2024 are more genuinely directory-like: DappRadar, a pornhosts blocklist, WebPulse, the Citizen Lab Block List, the AllSides Media Bias Chart, MediaBias/FactCheck, YouTube category labels, and one paper's "external political, government, media and issue-page sources". | - **The row does not contain open directories.** ''report_website_classification.mjs'' now prints every name behind it. Sorted, the twelve recent-window papers are three commercial vendors (Cloudflare Radar, SimilarWeb, Symantec SiteReview); four media-bias raters with closed editorial processes (AllSides, and Media Bias/Fact Check in three papers, one of which also cites Science Feedback); one model (Homepage2Vec); two taxonomies rather than label databases (IAB, NAICSlite); one paper's own rule set; and one that genuinely is an open, inspectable repository — DuckDuckGo's Tracker Radar Entity List — but of //tracker entities//, not website topics. **Not one is DMOZ, Curlie or any comparable open topic directory.** |
| | |
| | **The first version of this list, on this page and on the content page, named eleven of the twelve** — it silently dropped ''NDSS/2026/revealing-the-secret-power…'', whose resource string is "Media Bias/Fact Check (MBFC)", while the derived "four media-bias raters" counted it. Caught in re-review (§12.13, finding 2). A hand-written enumeration beside an embedded script output is the one kind of listing on this page that can still go stale, and this is the second time in one run it did. By contrast the eight in 2022–2024 are more genuinely directory-like: DappRadar, a pornhosts blocklist, WebPulse, the Citizen Lab Block List, the AllSides Media Bias Chart, MediaBias/FactCheck, YouTube category labels, and one paper's "external political, government, media and issue-page sources". |
| |
| So the openness half of this page's central argument has **no counter-evidence in the recent window**; the enum row that looked like counter-evidence is mostly mis-filed vendors. This is ''classification.method'''s 58% run-to-run stability doing exactly what the methodology bullet warns it does, on the one row where the page had built an argument on top of it. | So the openness half of this page's central argument has **no counter-evidence in the recent window**; the enum row that looked like counter-evidence is mostly mis-filed vendors. This is ''classification.method'''s 58% run-to-run stability doing exactly what the methodology bullet warns it does, on the one row where the page had built an argument on top of it. |
| ^ Family ^ Papers (of 177) ^ | ^ Family ^ Papers (of 177) ^ |
| | GPT-4o | 46 (26.0%) | | | GPT-4o | 46 (26.0%) | |
| | GPT-4 (non-4o) | 40 (22.6%) | | | GPT-4 (non-4o) | 44 (24.9%) | |
| | GPT-3.5 / GPT-3 / ChatGPT | 29 (16.4%) | | | GPT-3.5 / GPT-3 / ChatGPT | 29 (16.4%) | |
| | **UNMAPPED** | **23 (13.0%)** | | | **UNMAPPED** | **18 (10.2%)** | |
| | Llama family | 13 (7.3%) | | | Llama family | 13 (7.3%) | |
| | Gemini / PaLM | 12 (6.8%) | | | Gemini / PaLM | 12 (6.8%) | |
| | Unnamed LLM | 11 (6.2%) | | | Unnamed LLM | 11 (6.2%) | |
| | OpenAI reasoning / GPT-5 tier | 9 (5.1%) | | | OpenAI reasoning / GPT-5 tier | 10 (5.6%) | |
| | Qwen family | 8 (4.5%) | | | Qwen family | 8 (4.5%) | |
| | DeepSeek family | 7 (4.0%) | | | DeepSeek family | 7 (4.0%) | |
| | Encoder / seq2seq LM (not a chat LLM) | 2 (1.1%) | | | Encoder / seq2seq LM (not a chat LLM) | 2 (1.1%) | |
| |
| The 25-string unmapped residue is printed in full by the script (§12.9). **Four of those strings are a fold failure, not a long tail:** ''GPT 4.1'', ''GPT-4.0'', ''GPT-4.5'' and ''GPT-o1'' are OpenAI models the ''GPT-4 (non-4o)'' regex misses on a space or a decimal. They were left visible rather than folded, because patching a fold to absorb its own residue after seeing the output stops it being a documented rule; a task item exists to fold them and re-derive the three GPT rows. ''Grok-3'', ''GLM-4.5'', ''ChatGLM'' and ''Kimi'' are the genuine long tail — four vendors with no family in the list. | **The table above is the 2026-09-21 re-derivation.** Until then it read GPT-4 (non-4o) 40 (22.6%), UNMAPPED 23 (13.0%) and OpenAI reasoning 9 (5.1%), because four residue strings — ''GPT 4.1'', ''GPT-4.0'', ''GPT-4.5'' and ''GPT-o1'' — are OpenAI models that the fold missed on a space or a decimal. They were left visible on 2026-09-03 rather than quietly folded, because patching a fold to absorb its own residue after seeing the output stops it being a documented rule. The fold has now been extended **as a rule about how OpenAI model strings are spelled**, not as a list of those four: see §12.14 for the rule, the two further strings it moves, and the full diff. ''Grok-3'', ''GLM-4.5'', ''ChatGLM'' and ''Kimi'' stay in the now 20-string residue — they are the genuine long tail, four vendors with no family in the list, and no family was added for them. |
| |
| === 5b. The reproducibility buckets, which had to be rebuilt === | === 5b. The reproducibility buckets, which had to be rebuilt === |
| Rebuilt, split by hosting, which is the axis that actually decides reproducibility: | Rebuilt, split by hosting, which is the axis that actually decides reproducibility: |
| |
| ^ Bucket ^ Papers ^ Share ^ | ^ Bucket ^ Papers ^ Share of 175 ^ |
| | **A** dated **hosted** snapshot | 11 | 6.2% | | | **A** dated **hosted** snapshot | 10 | 5.7% | |
| | **B** **open-weight** checkpoint with a size | 25 | 14.1% | | | **B** **open-weight** checkpoint with a size | 24 | 13.7% | |
| | **A+B** resolvable to an artefact | **36** | **20.3%** | | | **A+B** resolvable to an artefact | **34** | **19.4%** | |
| | **C** family only, no version | 130 | 73.4% | | | **C** family only, no version | 131 | 74.9% | |
| | **D** no identifiable model | 11 | 6.2% | | | **D** no identifiable model | 10 | 5.7% | |
| | |
| | **The population is the 175 papers that used or produced LLM labels, not the 177 that mention one, and getting that wrong was the third bug in this table.** A re-review pass (§12.13, finding 1) found the bucket loop scanning //every// ''llm'' tuple including ''compared'' ones, which promoted two papers on the strength of a baseline they argued against — one of them a paper whose **only** ''llm'' tuple is ''compared'', counted in bucket A for a model it never ran. The script now buckets on ''llmUsed()'' and **throws if the buckets do not sum to the population**, so the two filters cannot diverge again silently. |
| |
| Every string in all four buckets is printed by the script (§12.9). **Two calls a reasonable person would make differently:** ''Mistral Large'' counts as B though it is a hosted API model, and ''FLAN-T5-XXL'' counts as B on a word-sized parameter count (''PARAM_SIZE'' accepts ''xxl'', ''large'', ''mini''). Both are visible in the printed list rather than buried in a share. | Every string in all four buckets is printed by the script (§12.9). **Two calls a reasonable person would make differently:** ''Mistral Large'' counts as B though it is a hosted API model, and ''FLAN-T5-XXL'' counts as B on a word-sized parameter count (''PARAM_SIZE'' accepts ''xxl'', ''large'', ''mini''). Both are visible in the printed list rather than buried in a share. |
| | PASS-CITE | 1 | present once inline citation markers are stripped — the extractor drops them, so ''Qwen3 [49]'' becomes ''Qwen3'' | | | PASS-CITE | 1 | present once inline citation markers are stripped — the extractor drops them, so ''Qwen3 [49]'' becomes ''Qwen3'' | |
| | PASS-NGRAM | 2 | ≥80% of the quote's word 5-grams present; the sentence is in the paper but the extraction reworded a word or two, or ''.cols'' interleaved a float into it | | | PASS-NGRAM | 2 | ≥80% of the quote's word 5-grams present; the sentence is in the paper but the extraction reworded a word or two, or ''.cols'' interleaved a float into it | |
| | **FAIL** | **6** | not present under any of the above | | | **RESCUED** | **6** | below every tier above against the rendering the extractor read, and at or above the n-gram tier against an independent ''pypdf'' rendering of the same ''paper.pdf'' | |
| | | **FAIL in both renderings** | **0** | not present under any tier in either rendering | |
| |
| **All six failures are ''privacy-policy'' or ''consent-notice'' tuples, and no page on this site quotes any of them.** Every tuple behind a figure on the three edited pages passed at some tier. | **Re-run 2026-09-21 with the PDF fallback, and the six FAILs are now six RESCUEDs: nothing in this pass fails in both renderings.** See //Quote-check refresh, 2026-09-21// at the foot of this page. Four of the six go to a **complete** n-gram match in the PDF (36/36, 14/14, 21/21, 11/11) and one to 19/19; the sixth, ''PETS/2026/word-level-annotation…'', goes 11/20 → 16/20, which clears the 80% tier. All six are ''privacy-policy'' or ''consent-notice'' tuples and no page on this site quotes any of them, so no published claim moves; every tuple behind a figure on the three edited pages passed at some tier already. |
| |
| **Two of the soft tiers were added after reading the failures, and that is worth admitting.** The first run reported 15 fail; reading them showed the extraction drops inline citation markers, and that ''paper.cols.txt'' interleaves table captions into sentences — ''PETS/2026/disclosure-divergence…'' failed because the caption //"Table 1: LLM backend comparison on 100-app validation."// lands inside the quoted sentence. Adding tiers for those is right, since the alternative is publishing "half the quotes are unlocatable", which is false. But it is also a checker relaxed until it agreed with a hypothesis. The mitigation is that every soft pass is itemised with its reason and, for PASS-NGRAM, the broken n-grams, so each call is checkable by hand. | **Two of the soft tiers were added after reading the failures, and that is worth admitting.** The first run reported 15 fail; reading them showed the extraction drops inline citation markers, and that ''paper.cols.txt'' interleaves table captions into sentences — ''PETS/2026/disclosure-divergence…'' failed because the caption //"Table 1: LLM backend comparison on 100-app validation."// lands inside the quoted sentence. Adding tiers for those is right, since the alternative is publishing "half the quotes are unlocatable", which is false. But it is also a checker relaxed until it agreed with a hypothesis. The mitigation is that every soft pass is itemised with its reason and, for PASS-NGRAM, the broken n-grams, so each call is checkable by hand. |
| |
| **A latent bug found in review changed nothing, and is recorded anyway.** 90 of the 5,869 ''paper.cols.txt'' files (1.5%) contain **NUL bytes** — 913 in the TGNN paper alone. They are not whitespace to ''\s'', invisible in a terminal, and they make shell ''grep'' treat the file as binary and suppress every match silently (use ''grep -a''). The checker did not strip them; it now does. Re-running gave **15 / 9 / 6 before and after**, because no NUL happened to land inside one of these 30 quotes. On a different sample it would have been a published false FAIL. | **A latent bug found in review changed nothing, and is recorded anyway.** 90 of the 5,869 ''paper.cols.txt'' files (1.5%) contain **NUL bytes** — 913 in the TGNN paper alone. They are not whitespace to ''\s'', invisible in a terminal, and they make shell ''grep'' treat the file as binary and suppress every match silently (use ''grep -a''). The checker did not strip them; it now does. Re-running gave **15 / 9 / 6 before and after**, because no NUL happened to land inside one of these 30 quotes. On a different sample it would have been a published false FAIL. ((The ''6'' in that sentence is the 2026-09-03 FAIL count. Under the PDF fallback added on 2026-09-04 the same six are ''RESCUED'' and the fail-in-both count is 0 — see //Quote-check refresh, 2026-09-21//. The NUL finding itself is unaffected: it is about ''.cols'' preprocessing, not about the fallback.)) |
| |
| **6 of 30 (20%) is far above the corpus-wide 0.9% unlocatable rate** on [[literature:corpus]]. It is not a re-measurement: 30 tuples, non-random, all 2024–2026, weighted to PETS 2026 whose PDFs are the newest and worst-rendered. Read it as a reason to check quotes from the 2025–2026 slice specifically. | **6 of 30 (20%) is far above the corpus-wide 0.9% unlocatable rate** on [[literature:corpus]]. It is not a re-measurement: 30 tuples, non-random, all 2024–2026, weighted to PETS 2026 whose PDFs are the newest and worst-rendered. Read it as a reason to check quotes from the 2025–2026 slice specifically. |
| * **The magnitude of the third-party-service drop.** §12.4. The 2026-alone column is 20 papers. Needs CCS 2026 and IMC 2026. | * **The magnitude of the third-party-service drop.** §12.4. The 2026-alone column is 20 papers. Needs CCS 2026 and IMC 2026. |
| * **Whether the curated-database row means anything at all.** §12.4. It is not significant and its contents are mis-filed. A stronger answer would need the twelve papers read and the enum corrected, not re-queried. | * **Whether the curated-database row means anything at all.** §12.4. It is not significant and its contents are mis-filed. A stronger answer would need the twelve papers read and the enum corrected, not re-queried. |
| * **Whether ''target == "other"'' hides more LLM website, script or tracker classification.** Three earlier drafts of these pages said "reading 116 papers would settle it". **That was lazy and the reviewer was right to say so** (§12.12, finding 10): ''other'' carries a ''targetDetail'', stated on all 157 such tuples, and a keyword probe over it now runs in the report script. It returns 16 tuples and **none is a website-topic, JavaScript, tracker or cookie classification** — they are IoT device categories, image content, decompiler clusters, MCP server categories, GDPR data categories. So the zeros survive the ''other'' bucket **at keyword recall**, which is what the pages now say. A probe is not a read of 116 papers and cannot rule out a ''targetDetail'' phrased in none of those words. | * **Whether ''target == "other"'' hides more LLM website, script or tracker classification.** Three earlier drafts of these pages said "reading 116 papers would settle it". **That was lazy and the reviewer was right to say so** (§12.12, finding 10): ''other'' carries a ''targetDetail'', stated on all 157 such tuples, and a keyword probe over it now runs in the report script. It returns **20** tuples and **none is a website-topic, JavaScript, tracker or cookie classification** — they are IoT device categories and control pages, image content, decompiler clusters, MCP server categories, GDPR data categories, phishing-personalisation page text and threat-intelligence page triage. So the zeros survive the ''other'' bucket **at keyword recall**, which is what the pages now say. |
| * **Whether the four GPT strings in the fold residue change a published share.** §12.5a. Not folded on purpose; a task item exists. | |
| | **The probe's width decided that answer, and the first width was wrong.** It used ''\bpage\b'', which does not match the compound "webpage" — there is no word boundary between "web" and "page" — and it had no bare ''web'' at all, so it returned 16 rather than 20 and silently dropped //"relevant person-specific webpage information"// and //"IOB presence and trustworthiness in web content"//. Found in re-review (§12.13, finding 3). Both were then read and neither changes the conclusion, which is the only reason the published claim survived a probe that was under-recalling by 20%. A probe is not a read of 116 papers, its hits must be read rather than counted, and its regex is a load-bearing part of the claim. |
| | * **Whether the four GPT strings in the fold residue change a published share.** §12.5a. **Settled on 2026-09-21 and no longer open** — the fold was extended as a rule and the three GPT rows re-derived; they do change one published share on this page and none anywhere else. See §12.14. |
| * **Whether the six quote failures are extraction paraphrase or ''.cols'' rendering.** One was read and was rendering. The other five were not, because no page quotes them. | * **Whether the six quote failures are extraction paraphrase or ''.cols'' rendering.** One was read and was rendering. The other five were not, because no page quotes them. |
| |
| ------------------------------------- --------------- ----- | ------------------------------------- --------------- ----- |
| GPT-4o 46 26.0% | GPT-4o 46 26.0% |
| GPT-4 (non-4o) 40 22.6% | GPT-4 (non-4o) 44 24.9% |
| GPT-3.5 / GPT-3 / ChatGPT 29 16.4% | GPT-3.5 / GPT-3 / ChatGPT 29 16.4% |
| UNMAPPED 23 13.0% | UNMAPPED 18 10.2% |
| Llama family 13 7.3% | Llama family 13 7.3% |
| Gemini / PaLM 12 6.8% | Gemini / PaLM 12 6.8% |
| Unnamed LLM 11 6.2% | Unnamed LLM 11 6.2% |
| OpenAI reasoning / GPT-5 tier 9 5.1% | OpenAI reasoning / GPT-5 tier 10 5.6% |
| Qwen family 8 4.5% | Qwen family 8 4.5% |
| DeepSeek family 7 4.0% | DeepSeek family 7 4.0% |
| Encoder / seq2seq LM (not a chat LLM) 2 1.1% | Encoder / seq2seq LM (not a chat LLM) 2 1.1% |
| |
| --- UNMAPPED residue: 25 distinct strings, printed in full --- | --- UNMAPPED residue: 20 distinct strings, printed in full --- |
| 2x Grok-3 | 2x Grok-3 |
| 2x local LLMs (custom prompts) | 2x local LLMs (custom prompts) |
| 1x BLIP2 | 1x BLIP2 |
| 1x Chat-GPT 3.5 and 4 | |
| 1x ChatGLM | 1x ChatGLM |
| 1x custom structured prompts with fine-tuned LLMs | 1x custom structured prompts with fine-tuned LLMs |
| 1x foundation LLMs | 1x foundation LLMs |
| 1x GLM-4.5 | 1x GLM-4.5 |
| 1x GPT 4.1 | |
| 1x GPT-4.0 | |
| 1x GPT-4.5 | |
| 1x GPT-o1 | |
| 1x HtmlLLM-Detector | 1x HtmlLLM-Detector |
| 1x Kimi | 1x Kimi |
| DOES THE PAPER NAME A MODEL YOU COULD RESOLVE? | DOES THE PAPER NAME A MODEL YOU COULD RESOLVE? |
| ======================================================================== | ======================================================================== |
| What the strongest thing the paper names is Papers (of 177) Share | Population: the 175 papers that USED or PRODUCED LLM labels, not the 177 that mention one. |
| | What the strongest thing the paper names is Papers (of 175) Share |
| ---------------------------------------------------------------- --------------- ----- | ---------------------------------------------------------------- --------------- ----- |
| A a DATED HOSTED SNAPSHOT (gpt-4-turbo-2024-04-09) 11 6.2% | A a DATED HOSTED SNAPSHOT (gpt-4-turbo-2024-04-09) 10 5.7% |
| B an OPEN-WEIGHT CHECKPOINT with a size (Llama-3.1-70B-Instruct) 25 14.1% | B an OPEN-WEIGHT CHECKPOINT with a size (Llama-3.1-70B-Instruct) 24 13.7% |
| A or B — resolvable to an artefact at all 36 20.3% | A or B — resolvable to an artefact at all 34 19.4% |
| C a FAMILY with no version (GPT-4, ChatGPT, Mistral) 130 73.4% | C a FAMILY with no version (GPT-4, ChatGPT, Mistral) 131 74.9% |
| D NO IDENTIFIABLE MODEL ("an LLM", "foundation LLMs") 11 6.2% | D NO IDENTIFIABLE MODEL ("an LLM", "foundation LLMs") 10 5.7% |
| |
| >>> B is the stronger kind of pin: an open-weight checkpoint stays | >>> B is the stronger kind of pin: an open-weight checkpoint stays |
| released artefact but not a revision. Every string is below. | released artefact but not a revision. Every string is below. |
| |
| --- bucket A — dated hosted snapshot: 11 distinct strings across 11 papers, printed in full --- | --- bucket A — dated hosted snapshot: 10 distinct strings across 10 papers, printed in full --- |
| 2x gpt-4-turbo-2024-04-09 [IMC/2024/analyzing-corporate-privacy-policies-using-ai-chatbots] [IMC/2024/beyond-the-guidelines-assessing-metas-political-ad-moderation-in-the-eu] | 2x gpt-4-turbo-2024-04-09 [IMC/2024/analyzing-corporate-privacy-policies-using-ai-chatbots] [IMC/2024/beyond-the-guidelines-assessing-metas-political-ad-moderation-in-the-eu] |
| 1x ChatGPT (gpt-3.5-turbo-0613) [USENIX/2024/llm-fuzzer-scaling-assessment-of-large-language-model-jailbreaks] | 1x ChatGPT (gpt-3.5-turbo-0613) [USENIX/2024/llm-fuzzer-scaling-assessment-of-large-language-model-jailbreaks] |
| 1x gpt-3.5-turbo-0125 [USENIX/2025/mind-the-inconspicuous-revealing-the-hidden-weakness-in-aligned-llms-refusal-bou] | 1x gpt-3.5-turbo-0125 [USENIX/2025/mind-the-inconspicuous-revealing-the-hidden-weakness-in-aligned-llms-refusal-bou] |
| 1x GPT-3.5-turbo-0125 [NDSS/2025/automated-expansion-of-privacy-data-taxonomy-for-compliant-data-breach-notification] | |
| 1x gpt-3.5-turbo-0613 [NDSS/2025/generating-api-parameter-security-rules-with-llm-for-api-misuse-detection] | 1x gpt-3.5-turbo-0613 [NDSS/2025/generating-api-parameter-security-rules-with-llm-for-api-misuse-detection] |
| 1x GPT-4 (gpt-4-0613) [USENIX/2024/llm-fuzzer-scaling-assessment-of-large-language-model-jailbreaks] | 1x GPT-4 (gpt-4-0613) [USENIX/2024/llm-fuzzer-scaling-assessment-of-large-language-model-jailbreaks] |
| 1x OpenAI gpt-4o-2024-05-13 [USENIX/2025/mbfuzzer-a-multi-party-protocol-fuzzer-for-mqtt-brokers] | 1x OpenAI gpt-4o-2024-05-13 [USENIX/2025/mbfuzzer-a-multi-party-protocol-fuzzer-for-mqtt-brokers] |
| |
| --- bucket B — open-weight checkpoint with a size: 27 distinct strings across 25 papers, printed in full --- | --- bucket B — open-weight checkpoint with a size: 25 distinct strings across 24 papers, printed in full --- |
| 2x Llama-3.1-8B-Instruct | 2x Llama-3.1-8B-Instruct |
| 1x FLAN-T5-XXL | 1x FLAN-T5-XXL |
| 1x Llama3-8B | 1x Llama3-8B |
| 1x Llama3:70b + GPT-4 hybrid | 1x Llama3:70b + GPT-4 hybrid |
| 1x Mistral Large | |
| 1x Mistral-7B | |
| 1x Mistral-7B-Instruct-v0.2 | 1x Mistral-7B-Instruct-v0.2 |
| 1x Qwen2-72B-Instruct | 1x Qwen2-72B-Instruct |
| 1x Vicuna-7b | 1x Vicuna-7b |
| |
| --- bucket C — family only, no version: 91 distinct strings across 130 papers, printed in full --- | --- bucket C — family only, no version: 88 distinct strings across 131 papers, printed in full --- |
| 25x GPT-4o | 25x GPT-4o |
| 15x GPT-4 | 14x GPT-4 |
| 6x GPT-4.1 | 6x GPT-4.1 |
| 6x GPT-4o-mini | 6x GPT-4o-mini |
| 5x ChatGPT | 5x ChatGPT |
| 4x GPT-3.5 | 4x GPT-3.5 |
| 4x GPT-5 | |
| 3x ChatGPT-4 | 3x ChatGPT-4 |
| 3x GPT-3.5-turbo | 3x GPT-3.5-turbo |
| 2x DeepSeek-V3 | 3x GPT-5 |
| 2x GPT-3.5 Turbo | 2x GPT-3.5 Turbo |
| 2x GPT-4 Turbo | 2x GPT-4 Turbo |
| 1x ChatGPT-4.0 | 1x ChatGPT-4.0 |
| 1x Claude Sonnet 3.5 | 1x Claude Sonnet 3.5 |
| 1x Claude-4.5-Sonnet | |
| 1x Claude, GPT-4o, Gemini, and DeepSeek (cross-LLM voting) | 1x Claude, GPT-4o, Gemini, and DeepSeek (cross-LLM voting) |
| 1x DeepSeek | 1x DeepSeek |
| 1x DeepSeek-R1 | 1x DeepSeek-R1 |
| 1x DeepSeek-R1 alert-verification module | 1x DeepSeek-R1 alert-verification module |
| | 1x DeepSeek-V3 |
| 1x DeepSeek-V3.2-Exp | 1x DeepSeek-V3.2-Exp |
| 1x fine-tuned GPT-4o | 1x fine-tuned GPT-4o |
| 1x Gemini 2.5 Pro | 1x Gemini 2.5 Pro |
| 1x Gemini 2.5-Flash | 1x Gemini 2.5-Flash |
| | 1x Gemini Ultra |
| 1x Gemini-2.0-Flash | 1x Gemini-2.0-Flash |
| 1x Gemini-2.5-Flash-Lite | |
| 1x Gemini-2.5-Flash-Lite and GPT-4o-mini ensemble | 1x Gemini-2.5-Flash-Lite and GPT-4o-mini ensemble |
| 1x Gemini-2.5-pro-preview-05-06 | 1x Gemini-2.5-pro-preview-05-06 |
| 1x Gemini-3.1-Pro evaluator | 1x Gemini-3.1-Pro evaluator |
| 1x GLM-4.5 | |
| 1x GPT 4.1 | 1x GPT 4.1 |
| 1x GPT-3 curie | 1x GPT-3 curie |
| 1x GPT-4-turbo | 1x GPT-4-turbo |
| 1x GPT-4-Turbo | 1x GPT-4-Turbo |
| 1x GPT-4, GPT-4 Turbo, and GPT-3.5 Turbo | |
| 1x GPT-4.0 | 1x GPT-4.0 |
| 1x GPT-4.1 (custom extraction prompt) | 1x GPT-4.1 (custom extraction prompt) |
| 1x Vertex AI text-bison | 1x Vertex AI text-bison |
| |
| --- bucket D — no identifiable model: 11 distinct strings across 11 papers, printed in full --- | --- bucket D — no identifiable model: 10 distinct strings across 10 papers, printed in full --- |
| 2x LLM judge (custom) [WWW/2026/inference-cost-attacks-for-retrieval-augmented-large-language-models] [NDSS/2026/when-cache-poisoning-meets-llm-systems-semantic-cache-poisoning-and-its-countermeasures] | 2x LLM judge (custom) [WWW/2026/inference-cost-attacks-for-retrieval-augmented-large-language-models] [NDSS/2026/when-cache-poisoning-meets-llm-systems-semantic-cache-poisoning-and-its-countermeasures] |
| 1x foundation LLMs [IEEE-SP/2025/code-speaks-louder-exploring-security-and-privacy-relevant-regional-variations-i] | 1x foundation LLMs [IEEE-SP/2025/code-speaks-louder-exploring-security-and-privacy-relevant-regional-variations-i] |
| 1x local LLMs (custom prompts) [USENIX/2026/a-large-scale-study-of-personalized-phishing-using-large-language-models] | 1x local LLMs (custom prompts) [USENIX/2026/a-large-scale-study-of-personalized-phishing-using-large-language-models] |
| 1x PhishLLM [USENIX/2025/unsafe-llm-based-search-quantitative-analysis-and-mitigation-of-safety-risks-in] | 1x PhishLLM [USENIX/2025/unsafe-llm-based-search-quantitative-analysis-and-mitigation-of-safety-risks-in] |
| 1x Prompt Instruct [WWW/2026/privsniffer-graph-based-contextual-privacy-leakage-detection-for-user-generated] | |
| 1x weighted multi-model ensemble (custom) [WWW/2026/webgeoinfer-structure-free-multi-stage-framework-for-geolocation-inference-from] | 1x weighted multi-model ensemble (custom) [WWW/2026/webgeoinfer-structure-free-multi-stage-framework-for-geolocation-inference-from] |
| |
| 157 of 157 (100.0%) state a targetDetail. | 157 of 157 (100.0%) state a targetDetail. |
| |
| --- targetDetail matching the probe (websit|domain|\burl\b|homepage|\bpage\b|script|tracker|cookie|\bsdk\b|first.part|third.part|categor) — 16 tuples, all printed --- | --- targetDetail matching the probe (websit|domain|\burl\b|homepage|page|\bweb\b|script|tracker|cookie|\bsdk\b|first.part|third.part|categor) — 20 tuples, all printed --- |
| IMC/2023/in-the-room-where-it-happens-characterizing-local-communication-and-threats-in-s | IMC/2023/in-the-room-where-it-happens-characterizing-local-communication-and-threats-in-s |
| IoT device vendors and categories | IoT device vendors and categories |
| IMC/2025/learning-as-to-organization-mappings-with-borges | IMC/2025/learning-as-to-organization-mappings-with-borges |
| favicon and associated final-URL groups | favicon and associated final-URL groups |
| | NDSS/2025/hidden-and-lost-control-on-security-design-risks-in-iot-user-facing-matter-controller |
| | UMCCI flaws and user-facing Matter control pages |
| USENIX/2025/evaluating-privacy-policies-under-modern-privacy-laws-at-scale-an-llm-based-auto | USENIX/2025/evaluating-privacy-policies-under-modern-privacy-laws-at-scale-an-llm-based-auto |
| personal information categories, purposes, and third-party recipients | personal information categories, purposes, and third-party recipients |
| WWW/2026/adaptive-location-hierarchy-learning-for-long-tailed-mobility-prediction | WWW/2026/adaptive-location-hierarchy-learning-for-long-tailed-mobility-prediction |
| hierarchical mappings between location categories, activities, and needs | hierarchical mappings between location categories, activities, and needs |
| | USENIX/2026/a-large-scale-study-of-personalized-phishing-using-large-language-models |
| | relevant person-specific webpage information |
| | NDSS/2026/indicator-of-benignity-an-industry-view-of-false-positive-in-malicious-domain-detection-and-its-mitigation |
| | IOB presence and trustworthiness in web content |
| PETS/2026/operationalizing-the-motivated-intruder-a-codebook-guided-inference-framework-fo | PETS/2026/operationalizing-the-motivated-intruder-a-codebook-guided-inference-framework-fo |
| GDPR personal data, special-category data, and trade-secret sensitivity | GDPR personal data, special-category data, and trade-secret sensitivity |
| WWW/2026/opendigger-a-practical-framework-for-assessing-community-health-and-sustainabili | WWW/2026/opendigger-a-practical-framework-for-assessing-community-health-and-sustainabili |
| repository technical domain | repository technical domain |
| | WWW/2026/bridging-expert-reasoning-and-llm-detection-a-knowledge-driven-framework-for-mal |
| | threat-intelligence pages containing actionable malicious-code analysis |
| |
| >>> Read the list, do not trust the count: the question is WHAT these are, | >>> Read the list, do not trust the count: the question is WHAT these are, |
| FAIL = not present under any of the above. | FAIL = not present under any of the above. |
| |
| FAIL IMC/2024/analyzing-corporate-privacy-policies-using-ai-chatbots [privacy-policy] 26/36 5-grams present | RESCUED IMC/2024/analyzing-corporate-privacy-policies-using-ai-chatbots [privacy-policy] 26/36 -> 36/36 PDF "we design a set of task prompts for an AI chatbot to split scraped content into …" |
| quote: we design a set of task prompts for an AI chatbot to split scraped content into sections, and then extract and label mentions of collected data types, data collection purposes, data retention and protection practices, and user rights and choices. | |
| longest verbatim prefix (19 of 246 chars): we design a set of | |
| PASS USENIX/2024/less-defined-knowledge-and-more-true-alarms-reference-based-phishing-detection-w [website-category] "This prompt is fed into a language model using a chain-of-thought approach, enfo…" | PASS USENIX/2024/less-defined-knowledge-and-more-true-alarms-reference-based-phishing-detection-w [website-category] "This prompt is fed into a language model using a chain-of-thought approach, enfo…" |
| PASS-NGRAM IMC/2025/an-in-depth-investigation-of-data-collection-in-llm-app-ecosystems [privacy-policy] 10/10 5-grams "We develop an LLM-based framework to check the consistency of data collection di…" | PASS-NGRAM IMC/2025/an-in-depth-investigation-of-data-collection-in-llm-app-ecosystems [privacy-policy] 10/10 5-grams "We develop an LLM-based framework to check the consistency of data collection di…" |
| PASS PETS/2026/overcoming-language-barriers-multilingual-analysis-of-the-2023-swiss-privacy-law [privacy-policy] "For each policy, we issue a single inference request to the model and require it…" | PASS PETS/2026/overcoming-language-barriers-multilingual-analysis-of-the-2023-swiss-privacy-law [privacy-policy] "For each policy, we issue a single inference request to the model and require it…" |
| PASS WWW/2025/harmful-terms-and-where-to-find-them-measuring-and-modeling-unfavorable-financia [website-category] "To evaluate our classification methods, we manually annotated a sample of 500 we…" | PASS WWW/2025/harmful-terms-and-where-to-find-them-measuring-and-modeling-unfavorable-financia [website-category] "To evaluate our classification methods, we manually annotated a sample of 500 we…" |
| FAIL PETS/2026/word-level-annotation-of-gdpr-transparency-compliance-in-privacy-policies-using [privacy-policy] 11/20 5-grams present | RESCUED PETS/2026/word-level-annotation-of-gdpr-transparency-compliance-in-privacy-policies-using [privacy-policy] 11/20 -> 16/20 PDF "On a manually labelled sample of 340 randomly selected documents ... using GPT 4…" |
| quote: On a manually labelled sample of 340 randomly selected documents ... using GPT 4.1 as the classifying LLM ... achieved an accuracy of 99.7%. | |
| longest verbatim prefix (59 of 140 chars): On a manually labelled sample of 340 randomly selected docu | |
| PASS-LOOSE PETS/2026/word-level-annotation-of-gdpr-transparency-compliance-in-privacy-policies-using [privacy-policy] "This LLM-based classifier predicts the set of labels Lp relevant to each passage" | PASS-LOOSE PETS/2026/word-level-annotation-of-gdpr-transparency-compliance-in-privacy-policies-using [privacy-policy] "This LLM-based classifier predicts the set of labels Lp relevant to each passage" |
| FAIL CCS/2025/whispertest-a-voice-control-based-library-for-ios-ui-automation [consent-notice] 9/14 5-grams present | RESCUED CCS/2025/whispertest-a-voice-control-based-library-for-ios-ui-automation [consent-notice] 9/14 -> 14/14 PDF "we used a more efficient text-only model (Qwen2.5-7B) to detect the presence of …" |
| quote: we used a more efficient text-only model (Qwen2.5-7B) to detect the presence of consent dialogs | RESCUED PETS/2025/automating-governing-knowledge-commons-and-contextual-integrity-gkc-ci-privacy-p [privacy-policy] 13/21 -> 21/21 PDF "We randomly reserved 70% of the manual annotations to constitute our training da…" |
| longest verbatim prefix (25 of 95 chars): we used a more efficient | |
| FAIL PETS/2025/automating-governing-knowledge-commons-and-contextual-integrity-gkc-ci-privacy-p [privacy-policy] 13/21 5-grams present | |
| quote: We randomly reserved 70% of the manual annotations to constitute our training data (21,588 examples), while the other 30% (9252 examples) were testing data. | |
| longest verbatim prefix (54 of 156 chars): We randomly reserved 70% of the manual annotations to | |
| PASS PETS/2025/automating-governing-knowledge-commons-and-contextual-integrity-gkc-ci-privacy-p [privacy-policy] "For the prompted non-fine-tuned LLMs, we used GPT-4, GPT-4 Turbo, and GPT-3.5 Tu…" | PASS PETS/2025/automating-governing-knowledge-commons-and-contextual-integrity-gkc-ci-privacy-p [privacy-policy] "For the prompted non-fine-tuned LLMs, we used GPT-4, GPT-4 Turbo, and GPT-3.5 Tu…" |
| PASS PETS/2025/behavr-user-identification-based-on-vr-sensor-data [privacy-policy] "We also use simple string matching to search for relevant content." | PASS PETS/2025/behavr-user-identification-based-on-vr-sensor-data [privacy-policy] "We also use simple string matching to search for relevant content." |
| PASS PETS/2026/designing-reflective-thinking-based-contextual-privacy-policy-for-mobile-applica [privacy-policy] "We evaluated privacy policy segment extraction accuracy on three state-of-the-ar…" | PASS PETS/2026/designing-reflective-thinking-based-contextual-privacy-policy-for-mobile-applica [privacy-policy] "We evaluated privacy policy segment extraction accuracy on three state-of-the-ar…" |
| PASS PETS/2026/designing-reflective-thinking-based-contextual-privacy-policy-for-mobile-applica [privacy-policy] "We evaluated privacy policy segment extraction accuracy on three state-of-the-ar…" | PASS PETS/2026/designing-reflective-thinking-based-contextual-privacy-policy-for-mobile-applica [privacy-policy] "We evaluated privacy policy segment extraction accuracy on three state-of-the-ar…" |
| FAIL PETS/2026/personal-data-flows-and-privacy-policy-traceability-in-third-party-llm-apps-in-t [privacy-policy] 15/19 5-grams present | RESCUED PETS/2026/personal-data-flows-and-privacy-policy-traceability-in-third-party-llm-apps-in-t [privacy-policy] 15/19 -> 19/19 PDF "A researcher manually verified whether each LLM classification matched the corre…" |
| quote: A researcher manually verified whether each LLM classification matched the correct taxonomy label. GPT-4o-mini achieved an overall accuracy of 87.83%. | |
| longest verbatim prefix (31 of 150 chars): A researcher manually verified | |
| PASS-NGRAM PETS/2026/disclosure-divergence-measuring-privacy-policy-and-data-safety-misalignment-at-s [privacy-policy] 19/20 5-grams "the system outputs two sets of data types, C data (collected) and S data (shared…" | PASS-NGRAM PETS/2026/disclosure-divergence-measuring-privacy-policy-and-data-safety-misalignment-at-s [privacy-policy] 19/20 5-grams "the system outputs two sets of data types, C data (collected) and S data (shared…" |
| FAIL PETS/2026/disclosure-divergence-measuring-privacy-policy-and-data-safety-misalignment-at-s [privacy-policy] 7/11 5-grams present | RESCUED PETS/2026/disclosure-divergence-measuring-privacy-policy-and-data-safety-misalignment-at-s [privacy-policy] 7/11 -> 11/11 PDF "both models were evaluated using the same preprocessing pipeline, chunking strat…" |
| quote: both models were evaluated using the same preprocessing pipeline, chunking strategy, prompts, and output schema. | |
| longest verbatim prefix (66 of 112 chars): both models were evaluated using the same preprocessing pipeline, | |
| PASS USENIX/2025/darkgram-a-large-scale-analysis-of-cybercriminal-activity-channels-on-telegram [website-category] "GPT-4 added four new categories, bringing the total number of categories to 19. …" | PASS USENIX/2025/darkgram-a-large-scale-analysis-of-cybercriminal-activity-channels-on-telegram [website-category] "GPT-4 added four new categories, bringing the total number of categories to 19. …" |
| PASS USENIX/2025/darkgram-a-large-scale-analysis-of-cybercriminal-activity-channels-on-telegram [website-category] "We evaluated 100 randomly selected cases and found GPT-4's predictions to be acc…" | PASS USENIX/2025/darkgram-a-large-scale-analysis-of-cybercriminal-activity-channels-on-telegram [website-category] "We evaluated 100 randomly selected cases and found GPT-4's predictions to be acc…" |
| PETS/2026/ai-in-the-loop-privacy-preserving-real-time-scam-detection-and-conversational-sc [consent-notice] elided into 2 fragments | PETS/2026/ai-in-the-loop-privacy-preserving-real-time-scam-detection-and-conversational-sc [consent-notice] elided into 2 fragments |
| |
| PART A: 15 verbatim, 9 soft (elided or punctuation/glyph), 6 fail, 0 with no fulltext. | PART A: 15 verbatim, 9 soft (elided or punctuation/glyph), 6 rescued from the PDF, 0 fail in both renderings, 0 with no fulltext. |
| |
| |
| | 2 | "CCS 2026 and IMC 2026 … the two whose 2022–2024 populations lean hardest on third-party services" is **false for CCS**, which leans least (8.3%, lowest of seven). | **ACCEPTED — wrong** | Verified (§12.4). The sentence is rewritten and the by-venue table is now printed. It was the only sentence in the section with no query behind it. | | | 2 | "CCS 2026 and IMC 2026 … the two whose 2022–2024 populations lean hardest on third-party services" is **false for CCS**, which leans least (8.3%, lowest of seven). | **ACCEPTED — wrong** | Verified (§12.4). The sentence is rewritten and the by-venue table is now printed. It was the only sentence in the section with no query behind it. | |
| | 3 | "11 papers name a model you could re-run" measures OpenAI date-strings, penalises open-weight papers, and contradicts the paragraph two sentences later. ~25 papers naming ''Llama-3.1-70B-Instruct''-class checkpoints were counted as unpinned. | **ACCEPTED — rebuilt** | The bucket is now split by hosting: A 11, B 25, **A+B 36 (20.3%)**, C 130, D 11 (§12.5b). The page's headline changed from "fewer than one in thirteen pins the model" to "one paper in five names something you could resolve", and the advice split by hosted vs open-weight. | | | 3 | "11 papers name a model you could re-run" measures OpenAI date-strings, penalises open-weight papers, and contradicts the paragraph two sentences later. ~25 papers naming ''Llama-3.1-70B-Instruct''-class checkpoints were counted as unpinned. | **ACCEPTED — rebuilt** | The bucket is now split by hosting: A 11, B 25, **A+B 36 (20.3%)**, C 130, D 11 (§12.5b). The page's headline changed from "fewer than one in thirteen pins the model" to "one paper in five names something you could resolve", and the advice split by hosted vs open-weight. | |
| | 4 | [[privacy:javascript]] said "top 7%" while [[design:website_classification]] said 6.2% — the earlier split was not propagated to the page linking to that heading. | **ACCEPTED** | Both now read 20.3%, from the rebuilt buckets. The cross-page claim check in ''report_llm_currency.mjs'' does **not** cover this kind of sentence, and that is a real limit of the guard: it checks the corpus claims, not prose that quotes another page. | | | 4 | [[privacy:javascript]] said "top 7%" while [[design:website_classification]] said 6.2% — the earlier split was not propagated to the page linking to that heading. | **ACCEPTED** | Both were brought to 20.3% from the rebuilt buckets — and then both to **19.4%** when the re-review found those buckets still counting compared-against models (§12.13, finding 1), so this one figure was propagated twice across two pages in a single run. The cross-page claim check in ''report_llm_currency.mjs'' does **not** cover this kind of sentence, and that is a real limit of the guard: it checks the corpus claims, not prose that quotes another page. | |
| | 5 | "Direction survives both controls" overstates two nested subsamples. Fisher's exact: 2025-alone p = 0.27; curated-database p = 0.08–0.23 on every cut. The two "controls" are not independent tests. | **ACCEPTED** | Fisher's exact is now computed in ''report_website_classification.mjs'' and every //p// and //n// is on the page. 2025-alone is demoted from control to description. The curated row is retracted (finding 1). The third-party drop keeps its direction on p = 0.006 (five-venue) and 0.047 (full window). | | | 5 | "Direction survives both controls" overstates two nested subsamples. Fisher's exact: 2025-alone p = 0.27; curated-database p = 0.08–0.23 on every cut. The two "controls" are not independent tests. | **ACCEPTED** | Fisher's exact is now computed in ''report_website_classification.mjs'' and every //p// and //n// is on the page. 2025-alone is demoted from control to description. The curated row is retracted (finding 1). The third-party drop keeps its direction on p = 0.006 (five-venue) and 0.047 (full window). | |
| | 6 | §12.5 listed 9 residue strings from the **pre-fix** output while §12.5 twelve lines later said 11, and the content page promised all 11 were listed here. | **ACCEPTED** | §12.5 rewritten; all bucket listings are now generated by the script and embedded in §12.9 rather than transcribed, so they cannot go stale independently. | | | 6 | §12.5 listed 9 residue strings from the **pre-fix** output while §12.5 twelve lines later said 11, and the content page promised all 11 were listed here. | **ACCEPTED** | §12.5 rewritten; all bucket listings are now generated by the script and embedded in §12.9 rather than transcribed, so they cannot go stale independently. | |
| |
| **The pattern across all four reviewers.** Every finding that changed a published figure — findings 1, 2, 3, 5, 9 here, and finding 1 in §12.10 — was in a place where **a number or a sentence had no printed list behind it**: an enum row nobody had itemised, a claim about a missing venue-year, a regex bucket reported only as a size, a share reported without an //n//. None was in a figure the report script printed with its denominator. **That is the whole finding of this review pass:** the guard the site already has works, and it only covers what a script prints. Everything else is prose, and prose is where all six defects were. | **The pattern across all four reviewers.** Every finding that changed a published figure — findings 1, 2, 3, 5, 9 here, and finding 1 in §12.10 — was in a place where **a number or a sentence had no printed list behind it**: an enum row nobody had itemised, a claim about a missing venue-year, a regex bucket reported only as a size, a share reported without an //n//. None was in a figure the report script printed with its denominator. **That is the whole finding of this review pass:** the guard the site already has works, and it only covers what a script prints. Everything else is prose, and prose is where all six defects were. |
| | |
| | |
| | ==== 12.13 Re-review of the figures, 2026-09-03 ==== |
| | |
| | The figures-versus-script reviewer was re-run on the final state, because both the pages and the scripts had changed substantially since its first pass — the model-version table had been rebuilt from scratch, and Fisher's exact, the composition listings and the ''targetDetail'' probe were all new code it had never seen. It re-ran all five scripts (byte-identical to the committed outputs), **verified the hand-implemented Fisher's exact test independently against exact rational arithmetic in Python** — all eight cells agree to four decimal places — and confirmed the like-for-like growth table, the by-venue table, the ''change (pp)'' column, the curated-database counts, and every figure in §12.2, §12.4, §12.5a and §12.5b. Three findings: |
| | |
| | ^ # ^ Finding ^ Verdict ^ Action ^ |
| | | 1 | The reproducibility buckets scanned **every** ''llm'' tuple, not just ''used''/''produced'' ones, so a paper could be promoted by a model it only //compared against//. ''NDSS/2025/automated-expansion-of-privacy-data-taxonomy…'' was in bucket A for ''GPT-3.5-turbo-0125'' although its only ''llm'' tuple is ''compared'' — it is one of the report's own "2 compared-only" papers. ''CCS/2024/airgapagent…'' was promoted to bucket B by ''Mistral Large'', which it compared against, having //used// ''Gemini Ultra''. | **ACCEPTED, real bug in a live figure** | Population changed to the 175 used/produced papers and the loop to ''llmUsed()''. Published figures moved: A **11 → 10**, B **25 → 24**, **A+B 36 (20.3%) → 34 (19.4%)**, C 130 → 131, D 11 → 10. A ''throw'' now fires if the buckets do not sum to the population. Propagated to [[design:website_classification]] and [[privacy:javascript]], both of which had 20.3% saved. | |
| | | 2 | The hand-written enumeration of the twelve ''curated-database'' papers named only **eleven**, omitting ''NDSS/2026/revealing-the-secret-power…'' ("Media Bias/Fact Check (MBFC)"), while the derived "four media-bias raters" counted it — so a reader could not verify "which four". | **ACCEPTED** | Both copies rewritten and sorted by kind, with counts that sum to twelve. The re-check also improved the claim: **one of the twelve, DuckDuckGo's Tracker Radar Entity List, //is// an open inspectable repository** (verified on GitHub, public and active), just of tracker entities rather than website topics. Saying "the row does not contain open directories" was therefore slightly too strong and now reads "no comparable open //topic// directory". | |
| | | 3 | The ''targetDetail'' probe's ''\bpage\b'' cannot match "webpage", and there was no bare ''web'' — two tuples were silently dropped and the published count was 16, not 20. | **ACCEPTED** | Probe widened; it now returns 20. Both new hits were read: //"relevant person-specific webpage information"// (phishing personalisation) and //"IOB presence and trustworthiness in web content"// (malicious-domain false positives). Neither changes the conclusion, which is luck rather than method — the claim rested on a probe under-recalling by 20%. Both pages updated with the new count and with the fact that the width had to be corrected. | |
| | |
| | **Declared clean by this pass:** the Fisher implementation, the like-for-like growth section, the by-venue third-party-service table, the ''change (pp)'' column, the decision box, the four new //What to Report// items, the trimmed Topics API entry, the folded-services table, and the cross-page claims on [[privacy:javascript]] and [[design:ip_classification]]. |
| | |
| | **Both of the two bugs that reached a live page in this run were of one kind:** a query in a new script that did not apply a filter the same script applies elsewhere — the ''validation'' allowlist against ''report_website_classification.mjs'' (§12.2), and the ''usedOrMentioned'' filter against this script's own population section (finding 1). Neither was visible in the output. Both are now guarded: the allowlist is shared, and the bucket sum ''throw''s. **The cheap general check is: for every filter a script defines, grep the script for the places that should use it and do not.** |
| |
| |
| [[design:website_classification|← back to the content page]] · [[literature:corpus|corpus-level provenance]] | [[design:website_classification|← back to the content page]] · [[literature:corpus|corpus-level provenance]] |
| | |
| | ==== 12.14 The model-family fold extended, 2026-09-21 ==== |
| | |
| | **What was wrong.** §12.5a's residue contained four strings that are OpenAI models the ''GPT-4 (non-4o)'' and reasoning-tier regexes should have absorbed: ''GPT 4.1'' (a space instead of a hyphen), ''GPT-4.0'' and ''GPT-4.5'' (a decimal, which the ''%%(?![.\do])%%'' lookahead rejected outright although it was written only to keep ''4o'' out), and ''GPT-o1'' (the o-series, which the fold matched only in its bare ''o1-mini'' form). One paper each. They were **deliberately left in the residue on 2026-09-03** rather than patched after the output was read. |
| | |
| | **Why it was left, and what changed.** Patching a fold to absorb the residue you have just looked at is how a documented rule stops being one: the next reader cannot tell a rule from a list of the strings that embarrassed it. So the fold has been extended as a **rule about how OpenAI writes model names** — the separator after ''GPT'' may be a hyphen, a space or nothing; a version may carry a decimal; the o-series is written both bare and ''GPT''-prefixed — and the rule is stated in the script, above the table it feeds: |
| | |
| | <code> |
| | [/gpt[- ]?4o|gpt4o/i, 'GPT-4o'], |
| | [/gpt[- ]?4-turbo|gpt[- ]?4(\.\d+)?(?![\do])/i, 'GPT-4 (non-4o)'], |
| | [/gpt[- ]?3\.5|chat-?gpt|text-davinci|gpt[- ]?3(?!\.5)/i, 'GPT-3.5 / GPT-3 / ChatGPT'], |
| | [/gpt[- ]?5|gpt[- ]?o[1345]\b|\bo[134]-(mini|preview|pro)\b|\bo4-mini\b/i, 'OpenAI reasoning / GPT-5 tier'], |
| | </code> |
| | |
| | **The rule moves two strings that were never in the residue, and that is the point.** Applied to all 154 distinct ''resourceName'' strings the fold sees, it changes six: the four above, plus ''Chat-GPT 3.5 and 4'' (hyphenated "Chat-GPT", so ''chatgpt'' never matched it — UNMAPPED → GPT-3.5) and ''ChatGPT-4.0'' (which the old decimal lookahead pushed past the GPT-4 row into the ChatGPT row — GPT-3.5 → GPT-4). A rule fitted to the residue would have moved exactly four. Both were verified string by string before the fold was changed, not after. |
| | |
| | ^ Row ^ Was ^ Is ^ Which papers moved ^ |
| | | GPT-4o | 46 (26.0%) | 46 (26.0%) | none — the ''GPT-o1'' paper is IMC/2025 //an-in-depth-investigation-of-data-collection…//, which already counted here for its ''GPT-4o'' string | |
| | | **GPT-4 (non-4o)** | **40 (22.6%)** | **44 (24.9%)** | +4: ''GPT 4.1'' (PETS/2026), ''GPT-4.5'' (CCS/2025), ''GPT-4.0'' (NDSS/2025), ''ChatGPT-4.0'' (USENIX/2024) | |
| | | GPT-3.5 / GPT-3 / ChatGPT | 29 (16.4%) | 29 (16.4%) | **net zero, not "unchanged"**: the ''ChatGPT-4.0'' paper leaves, the ''Chat-GPT 3.5 and 4'' paper (NDSS/2025) arrives | |
| | | **OpenAI reasoning / GPT-5 tier** | **9 (5.1%)** | **10 (5.6%)** | +1: the ''GPT-o1'' paper | |
| | | **UNMAPPED** | **23 (13.0%)** | **18 (10.2%)** | −5 papers; the residue falls from **25 distinct strings to 20** | |
| | |
| | **The three-bucket reproducibility table (§12.5b) does not move**, and it was checked rather than assumed: ''A 10 / B 24 / C 131 / D 10'' before and after, and every bucket's printed string list is byte-identical. The buckets are built from ''HOSTED_SNAPSHOT'', ''OPEN_FAMILY''/''PARAM_SIZE'' and ''NAMED'', none of which the fold touches; ''GPT-o1'' was already in bucket C, because ''NAMED'' contains a bare ''gpt''. The fold and the buckets answer different questions and are deliberately separate regexes. |
| | |
| | **What stays in the residue, and why.** ''Grok-3'' (2 papers), ''GLM-4.5'', ''ChatGLM'' and ''Kimi'' are four vendors with no family in the list. They are a genuine long tail, not a fold failure, and **no family was added for them** — adding one would be the post-hoc patch this section exists to avoid. The remaining 16 strings are descriptions rather than models (''local LLMs (custom prompts)'', ''weighted multi-model ensemble (custom)''), systems built on a model (''PhishLLM'', ''UGCG-GUARD'', ''YouthSafe'', ''RFCGPT''), or non-OpenAI multimodal models (''LLaVA'', ''BLIP2'', ''text-bison''). All 20 are printed in full in §12.9. |
| | |
| | **Where this is published.** Only here. [[:design:website_classification]] publishes the §12.5b bucket table, not the family fold, so the content page needed no edit for this. Checked by grepping the raw source of all **189** pages on the wiki: ''GPT-4 (non-4o)'' and the other family labels return **this page alone**, and the four residue strings return this page and [[:provenance:statistics:annotation]] — where ''GPT 4.1'' appears as a raw ''resourceName'' inside a per-paper quote-check listing, not as a folded row, and so is unaffected by the fold. Same for ''ChatGLM'' on that page. |
| | |
| | <code> |
| | $ node scripts/report_llm_currency.mjs > scripts/report_llm_currency-output.txt |
| | $ diff <old> <new> |
| | 178c178 |
| | < GPT-4 (non-4o) 40 22.6% |
| | > GPT-4 (non-4o) 44 24.9% |
| | 180c180 |
| | < UNMAPPED 23 13.0% |
| | > UNMAPPED 18 10.2% |
| | 184c184 |
| | < OpenAI reasoning / GPT-5 tier 9 5.1% |
| | > OpenAI reasoning / GPT-5 tier 10 5.6% |
| | 192c192 |
| | < --- UNMAPPED residue: 25 distinct strings, printed in full --- |
| | > --- UNMAPPED residue: 20 distinct strings, printed in full --- |
| | 196d195 (Chat-GPT 3.5 and 4) 201,204d199 (GPT 4.1 / GPT-4.0 / GPT-4.5 / GPT-o1) |
| | </code> |
| | |
| | **Those are the only lines that changed in a 446-line report.** Every population, year, venue, target, validation and bucket figure is identical, and the ''compared''-only assertion and the bucket-sum assertion both still pass. |
| | |
| | **One unrelated repair in the same save.** The §12.9 block is regenerated from the committed output file, and the published copy had lost the ''%%\b%%'' escapes from the ''targetDetail'' probe's printed regex (it read ''url|…|web|…'' where the script prints ''%%\burl\b|…|\bweb\b|…%%''). The block now matches the file byte for byte. The probe itself never changed; only the copy on this page was wrong, and it is the width of that probe that decides the claim in §12.7. |
| | |
| | ^ Item ^ Value ^ |
| | | Date | 2026-09-21, unsupervised | |
| | | Script changes | ''scripts/report_llm_currency.mjs'' — the four OpenAI ''FAMILY'' rows, with the rule stated in a comment above them; committed output regenerated | |
| | | Reviewers | one ''sonnet'' figures-vs-script pass; one ''sonnet'' citations/quotes pass | |
| | | Pages saved | this page only | |
| | | Not edited | [[:design:website_classification]], [[:design:ip_classification]], [[:privacy:javascript]], [[:privacy:cookies]] — none publishes a model-family row | |
| | |
| | ===== Markup sweep, 2026-09-17 ===== |
| | |
| | Mechanical rendering repair only: a fresh live raw/XHTML export of 188 pages was checked with ''check_wrap.mjs'' and ''check_typography.mjs''. Affected plugin tags, CLI flags and heading markup were repaired; no figures or substantive prose were changed. The resulting source and rendered DOM were re-checked after saving. |
| | |
| | ===== Quote-check refresh, 2026-09-21 ===== |
| | |
| | The 2026-09-04 ''cols''-vs-PDF audit on [[:provenance:literature:corpus]] showed that 73.1% of evidence quotes that cannot be located in ''paper.cols.txt'' **are** present in an independent ''pypdf'' rendering of the same ''paper.pdf''. ''llm_currency_quotecheck.mjs'' already carried the fallback; what was stale was this page's copy of its output, taken on 2026-09-03. ''scripts/llm_currency_quotecheck-output.txt'' was regenerated by re-running the checker and §12.9's block replaced from it — the block was byte-identical to the old artifact before the re-run, and is byte-identical to the new one after it. |
| | |
| | <code> |
| | $ node scripts/llm_currency_quotecheck.mjs |
| | PART A: 15 verbatim, 9 soft (elided or punctuation/glyph), 6 rescued from the PDF, 0 fail in both renderings, 0 with no fulltext. |
| | </code> |
| | |
| | ^ Figure ^ Was ^ Is ^ Why ^ |
| | | PART A tuples | 30 | 30 | population unchanged | |
| | | verbatim (PASS) | 15 | 15 | unchanged | |
| | | soft (PASS-ELID / LOOSE / CITE / NGRAM) | 9 | 9 | unchanged | |
| | | rescued from the PDF | — | **6** | these were the old FAILs | |
| | | FAIL | **6** | **0** (in both renderings) | 6 = 6 + 0 | |
| | |
| | **This is the strongest vindication of the "relaxed checker" worry above, and the strongest reason to keep the worry.** The paragraph //Two of the soft tiers were added after reading the failures// admits the tiers were added until the checker agreed with a hypothesis. An **independent** rendering, produced by a different extractor and introduced for a different page, now locates all six of the quotes that survived even the relaxed tiers — so the hypothesis was right. What it does not establish is that the soft tiers are sound in general: they were still fitted after seeing the data, and a reader should keep treating each soft pass as a hand-checkable call rather than as a measurement. |
| | |
| | **Zero is a figure that needs its denominator stated.** "0 fail in both renderings" is 0 of **30** LLM tuples across the targets the three classification pages make claims about. It is not a statement about the corpus's 5,859 papers, about the 177 papers that classify something with an LLM, or about any other page's population. |
| | |
| | **Scope of this edit.** §12.9's output block, the tier table and the paragraph under it, the NUL-bug paragraph's footnote, and the content page's quote bullet. ''report_llm_currency.mjs'' and ''report_website_classification.mjs'' were **not** re-run in this pass — the 2026-09-13 significance figures, the venue-composition control, PART B's 49 figure checks and every citation stand as published. |
| | |
| | ^ Item ^ Value ^ |
| | | Date | 2026-09-21, unsupervised | |
| | | Command | ''%%node scripts/llm_currency_quotecheck.mjs > scripts/llm_currency_quotecheck-output.txt%%'' | |
| | | Artifacts | ''scripts/llm_currency_quotecheck-output.txt'' (regenerated; the 2026-09-03 copy kept as ''…-output.txt.bak0921'') | |
| | | Script changes | none | |
| | | Reviewers | one ''sonnet'' figures-vs-script pass over this page and [[:design:website_classification]] | |
| | | Pages saved | this page, [[:design:website_classification]] | |
| | | Not edited | [[:provenance:privacy:javascript]] mentions this checker but publishes none of its counts | |
| |