| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| provenance:design:website_classification [2026/09/21 13:45] – Quote-check refresh 2026-09-21: regenerated llm_currency_quotecheck output block; 6 FAIL -> 6 rescued from the PDF, 0 fail in both renderings. Authored by Claude karel.kubicek.claude | provenance:design:website_classification [2026/09/21 14:40] (current) – Review follow-ups: 12.8's 'task item exists' bullet closed by 12.14, and 12.14's page-sweep claim corrected (provenance:statistics:annotation carries the raw string, not a fold row). Authored by Claude karel.kubicek.claude |
|---|
| ^ Family ^ Papers (of 177) ^ | ^ Family ^ Papers (of 177) ^ |
| | GPT-4o | 46 (26.0%) | | | GPT-4o | 46 (26.0%) | |
| | GPT-4 (non-4o) | 40 (22.6%) | | | GPT-4 (non-4o) | 44 (24.9%) | |
| | GPT-3.5 / GPT-3 / ChatGPT | 29 (16.4%) | | | GPT-3.5 / GPT-3 / ChatGPT | 29 (16.4%) | |
| | **UNMAPPED** | **23 (13.0%)** | | | **UNMAPPED** | **18 (10.2%)** | |
| | Llama family | 13 (7.3%) | | | Llama family | 13 (7.3%) | |
| | Gemini / PaLM | 12 (6.8%) | | | Gemini / PaLM | 12 (6.8%) | |
| | Unnamed LLM | 11 (6.2%) | | | Unnamed LLM | 11 (6.2%) | |
| | OpenAI reasoning / GPT-5 tier | 9 (5.1%) | | | OpenAI reasoning / GPT-5 tier | 10 (5.6%) | |
| | Qwen family | 8 (4.5%) | | | Qwen family | 8 (4.5%) | |
| | DeepSeek family | 7 (4.0%) | | | DeepSeek family | 7 (4.0%) | |
| | Encoder / seq2seq LM (not a chat LLM) | 2 (1.1%) | | | Encoder / seq2seq LM (not a chat LLM) | 2 (1.1%) | |
| |
| The 25-string unmapped residue is printed in full by the script (§12.9). **Four of those strings are a fold failure, not a long tail:** ''GPT 4.1'', ''GPT-4.0'', ''GPT-4.5'' and ''GPT-o1'' are OpenAI models the ''GPT-4 (non-4o)'' regex misses on a space or a decimal. They were left visible rather than folded, because patching a fold to absorb its own residue after seeing the output stops it being a documented rule; a task item exists to fold them and re-derive the three GPT rows. ''Grok-3'', ''GLM-4.5'', ''ChatGLM'' and ''Kimi'' are the genuine long tail — four vendors with no family in the list. | **The table above is the 2026-09-21 re-derivation.** Until then it read GPT-4 (non-4o) 40 (22.6%), UNMAPPED 23 (13.0%) and OpenAI reasoning 9 (5.1%), because four residue strings — ''GPT 4.1'', ''GPT-4.0'', ''GPT-4.5'' and ''GPT-o1'' — are OpenAI models that the fold missed on a space or a decimal. They were left visible on 2026-09-03 rather than quietly folded, because patching a fold to absorb its own residue after seeing the output stops it being a documented rule. The fold has now been extended **as a rule about how OpenAI model strings are spelled**, not as a list of those four: see §12.14 for the rule, the two further strings it moves, and the full diff. ''Grok-3'', ''GLM-4.5'', ''ChatGLM'' and ''Kimi'' stay in the now 20-string residue — they are the genuine long tail, four vendors with no family in the list, and no family was added for them. |
| |
| === 5b. The reproducibility buckets, which had to be rebuilt === | === 5b. The reproducibility buckets, which had to be rebuilt === |
| |
| **The probe's width decided that answer, and the first width was wrong.** It used ''\bpage\b'', which does not match the compound "webpage" — there is no word boundary between "web" and "page" — and it had no bare ''web'' at all, so it returned 16 rather than 20 and silently dropped //"relevant person-specific webpage information"// and //"IOB presence and trustworthiness in web content"//. Found in re-review (§12.13, finding 3). Both were then read and neither changes the conclusion, which is the only reason the published claim survived a probe that was under-recalling by 20%. A probe is not a read of 116 papers, its hits must be read rather than counted, and its regex is a load-bearing part of the claim. | **The probe's width decided that answer, and the first width was wrong.** It used ''\bpage\b'', which does not match the compound "webpage" — there is no word boundary between "web" and "page" — and it had no bare ''web'' at all, so it returned 16 rather than 20 and silently dropped //"relevant person-specific webpage information"// and //"IOB presence and trustworthiness in web content"//. Found in re-review (§12.13, finding 3). Both were then read and neither changes the conclusion, which is the only reason the published claim survived a probe that was under-recalling by 20%. A probe is not a read of 116 papers, its hits must be read rather than counted, and its regex is a load-bearing part of the claim. |
| * **Whether the four GPT strings in the fold residue change a published share.** §12.5a. Not folded on purpose; a task item exists. | * **Whether the four GPT strings in the fold residue change a published share.** §12.5a. **Settled on 2026-09-21 and no longer open** — the fold was extended as a rule and the three GPT rows re-derived; they do change one published share on this page and none anywhere else. See §12.14. |
| * **Whether the six quote failures are extraction paraphrase or ''.cols'' rendering.** One was read and was rendering. The other five were not, because no page quotes them. | * **Whether the six quote failures are extraction paraphrase or ''.cols'' rendering.** One was read and was rendering. The other five were not, because no page quotes them. |
| |
| ------------------------------------- --------------- ----- | ------------------------------------- --------------- ----- |
| GPT-4o 46 26.0% | GPT-4o 46 26.0% |
| GPT-4 (non-4o) 40 22.6% | GPT-4 (non-4o) 44 24.9% |
| GPT-3.5 / GPT-3 / ChatGPT 29 16.4% | GPT-3.5 / GPT-3 / ChatGPT 29 16.4% |
| UNMAPPED 23 13.0% | UNMAPPED 18 10.2% |
| Llama family 13 7.3% | Llama family 13 7.3% |
| Gemini / PaLM 12 6.8% | Gemini / PaLM 12 6.8% |
| Unnamed LLM 11 6.2% | Unnamed LLM 11 6.2% |
| OpenAI reasoning / GPT-5 tier 9 5.1% | OpenAI reasoning / GPT-5 tier 10 5.6% |
| Qwen family 8 4.5% | Qwen family 8 4.5% |
| DeepSeek family 7 4.0% | DeepSeek family 7 4.0% |
| Encoder / seq2seq LM (not a chat LLM) 2 1.1% | Encoder / seq2seq LM (not a chat LLM) 2 1.1% |
| |
| --- UNMAPPED residue: 25 distinct strings, printed in full --- | --- UNMAPPED residue: 20 distinct strings, printed in full --- |
| 2x Grok-3 | 2x Grok-3 |
| 2x local LLMs (custom prompts) | 2x local LLMs (custom prompts) |
| 1x BLIP2 | 1x BLIP2 |
| 1x Chat-GPT 3.5 and 4 | |
| 1x ChatGLM | 1x ChatGLM |
| 1x custom structured prompts with fine-tuned LLMs | 1x custom structured prompts with fine-tuned LLMs |
| 1x foundation LLMs | 1x foundation LLMs |
| 1x GLM-4.5 | 1x GLM-4.5 |
| 1x GPT 4.1 | |
| 1x GPT-4.0 | |
| 1x GPT-4.5 | |
| 1x GPT-o1 | |
| 1x HtmlLLM-Detector | 1x HtmlLLM-Detector |
| 1x Kimi | 1x Kimi |
| 157 of 157 (100.0%) state a targetDetail. | 157 of 157 (100.0%) state a targetDetail. |
| |
| --- targetDetail matching the probe (websit|domain|url|homepage|page|web|script|tracker|cookie|sdk|first.part|third.part|categor) — 20 tuples, all printed --- | --- targetDetail matching the probe (websit|domain|\burl\b|homepage|page|\bweb\b|script|tracker|cookie|\bsdk\b|first.part|third.part|categor) — 20 tuples, all printed --- |
| IMC/2023/in-the-room-where-it-happens-characterizing-local-communication-and-threats-in-s | IMC/2023/in-the-room-where-it-happens-characterizing-local-communication-and-threats-in-s |
| IoT device vendors and categories | IoT device vendors and categories |
| |
| [[design:website_classification|← back to the content page]] · [[literature:corpus|corpus-level provenance]] | [[design:website_classification|← back to the content page]] · [[literature:corpus|corpus-level provenance]] |
| | |
| | ==== 12.14 The model-family fold extended, 2026-09-21 ==== |
| | |
| | **What was wrong.** §12.5a's residue contained four strings that are OpenAI models the ''GPT-4 (non-4o)'' and reasoning-tier regexes should have absorbed: ''GPT 4.1'' (a space instead of a hyphen), ''GPT-4.0'' and ''GPT-4.5'' (a decimal, which the ''%%(?![.\do])%%'' lookahead rejected outright although it was written only to keep ''4o'' out), and ''GPT-o1'' (the o-series, which the fold matched only in its bare ''o1-mini'' form). One paper each. They were **deliberately left in the residue on 2026-09-03** rather than patched after the output was read. |
| | |
| | **Why it was left, and what changed.** Patching a fold to absorb the residue you have just looked at is how a documented rule stops being one: the next reader cannot tell a rule from a list of the strings that embarrassed it. So the fold has been extended as a **rule about how OpenAI writes model names** — the separator after ''GPT'' may be a hyphen, a space or nothing; a version may carry a decimal; the o-series is written both bare and ''GPT''-prefixed — and the rule is stated in the script, above the table it feeds: |
| | |
| | <code> |
| | [/gpt[- ]?4o|gpt4o/i, 'GPT-4o'], |
| | [/gpt[- ]?4-turbo|gpt[- ]?4(\.\d+)?(?![\do])/i, 'GPT-4 (non-4o)'], |
| | [/gpt[- ]?3\.5|chat-?gpt|text-davinci|gpt[- ]?3(?!\.5)/i, 'GPT-3.5 / GPT-3 / ChatGPT'], |
| | [/gpt[- ]?5|gpt[- ]?o[1345]\b|\bo[134]-(mini|preview|pro)\b|\bo4-mini\b/i, 'OpenAI reasoning / GPT-5 tier'], |
| | </code> |
| | |
| | **The rule moves two strings that were never in the residue, and that is the point.** Applied to all 154 distinct ''resourceName'' strings the fold sees, it changes six: the four above, plus ''Chat-GPT 3.5 and 4'' (hyphenated "Chat-GPT", so ''chatgpt'' never matched it — UNMAPPED → GPT-3.5) and ''ChatGPT-4.0'' (which the old decimal lookahead pushed past the GPT-4 row into the ChatGPT row — GPT-3.5 → GPT-4). A rule fitted to the residue would have moved exactly four. Both were verified string by string before the fold was changed, not after. |
| | |
| | ^ Row ^ Was ^ Is ^ Which papers moved ^ |
| | | GPT-4o | 46 (26.0%) | 46 (26.0%) | none — the ''GPT-o1'' paper is IMC/2025 //an-in-depth-investigation-of-data-collection…//, which already counted here for its ''GPT-4o'' string | |
| | | **GPT-4 (non-4o)** | **40 (22.6%)** | **44 (24.9%)** | +4: ''GPT 4.1'' (PETS/2026), ''GPT-4.5'' (CCS/2025), ''GPT-4.0'' (NDSS/2025), ''ChatGPT-4.0'' (USENIX/2024) | |
| | | GPT-3.5 / GPT-3 / ChatGPT | 29 (16.4%) | 29 (16.4%) | **net zero, not "unchanged"**: the ''ChatGPT-4.0'' paper leaves, the ''Chat-GPT 3.5 and 4'' paper (NDSS/2025) arrives | |
| | | **OpenAI reasoning / GPT-5 tier** | **9 (5.1%)** | **10 (5.6%)** | +1: the ''GPT-o1'' paper | |
| | | **UNMAPPED** | **23 (13.0%)** | **18 (10.2%)** | −5 papers; the residue falls from **25 distinct strings to 20** | |
| | |
| | **The three-bucket reproducibility table (§12.5b) does not move**, and it was checked rather than assumed: ''A 10 / B 24 / C 131 / D 10'' before and after, and every bucket's printed string list is byte-identical. The buckets are built from ''HOSTED_SNAPSHOT'', ''OPEN_FAMILY''/''PARAM_SIZE'' and ''NAMED'', none of which the fold touches; ''GPT-o1'' was already in bucket C, because ''NAMED'' contains a bare ''gpt''. The fold and the buckets answer different questions and are deliberately separate regexes. |
| | |
| | **What stays in the residue, and why.** ''Grok-3'' (2 papers), ''GLM-4.5'', ''ChatGLM'' and ''Kimi'' are four vendors with no family in the list. They are a genuine long tail, not a fold failure, and **no family was added for them** — adding one would be the post-hoc patch this section exists to avoid. The remaining 16 strings are descriptions rather than models (''local LLMs (custom prompts)'', ''weighted multi-model ensemble (custom)''), systems built on a model (''PhishLLM'', ''UGCG-GUARD'', ''YouthSafe'', ''RFCGPT''), or non-OpenAI multimodal models (''LLaVA'', ''BLIP2'', ''text-bison''). All 20 are printed in full in §12.9. |
| | |
| | **Where this is published.** Only here. [[:design:website_classification]] publishes the §12.5b bucket table, not the family fold, so the content page needed no edit for this. Checked by grepping the raw source of all **189** pages on the wiki: ''GPT-4 (non-4o)'' and the other family labels return **this page alone**, and the four residue strings return this page and [[:provenance:statistics:annotation]] — where ''GPT 4.1'' appears as a raw ''resourceName'' inside a per-paper quote-check listing, not as a folded row, and so is unaffected by the fold. Same for ''ChatGLM'' on that page. |
| | |
| | <code> |
| | $ node scripts/report_llm_currency.mjs > scripts/report_llm_currency-output.txt |
| | $ diff <old> <new> |
| | 178c178 |
| | < GPT-4 (non-4o) 40 22.6% |
| | > GPT-4 (non-4o) 44 24.9% |
| | 180c180 |
| | < UNMAPPED 23 13.0% |
| | > UNMAPPED 18 10.2% |
| | 184c184 |
| | < OpenAI reasoning / GPT-5 tier 9 5.1% |
| | > OpenAI reasoning / GPT-5 tier 10 5.6% |
| | 192c192 |
| | < --- UNMAPPED residue: 25 distinct strings, printed in full --- |
| | > --- UNMAPPED residue: 20 distinct strings, printed in full --- |
| | 196d195 (Chat-GPT 3.5 and 4) 201,204d199 (GPT 4.1 / GPT-4.0 / GPT-4.5 / GPT-o1) |
| | </code> |
| | |
| | **Those are the only lines that changed in a 446-line report.** Every population, year, venue, target, validation and bucket figure is identical, and the ''compared''-only assertion and the bucket-sum assertion both still pass. |
| | |
| | **One unrelated repair in the same save.** The §12.9 block is regenerated from the committed output file, and the published copy had lost the ''%%\b%%'' escapes from the ''targetDetail'' probe's printed regex (it read ''url|…|web|…'' where the script prints ''%%\burl\b|…|\bweb\b|…%%''). The block now matches the file byte for byte. The probe itself never changed; only the copy on this page was wrong, and it is the width of that probe that decides the claim in §12.7. |
| | |
| | ^ Item ^ Value ^ |
| | | Date | 2026-09-21, unsupervised | |
| | | Script changes | ''scripts/report_llm_currency.mjs'' — the four OpenAI ''FAMILY'' rows, with the rule stated in a comment above them; committed output regenerated | |
| | | Reviewers | one ''sonnet'' figures-vs-script pass; one ''sonnet'' citations/quotes pass | |
| | | Pages saved | this page only | |
| | | Not edited | [[:design:website_classification]], [[:design:ip_classification]], [[:privacy:javascript]], [[:privacy:cookies]] — none publishes a model-family row | |
| |
| ===== Markup sweep, 2026-09-17 ===== | ===== Markup sweep, 2026-09-17 ===== |