User Tools

Site Tools


provenance:design:website_classification

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
provenance:design:website_classification [2026/09/21 13:45] – Quote-check refresh 2026-09-21: regenerated llm_currency_quotecheck output block; 6 FAIL -> 6 rescued from the PDF, 0 fail in both renderings. Authored by Claude karel.kubicek.claudeprovenance:design:website_classification [2026/09/21 14:40] (current) – Review follow-ups: 12.8's 'task item exists' bullet closed by 12.14, and 12.14's page-sweep claim corrected (provenance:statistics:annotation carries the raw string, not a fold row). Authored by Claude karel.kubicek.claude
Line 285: Line 285:
 ^ Family ^ Papers (of 177) ^ ^ Family ^ Papers (of 177) ^
 | GPT-4o | 46 (26.0%) | | GPT-4o | 46 (26.0%) |
-| GPT-4 (non-4o) | 40 (22.6%) |+| GPT-4 (non-4o) | 44 (24.9%) |
 | GPT-3.5 / GPT-3 / ChatGPT | 29 (16.4%) | | GPT-3.5 / GPT-3 / ChatGPT | 29 (16.4%) |
-| **UNMAPPED** | **23 (13.0%)** |+| **UNMAPPED** | **18 (10.2%)** |
 | Llama family | 13 (7.3%) | | Llama family | 13 (7.3%) |
 | Gemini / PaLM | 12 (6.8%) | | Gemini / PaLM | 12 (6.8%) |
 | Unnamed LLM | 11 (6.2%) | | Unnamed LLM | 11 (6.2%) |
-| OpenAI reasoning / GPT-5 tier | 9 (5.1%) |+| OpenAI reasoning / GPT-5 tier | 10 (5.6%) |
 | Qwen family | 8 (4.5%) | | Qwen family | 8 (4.5%) |
 | DeepSeek family | 7 (4.0%) | | DeepSeek family | 7 (4.0%) |
Line 299: Line 299:
 | Encoder / seq2seq LM (not a chat LLM) | 2 (1.1%) | | Encoder / seq2seq LM (not a chat LLM) | 2 (1.1%) |
  
-The 25-string unmapped residue is printed in full by the script (§12.9). **Four of those strings are a fold failure, not a long tail:** ''GPT 4.1'', ''GPT-4.0'', ''GPT-4.5'' and ''GPT-o1'' are OpenAI models the ''GPT-4 (non-4o)'' regex misses on a space or a decimal. They were left visible rather than folded, because patching a fold to absorb its own residue after seeing the output stops it being a documented rule; a task item exists to fold them and re-derive the three GPT rows. ''Grok-3'', ''GLM-4.5'', ''ChatGLM'' and ''Kimi'' are the genuine long tail — four vendors with no family in the list.+**The table above is the 2026-09-21 re-derivation.** Until then it read GPT-4 (non-4o) 40 (22.6%), UNMAPPED 23 (13.0%) and OpenAI reasoning 9 (5.1%), because four residue strings — ''GPT 4.1'', ''GPT-4.0'', ''GPT-4.5'' and ''GPT-o1'' — are OpenAI models that the fold missed on a space or a decimal. They were left visible on 2026-09-03 rather than quietly folded, because patching a fold to absorb its own residue after seeing the output stops it being a documented rule. The fold has now been extended **as a rule about how OpenAI model strings are spelled**, not as a list of those four: see §12.14 for the rule, the two further strings it moves, and the full diff. ''Grok-3'', ''GLM-4.5'', ''ChatGLM'' and ''Kimi'' stay in the now 20-string residue — they are the genuine long tail, four vendors with no family in the list, and no family was added for them.
  
 === 5b. The reproducibility buckets, which had to be rebuilt === === 5b. The reproducibility buckets, which had to be rebuilt ===
Line 372: Line 372:
  
 **The probe's width decided that answer, and the first width was wrong.** It used ''\bpage\b'', which does not match the compound "webpage" — there is no word boundary between "web" and "page" — and it had no bare ''web'' at all, so it returned 16 rather than 20 and silently dropped //"relevant person-specific webpage information"// and //"IOB presence and trustworthiness in web content"//. Found in re-review (§12.13, finding 3). Both were then read and neither changes the conclusion, which is the only reason the published claim survived a probe that was under-recalling by 20%. A probe is not a read of 116 papers, its hits must be read rather than counted, and its regex is a load-bearing part of the claim. **The probe's width decided that answer, and the first width was wrong.** It used ''\bpage\b'', which does not match the compound "webpage" — there is no word boundary between "web" and "page" — and it had no bare ''web'' at all, so it returned 16 rather than 20 and silently dropped //"relevant person-specific webpage information"// and //"IOB presence and trustworthiness in web content"//. Found in re-review (§12.13, finding 3). Both were then read and neither changes the conclusion, which is the only reason the published claim survived a probe that was under-recalling by 20%. A probe is not a read of 116 papers, its hits must be read rather than counted, and its regex is a load-bearing part of the claim.
-  * **Whether the four GPT strings in the fold residue change a published share.** §12.5a. Not folded on purpose; a task item exists.+  * **Whether the four GPT strings in the fold residue change a published share.** §12.5a. **Settled on 2026-09-21 and no longer open** — the fold was extended as a rule and the three GPT rows re-derived; they do change one published share on this page and none anywhere else. See §12.14.
   * **Whether the six quote failures are extraction paraphrase or ''.cols'' rendering.** One was read and was rendering. The other five were not, because no page quotes them.   * **Whether the six quote failures are extraction paraphrase or ''.cols'' rendering.** One was read and was rendering. The other five were not, because no page quotes them.
  
Line 557: Line 557:
 -------------------------------------  ---------------  ----- -------------------------------------  ---------------  -----
 GPT-4o                                 46               26.0% GPT-4o                                 46               26.0%
-GPT-4 (non-4o)                         40               22.6%+GPT-4 (non-4o)                         44               24.9%
 GPT-3.5 / GPT-3 / ChatGPT              29               16.4% GPT-3.5 / GPT-3 / ChatGPT              29               16.4%
-UNMAPPED                               23               13.0%+UNMAPPED                               18               10.2%
 Llama family                           13               7.3% Llama family                           13               7.3%
 Gemini / PaLM                          12               6.8% Gemini / PaLM                          12               6.8%
 Unnamed LLM                            11               6.2% Unnamed LLM                            11               6.2%
-OpenAI reasoning / GPT-5 tier          9                5.1%+OpenAI reasoning / GPT-5 tier          10               5.6%
 Qwen family                            8                4.5% Qwen family                            8                4.5%
 DeepSeek family                        7                4.0% DeepSeek family                        7                4.0%
Line 571: Line 571:
 Encoder / seq2seq LM (not a chat LLM)  2                1.1% Encoder / seq2seq LM (not a chat LLM)  2                1.1%
  
---- UNMAPPED residue: 25 distinct strings, printed in full ---+--- UNMAPPED residue: 20 distinct strings, printed in full ---
       2x  Grok-3       2x  Grok-3
       2x  local LLMs (custom prompts)       2x  local LLMs (custom prompts)
       1x  BLIP2       1x  BLIP2
-      1x  Chat-GPT 3.5 and 4 
       1x  ChatGLM       1x  ChatGLM
       1x  custom structured prompts with fine-tuned LLMs       1x  custom structured prompts with fine-tuned LLMs
       1x  foundation LLMs       1x  foundation LLMs
       1x  GLM-4.5       1x  GLM-4.5
-      1x  GPT 4.1 
-      1x  GPT-4.0 
-      1x  GPT-4.5 
-      1x  GPT-o1 
       1x  HtmlLLM-Detector       1x  HtmlLLM-Detector
       1x  Kimi       1x  Kimi
Line 766: Line 761:
 157 of 157 (100.0%) state a targetDetail. 157 of 157 (100.0%) state a targetDetail.
  
---- targetDetail matching the probe (websit|domain|url|homepage|page|web|script|tracker|cookie|sdk|first.part|third.part|categor) — 20 tuples, all printed ---+--- targetDetail matching the probe (websit|domain|\burl\b|homepage|page|\bweb\b|script|tracker|cookie|\bsdk\b|first.part|third.part|categor) — 20 tuples, all printed ---
     IMC/2023/in-the-room-where-it-happens-characterizing-local-communication-and-threats-in-s     IMC/2023/in-the-room-where-it-happens-characterizing-local-communication-and-threats-in-s
         IoT device vendors and categories         IoT device vendors and categories
Line 1571: Line 1566:
  
 [[design:website_classification|← back to the content page]] · [[literature:corpus|corpus-level provenance]] [[design:website_classification|← back to the content page]] · [[literature:corpus|corpus-level provenance]]
 +
 +==== 12.14 The model-family fold extended, 2026-09-21 ====
 +
 +**What was wrong.** §12.5a's residue contained four strings that are OpenAI models the ''GPT-4 (non-4o)'' and reasoning-tier regexes should have absorbed: ''GPT 4.1'' (a space instead of a hyphen), ''GPT-4.0'' and ''GPT-4.5'' (a decimal, which the ''%%(?![.\do])%%'' lookahead rejected outright although it was written only to keep ''4o'' out), and ''GPT-o1'' (the o-series, which the fold matched only in its bare ''o1-mini'' form). One paper each. They were **deliberately left in the residue on 2026-09-03** rather than patched after the output was read.
 +
 +**Why it was left, and what changed.** Patching a fold to absorb the residue you have just looked at is how a documented rule stops being one: the next reader cannot tell a rule from a list of the strings that embarrassed it. So the fold has been extended as a **rule about how OpenAI writes model names** — the separator after ''GPT'' may be a hyphen, a space or nothing; a version may carry a decimal; the o-series is written both bare and ''GPT''-prefixed — and the rule is stated in the script, above the table it feeds:
 +
 +<code>
 +  [/gpt[- ]?4o|gpt4o/i, 'GPT-4o'],
 +  [/gpt[- ]?4-turbo|gpt[- ]?4(\.\d+)?(?![\do])/i, 'GPT-4 (non-4o)'],
 +  [/gpt[- ]?3\.5|chat-?gpt|text-davinci|gpt[- ]?3(?!\.5)/i, 'GPT-3.5 / GPT-3 / ChatGPT'],
 +  [/gpt[- ]?5|gpt[- ]?o[1345]\b|\bo[134]-(mini|preview|pro)\b|\bo4-mini\b/i, 'OpenAI reasoning / GPT-5 tier'],
 +</code>
 +
 +**The rule moves two strings that were never in the residue, and that is the point.** Applied to all 154 distinct ''resourceName'' strings the fold sees, it changes six: the four above, plus ''Chat-GPT 3.5 and 4'' (hyphenated "Chat-GPT", so ''chatgpt'' never matched it — UNMAPPED → GPT-3.5) and ''ChatGPT-4.0'' (which the old decimal lookahead pushed past the GPT-4 row into the ChatGPT row — GPT-3.5 → GPT-4). A rule fitted to the residue would have moved exactly four. Both were verified string by string before the fold was changed, not after.
 +
 +^ Row ^ Was ^ Is ^ Which papers moved ^
 +| GPT-4o | 46 (26.0%) | 46 (26.0%) | none — the ''GPT-o1'' paper is IMC/2025 //an-in-depth-investigation-of-data-collection…//, which already counted here for its ''GPT-4o'' string |
 +| **GPT-4 (non-4o)** | **40 (22.6%)** | **44 (24.9%)** | +4: ''GPT 4.1'' (PETS/2026), ''GPT-4.5'' (CCS/2025), ''GPT-4.0'' (NDSS/2025), ''ChatGPT-4.0'' (USENIX/2024) |
 +| GPT-3.5 / GPT-3 / ChatGPT | 29 (16.4%) | 29 (16.4%) | **net zero, not "unchanged"**: the ''ChatGPT-4.0'' paper leaves, the ''Chat-GPT 3.5 and 4'' paper (NDSS/2025) arrives |
 +| **OpenAI reasoning / GPT-5 tier** | **9 (5.1%)** | **10 (5.6%)** | +1: the ''GPT-o1'' paper |
 +| **UNMAPPED** | **23 (13.0%)** | **18 (10.2%)** | −5 papers; the residue falls from **25 distinct strings to 20** |
 +
 +**The three-bucket reproducibility table (§12.5b) does not move**, and it was checked rather than assumed: ''A 10 / B 24 / C 131 / D 10'' before and after, and every bucket's printed string list is byte-identical. The buckets are built from ''HOSTED_SNAPSHOT'', ''OPEN_FAMILY''/''PARAM_SIZE'' and ''NAMED'', none of which the fold touches; ''GPT-o1'' was already in bucket C, because ''NAMED'' contains a bare ''gpt''. The fold and the buckets answer different questions and are deliberately separate regexes.
 +
 +**What stays in the residue, and why.** ''Grok-3'' (2 papers), ''GLM-4.5'', ''ChatGLM'' and ''Kimi'' are four vendors with no family in the list. They are a genuine long tail, not a fold failure, and **no family was added for them** — adding one would be the post-hoc patch this section exists to avoid. The remaining 16 strings are descriptions rather than models (''local LLMs (custom prompts)'', ''weighted multi-model ensemble (custom)''), systems built on a model (''PhishLLM'', ''UGCG-GUARD'', ''YouthSafe'', ''RFCGPT''), or non-OpenAI multimodal models (''LLaVA'', ''BLIP2'', ''text-bison''). All 20 are printed in full in §12.9.
 +
 +**Where this is published.** Only here. [[:design:website_classification]] publishes the §12.5b bucket table, not the family fold, so the content page needed no edit for this. Checked by grepping the raw source of all **189** pages on the wiki: ''GPT-4 (non-4o)'' and the other family labels return **this page alone**, and the four residue strings return this page and [[:provenance:statistics:annotation]] — where ''GPT 4.1'' appears as a raw ''resourceName'' inside a per-paper quote-check listing, not as a folded row, and so is unaffected by the fold. Same for ''ChatGLM'' on that page.
 +
 +<code>
 +$ node scripts/report_llm_currency.mjs > scripts/report_llm_currency-output.txt
 +$ diff <old> <new>
 +178c178
 +< GPT-4 (non-4o)                         40               22.6%
 +> GPT-4 (non-4o)                         44               24.9%
 +180c180
 +< UNMAPPED                               23               13.0%
 +> UNMAPPED                               18               10.2%
 +184c184
 +< OpenAI reasoning / GPT-5 tier          9                5.1%
 +> OpenAI reasoning / GPT-5 tier          10               5.6%
 +192c192
 +< --- UNMAPPED residue: 25 distinct strings, printed in full ---
 +> --- UNMAPPED residue: 20 distinct strings, printed in full ---
 +196d195  (Chat-GPT 3.5 and 4)   201,204d199  (GPT 4.1 / GPT-4.0 / GPT-4.5 / GPT-o1)
 +</code>
 +
 +**Those are the only lines that changed in a 446-line report.** Every population, year, venue, target, validation and bucket figure is identical, and the ''compared''-only assertion and the bucket-sum assertion both still pass.
 +
 +**One unrelated repair in the same save.** The §12.9 block is regenerated from the committed output file, and the published copy had lost the ''%%\b%%'' escapes from the ''targetDetail'' probe's printed regex (it read ''url|…|web|…'' where the script prints ''%%\burl\b|…|\bweb\b|…%%''). The block now matches the file byte for byte. The probe itself never changed; only the copy on this page was wrong, and it is the width of that probe that decides the claim in §12.7.
 +
 +^ Item ^ Value ^
 +| Date | 2026-09-21, unsupervised |
 +| Script changes | ''scripts/report_llm_currency.mjs'' — the four OpenAI ''FAMILY'' rows, with the rule stated in a comment above them; committed output regenerated |
 +| Reviewers | one ''sonnet'' figures-vs-script pass; one ''sonnet'' citations/quotes pass |
 +| Pages saved | this page only |
 +| Not edited | [[:design:website_classification]], [[:design:ip_classification]], [[:privacy:javascript]], [[:privacy:cookies]] — none publishes a model-family row |
  
 ===== Markup sweep, 2026-09-17 ===== ===== Markup sweep, 2026-09-17 =====
provenance/design/website_classification.1789998321.txt.gz · Last modified: by karel.kubicek.claude