User Tools

Site Tools


privacy:javascript

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
privacy:javascript [2026/08/12 10:05] – Review pass (Fable): fix stale figures that sat OUTSIDE the corpus section and so were missed by the windowed staleness guard. The embedded js_fold.mjs docstring still quoted the 4,322-corpus counts (959/1,063, LLVM 66, Esprima 15) two screens below the t karel.kubicek.claudeprivacy:javascript [2026/09/03 22:34] (current) – Re-review fix: the resolvable-model figure moved to 19.4% of 175 after the buckets were restricted to papers that used an LLM rather than compared against one. Authored by Claude karel.kubicek.claude
Line 147: Line 147:
 ==== What the corpus cannot tell you, and what we found instead ==== ==== What the corpus cannot tell you, and what we found instead ====
  
-**LLM-based classification of web scripts is, as of 2026-08-06essentially absent from the peer-reviewed literature.** A targeted search across PETS 2025/2026, USENIX Security 2025, NDSS 2025/2026, IMC 2025, TheWebConf 2025/2026, CCS 2025 and arXiv found no paper that classifies web scripts as trackers with a language model, or that uses one to summarise a script's privacy-relevant behaviour. The nearest work is adjacent rather than on-point: LLM-aided **deobfuscation** feeding a graph classifier for JavaScript //malware//,((//Breaking Obfuscation: Cluster-Aware Graph with LLM-Aided Recovery for Malicious JavaScript Detection//, [[https://arxiv.org/abs/2507.22447|arXiv:2507.22447]], 2025.)) LLM screening of malicious npm packages, and ''humanify'', which uses a model only to //suggest identifier names// during de-minification.(([[https://github.com/jehna/humanify|github.com/jehna/humanify]], v3.1.1, checked 2026-08-06. The AST rewrite is done by ''oxc''; the model only proposes names.))+**LLM-based classification of web scripts is, as of 2026-09-03still absent from the peer-reviewed literature in these seven venues — and this is now a measured zero rather than the result of a keyword search.** ''classification.method == "llm"'' fires on **177 of the 5,859 corpus papers**, but on **0 of the 44 papers that classify a ''javascript'' target and 0 of the 31 that classify a ''fingerprinting-script'' target**. A targeted search outside the corpus — PETS 2025/2026, USENIX Security 2025/2026, NDSS 2025/2026, IMC 2025, TheWebConf 2025/2026, CCS 2025 and arXiv — found no paper that classifies web scripts as trackers with a language model, or that uses one to summarise a script's privacy-relevant behaviour.
  
-<wrap todo>Treat this as an opportunity, not a settled answer. If you are planning an LLM-based script classifier, you are not late — but you also have no baseline to cite, so budget for building one, and for the reviewer question about cost, reproducibility and prompt/version drift that this page cannot yet answer for you.</wrap>+**Read that against the corpus-wide curve before concluding the field is not interested.** LLM classification went from 2 papers in 2023 to 71 in the partial 2026 (17.1% of that year), at all seven venues, so the zero above is specific to this target rather than a statement about the method. The per-target table on [[design:website_classification#Where LLMs actually appear]] gives the full ranking; ''javascript'' and ''fingerprinting-script'' are **zero rows** in it, alongside ''malware'', ''sdk-or-library'' and ''website-popularity''. The lowest //non-zero// rows are ''ip-address'' (1 of 295) and ''web-request'' (1 of 258). 
 + 
 +**The nearest peer-reviewed work is one layer down, at the request.** TGNN {[xiong2026_tgnn]} (TheWebConf 2026) uses Qwen3 to label HTTP request/response quadruples as tracking or not, and reports an annotation F1 of **98.17%** against expert labels where filter lists reach **55.14%** on the same ground truth.((The paper reports this figure twice as **98.17%** — in §4.1.3 //LLM-based Labeling// ("𝑀𝜆 performs well (𝐹1-score of 98.17%)") and again beside Figure 5 — and once as **98.19%**, in the contributions paragraph of its Introduction. Its abstract states no F1 for the annotation component at all. The discrepancy is the paper's, not ours; the body figure is quoted here and on [[privacy:requests]], so the two pages agree. The 55.14% filter-list comparison appears only in the Introduction. Located in ''paper.cols.txt'' on 2026-09-03; note that this file contains NUL bytes, so shell ''grep'' reports nothing without ''-a''.)) It is not a script classifier — it explicitly contrasts its approach with prior methods that do "single-domain analysis-such as string matching against domain lists or scrutinizing JavaScript execution within a page" — but it is the closest thing to a citable precedent for using a model to //manufacture tracker labels//, which is this page's weakest link. [[privacy:requests]] treats it in full. Beyond it the adjacent work is not peer-reviewed and not on-point: LLM-aided **deobfuscation** feeding a graph classifier for JavaScript //malware//,((//Breaking Obfuscation: Cluster-Aware Graph with LLM-Aided Recovery for Malicious JavaScript Detection//, [[https://arxiv.org/abs/2507.22447|arXiv:2507.22447]], 2025.)) LLM screening of malicious npm packages, and ''humanify'', which uses a model only to //suggest identifier names// during de-minification.(([[https://github.com/jehna/humanify|github.com/jehna/humanify]], v3.1.1, checked 2026-08-06. The AST rewrite is done by ''oxc''; the model only proposes names.)) 
 + 
 +<WRAP todo>Treat this as an opportunity, not a settled answer. If you are planning an LLM-based script classifier, you are not late — but you also have no baseline to cite, so budget for building one, and for the reviewer question about cost, reproducibility and prompt/version drift that this page cannot yet answer for you. Two things you can borrow rather than invent: TGNN's annotation-versus-filter-list comparison {[xiong2026_tgnn]} is the experimental design a reviewer will expect, and the model-reporting figures on [[design:website_classification#And almost nobody names a model you could resolve]] show that only **19.4%** of the 175 corpus papers that actually use an LLM name a model resolvable to an actual artefact — so naming yours to the checkpoint is cheap novelty.</WRAP>
  
 ==== Two 2025 results that change how you design a crawl ==== ==== Two 2025 results that change how you design a crawl ====
Line 175: Line 179:
 | **llm** | **2** | **1.0%** | | **llm** | **2** | **1.0%** |
  
-The ''llm'' row is new: on the 4,322-paper corpus this enum never fired for a JavaScript-classification task at allTwo papers is not a trend, and it does not contradict the finding below that no peer-reviewed paper yet classifies web scripts //as trackers// with a language model.+**The ''llm'' row does not mean what it looks like it means, and it is worth being precise because two other pages depend on the same field.** The enum is per //paper//, not per JavaScript task: these are two papers in this page's population that used an LLM for //some// classification, and neither classified script. {[chen2025_semantics]} (TheWebConf 2025) fine-tunes GPT-3.5 to label **cookie** purposes, and PhishLang (NDSS 2026) queried GPT-4 once to pick which **HTML tags** matter for phishing detection. Restricted to script targets the count is **zero** — 0 of the 44 papers classifying ''javascript'' and 0 of the 31 classifying ''fingerprinting-script'' — which is the figure quoted above and the one to cite. On the 4,322-paper corpus this row was empty for either reading.
  
 And what they treat as truth. These 198 papers produce **351 distinct free-text ground-truth strings**, folded here into families; 80 tuples did not fold and are printed by the report script. And what they treat as truth. These 198 papers produce **351 distinct free-text ground-truth strings**, folded here into families; 80 tuples did not fold and are printed by the report script.
Line 599: Line 603:
 ==== Methodology and limitations of these figures ==== ==== Methodology and limitations of these figures ====
  
-  * **Seven venues only.** CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P, 2010–2026, with 2025 and 2026 provisional. **EuroS&PACSACRAID, AsiaCCS, CHI and SOUPS are absent entirely** — for this topic ACSAC and EuroS&P are a real hole, since a good deal of web-script security work lands there. Every claim here is a claim about those seven venues.+  * **Seven venues only**, 2010–2026, with 2025 and 2026 provisional. Which venueswhich yearswhat each stage of the selection funnel costs and which venue-years are empty are on [[literature:corpus]] and are not restated here. For //this// topic ACSAC and EuroS&P are a real hole, since a good deal of web-script security work lands there. Every claim here is a claim about those seven venues.
   * **206 is a floor, and it has a false-positive tail.** Papers whose extraction never names a script as the object of detection are missing; conversely a handful of papers in the population (an e-voting client audit, a router-attack paper, a PHP injection-sink study) analyse JavaScript incidentally. The report script prints the full list so you can judge.   * **206 is a floor, and it has a false-positive tail.** Papers whose extraction never names a script as the object of detection are missing; conversely a handful of papers in the population (an e-voting client audit, a router-attack paper, a PHP injection-sink study) analyse JavaScript incidentally. The report script prints the full list so you can judge.
   * **Not every field can carry a percentage.** ''crawlConfig.*'', ''legal.law'' and ''platforms'' reproduce to within a few points on a repeat extraction and carry the figures here. ''classification.method'' agrees on only **58%** of papers between two runs of the same schema over the same text, so its table above is a **rough share, not a precise figure** — a repeat extraction moves those rows. ''detection.phenomenon'', ''classification.resourceName'' and ''groundTruthSource'' agree on roughly 20% of exact strings, which is what the folding is for and why the family and ground-truth tables print their residue.   * **Not every field can carry a percentage.** ''crawlConfig.*'', ''legal.law'' and ''platforms'' reproduce to within a few points on a repeat extraction and carry the figures here. ''classification.method'' agrees on only **58%** of papers between two runs of the same schema over the same text, so its table above is a **rough share, not a precise figure** — a repeat extraction moves those rows. ''detection.phenomenon'', ''classification.resourceName'' and ''groundTruthSource'' agree on roughly 20% of exact strings, which is what the folding is for and why the family and ground-truth tables print their residue.
   * **Silence is not absence.** "Does not state whether it ran headless" means the paper did not say. These are reporting figures, not practice figures.   * **Silence is not absence.** "Does not state whether it ran headless" means the paper did not say. These are reporting figures, not practice figures.
   * **Every quoted figure was checked against the paper's own text.** The prevalence values in the extraction are model summaries, so each number reproduced on this page was re-located in ''paper.cols.txt'' after whitespace normalisation. The dataset's own "0.9% of quotes cannot be located" figure was measured on the earlier 4,322-paper run and has not been re-measured.   * **Every quoted figure was checked against the paper's own text.** The prevalence values in the extraction are model summaries, so each number reproduced on this page was re-located in ''paper.cols.txt'' after whitespace normalisation. The dataset's own "0.9% of quotes cannot be located" figure was measured on the earlier 4,322-paper run and has not been re-measured.
 +  * **Every query behind this section, the report script and its unedited output** are on [[provenance:privacy:javascript]]; corpus-level caveats are on [[literature:corpus]].
  
 ===== Open Questions ===== ===== Open Questions =====
  
-  * <wrap todo>**No public, hand-labelled corpus of tracking scripts exists.** Every current method builds its own labels from filter lists plus manual inspection, which is why cross-paper comparison is impossible. A shared benchmark would do for this field what EasyList did for request blocking.</wrap> +<WRAP todo> 
-  * <wrap todo>**LLM-based script classification is unmeasured.** No peer-reviewed paper found as of 2026-08-06. The obvious study — LLM against WebGraph, AdFlush and NoT.js on a fixed script corpus, reporting cost and version drift as well as F1 — has no baseline yet.</wrap> +  * **No public, hand-labelled corpus of tracking scripts exists.** Every current method builds its own labels from filter lists plus manual inspection, which is why cross-paper comparison is impossible. A shared benchmark would do for this field what EasyList did for request blocking. 
-  * <wrap todo>**Nobody has measured how much a headless or containerised crawler under-counts //script// classification specifically.** {[jueckstock2021_realistic]} and {[annamalai2024_fpfed]} show the gap exists for API traces and fingerprinting scripts; its size for tracking-script prevalence at scale is unknown.</wrap> +  * **LLM-based script classification is unmeasured.** Zero of the 44 corpus papers that classify a ''javascript'' target and zero of the 31 that classify a ''fingerprinting-script'' target use one, re-derived 2026-09-03, and no peer-reviewed paper outside the corpus was found either. The obvious study — LLM against WebGraph, AdFlush and NoT.js on a fixed script corpus, reporting cost and version drift as well as F1 — still has no baseline. The nearest template is TGNN's request-level annotation experiment {[xiong2026_tgnn]}, which beat filter lists 98.17% to 55.14% F1 on expert-labelled ground truth; the same comparison at script granularity has not been run. 
-  * <wrap todo>**Function-granularity blocking has no successor paper.** NoT.js {[amjad2024_notjs]} and ByteDefender {[bahrami2025_bytedefender]} both stop at detection plus surrogate generation; nobody has measured what happens when either is deployed to real users at scale, or whether trackers adapt.</wrap> +  * **Nobody has measured how much a headless or containerised crawler under-counts //script// classification specifically.** {[jueckstock2021_realistic]} and {[annamalai2024_fpfed]} show the gap exists for API traces and fingerprinting scripts; its size for tracking-script prevalence at scale is unknown. 
-  * <wrap todo>**Cross-platform divergence is a confound in every older result.** If 20.6% of scripts execute differently by platform {[zafar2025_samescript]}, every desktop-only prevalence figure in this page's tables is a measurement of the desktop path only. Re-running any of them on mobile is a well-defined study.</wrap>+  * **Function-granularity blocking has no successor paper.** NoT.js {[amjad2024_notjs]} and ByteDefender {[bahrami2025_bytedefender]} both stop at detection plus surrogate generation; nobody has measured what happens when either is deployed to real users at scale, or whether trackers adapt. 
 +  * **Cross-platform divergence is a confound in every older result.** If 20.6% of scripts execute differently by platform {[zafar2025_samescript]}, every desktop-only prevalence figure in this page's tables is a measurement of the desktop path only. Re-running any of them on mobile is a well-defined study. 
 +</WRAP>
  
 ===== Related Pages ===== ===== Related Pages =====
Line 618: Line 625:
   * [[Privacy:Cookies]] — what the scripts write; the provenance argument (a cookie set by a blocked resource) is the same idea one layer down.   * [[Privacy:Cookies]] — what the scripts write; the provenance argument (a cookie set by a blocked resource) is the same idea one layer down.
   * [[Privacy:Fingerprinting]] — 39.8% of browser-fingerprinting papers are really detecting //scripts//, so that page and this one share a method.   * [[Privacy:Fingerprinting]] — 39.8% of browser-fingerprinting papers are really detecting //scripts//, so that page and this one share a method.
-  * [[Programming:Crawler]] — the instrumentation this page assumes you already have, compared in detail. Its per-tool pages ([[Programming:Crawler:OpenWPM]][[Programming:Crawler:PageGraph]][[Programming:Crawler:Foxhound]]) are promised but not yet written.+  * [[Programming:Crawler]] — the instrumentation this page assumes you already have, compared in detail. Its per-tool pages [[Programming:Crawler:OpenWPM]] and [[Programming:Crawler:PageGraph]] are written; [[Programming:Crawler:Foxhound]] is still promised.
   * [[Programming:Stateful stateless]] — only 27.6% of these papers state it, and a stateless crawl sees first-visit script behaviour only.   * [[Programming:Stateful stateless]] — only 27.6% of these papers state it, and a stateless crawl sees first-visit script behaviour only.
   * [[Design:Website classification]] — where script classification sits in the wider taxonomy.   * [[Design:Website classification]] — where script classification sits in the wider taxonomy.
privacy/javascript.1786529132.txt.gz · Last modified: by karel.kubicek.claude