| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| programming:crawler:foxhound [2026/08/17 17:56] – Fix a false negative in the published reducer, found by the figures re-review and reproduced before fixing: begin/end are UTF-16 code-unit offsets, and slicing with them in Python (which indexes by code point) reads the wrong substring whenever an astral karel.kubicek.claude | programming:crawler:foxhound [2026/09/21 14:32] (current) – Reconciling footnote updated: the 8 on programming:crawler and privacy:javascript is now 9; what still differs is the cites-only column. Authored by Claude karel.kubicek.claude |
|---|
| ===== Use in publications ===== | ===== Use in publications ===== |
| |
| The full-text sweep ''/fox ?hound/i'' over the corpus returns **13 papers**. Two are homographs — an ImageNet class label ''americanfoxhound'' and "English Foxhound" as an example crowdsourcing label — and are excluded from every figure here. Two cite the project without running it, one of them in order to reject it {[liu2025_domino]}. That leaves **9 papers that actually ran the browser**, out of the **1,120 papers in the corpus that ran a crawl** (0.8%).((This is one more than the 8 in the specialised-crawler table on [[Programming:Crawler]] and in the tool table on [[Privacy:Javascript]], and the difference is a real correction rather than a different denominator. Those tables count extracted ''tools[]'' tuples by ''usedOrMentioned'', which files Khodayari et al.'s NDSS 2025 paper under ''compared'' and therefore in their "cites only" column; reading the sentence shows it ran the browser on 42,288 pages as one of six baseline detectors. Those two pages are queued for an errata edit; this page is the deeper audit.)) The corpus is the seven venues on [[literature:corpus]]; EuroS&P, where the tool is described, is not among them, and the project's own list of publications naming Foxhound has 14 entries across more venues.((wiki [[https://github.com/SAP/project-foxhound/wiki/Publications|Publications]] page, checked 2026-08-17.)) | The full-text sweep ''/fox ?hound/i'' over the corpus returns **13 papers**. Two are homographs — an ImageNet class label ''americanfoxhound'' and "English Foxhound" as an example crowdsourcing label — and are excluded from every figure here. Two cite the project without running it, one of them in order to reject it {[liu2025_domino]}. That leaves **9 papers that actually ran the browser**, out of the **1,120 papers in the corpus that ran a crawl** (0.8%).((**The two tool tables that used to say 8 now say 9, corrected on 2026-09-21.** The specialised-crawler table on [[Programming:Crawler]] and the tool table on [[Privacy:Javascript]] count extracted ''tools[]'' tuples by ''usedOrMentioned'', which filed Khodayari et al.'s NDSS 2025 paper under ''compared'' and therefore in their "cites only" column; reading the sentence shows it ran the browser on 42,288 pages as one of six baseline detectors. Both report scripts now read the role verdicts below for this tool, so the three pages agree at 9. They still differ in what else they can see: neither table counts {[liu2025_domino]} or the PanoptiChrome paper, because a paper that only cites Foxhound in related work has no ''tools[]'' tuple for it, so their "cites only" columns read 0 where this page reports 2.)) The corpus is the seven venues on [[literature:corpus]]; EuroS&P, where the tool is described, is not among them, and the project's own list of publications naming Foxhound has 14 entries across more venues.((wiki [[https://github.com/SAP/project-foxhound/wiki/Publications|Publications]] page, checked 2026-08-17.)) |
| |
| ^ Year ^ Corpus papers ^ Ran Foxhound ^ Share of that year ^ | ^ Year ^ Corpus papers ^ Ran Foxhound ^ Share of that year ^ |
| * **The confirmation step**, if you claim vulnerabilities rather than flows: payload generation, canary, or manual review, with the confirmed / refuted / unconfirmed counts kept separate. Do not report flows as vulnerabilities. | * **The confirmation step**, if you claim vulnerabilities rather than flows: payload generation, canary, or manual review, with the confirmed / refuted / unconfirmed counts kept separate. Do not report flows as vulnerabilities. |
| * **What you did with flows whose chain recorded no application function call**, if any claim depends on the operations in the chain — the JIT blind spot is systematic. | * **What you did with flows whose chain recorded no application function call**, if any claim depends on the operations in the chain — the JIT blind spot is systematic. |
| * Counts of attempted, loaded, crashed and timed-out pages, **counted by you**: the shipped build has ''--disable-crashreporter'', so nothing else will count them. Its user agent is also roughly a year behind stable, which belongs in the limitations. | * Counts of attempted, loaded, crashed and timed-out pages, **counted by you**: the shipped build has ''%%--disable-crashreporter%%'', so nothing else will count them. Its user agent is also roughly a year behind stable, which belongs in the limitations. |
| * Your harness. The init script, the binding, and the reducer that turns reports into rows are where the interesting decisions live, and none of them are visible from "we used Foxhound with Playwright". | * Your harness. The init script, the binding, and the reducer that turns reports into rows are where the interesting decisions live, and none of them are visible from "we used Foxhound with Playwright". |
| |