Table of Contents
Provenance: security:online_scams
Working log behind Online scams: fake shops, support scams and crypto scams. Corpus-wide caveats (venue coverage, the selection funnel, provisional venue-years) are on corpus and are not restated here. Citations use the shared bibliography; the 19 entries this run added are listed under Bibliography below and belong to the content page, not to this log. No ~~DISCUSSION~~ block: comments belong on the content page.
Voice here is a working log. It is read by somebody checking a number.
The run
| Item | Value |
|---|---|
| Date | 2026-09-24 (single sitting, unsupervised drain run) |
| Item | security:online_scams (new) — fake shops, support scams, crypto scams, chosen by Karel in the 2026-09-22 gap pass; not on the roadmap's Queued table |
| Corpus | data/extract/run1/extractions.jsonl, 5,859 papers, 5,855 with paper.cols.txt; 7 venues, 2010–2026 |
| Page status | New page. sitemap.mjs on 2026-09-24: 189 pages, 0 promised-but-missing, no page or red link for scams. Exported and read before writing: security, security:phishing and its provenance, security:virustotal, design:platforms, programming:crawler_detection, design:crawling_location, roadmap, start. Only design:platforms touches the topic, as one row (“spam / abuse / fraud / fake accounts”, 79 papers) of a phenomenon fold. Decision: create, as the eighth child of security:; no page to broaden. Overlap: security:phishing's hand-mapped population already counts three of this page's 27 papers (SCAMMAGNIFIER and LOKI as live-site, Beyond Phish as detector) under its reading of “phishing”; that population was not re-derived here (out of item). A one-line disambiguator was prepared for both pages and saved with the page — see Publication sequence |
| Models | Page, scripts, verdicts, hand codes and this log: Claude (Opus 5.5). Reading notes on the 27 IN papers: three sonnet sub-agents, nine papers each, told to quote verbatim from paper.cols.txt and write NOT STATED otherwise (notes/scams_papers_{A,B,C}.md, kept in the workdir; see Hand codes for why they are not published). External facts: one sonnet sub-agent told to fetch rather than recall and to log rejected sources (notes/scams_external.md); every fact that reached the page was re-fetched by external_checks_online_scams.sh. Review layer: see Review |
| Scripts added | scripts/scam_fold.mjs (probes, inclusion rule, 183 verdicts, hand codes for the 27), scripts/report_online_scams.mjs (+ -output.txt), scripts/verify_scam_figures.mjs (+ -output.txt), scripts/external_checks_online_scams.sh (+ -output.txt), scripts/pw_fetch_text.mjs (Playwright fetch for Cloudflare-walled vendor blogs), scripts/scams_notes_quotecheck.mjs (+ -output.txt), scripts/scams_recall_probe.mjs (+ -output.txt), scripts/bib_additions_online_scams.bib, scripts/build_provenance_online_scams.py, scripts/provenance_online_scams_prose.txt |
| Guards run before publication | check_wrap.mjs, check_tables.mjs, check_page_numbers.mjs whole-page against the concatenation of the report and verifier outputs, check_attributions.mjs (8 table attributions; prose attributions checked by hand, the prose mode matched none), bib_dedup_scan.py on a fresh export plus the additions (every candidate pair touching a new key is two different papers: same surname and year), citekey resolution (33 keys, all resolve) |
| Write path | node scripts/dw.mjs put (JSON-RPC) with --if-rev on every save |
| Accidental exposure | None. Credentials stayed in .env; GH_TOKEN only passed through to the GitHub API by the external-check script |
Scope and judgement calls
| Decision | Why | What a reasonable person might have done instead |
|---|---|---|
The boundary with security:phishing is what the victim hands over. Credentials, a seed phrase or a signature that gives authority → phishing; money knowingly sent for goods, support or returns → here | It is a rule a reader can apply to a new paper without reading ours. It sends wallet drainers and “transaction phishing” to phishing, where the literature already files them, and keeps giveaway and investment scams (the victim sends the coins) here | Draw the line by the literature's own label (“phishing” in the title). Rejected: CtPhishCapture and TxPhishScope would split by title wording, not by mechanism |
| Platform papers are IN only when the off-platform scam site or its wallet is measured | design:platforms owns accounts, comments and posts. Give and Take, Like-Comment-Get-Scammed and Evolving Bots follow the scam to a website, domain or wallet and are IN; The Imitation Game, Pirates of Charity, TBTrackerX, Chameleon Channels stay CONTEXT | Put every platform-scam paper here. Rejected: the page would duplicate design:platforms's access-route material |
| The seo_spam family (124 papers in the gap pass) was read, and search poisoning enters only when the destinations measured are scams or fraudulent storefronts | Leontiadis 2011 and 2014, Search + Seizure, the tech-support search/ads paper, NOKEScam and LOKI are IN. Generic SEO and cloaking (Juice, deSEO, SURF, Cloak and dagger, long-tail SEO spam, Klingon black keywords, promotional infections, linguistic collisions, local-business drug poisoning, EvilSeed) are CONTEXT search-abuse. Cloak of Visibility and the phishing cloaking papers are cited for method only | A separate search-poisoning page. Not proposed here; the 12 CONTEXT papers are the start of that list |
| The 2011–2014 spam-value-chain storefront papers are IN (Click Trajectories, Show Me the Money, Priceless, Taster's Choice, PharmaLeaks, and the two pharmacy search-poisoning papers) | They are where payment-identifier grouping, test purchases, order-number sampling and feed-overlap measurement come from, and the page dates them as historical. With Search + Seizure they are the eight 2011–2014 storefront papers of the 27 | Exclude them because many of the goods shipped (counterfeit rather than non-delivery). Reasonable; the page would lose its only account of why test purchases stopped |
| Social-engineering ad papers are IN (Vadrevu 2019, Subramani 2020, Szurdi 2021, TRIDENT 2023) | Their landing pages are tech-support, survey and fake-prize scams, and they carry the vantage and cloaking experiments the page needs | Treat them as malvertising and send them elsewhere. There is no malvertising page; they would be orphaned |
| Reshipping-mule, scam-baiting and back-office studies are CONTEXT (Drops for Stuff, Scambaiter, cybercriminal-minds, mobile gambling scams, the revenue-estimation study) | None measures a scam website as the object. Two are cited for method: [1Park, Youngsam; Jones, Jackie; McCoy, Damon; Shi, Elaine; Jakobsson, Markus (2014): "Scambaiter: Understanding Targeted Nigerian Scams on Craigslist", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] and [2Gómez, Gibran; Liebergen, Kevin van; Caballero, Juan (2023): "Cybercrime Bitcoin Revenue Estimations: Quantifying the Impact of Methodology and Coverage", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] | Count Drops for Stuff IN (its data are databases of seven scam websites). Reasonable; declined because the victims it measures are mules recruited by job scams, not site visitors |
| A named lineage exception rather than an object rule for storefronts (written after the generic review) | The verdicts kept the 2011–2014 pharmacy and counterfeit storefront papers IN and filed the 2022 illicit-drug local-search-poisoning paper, SEISE and the black-keywords paper as CONTEXT. The written rule's “counterfeit or unlicensed” admitted both; the criterion actually applied was lineage. The rule now says so | Move the 2022 paper IN (28 papers; the search channel would then have a 2022–2024 paper). Reasonable; declined because its object is the promotion of openly illicit trade, not a buyer being defrauded |
| Hand codes by one reader, from structured notes, not double-coded | 27 papers × 7 fields in one sitting. The notes carry verbatim quotes per field; the codes are in scam_fold.mjs and every figure the page prints from a single paper is re-checked by a needle | Double-code a sample. Not done; the page says so in its limitations paragraph |
| The page publishes no loss total of its own | Every harm figure is one paper's, with its instrument and bias. Two papers' “>100M” extrapolations are quoted with their assumption | Sum the wallet inflows across papers. Rejected: different windows, coins, prices and attribution rules |
| LLM classification dated as emerging on 1 of 27 papers, with the corpus-wide LLM counts beside it | The spec asks for year-by-year counts behind any “current” claim. The scam literature's one LLM classifier paper is 2025, provisional | Call LLM classification current practice, on the strength of Chrome's deployment and outside papers. Rejected for the corpus claim; the outside evidence is on the page as footnotes |
| Outside-corpus papers only as footnotes (ScamFerret, DIMVA 2025; toll-scam domains, eCrime 2025) | Both verified from arXiv abstract pages; neither is in any count | Add them to the bibliography. Reasonable; footnotes keep the bibliography to the corpus plus cited method papers |
Probes and the candidate set
All run over paper.cols.txt, whitespace-collapsed, paper-counted; the gap-pass regex is also run exactly as the gap pass ran it (paper.norm.txt, raw). Regexes are in scam_fold.mjs below.
| Probe | Threshold | Papers |
|---|---|---|
GAP (2026-09-22 gap-pass family scams), as run on paper.norm.txt | ≥ 5 | 43 (27 web, 21 from 2024–2026) — reproduces the gap pass |
GAP, same regex on paper.cols.txt | ≥ 5 | 45 |
| SCAM, the bare word | ≥ 10 | 127 |
| SPEC, named scam types (pharmacy, counterfeit, giveaway, rug pull, reshipping, …) | ≥ 5 | 79 |
| LIN, 2011–2016 vocabulary (affiliate program, spam-advertised, scareware, social-engineering ads, test purchase) | ≥ 5 | 49 |
| SCHEMA, the extraction's detection / classification / population free text | any | 57 |
| Union = candidate set | 183 |
Recall checks outside the union:
- a title sweep of all 5,859 papers for scam, fraud, fake, counterfeit, spam, cloak, swindl, deceptive, scareware, giveaway, pharmac, extort, ponzi that are not candidates added no IN paper; it surfaced two method papers cited as CONTEXT outside the set (
EXTRA_CONTEXTin the fold); - after the generic review, a title-plus-summary probe over empirical web-platform papers outside the candidates (fraud, fake, counterfeit, decepti, scareware, illicit, malvertis, rogue, swindl, scam, pharmac, storefront, affiliate, extort, ponzi, tech support) returned 36 papers — click, ad, view and review fraud, deceptive patterns, cache deception, deepfakes, domain parking, one phishing life-cycle paper — and no scam-website paper (
scripts/scams_recall_probe-output.txt, printed below with the other outputs); - no paper in the corpus combines an ad library, archive or repository with scam or fraud in its title, summary or tool names (0), which is what the channel table's “untested” rests on;
- a schema-only probe added one paper to the full-text union (PharmaLeaks), which is why SCHEMA is a probe.
Queries, with their denominators
Every corpus figure on the page, the query behind it, and its population. The printed values are in the report output below.
| Figure on the page | Query | Population / denominator |
|---|---|---|
| 183 candidates; IN 27 (14.8%), CONTEXT 67 by code, OUT 89 | PASS A union; PASS B verdicts | 183 candidates |
| probe precision and recall table | PASS B | hits of each probe; recall of the 27 |
| gap pass 43 / 27 web / 21 recent; 23 recent GAP hits of which 7 IN; 9 IN missed | PASS A, PASS B | 5,859 papers (norm); 45 GAP hits (cols) |
| 26 web, 24 crawled | PASS C: platforms includes web; crawlConfig non-null or automated-web-crawl | 27 IN |
| year buckets and types; venues | PASS C | 27 IN; last bucket starred |
| channel, grouping, label, cloaking, harm, interaction counts; 19 grouped; 14 measured blocklist coverage; 7 paid/phoned/messaged | PASS D over HAND | 27 IN, hand-coded |
| “four of the six 2022–2024 papers” did not group | PASS D per-era column | 6 IN papers from 2022–2024 |
| ethics outcomes 4 / 14 / 2 / 2 / 1 / 4 no record | PASS E: ethics.reviewOutcome, null counted separately | 27 IN |
| released code or data 11 / 2 / 1 / 13 | PASS E: artifacts.availability | 27 IN |
LLM: 1 of 27 classifies with an LLM; 7 CONTEXT papers with an llm tuple; corpus-wide 2 / 27 / 77 / 71 by year | PASS E: classification[].method == 'llm' | 27 IN; 67 CONTEXT; 5,859 papers |
| 4 papers from 2025, 1 from 2026 | PASS C by year | 27 IN |
The llm enum marks two IN papers; the page says one classifies scam sites with an LLM because LOKI's LLM tuple is a branded-keyword filter (FLAN-T5), which the extraction filed under target: mobile-app — an extraction error, recorded here and not patched.
Hand verdicts and hand codes
The 183 verdicts carry a reason each (VERDICTS in the fold; also printed in PASS G for CONTEXT). Verdicts were made from title, summary and full-text snippets around the probe hits; the borderline ones were read further (Evolving Bots, the Bing advertiser-fraud study, cybercriminal-minds, Fashion crimes, Knowing your enemy, the WebSocket paper).
Hand codes (HAND in the fold) were made by the page author from three sets of reading notes written by sonnet sub-agents, one record per IN paper, with a verbatim quote per field. The notes are not published: a machine check of their quotes (scams_notes_quotecheck.mjs, output below) located 586 of 777 quotes verbatim (75.4%) in the paper's renderings. Two of the misses were read by hand and both sentences are in the paper, split by a column break (paper.cols.txt still splices about a quarter of pages); the rest were not individually checked, so the notes are treated as reading aids, not as evidence. The evidence for any figure on the page is its needle in verify_scam_figures.mjs.
Codes that were changed: Priceless's cloaking code from measured to addressed while coding (merchants refusing undercover buyers is counter-intelligence, not a two-vantage comparison); PharmaLeaks's interaction from purchases to none after the generic review (its orders were Click Trajectories' and Priceless's, reused to authenticate the leak), which moved “paid, phoned or messaged” from 7 to 6. SCAMMAGNIFIER's harm is coded leaked-revenue because the field's definition covers partner transaction data. After the generic review the mobile gambling-scams paper moved from CONTEXT ecosystem to a new mobile-app code (back office 5 → 4), and the rule gained two written clauses — the 2011–2014 storefront lineage exception and the scareware purchase-page criterion — that describe verdicts already made rather than changing any.
How far the codes were checked. The figures reviewer compared 10 rows with the reading notes and re-derived two figures from the papers; the generic reviewer checked six codes directly against full text (five held; the sixth was the PharmaLeaks recode above). So about a quarter of the 27 rows have been checked against the papers themselves; the rest rest on the notes, which this log does not treat as evidence.
Quotes and figures checked against the papers
verify_scam_figures.mjs: 187 needles, each a sentence or clause carrying a figure or a quote on the page, looked up in paper.cols.txt → paper.norm.txt → paper.txt → a pypdf extraction. All 187 located (166 in .cols, 21 only in the PDF). Three mutation controls — a real needle with one digit changed — are not located, so the lookup can fail. No needle is under 20 characters; the 14 that were (bare percentages and dollar amounts) were lengthened with their sentence after the first run, because a bare “2.34%” passes against the wrong sentence.
Needles that needed care: Miramirkhani et al.'s blacklist sentence is spliced in .cols (“only 108 (7%) were each word is correlated…”), so it is checked as three fragments; Vadrevu and Perdisci's “$4.8” is split by a column break and checked as two; Beyond Phish's false-negative rate sits before a table caption and is checked as “false positive rate (FPR) and 1.63%”. Li et al.'s 21-hours figure has a denominator the first draft got wrong (see Mistakes).
A second check pulled every double-quoted span of 12 or more characters out of the page source and required it to be inside a needle; the misses were quotation marks around the page's own phrases, a nested quote changed from double to single ('blind'), and external quotes checked by the external script.
External sources
| Source | Used for | How verified |
|---|---|---|
| Let's Encrypt blog, 2025-08-14, RFC 6962 logs end of life | CT box: read-only 2025-11-30, shut down 2026-02-28 | re-fetched by the external script |
certstream.calidog.io and CaliDog/certstream-server | CT box: front end answers; last commit 2025-09-04; issue #143 (websocket down) open since 2026-02-09 | HTTP status, GitHub commits and issues API. The live WSS stream was not tested (no client here). Issue and successor found by the external-currency reviewer |
reloading01/certstream-server-rust README | CT box: a Certstream-compatible server that reads RFC 6962 and static-CT logs | README re-fetched; its claim is the maintainer's, not tested here |
| Google blog, 2025-05-08, Gemini Nano in Enhanced Protection | cloaking section: tech-support scam pages | Playwright fetch; the claim is limited to tech-support scams, as Google's post is |
| Google blog, 2025-09-18, new AI features for Chrome | cloaking section: announced extension to “fake viruses or fake giveaways” | Playwright fetch. Whether it shipped: not established |
| Microsoft Edge blog, 2025-01-27 (preview) and 2025-10-31 (default on), scareware blocker | cloaking section | Playwright fetch of both posts (curl gets 403) |
Safe Browsing v4 ThreatType reference and v4 overview | label section: scams fall under SOCIAL_ENGINEERING; v4 deprecated | both re-fetched; end date and v5 migration left to security:phishing |
| Online Safety Act 2023 ss. 38–39 (legislation.gov.uk) | channel table: fraudulent-advertising duties | statute text re-fetched |
| FBI IC3 2025 Internet Crime Report PDF | harm section: $20.877 billion reported | PDF re-fetched, text extracted, two phrases matched |
| GASA Global State of Scams 2025 | harm section: “$442 Billion (estimated…)”, 46,000 people, 42 markets | landing page on gasa.org re-fetched; the PDF read is a third-party mirror (uploads.agencialupa.org), flagged in the footnote |
| GOV.UK, Report Fraud replaces Action Fraud, 2025-12-04 | harm section | re-fetched |
| arXiv 2502.10110 (ScamFerret, DIMVA 2025) and 2510.14198 (toll scams, eCrime 2025) | currency and open questions, footnotes | abstract pages re-fetched; venue from the arXiv comments field |
LOKI Zenodo record and codepujan/ndss-loki-artifact; pragseclab/Crimson; mbitaab/beyondphish; phani-vadrevu/seacma; ian7yang/trident; jszurdi/ODIN; Double and Nothing dataset CSV; nokescam.com DNS | read-first table and open questions | GitHub commits API, HTTP fetch, DNS lookup in the external script |
Rejected or not used, and why:
| Source | Why not |
|---|---|
| FTC Consumer Sentinel Network Data Book 2025 | not published as of 2026-09-24 (the 2025 URL is 404); the 2024 book exists but the page needs only one official series, and IC3's is current |
| CryptoScamDB | API returns 502, last commit 2020-12-17: dead; not recommended as a label source |
| Chainabuse (TRM Labs) | live, but the free API is 10 calls a month; a researcher cannot draw a population from it without a partnership. Not on the page rather than half-described |
| ScamSniffer 2025 loss report | a vendor blog about signature phishing — the phishing side of the boundary |
MetaMask eth-phishing-detect | actively maintained, but a phishing blocklist; Muzammil et al.'s coverage table already includes MetaMask |
| Australian Scamwatch dashboard | operator verified (ACCC); figures sit in an interactive dashboard not fetched here |
| eCrime 2025 Contextual Classification of Cybercriminal Posts … Tech Support Scam Marketplaces | exists, but no verbatim abstract obtainable (IEEE Xplore and ResearchGate refused automated fetches); not cited |
| eCrime 2025 LLM scam-baiting system; arXiv scam-scenario and fake-shop crawling preprints | on-topic but about conversations or one country's shops; the two footnotes already make the outside-corpus point |
| Ofcom draft Fraudulent Advertising Codes (consultation to 2026-10-02) | the external reviewer reached it only via a trade-association summary; Ofcom's own site refused automated fetches. Not on the page |
| socialmediatransparency.org (DSA ad-repository tracker) | a secondary tracker; the page points to design:platforms:ad_archives instead |
What could not be established
- Whether CertStream carries static-CT logs. The front end answers; the stream was not tested; an open issue says the websocket is down. A CT-based study in 2026 should check its client against the current log list.
- Prevalence of scam sites. No paper here estimates it; the only capture–recapture attempt (Leontiadis et al.) gives two pharmacy-population estimates a factor of three apart.
- Double-coding. The hand codes have one reader. Single-paper differences are soft.
- Artifacts for three papers. SCAMMAGNIFIER, Give and Take and Ctrl+Alt+Deceive print no artifact URL in their full text; an artifact appendix outside
paper.cols.txtmay exist. NOKEScam's release domain does not resolve. - Whether Chrome's announced extension to fake-virus and fake-giveaway pages has shipped. Announced 2025-09-18; no Google page confirming it was found.
- eCrime coverage. The corpus lacks the venue where much scam work appears; the outside search in this sitting was a keyword pass, not a sweep of eCrime, EuroS&P, ACSAC, RAID and WEIS proceedings.
Mistakes caught during the run
- The gap-pass slope. The item described the family as “the steepest recent slope in the pass”. Re-running the regex reproduces 43 / 27 / 21 exactly, but only 7 of the 23 recent hits (in
.cols) are scam-website measurement; the slope is carried by platform, on-chain, user-study, telephony and LLM-conversation papers. The page says so. - “Flat at five to seven papers per bucket” — the 2014–2017 bucket has 3. Rewritten as “three to seven”.
- “Every paper here that grouped sites found a steep head” — not checked for all 19. Rewritten as “the papers that report how sites distribute over operators”.
- “108 (7%) were ever blacklisted during the study” — the paper says “were blacklisted”; “ever … during the study” was an inference. Rewritten.
- Li et al.'s 21 hours — first drafted with 2,266 wallets as the denominator; the paper computes it over “that small overlap”, the 300 addresses BitcoinAbuse also listed.
- Srinivasan et al.'s crawler “was eventually detected” — the paper offers detection as one of two explanations. Rewritten to its own warning that PhantomJS “can, in principle, be detected by scammers”.
- The 2% conversion rate behind Dial One for Scam's $9.7M — first drafted as “taken from earlier fake-antivirus work”; the paper says it assumes tech-support victims convert like fake-antivirus buyers of the “full version”. Rewritten to say that.
- Two quotes with changed case (“any work…”, “if an attacker…”) and one changed verb (“belong” for the paper's “belongs”) — restored to the paper's wording or moved outside the quotation marks.
- A pig-butchering open question claimed no paper measured the website side; Like-Comment-Get-Scammed follows victims from comments to fake investment platforms to wallets. The bullet was deleted.
- “No tech-support scam paper since 2018” — the social-engineering ad papers through 2023 include tech-support scams as a category. Qualified.
Publication sequence
Every save used dw.mjs put with --if-rev set to the revision read just before it; times are UTC, 2026-09-24.
| Step | Page | Revision before → after |
|---|---|---|
| 1 | bibliography — 19 entries appended to a fresh export, before </bibtex>; cache purged | 1790126248 → 1790286292 |
| 2 | online_scams — new page (21:45) | none → 1790286299 |
| 3 | online_scams — this page, first save (21:45) | none → 1790286308 |
| 4 | security — eighth row; “seven” → “eight” in box, heading and methodology; “Re-derived 2026-09-24” | 1790032323 → 1790286323 |
| 5 | phishing — one-line disambiguator in the scope paragraph | 1789632761 → 1790286325 |
| 6 | platforms — Related pages bullet | 1790040724 → 1790286327 |
| 7 | roadmap — Assessed row | 1790127423 → 1790286329 |
| 8 | roadmap — section 3e | 1790127424 → 1790286330 |
| 9 | security — Eighth child, 2026-09-24 log entry with the re-run set check | 1790032596 → 1790286383 |
| 10 | this page, second save with this table complete | see the page history |
Checks after the saves: ?purge=true on the bibliography and on every page above; the rendered content page shows 33 reference entries for 33 distinct citekeys (177 markers) and no wikilink2 (red link); the same holds for this page, security, phishing, platforms, roadmap, roadmap and security. Cross-page anchors resolve: design:crawling_location#residential_and_mobile_proxies, security:phishing#cloaking, security:phishing#google_safe_browsing. report_namespace_overviews.mjs exited 0 with security 8 8 equal. sitemap.mjs still listed 189 pages because ?do=sitemap is cached; its exit 1 comes from 21 bare-link defects on other provenance: pages that were there before this sitting and are not touched here.
Other pages touched in this sitting
- security — new table row; “seven” → “eight” in the box, the table heading and the methodology paragraph; the namespace set check in
report_namespace_overviews.mjsre-run after the save (see Publication sequence). - phishing — one-line disambiguator in the scope paragraph.
- platforms — one-line disambiguator in Related pages.
- bibliography — 19 entries appended before
</bibtex>.
Review
Four reviewers, each told that the author's context might not be exhaustive and each handed the page, this log's draft, the report script with its unedited output, the verifier with its output, the external-check script with its output, the fold module and the reading notes (brief: scams/review_brief.md). The three focused passes ran in parallel on the first complete draft, before publication; every accepted fix was applied and the report, verifier, external checks and guards re-run before the generic pass. Findings files: notes/scams_review_figures.md, notes/scams_review_citations.md, notes/scams_review_external.md, notes/scams_review_generic.md.
| Reviewer | Finding | Action |
|---|---|---|
| figures-vs-script (sonnet) | Re-ran the report and the verifier: both byte-identical to the committed outputs. Compared 10 of the 27 hand-code rows with the reading notes (and re-derived two figures from the papers): no code contradicted. Every table cell and all arithmetic reproduced | — |
| figures-vs-script (sonnet) | Imprecision. “the rest are platform, on-chain, user-study, telephony and LLM-conversation papers” covers 13 of the 16 non-IN recent GAP hits; one is phishing-boundary and two are OUT | Accepted. Sentence now gives 13 / 1 / 2 |
| figures-vs-script (sonnet) | Imprecision. Surveylance and Szurdi et al. carry two channel codes but were listed under one channel each, unlike Srinivasan et al. and LOKI | Accepted. Both now listed under both rows |
| citations-and-quotes (sonnet) | All 33 keys resolve; no collisions; no duplicate papers under other keys; the 19 new entries' metadata verified against NDSS and USENIX pages and Crossref, including the two hand-filled USENIX author lists; ~40 attributions correct; 18 substantive claims re-read in context | — |
| citations-and-quotes (sonnet) | Defect. “transaction phishing” in quotation marks is not [3He, Bowen; Chen, Yuan; Chen, Zhuo; Hu, Xiaohui; Hu, Yufeng; Wu, Lei; Chang, Rui; Wang, Haoyu; Zhou, Yajin (2023): "TxPhishScope: Towards Detecting and Understanding Transaction-based Phishing on Ethereum", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)]'s term; the paper says “transaction-based phishing” | Accepted. Quote corrected; the paper's phrase is now a needle |
| citations-and-quotes (sonnet) | Imprecision. The Edge footnote gave a 2025-01-27 preview date but only the 2025-10-31 URL | Accepted. Both posts cited; the preview post is now fetched by the external script |
| citations-and-quotes (sonnet) | Nit. kotzias2025_ctrl author “Iuit, Javier Aldana” may split a two-part surname | Rejected. Crossref for 10.14722/ndss.2025.241816 gives given “Javier Aldana”, family “Iuit”; the entry follows the publisher's metadata |
| citations-and-quotes (sonnet) | Nit. “Li et al.” names two different papers | Accepted where they sit close together (grouping pitfall, harm extrapolation): venue-year added |
| external-currency (sonnet) | Defect. Safe Browsing v4 is deprecated; the page cited the v4 ThreatType reference without saying so | Accepted and re-verified from Google's own v4 overview (“The Safe Browsing APIs (v4) are deprecated”); the end date and v5 migration are pointed to security:phishing, which already carries them |
| external-currency (sonnet) | Imprecision. Google said on 2025-09-18 that it would extend Chrome's Gemini Nano check to “fake viruses or fake giveaways” | Accepted and re-verified (Playwright fetch of Google's post). The page says it was announced and that shipping was not established |
| external-currency (sonnet) | Defect (dating). CertStream: no commit since 2025-09-04, an open issue since 2026-02-09 reporting the websocket down; a maintained Certstream-compatible server reads static-CT logs | Accepted and re-verified (GitHub API for the issue state; the certstream-server-rust README). The CT box now says to treat the public server as unmaintained across the transition |
| external-currency (sonnet) | Omission. EU DSA Art. 39 ad repositories as a scam-ad discovery channel; the X €120m decision | Accepted in part. A channel-table row notes that no paper in the population uses a platform ad repository and points to design:platforms:ad_archives, which already covers the repositories and the X decision. The reviewer's source (a transparency-tracking site) was not used |
| external-currency (sonnet) | Omission. UK Online Safety Act fraudulent-advertising duty; Ofcom's draft codes out for consultation until 2026-10-02 | Accepted in part. The statutory duties (ss. 38–39) are cited from legislation.gov.uk and re-fetched by the external script. The Ofcom consultation is not on the page: the reviewer reached it only through a trade-association summary (Ofcom's site refused automated fetches) |
| external-currency (sonnet) | Everything else re-fetched and reproduced (CT dates, Edge default-on, IC3, GASA, Report Fraud, arXiv papers, artifacts, nokescam.com DNS). GASA has 2026 country reports but no newer global edition; ScamAdviser is owned by Gogolook since 2024 | No change; noted |
| generic (fable) | F1 (defect). The blocklist row called the method “superseded” in a column defined as usage; three 2023–2025 papers use a list as the label | Accepted. Column now says it records use and marks judgements; the row reads “current in use … judgement: do not use it as ground truth” |
| generic (fable) | F7 (defect). “all validated against people looking at sites” — 19 of 27; Starov 2018 labels with VirusTotal only | Accepted. “19 of the 27 had researchers look at sites” |
| generic (fable) | F2: “Search-query discovery: current” rests on one 2026 paper after none in 2022–2024 | Accepted in both tables |
| generic (fable) | F3: the written rule admits the 2022 illicit-drug search-poisoning paper, SEISE and the black-keywords paper that the verdicts file as CONTEXT; the criterion in use is lineage | Accepted as a rule clause, verdicts unchanged; the page says a reader could move the 2022 paper IN; recorded under Scope and judgement calls with the alternative |
| generic (fable) | F4: scareware is an IN type in one place and CONTEXT (download payload) in another | Accepted as a rule clause (purchase page vs download) |
| generic (fable) | F5: the mobile gambling-scams paper sat under a code (back office) that does not describe it | Accepted. New CONTEXT code mobile-app; back office 5 → 4 |
| generic (fable) | F6: “None measures losses directly” contradicted by leaked revenue and victim-report rows | Accepted. Reworded to “none gives a loss total …” |
| generic (fable) | F8: harm and interaction tables list subsets under “Papers” | Accepted. Columns renamed “Example papers”/“Example paper”, with a pointer to the full coding |
| generic (fable) | F9: PharmaLeaks coded as buying; its orders were earlier studies' | Accepted. Recoded; “paid, phoned or messaged” 7 → 6 |
| generic (fable) | F10: “not behind” argued from counts with no denominators | Accepted. Denominators added to the report (2/719, 27/690, 77/770, 71/415) and printed on the page with the 1-of-5 figure |
| generic (fable) | F11: site lifetime cited four times, defined three ways, never treated as a method | Accepted. The three definitions (each a needle) now preface the results table; a What to report bullet asks for the liveness definition, cadence and right-censoring |
| generic (fable) | F12: no guidance on labelling the scam type | Accepted. A paragraph under Ground truth (Kotzias's seven types, LOKI's ten categories, the page's own seven) and a What to report bullet |
| generic (fable) | F13: this log said the hand codes were checked “against the notes and the papers” | Accepted. The first row of this table now says what was done; Hand verdicts and hand codes states how many rows were checked against papers |
| generic (fable) | F14: this log described sibling edits and post-publication re-runs as done while the draft was unpublished | Accepted. A Publication sequence section records each save with its revision, written as the saves happened |
| generic (fable) | F15, F16, F17, F19: “two channels” for two query lists; “held-out test”; the intro's harm trio; a negative universal without “in this population” | Accepted, all four |
| generic (fable) | F18: record a title-plus-summary recall probe, which agrees | Accepted. Re-run here (36 hits, none a scam-website paper; 0 ad-library papers) and published below |
After the generic pass's fixes the report, verifier (193 needles, all located) and external checks were re-run, and check_page_numbers.mjs, check_wrap.mjs, check_tables.mjs and check_wrapped_lists.py re-passed. The focused reviewers were not re-run on the post-generic revision; every change since their pass is listed in the rows above and was re-checked by the author against the scripts' outputs.
Report output, unedited
node scripts/report_online_scams.mjs > scripts/report_online_scams-output.txt — deterministic, about 20 s.
- report_online_scams-output.txt
=== PASS A — candidate probes (full text, paper-counted) === papers: 5859; with paper.norm.txt: 5859; with paper.cols.txt: 5855 probe rendering threshold papers ---------------------------------------------------------- ------------------------------------ --------- ------ GAP (2026-09-22 gap pass, as run) paper.norm.txt, raw >= 5 43 GAP (same regex) paper.cols.txt, whitespace-collapsed >= 5 45 SCAM (bare word scam/scammer/…) paper.cols.txt, collapsed >= 10 127 SPEC (named scam types) paper.cols.txt, collapsed >= 5 79 LIN (spam value chain, scareware, SE ads) paper.cols.txt, collapsed >= 5 49 SCHEMA (detection / classification / population free text) extraction any 57 UNION = candidate set 183 gap pass as run: 43 papers, 27 include the web platform, 21 from 2024–2026 === PASS B — hand verdicts over the candidate set === INCLUSION RULE (written before the verdicts were counted): IN — the paper's measured objects include web-delivered scams: websites, landing pages, or the ad / search / redirect chains that deliver them, whose purpose is to make the visitor PAY or TRANSFER money under false pretences (goods that never arrive or are counterfeit or unlicensed, "support", doubled or invested crypto, survey prizes, one-click billing). The paper collects instances (crawl, feed, telemetry, leaked or seized back end) and measures, clusters, labels or detects them. CONTEXT — scams, but the object is not a scam website: scam accounts/comments on a platform (design:platforms), on-chain scam contracts/tokens with no website, phone/SMS scams, user studies of scam susceptibility, credential/seed/signature theft that the literature calls phishing (security:phishing), search abuse or scam ads whose destinations are not primarily scams, and back-office or chat studies of fraud crews. OUT — scams are incidental, or "fraud" is against a platform/merchant/operator (ad fraud, telephony fraud, concession abuse), malware delivery, or a homonym (Scamper, Scamalytics). Boundary with security:phishing: WHAT THE VICTIM HANDS OVER. Credentials, seed phrases or a signature that gives the attacker authority -> phishing. Money the victim knowingly sends -> here. Named lineage exception (added after the generic review): the 2011–2014 spam- and SEO-advertised pharmacy and counterfeit storefront papers are IN for the methods they introduced (payment grouping, test purchases, order numbers, feed overlap). Later search poisoning that promotes openly illicit trade (controlled drugs, gambling, black-market goods) is CONTEXT search-abuse: the criterion in use is lineage, not the object, and the page says so. Scareware counts when the paper measures the purchase page; a paper that measures the download (the software payload) is CONTEXT. candidates 183: IN 27 (14.8%), CONTEXT 67 (36.6%), OUT 89 (48.6%) CONTEXT, by code (papers): code papers meaning ----------------- ------ --------------------------------------------------------------------------------------------------------------------------------- onchain 12 scam contracts, tokens or transactions with no website in the object search-abuse 12 search poisoning / cloaking / SEO whose destinations are generic, or openly illicit trade outside the 2011–2014 lineage exception platform 11 scam accounts, comments or posts on a platform are the object (design:platforms) telephony 10 phone or SMS scams user-study 5 people's exposure, susceptibility or discussion of scams phishing-boundary 5 credential, seed or signature theft called phishing (security:phishing) ads 5 scam ads measured as ad exposure (privacy:advertising) ecosystem 4 back-office, dark-web or chat studies of fraud crews; revenue-estimation method llm 2 LLMs in scam conversations (detection, scam-baiting, scam operations) mobile-app 1 scam apps (not websites) measured as the object OUT, by code (papers): code papers meaning ------------- ------ ---------------------------------------------------------------------------------- user-general 21 general security/privacy user study in which scams come up mention 19 scams incidental to another topic other-fraud 14 fraud against a platform, merchant or operator; account abuse; credential spoofing onchain-other 13 blockchain security or tracing that is not a scam measurement malware 9 malware or unwanted-software delivery infra 9 network measurement; "Scamper"/"Scamalytics" homonyms domain-abuse 4 typo/bit-squatting, parking Precision and recall of each probe against the IN set: probe hits IN precision CONTEXT recall of IN --------------- ---- -- --------- ------- ----------------- GAP >= 5 (cols) 45 18 40.0% 20 18 of 27 (66.7%) SCAM >= 10 127 19 15.0% 52 19 of 27 (70.4%) SPEC >= 5 79 23 29.1% 43 23 of 27 (85.2%) LIN >= 5 49 11 22.4% 14 11 of 27 (40.7%) SCHEMA 57 16 28.1% 28 16 of 27 (59.3%) UNION 183 27 14.8% 67 27 of 27 (100.0%) IN papers the GAP probe misses: CCS/2010/dissecting-one-click-frauds (gap hits 4) CCS/2012/priceless-the-role-of-payments-in-abuse-advertised-goods (gap hits 0) CCS/2014/a-nearly-four-year-longitudinal-study-of-search-engine-poisoning (gap hits 0) IEEE-SP/2011/click-trajectories-end-to-end-analysis-of-the-spam-value-chain (gap hits 1) IMC/2012/tasters-choice-a-comparative-analysis-of-spam-feeds (gap hits 1) IMC/2014/search-seizure-the-effectiveness-of-interventions-on-seo-campaigns (gap hits 1) USENIX/2011/measuring-and-analyzing-search-redirection-attacks-in-the-illicit-online-prescri (gap hits 1) USENIX/2011/show-me-the-money-characterizing-spam-advertised-revenue (gap hits 0) USENIX/2012/pharmaleaks-understanding-the-business-of-online-pharmaceutical-affiliate-progra (gap hits 0) GAP >= 5 (cols) and IN: 18; of the GAP set from 2024–2026: 23, of which IN 7 GAP set, 2024–2026, by verdict/code: IN:general 4 CONTEXT:platform 4 CONTEXT:onchain 3 CONTEXT:user-study 3 IN:crypto 2 CONTEXT:llm 2 OUT:other-fraud 1 CONTEXT:telephony 1 IN:shop 1 OUT:user-general 1 CONTEXT:phishing-boundary 1 === PASS C — the IN population === IN: 27 papers; include the web platform: 26; crawled (crawlConfig or automated-web-crawl): 24 By year bucket (2025–2026* provisional): bucket IN CONTEXT types (IN) ---------- -- ------- ----------------------------------------------- 2010–2013 7 10 oneclick 1, shop 6 2014–2017 3 9 shop 2, techsupport 1 2018–2021 6 8 general 1, socialeng 3, survey 1, techsupport 1 2022–2024 6 16 crypto 2, general 2, shop 1, socialeng 1 2025–2026* 5 24 crypto 1, general 3, shop 1 By year (IN): 2010: 1; 2011: 3; 2012: 3; 2014: 2; 2017: 1; 2018: 3; 2019: 1; 2020: 1; 2021: 1; 2023: 4; 2024: 2; 2025*: 4; 2026*: 1 By venue (IN): CCS 3; IMC 6; NDSS 6; PETS 0; USENIX 5; WWW 4; IEEE-SP 3 By scam type (IN): type papers years meaning ----------- ------ ------------------------------------------------- -------------------------------------------------------------------------------------------- shop 10 2011,2011,2011,2012,2012,2012,2014,2014,2023,2025 fake, counterfeit or unlicensed storefronts (incl. spam/SEO-advertised pharmacy) techsupport 2 2017,2018 technical-support scams crypto 3 2023,2024,2025 crypto giveaway / investment scam sites socialeng 4 2019,2020,2021,2023 social-engineering ad and redirect campaigns (tech support, scareware, surveys, fake prizes) survey 1 2018 survey scams oneclick 1 2010 one-click billing fraud general 6 2018,2023,2024,2025,2025,2026 several scam types, or scam as a class IN papers (key — type — note): 2010 CCS/2010/dissecting-one-click-frauds — oneclick — Japanese one-click billing-fraud websites; 2,140 reports, bank accounts and phone numbers as linking identifiers 2011 IEEE-SP/2011/click-trajectories-end-to-end-analysis-of-the-spam-value-chain — shop — spam value chain end to end: hosting, affiliate programs, payment, purchases 2011 USENIX/2011/measuring-and-analyzing-search-redirection-attacks-in-the-illicit-online-prescri — shop — unlicensed-pharmacy storefronts reached by search redirection; conversion estimate 2011 USENIX/2011/show-me-the-money-characterizing-spam-advertised-revenue — shop — spam-advertised pharmacy/software storefront revenue from sequential order numbers and purchases 2012 CCS/2012/priceless-the-role-of-payments-in-abuse-advertised-goods — shop — payment processing behind abuse-advertised storefronts; undercover purchases 2012 IMC/2012/tasters-choice-a-comparative-analysis-of-spam-feeds — shop — coverage of ten spam feeds for storefront domains: the seed-bias paper of the lineage 2012 USENIX/2012/pharmaleaks-understanding-the-business-of-online-pharmaceutical-affiliate-progra — shop — leaked affiliate-program databases: customers, revenue, payment 2014 CCS/2014/a-nearly-four-year-longitudinal-study-of-search-engine-poisoning — shop — four years of search poisoning toward unlicensed-pharmacy storefronts 2014 IMC/2014/search-seizure-the-effectiveness-of-interventions-on-seo-campaigns — shop — counterfeit-luxury storefronts via SEO; orders and supplier records 2017 NDSS/2017/dial-one-for-scam-a-large-scale-analysis-of-technical-support-scams — techsupport — technical support scams: malvertising crawl, infrastructure, calls to scammers 2018 IEEE-SP/2018/surveylance-automatically-detecting-online-survey-scams — survey — survey scam detector over search-derived crawl 2018 WWW/2018/betrayed-by-your-dashboard-discovering-malicious-campaigns-via-web-analytics — general — analytics IDs link malicious sites into campaigns; tech-support and other scams among them 2018 WWW/2018/exposing-search-and-advertisement-abuse-tactics-and-infrastructure-of-technical — techsupport — tech-support scams via search results and search ads 2019 IMC/2019/what-you-see-is-not-what-you-get-discovering-and-tracking-social-engineering-att — socialeng — social-engineering ad campaigns clustered by screenshot; tech-support, scareware, survey scams 2020 IMC/2020/when-push-comes-to-ads-measuring-the-rise-of-malicious-push-advertising — socialeng — malicious web-push ad campaigns; clustering; many are scams 2021 WWW/2021/where-are-you-taking-me-understanding-abusive-traffic-distribution-systems — socialeng — abusive traffic distribution systems; multi-profile, multi-vantage crawl; scam destinations 2023 IEEE-SP/2023/beyond-phish-toward-detecting-fraudulent-e-commerce-websites-at-scale — shop — fraudulent e-commerce website detector; Reddit-derived seeds 2023 IMC/2023/evolving-bots-the-new-generation-of-comment-bots-and-their-underlying-scam-campa — general — YouTube comment bots promoting scam domains; 72 scam campaigns 2023 NDSS/2023/double-and-nothing-understanding-and-detecting-cryptocurrency-giveaway-scams — crypto — crypto giveaway scam sites via CT logs; wallet flows 2023 USENIX/2023/trident-towards-detecting-and-mitigating-web-based-social-engineering-attacks — socialeng — social-engineering ad detector on crawled sites; tech-support, scareware 2024 IMC/2024/give-and-take-an-end-to-end-investigation-of-giveaway-scam-conversion-rates — crypto — giveaway scam conversion rates, landing pages and wallets 2024 NDSS/2024/like-comment-get-scammed-characterizing-comment-scams-on-media-platforms — general — YouTube comment scams followed to scam websites and wallets 2025* NDSS/2025/ctrlaltdeceive-quantifying-user-exposure-to-online-scams — general — vendor telemetry: user exposure to scam domains 2025* NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns — shop — fraudulent shopping campaigns via new registrations, automated checkout, merchant IDs 2025* USENIX/2025/nokescam-understanding-and-rectifying-non-sense-keywords-spear-scam-in-search-en — general — nonsense-keyword spear scams in Baidu; complaint data 2025* WWW/2025/the-poorest-man-in-babylon-a-longitudinal-study-of-cryptocurrency-investment-sca — crypto — crypto investment scam sites via CT logs; wallets; blocklist coverage 2026* NDSS/2026/loki-proactively-discovering-online-scams-by-mining-toxic-search-queries — general — scam discovery by mining toxic search queries === PASS D — hand-coded method attributes of the IN papers (read from full text) === denominator: the 27 IN papers; eras: 2010–2014 9, 2015–2021 7, 2022–2024 6, 2025–2026* 5 channel: channel papers of 27 2010–2014 2015–2021 2022–2024 2025–2026* meaning ----------------- ------ ----- --------- --------- --------- ---------- ----------------------------------------------------------------------------------- search 7 25.9% 3 3 0 1 organic search results for chosen queries search-ads 1 3.7% 0 1 0 0 sponsored search results ad-networks 6 22.2% 0 5 1 0 malvertising / low-tier ad networks / push ads / typosquat parking / URL shorteners social 3 11.1% 0 0 3 0 posts, comments or livestreams on a platform ct-logs 2 7.4% 0 0 1 1 Certificate Transparency stream new-registrations 1 3.7% 0 0 0 1 newly registered domain feed feeds 3 11.1% 1 1 0 1 blocklists, VirusTotal or commercial scam feeds user-reports 5 18.5% 2 0 1 2 victim or vigilante reports (forums, Reddit, complaints) spam-feeds 3 11.1% 3 0 0 0 email spam feeds telemetry 2 7.4% 0 0 0 2 security-vendor or platform telemetry / internal index leaked-backend 1 3.7% 1 0 0 0 leaked or seized operator databases prior-dataset 3 11.1% 1 1 1 0 another paper's scam list cluster: cluster papers of 27 2010–2014 2015–2021 2022–2024 2025–2026* meaning -------------- ------ ----- --------- --------- --------- ---------- ---------------------------------------------------------------- payment 6 22.2% 3 0 1 2 bank accounts, acquiring banks, merchant IDs or wallet addresses phone 4 14.8% 1 1 1 1 phone numbers or messenger contacts whois 6 22.2% 1 2 1 2 WHOIS registrant fields content 7 25.9% 4 2 0 1 HTML / text similarity or store templates screenshot 4 14.8% 0 2 1 1 perceptual hash of screenshots hosting 2 7.4% 0 1 0 1 IP / DNS / hosting infrastructure analytics-id 1 3.7% 0 1 0 0 shared analytics or tag-manager IDs redirect-graph 2 7.4% 2 0 0 0 connected components of redirect chains affiliate-id 1 3.7% 1 0 0 0 affiliate identifiers embedded in the page none 8 29.6% 1 1 4 2 no site-to-operator grouping gt: gt papers of 27 2010–2014 2015–2021 2022–2024 2025–2026* meaning ------------- ------ ----- --------- --------- --------- ---------- ----------------------------------------------------------------------- manual 19 70.4% 4 5 6 4 researchers looked at the sites classifier 9 33.3% 1 3 1 4 trained classifier (validated on a labelled sample) heuristic 11 40.7% 7 1 3 0 keyword / structural rules list-as-label 6 22.2% 0 3 2 1 a blocklist, VirusTotal or trust score used as the label source 4 14.8% 4 0 0 0 the source itself is the label (leaked back end, curated list, reports) llm 1 3.7% 0 0 0 1 LLM as the classifier cloak: cloak papers of 27 2010–2014 2015–2021 2022–2024 2025–2026* meaning --------- ------ ----- --------- --------- --------- ---------- ------------------------------------------------------------------------------ measured 7 25.9% 2 3 1 1 cloaking or victim-only serving measured (two vantages or identities compared) addressed 8 29.6% 3 3 1 1 crawler set up to look like a victim, cloaking not measured none 12 44.4% 4 1 4 3 not stated harm: harm papers of 27 2010–2014 2015–2021 2022–2024 2025–2026* meaning ------------------ ------ ----- --------- --------- --------- ---------- ----------------------------------------------------------------------------------- order-numbers 2 7.4% 2 0 0 0 sequential order numbers (purchase pairs) test-purchase 4 14.8% 4 0 0 0 completed purchases leaked-revenue 5 18.5% 4 0 0 1 leaked or partner revenue / transaction records traffic-stats 5 18.5% 3 1 0 1 visitor counts from exposed server stats or traffic panels, times a conversion rate wallet-flows 4 14.8% 0 0 3 1 inflows to scam wallet addresses telemetry-exposure 1 3.7% 0 0 0 1 users exposed / reaching checkout in vendor telemetry victim-reports 2 7.4% 1 0 0 1 losses stated in victim complaints or police records none 11 40.7% 1 6 3 1 not measured interact: interact papers of 27 2010–2014 2015–2021 2022–2024 2025–2026* meaning --------------- ------ ----- --------- --------- --------- ---------- ---------------------------------------- purchases 4 14.8% 4 0 0 0 bought from the scam calls 2 7.4% 1 1 0 0 phoned the scammers chats 1 3.7% 0 0 1 0 messaged the scammers checkout-no-pay 3 11.1% 2 0 0 1 went to the payment step and stopped accounts 1 3.7% 0 0 0 1 created victim accounts on the scam site ad-clicks 3 11.1% 0 2 1 0 clicked ads (cost to advertisers) form-filling 1 3.7% 0 1 0 0 completed survey forms automatically none 14 51.9% 4 3 4 3 no interaction grouped sites into operators/campaigns (cluster != none): 19 of 27 (70.4%) measured blocklist / feed coverage or lag against their own scam set: 14 of 27 (51.9%) measured some form of harm (harm != none): 16 of 27 (59.3%) cloaking measured or addressed: 15 of 27 (55.6%); measured: 7 paid, phoned or messaged the scammers: 6 of 27 per paper: 2010 dissecting-one-click-frauds ch=user-reports | cl=payment+phone+whois | gt=source+manual | bl=y | cloak=none | harm=victim-reports | int=none 2011 click-trajectories-end-to-end-analysis-of-the-sp ch=spam-feeds | cl=content+payment | gt=heuristic+manual | bl=n | cloak=addressed | harm=test-purchase | int=purchases 2011 measuring-and-analyzing-search-redirection-attac ch=search | cl=redirect-graph | gt=heuristic+source | bl=y | cloak=addressed | harm=traffic-stats | int=checkout-no-pay 2011 show-me-the-money-characterizing-spam-advertised ch=spam-feeds | cl=content | gt=heuristic | bl=n | cloak=none | harm=order-numbers+test-purchase+leaked-revenue+traffic-stats | int=purchases 2012 priceless-the-role-of-payments-in-abuse-advertis ch=user-reports+prior-dataset | cl=payment+content | gt=manual+heuristic | bl=n | cloak=addressed | harm=test-purchase | int=purchases+calls 2012 tasters-choice-a-comparative-analysis-of-spam-fe ch=spam-feeds+feeds | cl=affiliate-id | gt=heuristic | bl=y | cloak=none | harm=leaked-revenue | int=none 2012 pharmaleaks-understanding-the-business-of-online ch=leaked-backend | cl=none | gt=source | bl=n | cloak=none | harm=leaked-revenue | int=none 2014 a-nearly-four-year-longitudinal-study-of-search- ch=search | cl=redirect-graph | gt=source+heuristic | bl=n | cloak=measured | harm=none | int=none 2014 search-seizure-the-effectiveness-of-intervention ch=search | cl=content | gt=heuristic+classifier+manual | bl=n | cloak=measured | harm=order-numbers+test-purchase+leaked-revenue+traffic-stats | int=purchases+checkout-no-pay 2017 dial-one-for-scam-a-large-scale-analysis-of-tech ch=ad-networks | cl=phone+whois | gt=heuristic+manual | bl=y | cloak=measured | harm=traffic-stats | int=calls 2018 surveylance-automatically-detecting-online-surve ch=search+ad-networks | cl=whois+screenshot | gt=classifier+manual | bl=n | cloak=addressed | harm=none | int=form-filling 2018 betrayed-by-your-dashboard-discovering-malicious ch=feeds+prior-dataset | cl=analytics-id | gt=list-as-label | bl=y | cloak=none | harm=none | int=none 2018 exposing-search-and-advertisement-abuse-tactics- ch=search+search-ads | cl=hosting+content | gt=classifier | bl=y | cloak=addressed | harm=none | int=none 2019 what-you-see-is-not-what-you-get-discovering-and ch=ad-networks | cl=screenshot | gt=manual+list-as-label | bl=y | cloak=measured | harm=none | int=ad-clicks 2020 when-push-comes-to-ads-measuring-the-rise-of-mal ch=ad-networks | cl=content | gt=list-as-label+manual | bl=y | cloak=addressed | harm=none | int=ad-clicks 2021 where-are-you-taking-me-understanding-abusive-tr ch=ad-networks+search | cl=none | gt=manual+classifier | bl=y | cloak=measured | harm=none | int=none 2023 beyond-phish-toward-detecting-fraudulent-e-comme ch=user-reports | cl=none | gt=classifier+manual | bl=y | cloak=none | harm=none | int=none 2023 evolving-bots-the-new-generation-of-comment-bots ch=social | cl=none | gt=manual+list-as-label | bl=n | cloak=none | harm=none | int=none 2023 double-and-nothing-understanding-and-detecting-c ch=ct-logs | cl=whois+payment+screenshot | gt=heuristic+manual | bl=y | cloak=addressed | harm=wallet-flows | int=none 2023 trident-towards-detecting-and-mitigating-web-bas ch=ad-networks | cl=none | gt=manual+list-as-label | bl=n | cloak=none | harm=none | int=ad-clicks 2024 give-and-take-an-end-to-end-investigation-of-giv ch=social+prior-dataset | cl=none | gt=heuristic+manual | bl=n | cloak=measured | harm=wallet-flows | int=none 2024 like-comment-get-scammed-characterizing-comment- ch=social | cl=phone | gt=heuristic+manual | bl=y | cloak=none | harm=wallet-flows | int=chats 2025* ctrlaltdeceive-quantifying-user-exposure-to-onli ch=feeds+telemetry | cl=none | gt=list-as-label+classifier+manual | bl=y | cloak=addressed | harm=telemetry-exposure | int=none 2025* scammagnifier-piercing-the-veil-of-fraudulent-sh ch=new-registrations | cl=payment+whois | gt=classifier+manual | bl=n | cloak=none | harm=leaked-revenue | int=checkout-no-pay 2025* nokescam-understanding-and-rectifying-non-sense- ch=telemetry+user-reports | cl=content+whois | gt=classifier+manual | bl=n | cloak=measured | harm=victim-reports+traffic-stats | int=none 2025* the-poorest-man-in-babylon-a-longitudinal-study- ch=ct-logs | cl=hosting+screenshot+phone+payment | gt=llm+manual | bl=y | cloak=none | harm=wallet-flows | int=accounts 2026* loki-proactively-discovering-online-scams-by-min ch=search+user-reports | cl=none | gt=classifier | bl=n | cloak=none | harm=none | int=none === PASS E — extraction fields over the IN papers (enums; nulls and sentinels shown, never counted as answers) === ethics.reviewOutcome: none-mentioned 14; approved 4; NULL (no ethics object) 4; explicitly-discussed-no-review 2; not-required 2; sought-outcome-unstated 1 (of 27) ethics.notifiedAffectedParties: no 11; not-stated 6; NULL (no ethics object) 4; yes 3; partial 2; not-applicable 1 (of 27) artifacts.availability: none-mentioned 13; public 11; promised-not-yet-available 2; on-request 1 (of 27) classification[].method == 'llm': 2 of 27 IN — WWW/2025/the-poorest-man-in-babylon-a-longitudinal-study-of-cryptocurrency-investment-sca [website-category: Llama3:70b + GPT-4 hybrid] ; NDSS/2026/loki-proactively-discovering-online-scams-by-mining-toxic-search-queries [mobile-app: FLAN-T5-XXL] classification[].method == 'llm' among the 67 CONTEXT papers: 7 — 2025 sheeps-clothing-wolfish-intent-automated (ads); 2026 ai-in-the-loop-privacy-preserving-real-t (llm); 2026 love-lies-and-language-models-investigat (llm); 2026 ctphishcapture-uncovering-credential-the (phishing-boundary); 2026 chameleon-channels-measuring-youtube-acc (platform); 2025 pirates-of-charity-exploring-donation-ba (platform); 2025 fishing-for-smishing-understanding-sms-p (telephony) corpus-wide papers with an 'llm' classification tuple: 177 of 5859; by year: 2023 2 of 719 (0.3%), 2024 27 of 690 (3.9%), 2025* 77 of 770 (10.0%), 2026* 71 of 415 (17.1%) IN papers from 2025–2026*: 5; of them with an LLM classifying scam sites (hand: gt includes llm): 1; with any 'llm' tuple: 2 isEmpirical: true 27 (of 27) === PASS F — detection[] tuples of the IN papers (prevalence is a model summary: every figure the page uses is re-checked against paper.cols.txt by verify_scam_figures.mjs) === # CCS/2010/dissecting-one-click-frauds (5 tuples) - One Click Fraud incidents | metric: incident count | prevalence: 2,140 incident reports | quote: "All in all, we gathered 2,140 incident reports." - Miscreant groups | metric: connected clusters | prevalence: 105 clusters and 26 singletons | quote: "Excluding the 26 singletons, the whole graph G contains 105 connected subgraphs ("clusters")." - Malware on fraud websites | metric: share of downloaded fraud websites | prevalence: 14 websites contained malware | quote: "We found that a small number (14) of One Click Fraud websites contained some malware." - Blacklist associations | metric: share of IP addresses matching Google Safe Browsing Firefox 2 | prevalence: 16% of IP addresses | quote: "when comparing IP addresses of the servers that host One Click Frauds to the IP addresses of the servers in the blacklist, we see a significant (16%) number of hits." - Fraud concentration | metric: cumulative proportion of frauds | prevalence: Top 8 groups responsible for more than half the frauds | quote: "Taking into account similarities exhibited by WHOIS records, and malware deployment, we observe that the top 8 groups are responsible for more than half of the frauds we have collected." # IEEE-SP/2011/click-trajectories-end-to-end-analysis-of-the-spam-value-chain (7 tuples) - spam-advertised URL coverage | metric: share of received URLs covered | prevalence: 98.1% of received URLs | quote: "URLs covered 950,716,776 (98.1%)" - DNS infrastructure | metric: distinct DNS records and infrastructure | prevalence: null | quote: "The crawler periodically queries new records until it converges on a set of distinct results." - web redirection | metric: share of crawled URLs redirecting | prevalence: 32% of crawled URLs redirected at least once | quote: "To explain further, 32% of crawled URLs in our data redirected at least once." - affiliate-program storefronts | metric: distinct registered domains and affiliate programs | prevalence: 45 affiliate programs across 69,002 registered domains | quote: "We identify a total of 45 affiliate programs for the three categories combined, that are advertised via 69,002 distinct registered domains." - payment-bank concentration | metric: share of spam-advertised goods using concentrated acquirers | prevalence: Just three banks provided payment servicing for over 95% | quote: "This situation is dramatically reflected in Figure 5, which shows that just three banks provide the payment servicing for over 95% of the spam-advertised goods in our study." - purchase fulfillment | metric: authorized and settled purchases | prevalence: 120 attempted; 76 authorized and 56 settled | quote: "We attempted 120 purchases, of which 76 authorized and 56 settled." - supplier sharing | metric: number of physical-goods suppliers | prevalence: 13 different suppliers | quote: "Fulfillment for physical goods was sourced from 13 different suppliers (as determined by declared shipper and packaging)." # USENIX/2011/measuring-and-analyzing-search-redirection-attacks-in-the-illicit-online-prescri (8 tuples) - search-redirection attacks | metric: share of search-result URIs | prevalence: 44,503 URIs, 32% of results, actively redirected | quote: "We observed 44 503 of these URIs to be compromised websites (source infections) actively redirecting to pharmacies, 32% of the total." - infected source domains | metric: number of unique infected domains | prevalence: 7,298 source websites across both datasets | quote: "We identified 7 298 source websites from both data sets that had been infected to take part in search-redirection attacks" - infection persistence | metric: median infection lifetime | prevalence: 47 days; 16% remained in results throughout the 192-day sample | quote: "The median lifetime of infected websites is 47 days; this can be seen in the graph by observing where S(t) = 0.5." - redirection-chain connectivity | metric: share of infected domains in giant component | prevalence: 96% of infected domains were in the largest connected component | quote: "The largest connected component G0 contains 96% of all infected domains, 90% of the redirection domains and 92% of the pharmacy domains" - affiliate communities | metric: number of detected communities | prevalence: 73 communities; seven largest comprised more than half the nodes | quote: "The community detection algorithm identifies a total of 73 distinct communities." - blacklist coverage | metric: share of domains appearing on at least one blacklist | prevalence: 95% of source infections appeared on no blacklist; over two-thirds of pharmacies appeared on one | quote: "source infections are not widely reported by any of the blacklists (95% do not appear on a single blacklist), but around half of the redirects are found on at least one blacklist" - search-to-sale conversion | metric: estimated conversion rate | prevalence: Between 0.3% and 3%; calculated lower bound approximately 3.2% before conservative completion adjustment | quote: "Conversion ≈ = 3.2% ." - payment-processing concentration | metric: monthly visits to payment sites | prevalence: 94 pharmacies pointed to 21 payment-processing websites receiving 855,000 monthly visits | quote: "We found that 94 of these websites in fact pointed to one of 21 different payment processing websites." # USENIX/2011/show-me-the-money-characterizing-spam-advertised-revenue (4 tuples) - Spam-advertised order volume | metric: monthly attempted orders | prevalence: over 82,000 pharmaceutical orders and over 37,000 software orders monthly | quote: "Together, these reflect a monthly volume of over 82,000 pharmaceutical orders and over 37,000 software orders." - Customer basket contents | metric: distinct products and cart additions | prevalence: 289 distinct products; 38% of cart additions were outside the ED and sexually-related category | quote: "We observed 289 distinct products ... indeed, 38% of all items added to the cart were not in this category." - Cart-addition conversion | metric: visit-to-cart-addition conversion rate | prevalence: 0.5% overall on an IP basis | quote: "Based on the total number of visitors where we have referrer information, the conversion percentage on an IP basis is 0.5%." - Geographic customer distribution | metric: share of cart additions by region | prevalence: 80% originated from the U.S. and Canada; Europe contributed 6% | quote: "the vast majority of shopping cart insertions originate from the U.S. and Canada (80%) or Europe (6%)" # CCS/2012/priceless-the-role-of-payments-in-abuse-advertised-goods (6 tuples) - affiliate-program payment relationships | metric: number of acquiring banks and merchant relationships | prevalence: 40 affiliate programs and 30 acquiring banks | quote: "In aggregate, we executed 429 orders from 25 pharmaceutical and 15 software affiliate programs. These in turn were processed through 30 acquiring banks." - payment intervention effects | metric: outcome distribution after descriptor inactivity | prevalence: 69% of complaint cases moved to a new bank; nearly 21% had no successful subsequent purchase | quote: "Only 11% of subsequent purchases to programs that received complaints were processed on descriptors at the same bank, while 69% were processed on descriptors at a new bank and nearly 21% ... were not successful." - payment processing concentration | metric: share of purchases by bank | prevalence: Most pharmaceutical purchases went through twelve banks | quote: "When purchasing from all of the affiliate programs, most of the purchases go through just twelve banks with the remaining banks processing fewer than ten purchases." - order refusals | metric: successful authorizations versus refusals | prevalence: 429 successful attempts and 247 refusals | quote: "Together, this combined dataset includes 676 ordering attempts, of which 429 were successful, covering over two years of activity." - payment miscoding | metric: fraction of pharmaceutical transactions miscoded | prevalence: Almost 50% over the past eight months; closer to 70% in the last two months | quote: "Indeed, almost 50% of pharmaceutical transactions over the past eight months are miscoded (e.g, as Cosmetics, Grocery Stores, etc.), and in the last two months this fraction is closer to 70%." - undercover-purchase filtering | metric: number of programs requiring phone verification | prevalence: Fourteen pharmaceutical programs required phone confirmation over two years | quote: "As early as 2010, we experienced that some pharmaceutical programs (fourteen all told over the last two years) would hold an order until they had called and confirmed our order over the phone." # IMC/2012/tasters-choice-a-comparative-analysis-of-spam-feeds (7 tuples) - spam-advertised registered domains | metric: feed volume and unique registered domains | prevalence: over a billion messages distributed over three months | quote: "Using this corpus, corresponding to over a billion messages distributed over three months, we characterize the relationships between its constituent data sources." - feed purity | metric: fraction of feed domains meeting purity indicators | prevalence: DNS and HTTP success rates varied substantially across feeds | quote: "Table 2 shows several purity indicators for each feed. The first three (DNS, HTTP, and Tagged) are positive indicators." - spam-domain coverage | metric: fraction of aggregate domains covered | prevalence: 60% of live domains and 19% of tagged domains were exclusive to one feed | quote: "Across our feeds, 60% of all live domains and 19% of all tagged domains were exclusive to a single feed." - spam volume proportionality | metric: relative domain volume and rank | prevalence: null | quote: "The provider reported back to us the number of messages (normalized) containing each spam domain, as seen by their incoming mail servers over five days during the measurement period." - spam campaign timing | metric: relative first and last appearance time | prevalence: Hu saw over 75% of domains within one day and 95% within three days; dbl saw over 95% within one day | quote: "The Hu feed sees over 75% of the domains within a day after they appear in any feed, and 95% within three days; dbl is delayed even less, with over 95% appearing on the blacklist within a day." - affiliate-program coverage | metric: number and fraction of affiliate programs covered | prevalence: 45 leading affiliate programs | quote: "The prior Click Trajectories measurement effort identified 45 leading affiliate programs specializing in pharmaceutical sales, replica luxury goods, and “OEM” software." - RX-Promotion affiliate coverage | metric: distinct affiliate identifiers covered | prevalence: 846 distinct affiliate identifiers | quote: "This embedding allowed us to extract affiliate identifiers and map them to domains. In total, we were able to identify 846 distinct affiliate identifiers." # USENIX/2012/pharmaleaks-understanding-the-business-of-online-pharmaceutical-affiliate-progra (7 tuples) - customer orders and revenue | metric: counts of unique customers, billed orders, and revenue | prevalence: over 2M sales records and over $170M in settled revenue | quote: "over 2M sales records, with over 140 linked tables, coupled with several GB of related metadata" - repeat customer purchasing | metric: share of average program revenue | prevalence: repeat orders constituted 27% of GlavMed and 38% of SpamIt average revenue | quote: "repeat orders are an important part of the business, constituting 27% and 38% of average program revenue for GlavMed and SpamIt, respectively." - drug demand | metric: order volume and program revenue by drug category | prevalence: ED products generated 75% of GlavMed revenue and 82% of SpamIt revenue | quote: "ED and Related 580K (73%) $55M (75%)" - affiliate revenue concentration | metric: fraction of program revenue by affiliate percentile | prevalence: the top 10% of affiliates accounted for 75-90% of program revenue | quote: "The graph shows that just 10% of the highest-revenue affiliates account for 75-90% of total program revenue" - affiliate commissions | metric: annualized commission distribution | prevalence: median annualized commissions were $292, $3,320, and $428 | quote: "the median annualized affiliate commissions for GlavMed, SpamIt, and RX-Promotion are just $292, $3,320, and $428, respectively." - program costs and profit | metric: gross margin and net revenue as percentages of gross revenue | prevalence: RX-Promotion net revenue was 16.3% of gross revenue during March-September 2010 | quote: "Overall, the net revenue for this period-the profit returned to the affiliate program owners-is just 16.3% of gross revenue." - payment-provider concentration | metric: share of total revenue by payment provider | prevalence: three providers accounted for 84% of GlavMed and SpamIt revenue | quote: "Together these three providers are responsible for 84% of all revenue for GlavMed and SpamIt." # CCS/2014/a-nearly-four-year-longitudinal-study-of-search-engine-poisoning (5 tuples) - search-engine poisoning | metric: share of search results | prevalence: Search-redirection attacks rose from around 30% in late 2010 to nearly 60% in late 2012. | quote: "search-redirection attacks have steadily grown to take over a larger share of results (rising from around 30% in late 2010 to a peak of nearly 60% in late 2012)" - search redirection | metric: share of results | prevalence: 38.8% of 1,602,087 results were active search-redirections. | quote: "Out of those, more than 38% are active redirections; on any given day between 8.7% and 61.7% of the obtained results actively redirect." - cloaking | metric: share of unclassified results | prevalence: 19.5% of unclassified results appeared malicious according to VirusTotal. | quote: "When we observe a difference in the HTML returned between the two treatments, we infer there might have been cloaking." - source-infection persistence | metric: median survival time | prevalence: The median survival time for infections was 19 days; 1.7% remained infected at least two years. | quote: "One-third last five days or less, while the median survival time for infections is 19 days. ... 459 websites, 1.7% of the total, remain infected for at least two years!" - traffic-broker concentration | metric: number of autonomous systems hosting brokers | prevalence: Only seven autonomous systems supported more than ten traffic brokers daily. | quote: "It turns out that only 7 ASes (3 in the US, 3 in Germany, 1 in the Netherlands) support more than 10 traffic brokers every day." # IMC/2014/search-seizure-the-effectiveness-of-interventions-on-seo-campaigns (8 tuples) - poisoned search results | metric: number and percentage of PSRs | prevalence: 2.7M PSRs across 16 verticals | quote: "we detected 2.7M PSRs, across all verticals, using 27K doorways from unique domains and sending users to 7,404 different stores" - cloaking | metric: detected cloaked search results | prevalence: null | quote: "Dagger uses heuristics to detect cloaking by examining semantic differences between versions of the same page fetched first as a user and then as a search engine crawler" - iframe cloaking | metric: share of pages using iframe cloaking | prevalence: described as pervasive in counterfeit luxury | quote: "We found the use of iframe cloaking to be pervasive within the domain of counterfeit luxury" - counterfeit storefronts | metric: number of detected storefronts | prevalence: 7,484 stores | quote: "If either of the heuristics succeed, we treat the landing site as a counterfeit luxury store advertised through search poisoning." - SEO campaigns | metric: number of campaigns and held-out accuracy | prevalence: 52 campaigns; 86.8% average accuracy | quote: "The average accuracy on held-out data was 86.8% for multiway classification of Web pages into 52 different SEO campaigns." - order volume | metric: estimated orders over time | prevalence: 1,408 test orders from 290 stores | quote: "This technique exploits the fact that stores use monotonically increasing order numbers, where the difference between order numbers represents the total number of orders created over the time delta between orders." - user traffic | metric: visits and HTML page fetches | prevalence: 647 storefronts in 12 campaigns | quote: "AWStats is a Web analytics tool [1] that uses a Web site's server logs to report aggregated visitor information" - domain seizures | metric: number of observed seizures and seized domains | prevalence: 290 seizures directly observed during eight months | quote: "From the PSRs crawled, we directly observed 290 seizures over our eight-month period: 214 seized by GBC and 76 by SMGPA." # NDSS/2017/dial-one-for-scam-a-large-scale-analysis-of-technical-support-scams (8 tuples) - technical-support scam domains and phone numbers | metric: unique domains and phone numbers | prevalence: 8,698 unique domain names and 1,581 phone numbers | quote: "in a period of 250 days, we discover 8,698 unique domain names involved in technical support scams... urging them to call one of the 1,581 collected phone numbers." - scam-domain lifetime | metric: domain lifetime in days | prevalence: 43% of domains were reachable for up to three days | quote: "43% of the domains are reachable for up to three days." - scam campaigns | metric: campaign lifetime | prevalence: 69% of campaigns had a lifetime of less than 50 days | quote: "69% of the campaigns have a lifetime of less than 50 days." - blacklist coverage | metric: percentage of scam domains or IPs detected | prevalence: Only 108 of 1,524 domains and 28 of 685 IP addresses were already blacklisted | quote: "out of 1,524 scam domains, only 108 (7%) were blacklisted... From the 685 resolved IP addresses, only 28 (4%) were already present" - scam-page techniques | metric: percentage of collected scam pages | prevalence: 87% used HTML audio tags and 49% used very long alert messages | quote: "Lastly, we observed that 87% of the discovered scam pages were using HTML audio tags" - scammer social-engineering techniques | metric: percentage of 60 calls | prevalence: Stopped services/drivers appeared in 67% of calls | quote: "Stopped Services/Drivers 67" - call duration and requested price | metric: average duration and average requested price | prevalence: 17 minutes and $290.9 | quote: "The average duration of that interval is 17 minutes... The average support price across all support packages and all scammers is $290.9" - call-center size | metric: average number of reachable scammers | prevalence: Average call center size was 11 scammers | quote: "The average number of volunteers who were able to speak with a scammer across all ten studied phone numbers was 11" # IEEE-SP/2018/surveylance-automatically-detecting-online-survey-scams (4 tuples) - survey scam gateways | metric: true-positive and false-positive rates | prevalence: 8,623 survey gateways identified among 2,301,733 crawled URLs | quote: "S URVEYLANCE reported 8,623 survey gateways by crawling 2,301,733 URLs." - survey publisher exposure | metric: completed surveys | prevalence: 131,277 surveys completed from 318,219 survey publisher URLs | quote: "we were able to fill out 131,277 unique surveys using three different browsers (Chrome, Firefox, and Internet Explorer)." - malicious binary delivery | metric: unique binaries and distinct files | prevalence: 2,612 unique MD5s and 954 distinct binaries; 521 unknown to VirusTotal | quote: "we collected 2,612 unique binaries (unique MD5s) by visiting 22,057 URLs that delivered a binary, yielding 954 distinct polymorphic files." - browser-dependent information flow | metric: share interacting with WebStorage APIs | prevalence: 144 of 200 randomly selected gateways interacted with WebStorage APIs | quote: "From the 200 randomly selected survey gateways, we observed that 144 (72%) of them were interacting with browser WebStorage APIs" # WWW/2018/betrayed-by-your-dashboard-discovering-malicious-campaigns-via-web-analytics (8 tuples) - analytics identifiers on malicious web pages | metric: unique analytics IDs | prevalence: 9,395 IDs associated with malicious pages | quote: "we were able to extract 9,395 analytics IDs associated with malicious content." - malicious domains sharing analytics IDs | metric: newly discovered live websites | prevalence: 14,267 live websites, 76.5% previously unseen | quote: "Overall, we were able to discover 14,267 live websites containing malicious analytic IDs, 76.5% of which were new, previously unseen domains" - technical-support scam analytics campaigns | metric: unique Google Analytics IDs | prevalence: 872 IDs across 3,185 domains | quote: "we were able to extract 872 unique Google Analytics IDs across 3,185 domain names." - analytics identifiers in browser extensions | metric: extensions containing Google Analytics IDs | prevalence: 120 of 255 downloaded extensions | quote: "Out of 255 extensions that we could successfully download and unpack, we found Google Analytics IDs on 120 extensions." - analytics identifiers in malicious Android apps | metric: unique identifiers and APK samples | prevalence: 18,734 identifiers across 273,232 malware APKs | quote: "Overall, from 477,829 malicious APKs, we retrieved 18,734 unique analytics identifiers over 273,232 samples." - malware discovery through analytics-ID matching | metric: flagged APKs | prevalence: 100,379 newer APKs flagged as malware | quote: "By matching analytics IDs found on previous malicious samples, 100,379 unique APKs were flagged as malware." - phishing campaigns | metric: campaigns identified | prevalence: 13 phishing campaigns | quote: "After applying Algorithm 1 to the two-week collection of “unknown” websites ... we could identify 13 phishing campaigns" - analytics-based actor deanonymization | metric: malicious actors deanonymized | prevalence: 59 malicious actors behind VirusTotal URLs | quote: "we were able to deanonymize 59 malicious actors behind VirusTotal URLs by finding public WHOIS records for other domains sharing the same analytics IDs." # WWW/2018/exposing-search-and-advertisement-abuse-tactics-and-infrastructure-of-technical (5 tuples) - technical-support-scam domains | metric: unique FQDNs hosting TSS content | prevalence: 9,221 FQDNs, including 5,225 found through network-level amplification | quote: "In all, the total number of unique FQDNs hosting TSS content, |Ff −t ss | = 9,221, with 3,996 TSS FQDNs coming from the final-landing websites in search listings and 5,225 additional TSS FQDNs" - search-result and advertisement TSS pollution | metric: share of URIs leading to TSS websites | prevalence: 71.79% of advertisement URIs and 54.26% of search-result URIs | quote: "Among the AD URIs, 10,299 (71.79%) were observed as leading to TSS websites." - DNS-based TSS infrastructure amplification | metric: amplification factor and additional domains | prevalence: 5,225 additional TSS FQDNs; maximum amplification factor 275 | quote: "around 60% domains had A(d) ≤ 50 while the remaining 40% domains had A(d) > 50, with the maximum A(d) value equal to 275." - support-domain redirection infrastructure | metric: additional support domains | prevalence: 2,435 support domains | quote: "There were an additional, 2,435 support domains found." - blacklist coverage of TSS domains | metric: cumulative FQDN coverage | prevalence: 26.8% of discovered FQDNs were covered | quote: "Cumulatively, these lists cover only 26.8% FQDNs, that were found to be involved in TSS by our system." # IMC/2019/what-you-see-is-not-what-you-get-discovering-and-tracking-social-engineering-att (5 tuples) - SEACMA campaigns | metric: number of campaigns | prevalence: 108 of 130 clusters represented SE attack campaigns | quote: "Out of the 130 clusters we obtained, 108 clusters appeared to be clearly representing various SE attack campaigns." - SE attack instances | metric: number of attack instances | prevalence: 28,923 SE attack instances | quote: "Overall, these campaigns included 11,341 unique domains of publisher sites and 28,923 SE attack instances" - SEACMA publisher prevalence | metric: share of crawled websites | prevalence: 11,341 of 70,541 publisher websites (16%) | quote: "Among the 70,541 publisher websites that we crawled, only 11,341 (16%) were observed to host SEACMA ads." - Google Safe Browsing detection | metric: final detection rate | prevalence: 16.21% of attack domains after two months | quote: "even after two months time, GSB detected only a small percentage (16.2% overall) of SE attack domains." - Downloaded malicious files | metric: number marked malicious | prevalence: more than 9,000 of 9,476 milked files | quote: "In the end, we found that more than 9,000 of the milked files were marked as malicious." # IMC/2020/when-push-comes-to-ads-measuring-the-rise-of-malicious-push-advertising (5 tuples) - web-push notification collection | metric: number of WPN messages | prevalence: 21,541 push notification messages | quote: "we were able to collect a total of 21,541 push notification messages, including 12,441 notifications for the desktop environment and 9,100 for the mobile environment." - WPN advertising campaigns | metric: number of campaigns and ads | prevalence: 572 WPN ad campaigns and 5,143 WPN ads | quote: "PushAdMiner identified 572 WPN ad campaigns and a total of 5,143 WPN ads related to these campaigns." - malicious WPN advertising | metric: share of WPN ads confirmed malicious | prevalence: 51% of all WPN ads | quote: "Furthermore, PushAdMiner found 51% of all WPN ads as malicious." - URL-blacklist detection gaps | metric: share of URLs detected | prevalence: VirusTotal detected 11.31% after one month; GSB still flagged 1% | quote: "After one month, we submitted the same set of URLs once again, and we found that 1,388 (11.31%) of them were then detected by VT, though GSB still only flagged 1% of them." - ad-blocker effectiveness | metric: blocked Service Worker scripts and requests | prevalence: All tested blockers blocked 0 Service Worker scripts | quote: "both ad blocking mechanisms failed to block the registration of Service Worker scripts related to ad networks that support WPN ads" # WWW/2021/where-are-you-taking-me-understanding-abusive-traffic-distribution-systems (6 tuples) - user differentiation | metric: relative page counts and NRD scores | prevalence: 81% more malicious pages using six profiles than the best single profile | quote: "ODIN finds 81% more malicious and 96% more suspicious landing pages, compared to visiting pages only using the user profile which experienced the most malice." - IP-based cloaking | metric: malicious-page count | prevalence: More than twice as many malicious pages with multiple IP addresses | quote: "We find that using multiple IP addresses leads us to find more than twice as many malicious pages." - TDS infrastructure overlap | metric: traffic-broker domain overlap | prevalence: 19.2% to 44.1% overlap between non-pharmacy TDSs | quote: "we observe 19.2% to 44.1% traffic broker domains overlap between non-pharmacy TDSs." - malicious destination prevalence | metric: share of collected pages | prevalence: 26.5% after errors: 2.6% malicious, 4.0% suspicious, or 20.0% illicit | quote: "After removing errors, we find that 26.5% of all collected pages are malicious (2.6%), suspicious (4.0%) or illicit (20.0%)." - blacklist coverage | metric: coverage and detection delay | prevalence: GSB missed mobile-targeted malicious pages 76% of the time | quote: "GSB does not include malicious landing pages shown to mobile users 76% of the time." - malicious redirection prediction | metric: 10-fold accuracy and F1 score | prevalence: 99.0% accuracy and 92.7% F1 | quote: "Our classifier achieves an average 99.0% accuracy and 92.7% F 1 score (evaluated using 10-fold crossvalidation)" # IEEE-SP/2023/beyond-phish-toward-detecting-fraudulent-e-commerce-websites-at-scale (7 tuples) - Fraudulent e-commerce websites | metric: dataset prevalence | prevalence: 6,127 FCWs among 9,114 live URLs | quote: "We leverage social media to collect 6,127 FCWs that are actively luring victims." - FCW categories | metric: share of submissions | prevalence: Fake online shopping 60.38%; pet scams 20.52%; charity 6.04% | quote: "fake online shopping scams are the most common at 60.38% in our dataset, pet scams are 20.52%" - Existing blocklist coverage | metric: detection rate | prevalence: Google Safe Browsing detected 0.46% of FCWs | quote: "APWG and GSB detect only 25 and 10 FCWs within our dataset, respectively." - B EYOND P HISH detection | metric: detection rate and false-positive rate | prevalence: 98.34% detection rate and 1.34% false-positive rate in the user study | quote: "The model achieves a false positive rate of 2.46% and a 94.88% detection rate" - In-the-wild FCW detection | metric: accuracy | prevalence: 98.38% accuracy over 2,223 submissions | quote: "After examining the results, we find that B EYOND P HISH predicted the correct label 98.38% of the time." - Alexa-domain false positives | metric: false-positive rate | prevalence: 1.21% for ranks 10,000–20,000 and 0.92% for the bottom 10,000 | quote: "Comparing B EYOND P HISH's prediction to the experts' assigned labels yields a false positive rate of 1.21%." - Partner-dataset detection | metric: detection rate and false-positive rate | prevalence: 94.88% detection rate and 2.46% false-positive rate | quote: "Testing our model on the Palo Alto Networks data indicates a false positive rate of 2.46% and a high detection rate of 94.88%." # IMC/2023/evolving-bots-the-new-generation-of-comment-bots-and-their-underlying-scam-campa (5 tuples) - Social scam bots | metric: number of verified SSB accounts | prevalence: 1,134 SSBs | quote: "From our dataset of 45,322 videos ... we were able to obtain 1,134 active SSB accounts." - Scam campaigns | metric: number of confirmed scam campaigns | prevalence: 72 scam campaigns | quote: "From 74 SLDs, a total of 72 SLDs were confirmed to be scam domains." - Video infection by SSBs | metric: share of crawled videos | prevalence: 14,380 of 45,322 videos (31.73%) | quote: "From our dataset of 45,322 videos 14,380 (31.73%) were infected by one or more SSBs" - YouTube account termination | metric: share of SSBs terminated | prevalence: 47.9% over six months | quote: "Over the course of 7 monthly examinations (a period of 6 months), 47.9% of SSBs were able to be detected and terminated." - Self-engagement | metric: share of self-engagement instances that were first replies | prevalence: 99.56% | quote: "An overwhelming majority, specifically 99.56% of self-engagement instances, featured an SSB reply as the first reply to the comment." # NDSS/2023/double-and-nothing-understanding-and-detecting-cryptocurrency-giveaway-scams (6 tuples) - cryptocurrency giveaway scams | metric: number of confirmed scam websites | prevalence: 10,079 giveaway scam websites on 3,863 domains | quote: "In a six-month period starting from January 1, 2022, CryptoScamTracker recorded 10,079 giveaway scam websites hosted on a total of 3,863 domains." - scammer wallet transactions | metric: total cryptocurrency and USD value received | prevalence: $24.9M-$69.9M stolen across all studied cryptocurrencies | quote: "We recorded all transactions of each wallet and summed the incoming transactions where the transaction recipient is the scammer's wallet address." - blocklist coverage | metric: share of identified domains labeled by VirusTotal | prevalence: 16.75% of 3,610 domains appeared in VirusTotal blocklists | quote: "In total, out of the 3,610 domains discovered by CryptoScamTracker, only 16.75% domains appeared in VT's blocklists." - scam website lifespan | metric: lifespan in hours | prevalence: Half of giveaway websites had lifespans of 26.18 hours | quote: "We observe that half of the giveaway websites have a short lifespan of 26.18 hours." - visual template reuse | metric: number of clusters and screenshots | prevalence: 3,832 screenshots clustered into 1,198 clusters | quote: "Overall, we successfully clustered 3,832 webpage screenshots to 1,198 clusters." - wallet address reuse | metric: share of wallet addresses reused across domains | prevalence: 214 of 2,266 wallet addresses were reused | quote: "Out of a total 2,266 cryptocurrency wallet addresses, there are 214 (9.44%) wallet addresses that CryptoScamTracker encountered on more than one cryptocurrency scam domains." # USENIX/2023/trident-towards-detecting-and-mitigating-web-based-social-engineering-attacks (4 tuples) - SE-ad-related navigation | metric: accuracy, precision, recall, and F1 score | prevalence: 1,479 events resulting in SE attacks among 258,008 JavaScript-initiated navigation events | quote: "we crawled over 100K websites from October 2021 to January 2022 and collected 258,008 unique navigation events initiated by JavaScript (JS), including 1,479 events resulting in SE attacks." - Social-engineering attack types | metric: counts by attack category | prevalence: 857 unwanted-software downloads, 222 dating scams, 177 reward/lottery scams, 148 push notifications, 51 scareware, and 24 tech-support scams | quote: "We have six categories of SE attacks in total." - Concept drift | metric: accuracy, precision, and recall | prevalence: 97.37% accuracy, 98.25% precision, and 97.37% recall | quote: "we evaluated T RIDENT's accuracy over time by testing it on a dataset crawled in October 2022, almost one year after the initial model was trained." - Runtime overhead | metric: median page-load overhead | prevalence: 2.13% median overhead, corresponding to a 0.02-second increase | quote: "The median runtime overhead is 2.13% which results in a 0.02-second increase in the page load time" # IMC/2024/give-and-take-an-end-to-end-investigation-of-giveaway-scam-conversion-rates (6 tuples) - giveaway scam tweets | metric: number of scam tweets | prevalence: 457,248 scam tweets | quote: "Starting from this dataset, we identified all tweets that contained at least one known scam domain, totaling 457,248 tweets from 33,841 distinct accounts." - giveaway scam livestreams | metric: number of livestreams | prevalence: 2,069 livestreams from 1,632 channels | quote: "We ran our measurement pipeline from July 24, 2023 to January 21, 2024 (26 weeks), identifying a total of 2,069 livestreams from 1,632 different channels." - scam landing pages | metric: share of candidate sites classified as scams | prevalence: 4,611 sites manually examined; 343 YouTube-linked domains | quote: "We verified that every final clickthrough or landing page was related to a giveaway scam by (1) ensuring there was a valid cryptocurrency address published on the site." - victim payment conversion | metric: conversion rate | prevalence: 0.12% of tweets; 0.0039% of livestream views | quote: "For Twitter, the conversion rate of tweets (with associated cryptocurrency addresses) to victims was 0.12%-or roughly 1 in 1000 tweets netting a victim." - scam revenue | metric: USD revenue | prevalence: $2.7M Twitter and $1.9M YouTube | quote: "In total, we estimate that Twitter-based giveaway scams yielded $2.7M in revenue." - payment origins | metric: share of payments from centralized exchanges | prevalence: 755 of 1,309 payments (58%) | quote: "For the 1,309 payments across Twitter and YouTube, 755 (58%) came from centralized exchanges." # NDSS/2024/like-comment-get-scammed-characterizing-comment-scams-on-media-platforms (8 tuples) - scam comments | metric: share of captured comments | prevalence: 206,306 (2.34%) of 8,801,224 comments | quote: "Among those comments, we identified 206,306 (2.34%) comments that clearly belong to scammers." - visually similar symbols | metric: share of scam comments | prevalence: 168,938 (81.89%) scam comments | quote: "We discovered 168,938 (81.89%) scam comments contained at least one VSS" - channel-owner impersonation | metric: share of scam comments | prevalence: 28,118 (13.63%) | quote: "Through our Image-based filter, we discovered that 28,118 (13.63%) scammer accounts used the same or similar profile images as the corresponding channel owners." - username abuse | metric: share of scammer accounts | prevalence: 4,802 (45.56%) accounts | quote: "We found at least 4,802 (45.56%) scammer accounts abusing the username for scam activity." - deleted scam comments | metric: share of scam comments | prevalence: 123,506 (59.87%) | quote: "Throughout our dataset, we identified 123,506 (59.87%) scam comments that were deleted." - account deactivation | metric: share of scammer accounts | prevalence: 3,312 (31.42%) of 10,541 accounts | quote: "only 3,312 (31.42%) scam accounts out of the 10,541 scam accounts we captured during the six-month period were deactivated." - scam website blocklist coverage | metric: share of submitted URLs marked suspicious | prevalence: 1 of 24 URLs | quote: "In total, out of the 24 URLs we submitted, only 1 URL is marked as suspicious according to the online blocklist provided by VirusTotal." - cryptocurrency scam proceeds | metric: USD value of transactions | prevalence: $1.11M–$1.99M received by 31 scammers | quote: "Overall, scammers received a total of 67.64 BTC and 36.49 ETH, which is equivalent to $1.11M - $1.99M in USD value" # NDSS/2025/ctrlaltdeceive-quantifying-user-exposure-to-online-scams (7 tuples) - user exposure to scam domains | metric: daily exposed IP addresses and exposed-domain share | prevalence: 149K devices exposed daily; 25.1M IPs observed 415K scam SLDs | quote: "Each day, over 149K devices of the vendor's customers are exposed to online scams." - geographic scam exposure | metric: ratio of exposed daily IP addresses | prevalence: Philippines had 1.2% exposure versus Japan's 0.1% | quote: "The highest exposure happens in Philippines with 1.2% of daily devices in that country encountering scams" - scam domain lifetime | metric: median days | prevalence: 11-day median active time and 1-day median listing delay | quote: "Overall, the median active time for scams domains is 11 days" - advertised scam URLs | metric: share of observations, SLDs, and IPs | prevalence: 13.3% of observations followed advertisements; 59% of ads were social media | quote: "Overall, in 9.2M (13.3%) of all scam observations users followed an advertisement to reach the scam domain" - checkout-page visits | metric: share of shopping-scam IPs | prevalence: 411K desktop IPs, or 4%, reached checkout pages | quote: "over 411K desktop IPs (4% of all IPs visiting a shopping scam) reach a checkout page." - scam-type classification coverage | metric: classified-domain share | prevalence: 330,898 domains classified, representing 54.5% of all scam domains | quote: "The classified domains increase to 330,898 (54.5% ) when adding the domains from the shopping scam ML detector." - clustering-based label expansion | metric: unclassified-domain share | prevalence: Unclassified domains decreased from 59.0% to 24.7% | quote: "Using the expansion we reduce the unclassified domains more than half from 59.0% to 24.7%" # NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns (8 tuples) - fraudulent shopping websites | metric: number of identified websites | prevalence: 46,746 of 1,155,237 collected domains | quote: "we collected 1,155,237 domains, with 46,746 identified as potential fraudulent shopping websites using the ML-based classifier." - merchant-ID reuse across scam domains | metric: unique merchant IDs | prevalence: 5,278 merchant IDs from 41,863 completed checkouts | quote: "Ultimately, out of 41,863 completed checkouts, S CAM M AGNIFIER was able to extract merchant IDs for three different payment processors for 5,278 total." - historical merchant-domain linkage | metric: linked domains | prevalence: 14,394 domains linked to identified fraudulent merchants | quote: "Intriguingly, 14,394 domains are connected to these merchants." - checkout redirection | metric: fraudulent websites with intermediary redirects | prevalence: 263 fraudulent shopping websites | quote: "The results showed that 263 fraudulent shopping websites redirected to an intermediary domain, which we refer to as B." - advertising referral traffic | metric: share of users | prevalence: 28.78% Facebook, 21.10% Google, and 9.38% Bing | quote: "28.78% of users reached fraudulent shopping websites through advertisements on Facebook (including Instagram), 21.10% from Google (Ads or search results), and 9.38% from Bing." - rapid scam monetization | metric: share monetized within one year | prevalence: 97.73% monetized in less than a year | quote: "Specifically 20.17% of fraudulent shopping websites have transactions within 10 days after their creation date and 97.73% are monetized in less than a year." - browser-extension scam detection | metric: detection rate | prevalence: 76.74% (66 of 86 expert-labeled fraudulent websites) | quote: "The integrated approach achieved a much higher detection rate of 76.74% (66 detected fraudulent websites) compared to Beyond Phish's standalone performance of 59.30%" - shared website content | metric: number of clusters | prevalence: ten categories | quote: "We performed a simple clustering method on fraudulent shopping websites screenshots to cluster them into ten categories." # USENIX/2025/nokescam-understanding-and-rectifying-non-sense-keywords-spear-scam-in-search-en (6 tuples) - NOKEScam NSKeywords and domains | metric: counts of NSKeywords and domains | prevalence: 153,975 NSKeywords across 68,863 domains | quote: "Over seven months, we identified 153,975 NSKeywords across 68,863 domains." - Fraud-category prevalence | metric: share of detected NOKEScam pages | prevalence: online game transaction fraud 94.36%, pornography 4.81%, gambling 0.83% | quote: "NOKEScam primarily involves online game account transactions (94.36%), with a smaller number related to pornography (4.81%) and gambling (0.83%)." - Search-engine user impact | metric: average daily page views | prevalence: about 30,000 daily page views | quote: "Statistics revealed an average daily PV of about 30k for these fraudulent domains" - Governance effectiveness | metric: weekly complaint reduction | prevalence: 194-fold reduction, from 777 to 4 per week | quote: "complaints dropped precipitously, decreasing 194-fold from the most 777 to 4 per week" - Detection accuracy | metric: accuracy and false-positive rate | prevalence: 99.10% accuracy and 0.96% false-positive rate | quote: "Our method achieved 99.10% accuracy. Manual examination reveals that False Positives (FPs) ... with a 0.96% FP rate." - Resolved scam-domain infrastructure | metric: resolved IP count | prevalence: 68,863 domains resolved to 9,102 IPs | quote: "We sent DNS requests for these domains (with A records), revealing that they resolved to 9,102 IPs." # WWW/2025/the-poorest-man-in-babylon-a-longitudinal-study-of-cryptocurrency-investment-sca (8 tuples) - cryptocurrency investment scam websites | metric: unique scam websites | prevalence: 43,572 unique cryptocurrency investment scam websites during the first 8 months of 2024 | quote: "During that time, Crimson recorded 43,572 unique scam websites." - shared hosting infrastructure | metric: share of websites hosted on IP addresses | prevalence: More than half of scam websites were associated with 10% of IP addresses | quote: "We can therefore conclude that more than half of the scam websites are associated with merely 10% of the IP addresses in our dataset." - web-design reuse | metric: clustered homepage screenshots | prevalence: 17,285 screenshots, representing 40% of detected scam websites, grouped into 4,335 clusters | quote: "we grouped a total of 17,285 investment scam website home-page screenshots—representing 40% of all detected scam websites—into 4,335 clusters" - indicator reuse | metric: websites sharing indicators | prevalence: Email addresses were reused across 27,036 websites (58%) | quote: "Email Address 27,036 (58%)" - scam-site persistence | metric: active websites at observation end | prevalence: 22,983 (47%) continued hosting investment scams at the end | quote: "22,983 (47%) websites continued to be active and still hosted investment scams at the end of the observation period." - cryptocurrency scam losses | metric: estimated lower-bound loss | prevalence: 2.04M US dollars from 6.7% of detected scam websites | quote: "In total, transactions sent towards scam wallets sum up to 2.04M US dollars" - blocklist coverage | metric: intersection percentage | prevalence: VirusTotal covered 20%; MetaMask covered 2%; Google Safe Browsing covered 1% | quote: "VirusTotal [48] 20% Metamask [47] 2% ... Google Safe Browsing [50] 1%" - shared deposit wallet addresses | metric: sites showing identical wallet addresses | prevalence: All 15 sampled sites showed the same wallet address to both accounts | quote: "For all 15 sites, we observed the same wallet address being shown to our two different Crimson-generated accounts." # NDSS/2026/loki-proactively-discovering-online-scams-by-mining-toxic-search-queries (3 tuples) - scam website discovery | metric: number and proportion of discovered websites | prevalence: 52,493 websites, 19.3% of 271,161 collected websites | quote: "The collected websites are passed through the oracle classifier, which labels 52,493 (19.3%) of them as scams." - search-query toxicity | metric: toxicity and expansion | prevalence: 1,663 seed scam sites generated approximately 1.5 million keyword suggestions | quote: "Toxicity = # Websites flagged as scam by oracle / # Websites returned by a search query" - top-result scam exposure | metric: share of scam websites in top 20 results | prevalence: 13.98% Google; 22.22% Bing; 29.1% Naver; 45.3% Baidu | quote: "13.98% of scam websites appear in the top 20 Google Search results, whereas the ratio is much higher on Bing (22.22%), Naver (29.1%), and Baidu (45.3%)." 166 detection tuples across 27 IN papers; papers with none: 0 === PASS G — CONTEXT papers (reading list, not population), by code === platform (11): scam accounts, comments or posts on a platform are the object (design:platforms) 2010 CCS/2010/spam-the-underground-on-140-characters-or-less — Twitter spam; URLs crawled but the object is the account/tweet 2011 IMC/2011/suspended-accounts-in-retrospect-an-analysis-of-twitter-spam — Twitter spam accounts; affiliate programs named 2012 USENIX/2012/efficient-and-scalable-socware-detection-in-online-social-networks — Facebook socware (scam posts) 2017 WWW/2017/pinning-down-abuse-on-google-maps — fake Google Maps listings 2024 USENIX/2024/the-imitation-game-exploring-brand-impersonation-attacks-on-social-media-platfor — brand impersonation accounts 2025* CCS/2025/poster-longitudinal-analysis-of-romance-scam-infrastructure-evolution-evidence-o — romance-scam profiles (dating sites) 2025* IMC/2025/exploration-of-the-dynamics-of-buy-and-sale-of-social-media-accounts — account markets; scams among buyers/sellers 2025* USENIX/2025/please-dont-send-that-bot-anything-a-mixed-methods-study-of-personal-impersonati — payment impersonation on X/Bluesky 2025* WWW/2025/pirates-of-charity-exploring-donation-based-abuses-in-social-media-platforms — donation scams on social platforms 2026* NDSS/2026/tbtrackerx-fantastic-trigger-bots-and-where-to-find-malicious-campaigns-on-x — trigger bots on X 2026* USENIX/2026/chameleon-channels-measuring-youtube-accounts-repurposed-for-deception-and-profi — repurposed YouTube channels onchain (12): scam contracts, tokens or transactions with no website in the object 2018 WWW/2018/detecting-ponzi-schemes-on-ethereum-towards-healthier-blockchain-technology — Ponzi contracts 2019 USENIX/2019/the-anatomy-of-a-cryptocurrency-pump-and-dump-scheme — pump-and-dump via Telegram 2019 USENIX/2019/the-art-of-the-scam-demystifying-honeypots-in-ethereum-smart-contracts — honeypot contracts 2022 CCS/2022/understanding-security-issues-in-the-nft-ecosystem — NFT counterfeit/fraudulent trading 2022 IMC/2022/challenges-in-decentralized-name-management-the-case-of-ens — ENS records incl. scam sites 2023 USENIX/2023/token-spammers-rug-pulls-and-sniper-bots-an-analysis-of-the-ecosystem-of-tokens — rug pulls 2024 CCS/2024/characterizing-ethereum-address-poisoning-attack — address poisoning 2024 CCS/2024/tokenscout-early-detection-of-ethereum-scam-tokens-via-temporal-graph-learning — scam tokens 2024 WWW/2024/interface-illusions-uncovering-the-rise-of-visual-scams-in-cryptocurrency-wallet — wallet visual scams (token metadata) 2025* USENIX/2025/blockchain-address-poisoning — address poisoning 2025* WWW/2025/serial-scammers-and-attack-of-the-clones-how-scammers-coordinate-multiple-rug-pu — rug pulls 2026* USENIX/2026/a-midsummer-memes-dream-investigating-market-manipulations-in-the-meme-coin-ecos — meme-coin manipulation telephony (10): phone or SMS scams 2015 NDSS/2015/phoneypot-data-driven-understanding-of-telephony-threats — telephony honeypot 2018 IEEE-SP/2018/a-machine-learning-approach-to-prevent-malicious-calls-over-telephony-networks — malicious calls 2019 USENIX/2019/users-really-do-answer-telephone-scams — telephone-scam experiment 2020 USENIX/2020/whos-calling-characterizing-robocalls-through-audio-and-metadata-analysis — robocalls 2023 USENIX/2023/diving-into-robocall-content-with-snorcall — robocall content 2024 IMC/2024/poster-a-comprehensive-categorization-of-sms-scams — SMS scams 2025* IEEE-SP/2025/blind-users-really-do-heed-aural-telephone-scam-warnings — aural scam warnings 2025* IEEE-SP/2025/characterizing-robocalls-with-multiple-vantage-points — robocalls 2025* IMC/2025/fishing-for-smishing-understanding-sms-phishing-infrastructure-and-strategies-by — smishing from public user reports 2025* USENIX/2025/hey-mum-i-dropped-my-phone-down-the-toilet-investigating-hi-mum-and-dad-sms-scam — SMS impersonation scams user-study (5): people's exposure, susceptibility or discussion of scams 2024 USENIX/2024/i-experienced-more-than-10-defi-scams-on-defi-users-perception-of-security-breac — DeFi users' scam experiences 2025* CCS/2025/is-this-a-scam-the-nature-and-quality-of-reddit-discussion-about-scams — Reddit scam discussion (user reports as a source) 2025* NDSS/2025/the-kids-are-all-right-investigating-the-susceptibility-of-teens-and-adults-to-youtube-giveaway-scams — giveaway-scam susceptibility experiment 2025* USENIX/2025/scanned-and-scammed-insecurity-by-obsqrity-measuring-user-susceptibility-and-awa — QR-code scam susceptibility 2026* IEEE-SP/2026/international-students-and-scams-at-risk-abroad — students' scam exposure phishing-boundary (5): credential, seed or signature theft called phishing (security:phishing) 2023 CCS/2023/txphishscope-towards-detecting-and-understanding-transaction-based-phishing-on-e — transaction-based phishing (signing), security:phishing counts it 2024 NDSS/2024/drainclog-detecting-rogue-accounts-with-illegally-obtained-nfts-using-classifiers-learned-on-graphs — NFT drainer accounts 2025* IMC/2025/unmasking-the-shadow-economy-a-deep-dive-into-drainer-as-a-service-phishing-on-e — drainer-as-a-service 2025* NDSS/2025/dissecting-payload-based-transaction-phishing-on-ethereum — payload transaction phishing 2026* NDSS/2026/ctphishcapture-uncovering-credential-theft-based-phishing-scams-targeting-cryptocurrency-wallets — credential-theft wallet phishing (security:phishing) search-abuse (12): search poisoning / cloaking / SEO whose destinations are generic, or openly illicit trade outside the 2011–2014 lineage exception 2011 CCS/2011/cloak-and-dagger-dynamics-of-web-search-cloaking — search cloaking method (user vs crawler identities); destinations mixed 2011 CCS/2011/fashion-crimes-trending-term-exploitation-on-the-web — trending-term exploitation; fake AV is one of several monetisations 2011 CCS/2011/surf-detecting-and-measuring-search-poisoning — search poisoning detector; destinations not primarily scams 2011 USENIX/2011/deseo-combating-search-result-poisoning — SEO poisoning infrastructure; fake AV destinations among others 2012 IEEE-SP/2012/evilseed-a-guided-approach-to-finding-malicious-web-pages — guided discovery of malicious pages from seeds 2013 NDSS/2013/juice-a-longitudinal-study-of-an-seo-botnet — SEO botnet; redirection chains 2016 IEEE-SP/2016/seeking-nonsense-looking-for-trouble-efficient-promotional-infection-detection-t — promotional infections 2016 USENIX/2016/towards-measuring-and-mitigating-social-engineering-software-download-attacks — social-engineering download ads (fake updates, scareware); payload is software not a payment 2016 WWW/2016/characterizing-long-tail-seo-spam-on-cloud-web-hosting-services — SEO spam on cloud hosting 2017 IEEE-SP/2017/how-to-learn-klingon-without-a-dictionary-detection-and-measurement-of-black-key — black keywords (underground promotion) 2019 IEEE-SP/2019/measuring-and-analyzing-search-engine-poisoning-of-linguistic-collisions — misspelling-query poisoning 2022 NDSS/2022/auto-draft-261 — local-business search poisoning for illicit drugs mobile-app (1): scam apps (not websites) measured as the object 2022 IEEE-SP/2022/analyzing-ground-truth-data-of-mobile-gambling-scams — mobile gambling scam apps and their payment channels; apps, not websites ads (5): scam ads measured as ad exposure (privacy:advertising) 2012 CCS/2012/knowing-your-enemy-understanding-and-detecting-malicious-web-advertising — malvertising chains; fake-AV scams are one category 2017 IMC/2017/exploring-the-dynamics-of-search-advertiser-fraud — fraudulent advertisers seen from inside Bing; scam ads are one monetisation 2023 USENIX/2023/problematic-advertising-and-its-disparate-exposure-on-facebook — Facebook ads coded incl. scams (privacy:advertising) 2025* PETS/2025/more-and-scammier-ads-the-perils-of-youtubes-ad-privacy-settings — predatory/scam ad rates on YouTube 2025* PETS/2025/sheeps-clothing-wolfish-intent-automated-detection-and-evaluation-of-problematic — problematic acceptable ads incl. scams ecosystem (4): back-office, dark-web or chat studies of fraud crews; revenue-estimation method 2014 NDSS/2014/scambaiter-understanding-targeted-nigerian-scams-on-craigslist — Craigslist advance-fee scams via email; scam-baiting method and ethics 2015 CCS/2015/drops-for-stuff-an-analysis-of-reshipping-mule-scams — reshipping-scam back-office databases; mules recruited by job scams 2019 NDSS/2019/cybercriminal-minds-an-investigative-study-of-cryptocurrency-abuses-in-the-dark-web — dark-web Bitcoin addresses including investment-scam sites 2023 CCS/2023/cybercrime-bitcoin-revenue-estimations-quantifying-the-impact-of-methodology-and — how method and coverage change cybercrime revenue estimates from addresses (harm method) llm (2): LLMs in scam conversations (detection, scam-baiting, scam operations) 2026* PETS/2026/ai-in-the-loop-privacy-preserving-real-time-scam-detection-and-conversational-sc — LLM scam detection / scam-baiting on conversations 2026* USENIX/2026/love-lies-and-language-models-investigating-ais-role-in-romance-baiting-scams — LLM agents in romance-baiting scams (WhatsApp) Cited outside the candidate set (method citations, not verdicts): IEEE-SP/2016/cloak-of-visibility-detecting-when-machines-browse-a-different-web — cloaking measurement method (search and ads) NDSS/2020/into-the-deep-web-understanding-e-commerce-fraud-from-autonomous-chat-with-cybercriminals — automated chat with fraud crews; ethics of interacting === PASS H — the 2026-09-22 gap-pass claim, re-derived === Claim: "43 papers >= 5 hits, 27 web, 21 from 2024–2026". As run here on paper.norm.txt: 43 papers, 27 web, 21 from 2024–2026. IN papers from 2024–2026: 7 of 27 (25.9%); CONTEXT papers from 2024–2026: 31 of 67. The recent slope of the probe is carried by CONTEXT (platform, on-chain, user-study, telephony), not by scam-website measurement; see PASS B "GAP set, 2024–2026, by verdict/code".
Figure and quote verification, unedited
node scripts/verify_scam_figures.mjs 2>/dev/null > scripts/verify_scam_figures-output.txt. The font warnings pypdf prints go to stderr and are not included.
- verify_scam_figures-output.txt
control ok (mutated needle not located) 2025/NDSS/scammagnifier "54.65% of all collected fraudulent websites are managed by only 10 merchant IDs" control ok (mutated needle not located) 2017/NDSS/dial-one "95.8% of the domain names discovered by all three instances" control ok (mutated needle not located) 2024/IMC/give-and-take "the conversion rate of tweets (with associated cryptocurrency addresses) to victims was 0.13%" OK paper.cols.txt 2025/NDSS/scammagnifier | box: 54.55% / 10 merchant IDs "54.55% of all collected fraudulent websites are managed by only 10 merchant IDs" OK paper.cols.txt 2025/NDSS/scammagnifier | box: 974 domains for one merchant ID "associated with 974 domains" OK paper.cols.txt 2025/NDSS/scammagnifier | box: 14,394 linked domains "14,394 domains are connected to these merchants" OK paper.cols.txt 2025/NDSS/scammagnifier | channel table: NRD funnel "we collected 1,155,237 domains, with 46,746 identified as potential fraudulent shopping websites" OK paper.cols.txt 2025/NDSS/scammagnifier | channel table: window "from May 2023 to June 2024" OK paper.cols.txt 2025/NDSS/scammagnifier | results: 20.17% "20.17% of fraudulent shopping websites have transactions within 10 days after their creation date" OK paper.cols.txt 2025/NDSS/scammagnifier | ethics table "without finalizing transactions" OK pypdf 2025/NDSS/scammagnifier | harm table: processor withheld volumes "the axis are deliberately unspecified" OK paper.cols.txt 2025/NDSS/scammagnifier | ethics table: reported "We reported all merchant IDs found to the payment processors when possible" OK paper.cols.txt 2014/IMC/search-seizure | box: 58% / 11% "we ascribed more than half (58%) of all PSRs to their respective campaigns, these PSRs only account for 11% of all stores" OK paper.cols.txt 2014/IMC/search-seizure | box: 2.7M "and detected 2.7M PSRs, across all verticals" OK paper.cols.txt 2014/IMC/search-seizure | grouping table: 7,404 stores "sending users to 7,404 different stores" OK paper.cols.txt 2014/IMC/search-seizure | box: 52 campaigns "we identified 52 distinct SEO campaigns" OK paper.cols.txt 2014/IMC/search-seizure | grouping: IP features dropped (quote) "were ill-suited to differentiate SEO campaigns due to the growing popularity of shared hosting and reverse proxying infrastructure (e.g., CloudFlare)" OK paper.cols.txt 2014/IMC/search-seizure | channel table quote "Any work measuring search results is biased towards the search terms selected" OK paper.cols.txt 2014/IMC/search-seizure | results: 3.9% "these domain seizures represent just a small percentage (3.9%) of stores" OK paper.cols.txt 2014/IMC/search-seizure | results: 290 seizures "we directly observed 290 seizures" OK paper.cols.txt 2014/IMC/search-seizure | results: 7 and 15 days "within 7 and 15 days of the initial seizure" OK paper.cols.txt 2014/IMC/search-seizure | ethics: checkout then stop "We take the orders all the way to the payment processing page" OK paper.cols.txt 2014/IMC/search-seizure | cloaking: iframe "we have identified a new method of cloaking, which we call iframe cloaking" OK paper.cols.txt 2017/NDSS/dial-one | box: 108 of 1,524 "out of 1,524 scam domains, only 108 (7%) were" OK paper.cols.txt 2017/NDSS/dial-one | box: 16 same day "16 were already blacklisted on the day that ROBOVIC first detected them" OK paper.cols.txt 2017/NDSS/dial-one | box: 38 days "The rest were blacklisted, on average, 38 days" OK paper.cols.txt 2017/NDSS/dial-one | box: 95.7% "95.7% of the domain names discovered by all three instances" OK paper.cols.txt 2017/NDSS/dial-one | box: cloud IP filtering "evading crawlers located on popular commercial clouds" OK paper.cols.txt 2017/NDSS/dial-one | harm: 1,688,412 visitors "we discovered a total of 1,688,412 unique IP addresses visiting the monitored scam domains" OK paper.cols.txt 2017/NDSS/dial-one | harm: 142 domains "we were able to monitor the activity of 142 scam domains over a period of two months" OK paper.cols.txt 2017/NDSS/dial-one | harm: $9.7M "more than $9.7 million" OK paper.cols.txt 2017/NDSS/dial-one | harm: borrowed 2% rate "is the same as that of users paying for the "full version" of a fake antivirus (approximately 2%" OK paper.cols.txt 2017/NDSS/dial-one | ethics: did not pay "we did not pay any scammer" OK paper.cols.txt 2017/NDSS/dial-one | ethics: 60 calls "We interact in a controlled fashion with 60 scammers" OK paper.cols.txt 2017/NDSS/dial-one | ethics: IRB for calls "got permission to perform these recorded calls" OK paper.cols.txt 2017/NDSS/dial-one | ethics: consent waived "waive the requirement of consent" OK pypdf 2023/NDSS/double-and-nothing | box/coverage: 16.75% of 3,610 "out of the 3,610 domains discovered by CryptoScamTracker, only 16.75% domains appeared in VT" OK paper.cols.txt 2023/NDSS/double-and-nothing | grouping: 35.58% "These campaigns consist of 3,586 web pages, which covered 35.58% websites in our dataset" OK paper.cols.txt 2023/NDSS/double-and-nothing | harm: $24.9M-$69.9M "scammers stole a total of $24.9M-$69.9M" OK paper.cols.txt 2023/NDSS/double-and-nothing | channel table quote "If an attacker avoids using any of these keywords" OK paper.cols.txt 2023/NDSS/double-and-nothing | grouping pitfall quote "online exchanges (such as Coinbase), where multiple victims can share the same outgoing address" OK paper.cols.txt 2023/NDSS/double-and-nothing | label: ~4% false positives "false positives (approximately 4%)" OK paper.cols.txt 2023/NDSS/double-and-nothing | results: 21 hours "50% of the addresses were identified by CryptoScamTracker at least 21 hours before they were reported by users" OK paper.cols.txt 2023/NDSS/double-and-nothing | results: denominator 300 of 2,266 "300 (14.45%) of the wallet addresses appeared in Bitcoinabuse" OK paper.cols.txt 2023/NDSS/double-and-nothing | results: 14.14 hours "less than 14.14 hours" OK paper.cols.txt 2023/NDSS/double-and-nothing | ethics quote "we chose not to tamper with the ecosystem while studying it, to avoid measuring artifacts of our own intervention" OK pypdf 2023/NDSS/double-and-nothing | grouping/overlap: 10,079 / 3,863 "10,079 giveaway scam websites hosted on a total of 3,863 domains" OK paper.cols.txt 2023/NDSS/double-and-nothing | results: 2,266 wallets "we extracted 2,266 scammer wallet addresses" OK paper.cols.txt 2023/IEEE-SP/beyond-phish | box/coverage: 25 and 10 "APWG and GSB detect only 25 and 10 FCWs within our dataset" OK paper.cols.txt 2023/IEEE-SP/beyond-phish | coverage denominator 6,127 "We leverage social media to collect 6,127 FCWs" OK pypdf 2023/IEEE-SP/beyond-phish | channel table quote "are potentially biased because a human user has already decided that they might be FCWs" OK paper.cols.txt 2023/IEEE-SP/beyond-phish | label: 1.98% FP "1.98% false positive rate" OK paper.cols.txt 2023/IEEE-SP/beyond-phish | label: 1.63% FN (column splice after the number) "false positive rate (FPR) and 1.63%" OK paper.cols.txt 2021/WWW/where-are-you | box: 81% "finds 81% more malicious and 96% more suspicious landing pages" OK paper.cols.txt 2021/WWW/where-are-you | cloaking: IP pool "more than twice as many malicious pages" OK paper.cols.txt 2021/WWW/where-are-you | read-first: 240-IP pool "rotates through 240 different IP addresses" OK paper.cols.txt 2021/WWW/where-are-you | cloaking quote "we face cloaking 86% of the time and are explicitly blocked only 14% of the time" OK pypdf 2021/WWW/where-are-you | coverage: 92 of 3,746 "GSB only labels 92 out of the 3,746 malicious pages on the same day we detect them" OK paper.cols.txt 2021/WWW/where-are-you | ethics quote "to avoid ethical quandaries" OK paper.cols.txt 2021/WWW/where-are-you | cloaking: no emulator detection "cybercriminals do not attempt to detect phone emulation" OK paper.cols.txt 2021/WWW/where-are-you | cloaking: desktop tech support "are more often shown to desktop users" OK paper.cols.txt 2021/WWW/where-are-you | cloaking: phones survey scams "more often targeted by survey campaigns" OK paper.cols.txt 2021/WWW/where-are-you | read-first: six profiles "six different types of users" OK paper.cols.txt 2024/IMC/give-and-take | box: 0.12% "the conversion rate of tweets (with associated cryptocurrency addresses) to victims was 0.12%" OK paper.cols.txt 2024/IMC/give-and-take | box: 0.0039% "0.0039% of livestream views yielded a victim" OK paper.cols.txt 2024/IMC/give-and-take | channel overlap: 9% "only 361 (9%) appeared on Twitter" OK paper.cols.txt 2024/IMC/give-and-take | channel table quote "is biased towards tweets that are more discoverable" OK paper.cols.txt 2024/IMC/give-and-take | harm: 43% "only 43% of payments (695 of 1,633) fell within it" OK paper.cols.txt 2024/IMC/give-and-take | harm: $2.7M "Twitter-based giveaway scams yielded $2.7M in revenue" OK paper.cols.txt 2024/IMC/give-and-take | harm: $6.6M "the total jumps to $6.6M" OK pypdf 2024/IMC/give-and-take | cloaking: four types "We identified four types of cloaking behavior in our pilot study" OK paper.cols.txt 2024/IMC/give-and-take | ethics: retrospective "did not have the opportunity to report the identified scams" OK paper.cols.txt 2025/WWW/the-poorest-man | box: $2.04M "sum up to 2.04M US dollars" OK paper.cols.txt 2025/WWW/the-poorest-man | box: 6.7% "stem from only 6.7% of the total scam websites" OK paper.cols.txt 2025/WWW/the-poorest-man | box/results: 43,572 "43,572 unique cryptocurrency investment scam domain names" OK paper.cols.txt 2025/WWW/the-poorest-man | grouping: 67% / 26% "29,300 (67%) of them share only 4,900 (26% of all) IP addresses" OK paper.cols.txt 2025/WWW/the-poorest-man | grouping: 8 bits "Hamming distance of up to 8 bits" OK paper.cols.txt 2025/WWW/the-poorest-man | label: 90 / 88 / 87 (Table 1 cells) "GPT-4 90 1200 GPT-4 + Llama3:70b 88 130 Llama3:70b 87 0" OK paper.cols.txt 2025/WWW/the-poorest-man | label: 300 hand-labelled "We manually categorized a random sample of 300 websites" OK paper.cols.txt 2025/WWW/the-poorest-man | coverage quote "this low coverage in existing block-lists is due to the general nature of websites they aim to detect" OK paper.cols.txt 2025/WWW/the-poorest-man | coverage denominator "a random sample of one-third of the Crimson dataset" OK paper.cols.txt 2025/WWW/the-poorest-man | harm: extrapolation "we estimate the total financial loss to be more than 100M USD" OK pypdf 2025/WWW/the-poorest-man | ethics: not reported "refrained from reporting identified scam websites" OK paper.cols.txt 2025/WWW/the-poorest-man | CT box: local CertStream "We deploy a local server using Certstream" OK paper.cols.txt 2025/NDSS/ctrlaltdeceive | read-first: 607K "607K scam fully qualified domain names" OK paper.cols.txt 2025/NDSS/ctrlaltdeceive | read-first: 196.9 billion "196.9 billion URL visits" OK paper.cols.txt 2025/NDSS/ctrlaltdeceive | read-first: 10.2M "observed by a total of 10.2M IPs" OK paper.cols.txt 2025/NDSS/ctrlaltdeceive | read-first: 653K "observed by 653K IPs" OK paper.cols.txt 2025/NDSS/ctrlaltdeceive | read-first/harm: 4% checkout "411K (4%) reach the checkout page" OK paper.cols.txt 2025/NDSS/ctrlaltdeceive | channel table: 80% "80% of the devices are located in the United States, the European Union, Japan, and the United Kingdom" OK paper.cols.txt 2025/NDSS/ctrlaltdeceive | channel table "Shopping scams are over-represented since they come from two datasets" OK paper.cols.txt 2025/NDSS/ctrlaltdeceive | results: daily exposure "101K (0.8%) of desktop devices being exposed compared to 48K (0.3%) of mobile devices" OK paper.cols.txt 2025/NDSS/ctrlaltdeceive | results: 11 days "the median active time for scams domains is 11 days, although the mean is 38.7 days" OK paper.cols.txt 2025/NDSS/ctrlaltdeceive | results: 13.3% "In at least 9.2M (13.3%) of all scam observations users followed an advertisement" OK paper.cols.txt 2025/NDSS/ctrlaltdeceive | results: 59% "largely (59%) hosted on social media" OK paper.cols.txt 2025/NDSS/ctrlaltdeceive | label: 7 of 26 "only 7 out of 26 categories having high precision" OK paper.cols.txt 2025/NDSS/ctrlaltdeceive | label quote "this low accuracy is not specific to ScamAdviser, but plagues most commercial website classification services" OK paper.cols.txt 2025/NDSS/ctrlaltdeceive | harm quote "not all users who visit a checkout page will complete a purchase" OK paper.cols.txt 2025/NDSS/ctrlaltdeceive | label: trust score "with a trust score up to 10" OK paper.cols.txt 2025/NDSS/ctrlaltdeceive | label: two VT detections "received less than two detections" OK paper.cols.txt 2025/NDSS/ctrlaltdeceive | open questions: desktop-heavy feeds "mostly from desktop devices" OK paper.cols.txt 2018/WWW/exposing-search | overlap: 0 of 2,768 "0/2,768 FQDNs and 0/2,441 second-level domains" OK paper.cols.txt 2018/WWW/exposing-search | overlap: 92 of 1,994 "92/1,994 common IP addresses" OK paper.cols.txt 2018/WWW/exposing-search | overlap: 5 of 882 "5/882 common toll-free phone numbers" OK paper.cols.txt 2018/WWW/exposing-search | search ads: 71.79% "Among the AD URIs, 10,299 (71.79%) were observed as leading to TSS websites" OK paper.cols.txt 2018/WWW/exposing-search | search ads: 14,346 "14,346 distinct AD URIs" OK paper.cols.txt 2018/WWW/exposing-search | coverage: 26.8% "these lists cover only 26.8% FQDNs" OK paper.cols.txt 2018/WWW/exposing-search | coverage: support domains <1% "<1% of those were present in any of these lists" OK paper.cols.txt 2018/WWW/exposing-search | search ads: 50 ads validated "For a set of 50 fake technical support ADs" OK paper.cols.txt 2018/WWW/exposing-search | search ads: referrer instead of clicking "while maintaining the Referer to be the search engine displaying the AD" OK paper.cols.txt 2018/WWW/exposing-search | results: ~9 days "median lifetime of ~9 days" OK paper.cols.txt 2018/WWW/exposing-search | results: ~100 days "median lifetime of ~100 days" OK paper.cols.txt 2018/WWW/exposing-search | cloaking: PhantomJS "Our choice of using PhantomJS for crawling search results and ads can, in principle, be detected by scammers" OK paper.cols.txt 2019/IMC/what-you-see | coverage: 16.2% "GSB detected only a small percentage (16.2% overall) of SE attack domains" OK pypdf 2019/IMC/what-you-see | grouping: DBSCAN params "eps = 0.1 and MinPts = 3" OK paper.cols.txt 2019/IMC/what-you-see | coverage: denominator 2,042 milked domains (Table 4 total row) "Total 2042 1.42% 16.21%" OK paper.cols.txt 2019/IMC/what-you-see | results: >50% "for three of the ad networks more than 50% of their ads led to SE attacks" OK paper.cols.txt 2019/IMC/what-you-see | ethics: $4.8 (sentence spliced by a column break) "we estimate that the advertiser was charged about" OK paper.cols.txt 2019/IMC/what-you-see | ethics: $4.8 "USD $4.8 due to our experiments" OK paper.cols.txt 2019/IMC/what-you-see | cloaking: two networks "Propeller and Clickadu" OK pypdf 2019/IMC/what-you-see | cloaking: residential laptops "connected to residential networks" OK pypdf 2019/IMC/what-you-see | cloaking: eleven networks "11 different ad networks" OK paper.cols.txt 2020/IMC/when-push-comes | results: 51% "51% of all WPN ads we collected are malicious" OK paper.cols.txt 2020/IMC/when-push-comes | results: 572 / 5,143 "572 WPN ad campaigns, for a total of 5,143 WPN-based ads" OK paper.cols.txt 2020/IMC/when-push-comes | coverage: <1% "less than 1% of the URLs were detected as malicious by GSB or VT" OK paper.cols.txt 2020/IMC/when-push-comes | coverage: 11.31% "we found that 1,388 (11.31%) of them were then detected by VT" OK pypdf 2020/IMC/when-push-comes | cloaking quote "were much more likely to appear on real Android devices, rather than emulated environments" OK paper.cols.txt 2020/IMC/when-push-comes | ethics: $1.12 "The maximum cost per landing domain throughout our entire study was USD 1.12$" OK paper.cols.txt 2023/USENIX/trident | ethics: $1.5 "the cost to each advertiser would be USD $1.5 on average" OK paper.cols.txt 2011/IEEE-SP/click-trajectories | read-first/grouping: three banks "just three banks provide the payment servicing for over 95% of the spam-advertised goods" OK pypdf 2011/IEEE-SP/click-trajectories | ethics: 120/76/56 "We attempted 120 purchases, of which 76 authorized and 56 settled" OK paper.cols.txt 2011/IEEE-SP/click-trajectories | ethics: IRB declined "did not deem it appropriate for their review" OK paper.cols.txt 2011/IEEE-SP/click-trajectories | ethics: restrictions "non-prescription goods" OK paper.cols.txt 2011/IEEE-SP/click-trajectories | ethics: licensed software "already possessed a site license" OK paper.cols.txt 2011/IEEE-SP/click-trajectories | ethics: destroyed "scheduled to be destroyed" OK paper.cols.txt 2011/IEEE-SP/click-trajectories | channel table: 13M garbage domains "the 13M distinct domains produced by the Rustock bot are artifacts of a" OK paper.cols.txt 2011/USENIX/show-me-the-money | harm: order volumes "seven leading counterfeit pharmacies together have a total monthly order volume in excess of 82,000, while three counterfeit software stores process over 37,000 orders" OK paper.cols.txt 2011/USENIX/show-me-the-money | grouping pitfall quote "is simply using the SE2 engine, but is not in fact associated with the GlavMed operation" OK paper.cols.txt 2012/USENIX/pharmaleaks | harm: $170M "settled revenue totaling over US$170M" OK paper.cols.txt 2012/USENIX/pharmaleaks | harm: purchase-pair bias "true turnover is between 8% (low of GlavMed) and 35% (high of SpamIt) less than predicted" OK paper.cols.txt 2012/CCS/priceless | ethics: 676/429 "676 ordering attempts, of which 429 were successful" OK paper.cols.txt 2012/CCS/priceless | ethics quote "explicitly reviewed and approved by our institution" OK paper.cols.txt 2012/CCS/priceless | harm quote "some subset of our refusals may not be due to true payment processing problems but an active attempt to" OK paper.cols.txt 2012/IMC/tasters-choice | channel table quote "only tend to capture spam campaigns that are very broadly targeted" OK paper.cols.txt 2012/IMC/tasters-choice | overlap: 60% / 19% "60% of all live domains and 19% of all tagged domains were exclusive to a single feed" OK paper.cols.txt 2012/IMC/tasters-choice | overlap: Hu "the largest contributor of unique instances is the human-identified spam domain feed Hu, despite also being the smallest feed" OK paper.cols.txt 2011/USENIX/measuring-and-analyzing-search | channel table quote "The risk of starting from a single seed is to only identify a single unrepresentative campaign" OK paper.cols.txt 2011/USENIX/measuring-and-analyzing-search | coverage: 95% "95% do not appear on a single blacklist" OK paper.cols.txt 2011/USENIX/measuring-and-analyzing-search | coverage: two thirds "over two thirds of pharmacy websites show up on at least one blacklist" OK pypdf 2011/USENIX/measuring-and-analyzing-search | capture-recapture: 2,523 and 795 "predicting a population size of 2 523 and the other predicting 795" OK paper.cols.txt 2011/USENIX/measuring-and-analyzing-search | results: 47 days "The median lifetime of infected websites is 47 days" OK paper.cols.txt 2011/USENIX/measuring-and-analyzing-search | results: 4,652 "4 652 unique infected source domains" OK paper.cols.txt 2010/CCS/dissecting-one-click | grouping: top 8 "the top 8 groups are responsible for more than half of the frauds" OK paper.cols.txt 2010/CCS/dissecting-one-click | grouping: 112 groups "out of the 112 groups" OK pypdf 2010/CCS/dissecting-one-click | coverage: none "None of the URLs we find in our database is listed in the Google Safe Browsing database" OK paper.cols.txt 2010/CCS/dissecting-one-click | channel table quote "we probably captured the most successful frauds" OK paper.cols.txt 2018/WWW/betrayed-by | grouping: 7.6 / 480 "average campaign size of 7.6 domains, with the largest campaign including 480 domains" OK pypdf 2018/WWW/betrayed-by | grouping: 3.6 / 293 "the average campaign size was 3.6 with the largest campaign including 293 domains" OK pypdf 2018/WWW/betrayed-by | grouping pitfall "error pages shown by hosting providers when they have suspended a user" OK paper.cols.txt 2023/IMC/evolving-bots | results: 31.73% "From our dataset of 45,322 videos 14,380 (31.73%) were infected" OK paper.cols.txt 2023/IMC/evolving-bots | results: 45,322 "From our dataset of 45,322 videos" OK paper.cols.txt 2023/IMC/evolving-bots | channel table: top comments "We proceeded with the 'top comments' because it is the default comment presentation option" OK paper.cols.txt 2023/IMC/evolving-bots | results: top 1,000 US creators "top 1,000 United States YouTube creator list" OK pypdf 2024/NDSS/like-comment | results: 2.34% "we identified 206,306 (2.34%) comments that clearly belong to scammers" OK pypdf 2024/NDSS/like-comment | results: denominator "8,801,224 individual comments from 20 YouTube channels" OK pypdf 2024/NDSS/like-comment | harm: extrapolation "have potentially stolen more than 100 million US dollars" OK paper.cols.txt 2024/NDSS/like-comment | harm: 31 of 72 "31 out of 72 scammers we contacted" OK pypdf 2024/NDSS/like-comment | ethics "instead used polite excuses to leave the conversation" OK pypdf 2024/NDSS/like-comment | ethics quote "The IRB approved our approach of deception and waived the requirement of consent" OK paper.cols.txt 2024/NDSS/like-comment | coverage: 1 of 24 "out of the 24 URLs we submitted, only 1 URL is marked as suspicious" OK paper.cols.txt 2024/NDSS/like-comment | ethics: 74 / 50 "74 scammers and captured 50 complete conversations" OK paper.cols.txt 2026/NDSS/loki-proactively | channel table: 19.3% "labels 52,493 (19.3%) of them as scams" OK paper.cols.txt 2026/NDSS/loki-proactively | channel table: 980 queries "resulting in a total of 980 distinct queries" OK paper.cols.txt 2026/NDSS/loki-proactively | channel table: 271,161 "resulting in a total of 271,161 websites" OK paper.cols.txt 2026/NDSS/loki-proactively | channel table: four engines "Google, Baidu, Bing, and Naver" OK paper.cols.txt 2026/NDSS/loki-proactively | label: 103 features "leverages a total of 103 features" OK paper.cols.txt 2026/NDSS/loki-proactively | label: >90% "precision and recall exceed 90%" OK paper.cols.txt 2026/NDSS/loki-proactively | currency: LLM keyword filter "zero-shot text classification capabilities" OK paper.cols.txt 2026/NDSS/loki-proactively | label: corroboration "Since there is no ground truth attached to the new scam websites discovered by" OK paper.cols.txt 2025/USENIX/nokescam | grouping: 80.02% "top 10 campaigns accounting for 80.02%" OK paper.cols.txt 2025/USENIX/nokescam | grouping: 143 "we identified 143 NOKEScam campaigns" OK paper.cols.txt 2025/USENIX/nokescam | harm: $2,896 "The average fraud amount per NOKEScam victim was $2,896" OK paper.cols.txt 2025/USENIX/nokescam | harm: 20 of last 100 "20 complaints mentioned the amount of money defrauded" OK paper.cols.txt 2025/USENIX/nokescam | harm: assumed rate "rate of one-millionth of users" OK paper.cols.txt 2025/USENIX/nokescam | cloaking: crawler vs victim "timely content like the latest news" OK paper.cols.txt 2025/NDSS/ctrlaltdeceive | results preface: Kotzias lifetime definition "the activity time from first to last observations" OK paper.cols.txt 2018/WWW/exposing-search | results preface: Srinivasan lifetime definition "the difference between the earliest and most recent date that the domain was seen hosting TSS content" OK paper.cols.txt 2011/USENIX/measuring-and-analyzing-search | results preface: Leontiadis truncation "which is 192 days in our sample" OK paper.cols.txt 2025/NDSS/ctrlaltdeceive | type taxonomy: seven types "We examine seven popular scam types" OK paper.cols.txt 2026/NDSS/loki-proactively | type taxonomy: ten categories "across 10 major scam categories" OK paper.cols.txt 2026/NDSS/loki-proactively | type taxonomy: Trustpilot "Trustpilot categories" OK paper.cols.txt 2023/CCS/txphishscope | intro: the paper's own term "transaction-based phishing" OK paper.cols.txt 2023/CCS/cybercrime-bitcoin | grouping pitfall quote "the popular multi-input clustering fails to discover addresses for 40% of groups" OK paper.cols.txt 2023/CCS/cybercrime-bitcoin | harm quote "the revenue is not always underestimated. There exist methodologies that can introduce huge overestimation" OK paper.cols.txt 2023/CCS/cybercrime-bitcoin | harm: 30,424 "We collect 30,424 payment addresses" OK paper.cols.txt 2023/CCS/cybercrime-bitcoin | harm: giveaway scams among six "Ponzi schemes, giveaway scams, exchange scams" 193 needles; located by: paper.cols.txt 172, pypdf 21; weak (<20 chars): 0; not located: 0 === EXTERNAL FIGURES (non-corpus; primary source on the provenance page) === Let's Encrypt RFC 6962 logs: read-only 30 November 2025, shut down 28 February 2026 FBI IC3 2025: $20.877 billion reported losses GASA Global State of Scams 2025: "$442 billion", 46,000 people, 42 markets Report Fraud replaced Action Fraud: 4 December 2025 Chrome Gemini Nano scam detection announced 2025-05-08; Edge scareware blocker default-on 2025-10-31 ScamFerret (arXiv:2502.10110, DIMVA 2025): 0.972 accuracy, four English scam types, GPT-4 Toll scams (arXiv:2510.14198, eCrime 2025): 67,907 domains, 86.9% in five TLDs certstream-server last commit 2025-09-04; issue #143 opened 2026-02-09; certstream.calidog.io HTTP 200 on 2026-09-24 Chrome extension to fake viruses / fake giveaways announced 2025-09-18; Edge scareware blocker preview 2025-01-27 UK Online Safety Act 2023 ss. 38–39 === TEMPLATE LITERALS (the illustrative "What to Report" sentence; not data) === 2026-03-01 to 2026-05-31; a 40-keyword list; "We crawled 40,000 scam sites"
External checks, unedited
PLAYWRIGHT_BROWSERS_PATH=/workspace/.playwright bash scripts/external_checks_online_scams.sh > scripts/external_checks_online_scams-output.txt 2>&1
- external_checks_online_scams-output.txt
== 1. Certificate Transparency discovery infrastructure == OK [200] Let's Encrypt RFC 6962 logs read-only — "November 30, 2025 , we will make our RFC 6962 logs read-only" OK [200] Let's Encrypt RFC 6962 logs shut down — "On February 28, 2026 , we will entirely shut down our RFC 6962 logs" OK [200] Let's Encrypt: monitors need new client — "need to ensure that they have client software that's compatible with the new API" OK [200] certstream.calidog.io front end answers (live WSS stream NOT tested) OK CaliDog/certstream-server last commit on default branch: 2025-09-04T23:49:18Z OK CaliDog/certstream-server issue #143 still open: open 2026-02-09T12:43:03Z Certstream websocket down OK [200] certstream-server-rust README: RFC 6962 and static-CT — "compatible with existing Certstream clients and supports both RFC 6962 and static-CT logs" == 2. Browser-side scam defences == OK [200] Chrome: Gemini Nano against tech support scams (2025-05-08) — "protect users from remote tech support scams" OK [200] Edge scareware blocker default on (2025-10-31) — "is now enabled by default on most Windows and Mac devices" OK [200] Chrome: extension to fake viruses / fake giveaways announced (2025-09-18) — "Soon, we'll be expanding this protection to also stop sites that use fake viruses or fake giveaways to trick you" OK [200] Edge scareware blocker preview (2025-01-27) — "now available in preview in Microsoft Edge" OK [200] Safe Browsing v4 deprecated — "The Safe Browsing APIs (v4) are deprecated" OK [200] Safe Browsing v4 ThreatType SOCIAL_ENGINEERING — "SOCIAL_ENGINEERING" == 3. Loss statistics == OK IC3 2025 report — "$20.877 billion" OK IC3 2025 report — "Investment-related fraud was once again the largest component of these losses, followed by business email compromises and tech support scams" OK [200] GASA Global State of Scams 2025 landing page — "Global State of Scams" OK GASA 2025 PDF (third-party mirror) — "$442 Billion (estimated, as seen on methodology)" OK GASA 2025 PDF (third-party mirror) — "46,000" OK GASA 2025 PDF (third-party mirror) — "42 markets" OK GASA 2025 PDF (third-party mirror) — "figure reflects only the 42 markets surveyed" OK [200] GOV.UK: Report Fraud replaces Action Fraud — "From 4 December 2025, City of London Police is launching Report Fraud, a new service that replaces Action Fraud" OK [200] UK Online Safety Act s.39 fraudulent advertising (Category 2A) — "39 Duties about fraudulent advertising: Category 2A services" OK [200] UK Online Safety Act s.38 fraudulent advertising (Category 1) — "38 Duties about fraudulent advertising: Category 1 services" == 4. Outside-corpus papers cited in footnotes == OK [200] ScamFerret arXiv abstract — "ScamFerret achieves 0.972 accuracy in classifying four scam types in English" OK [200] ScamFerret venue — "Accepted for publication at DIMVA 2025" OK [200] Toll scam domains arXiv abstract — "a newly created dataset of 67,907 confirmed scam domains mostly registered in 2025" OK [200] Toll scam domains: five TLDs — "86.9% of domains concentrated in just five non-mainstream TLDs" OK [200] Toll scam domains venue — "accepted for presentation at eCrime 2025" == 5. Artifacts of papers in the population == OK github.com/codepujan/ndss-loki-artifact last commit: 2025-09-01T19:33:32Z OK github.com/pragseclab/Crimson last commit: 2025-05-31T18:31:34Z OK github.com/mbitaab/beyondphish last commit: 2023-05-24T16:21:08Z OK github.com/phani-vadrevu/seacma last commit: 2019-10-21T14:20:31Z OK github.com/ian7yang/trident last commit: 2023-09-19T02:03:30Z OK github.com/jszurdi/ODIN last commit: 2021-02-08T05:24:47Z OK [200] Double and Nothing dataset CSV (1416783 bytes) OK nokescam.com does not resolve (the page says so) external checks: 0 failed
Probes, inclusion rule, verdicts and hand codes
scripts/scam_fold.mjs as run.
- scam_fold.mjs
// scam_fold.mjs — probes, inclusion rule and hand verdicts behind // security:online_scams ("Measuring online scams"). Imported by // report_online_scams.mjs; nothing here is computed, it is the audit surface. // // The probes produce a CANDIDATE set. The population is the hand verdict // below, and report_online_scams.mjs throws if a candidate has no verdict or a // verdict names a paper that is not a candidate. // The 2026-09-22 gap pass's regex, unchanged (gap_probe_fulltext_20260922.mjs, family "scams"). export const GAP_RE = /fake (shop|store|webshop|pharmac)|tech(nical)? support scam|scam (site|website|page|domain)s?\b|online scam|crypto(currency)? scam|giveaway scam|pig[- ]butchering|fraudulent (shop|store|web ?site)/gi; // The bare word, counted with a high threshold so that a paper about scams qualifies and a mention does not. export const SCAM_RE = /\bscam(s|mer|mers|ming|med)?\b/gi; // Named scam types, so a paper that never says "scam" as often (pharmacy, counterfeit, giveaway, rug pull) still enters. export const SPEC_RE = /tech(nical)?[- ]support (scam|fraud)|fake (tech(nical)? )?support|fake (online )?(shop|store|webshop|e-?shop)s?|fraudulent (online )?(shop|store|shopping|e-?commerce)|counterfeit (goods|luxury|product|store|shop|merchandise)|(crypto(currency)?|bitcoin|NFT|token) (scam|fraud|giveaway|investment scam)s?|giveaway scam|pig[- ]butcher|rug[- ]?pull|ponzi|high[- ]yield investment|(online|illegal|illicit|rogue|counterfeit|unlicensed) pharmac(y|ies)|pharma(ceutical)? spam|survey scam|romance scam|advance[- ]fee|419 scam|reshipping|lottery scam|fake (job|recruit)|employment scam|refund scam|impersonation scam/gi; // The 2011–2016 vocabulary: spam value chain, scareware, social-engineering ads. export const LIN_RE = /scareware|fake anti-?virus|rogue (anti-?virus|security software)|affiliate program(me)?s?|spam-advertised|abuse-advertised|test purchase|social engineering (attack|campaign|ad|page)s?|scam ads?\b/gi; // Schema channel: the extraction's own free-text fields. export const SCHEMA_RE = /scam|fraudulent (shop|store|web|site|e-?commerce)|fake (shop|store|pharm)|counterfeit (goods|luxury|product|store|shop|pharm)|scareware|fake anti|survey scam|giveaway|ponzi|pharmaceutical products/i; export const THRESH = { GAP: 5, SCAM: 10, SPEC: 5, LIN: 5 }; export const INCLUSION_RULE = `INCLUSION RULE (written before the verdicts were counted): IN — the paper's measured objects include web-delivered scams: websites, landing pages, or the ad / search / redirect chains that deliver them, whose purpose is to make the visitor PAY or TRANSFER money under false pretences (goods that never arrive or are counterfeit or unlicensed, "support", doubled or invested crypto, survey prizes, one-click billing). The paper collects instances (crawl, feed, telemetry, leaked or seized back end) and measures, clusters, labels or detects them. CONTEXT — scams, but the object is not a scam website: scam accounts/comments on a platform (design:platforms), on-chain scam contracts/tokens with no website, phone/SMS scams, user studies of scam susceptibility, credential/seed/signature theft that the literature calls phishing (security:phishing), search abuse or scam ads whose destinations are not primarily scams, and back-office or chat studies of fraud crews. OUT — scams are incidental, or "fraud" is against a platform/merchant/operator (ad fraud, telephony fraud, concession abuse), malware delivery, or a homonym (Scamper, Scamalytics). Boundary with security:phishing: WHAT THE VICTIM HANDS OVER. Credentials, seed phrases or a signature that gives the attacker authority -> phishing. Money the victim knowingly sends -> here. Named lineage exception (added after the generic review): the 2011–2014 spam- and SEO-advertised pharmacy and counterfeit storefront papers are IN for the methods they introduced (payment grouping, test purchases, order numbers, feed overlap). Later search poisoning that promotes openly illicit trade (controlled drugs, gambling, black-market goods) is CONTEXT search-abuse: the criterion in use is lineage, not the object, and the page says so. Scareware counts when the paper measures the purchase page; a paper that measures the download (the software payload) is CONTEXT.`; export const CONTEXT_CODES = { platform: 'scam accounts, comments or posts on a platform are the object (design:platforms)', onchain: 'scam contracts, tokens or transactions with no website in the object', telephony: 'phone or SMS scams', 'user-study': "people's exposure, susceptibility or discussion of scams", 'phishing-boundary': 'credential, seed or signature theft called phishing (security:phishing)', 'search-abuse': 'search poisoning / cloaking / SEO whose destinations are generic, or openly illicit trade outside the 2011–2014 lineage exception', 'mobile-app': 'scam apps (not websites) measured as the object', ads: 'scam ads measured as ad exposure (privacy:advertising)', ecosystem: 'back-office, dark-web or chat studies of fraud crews; revenue-estimation method', llm: 'LLMs in scam conversations (detection, scam-baiting, scam operations)', }; export const OUT_CODES = { mention: 'scams incidental to another topic', 'user-general': 'general security/privacy user study in which scams come up', 'other-fraud': 'fraud against a platform, merchant or operator; account abuse; credential spoofing', malware: 'malware or unwanted-software delivery', 'domain-abuse': 'typo/bit-squatting, parking', infra: 'network measurement; "Scamper"/"Scamalytics" homonyms', 'onchain-other': 'blockchain security or tracing that is not a scam measurement', }; export const IN_TYPES = { shop: 'fake, counterfeit or unlicensed storefronts (incl. spam/SEO-advertised pharmacy)', techsupport: 'technical-support scams', crypto: 'crypto giveaway / investment scam sites', socialeng: 'social-engineering ad and redirect campaigns (tech support, scareware, surveys, fake prizes)', survey: 'survey scams', oneclick: 'one-click billing fraud', general: 'several scam types, or scam as a class', }; // [key, verdict, code (IN: scam type; CONTEXT/OUT: code), note] export const VERDICTS = [ ["IMC/2024/give-and-take-an-end-to-end-investigation-of-giveaway-scam-conversion-rates", "IN", "crypto", "giveaway scam conversion rates, landing pages and wallets"], ["NDSS/2023/double-and-nothing-understanding-and-detecting-cryptocurrency-giveaway-scams", "IN", "crypto", "crypto giveaway scam sites via CT logs; wallet flows"], ["WWW/2025/the-poorest-man-in-babylon-a-longitudinal-study-of-cryptocurrency-investment-sca", "IN", "crypto", "crypto investment scam sites via CT logs; wallets; blocklist coverage"], ["IMC/2023/evolving-bots-the-new-generation-of-comment-bots-and-their-underlying-scam-campa", "IN", "general", "YouTube comment bots promoting scam domains; 72 scam campaigns"], ["NDSS/2024/like-comment-get-scammed-characterizing-comment-scams-on-media-platforms", "IN", "general", "YouTube comment scams followed to scam websites and wallets"], ["NDSS/2025/ctrlaltdeceive-quantifying-user-exposure-to-online-scams", "IN", "general", "vendor telemetry: user exposure to scam domains"], ["NDSS/2026/loki-proactively-discovering-online-scams-by-mining-toxic-search-queries", "IN", "general", "scam discovery by mining toxic search queries"], ["USENIX/2025/nokescam-understanding-and-rectifying-non-sense-keywords-spear-scam-in-search-en", "IN", "general", "nonsense-keyword spear scams in Baidu; complaint data"], ["WWW/2018/betrayed-by-your-dashboard-discovering-malicious-campaigns-via-web-analytics", "IN", "general", "analytics IDs link malicious sites into campaigns; tech-support and other scams among them"], ["CCS/2010/dissecting-one-click-frauds", "IN", "oneclick", "Japanese one-click billing-fraud websites; 2,140 reports, bank accounts and phone numbers as linking identifiers"], ["CCS/2012/priceless-the-role-of-payments-in-abuse-advertised-goods", "IN", "shop", "payment processing behind abuse-advertised storefronts; undercover purchases"], ["CCS/2014/a-nearly-four-year-longitudinal-study-of-search-engine-poisoning", "IN", "shop", "four years of search poisoning toward unlicensed-pharmacy storefronts"], ["IEEE-SP/2011/click-trajectories-end-to-end-analysis-of-the-spam-value-chain", "IN", "shop", "spam value chain end to end: hosting, affiliate programs, payment, purchases"], ["IEEE-SP/2023/beyond-phish-toward-detecting-fraudulent-e-commerce-websites-at-scale", "IN", "shop", "fraudulent e-commerce website detector; Reddit-derived seeds"], ["IMC/2012/tasters-choice-a-comparative-analysis-of-spam-feeds", "IN", "shop", "coverage of ten spam feeds for storefront domains: the seed-bias paper of the lineage"], ["IMC/2014/search-seizure-the-effectiveness-of-interventions-on-seo-campaigns", "IN", "shop", "counterfeit-luxury storefronts via SEO; orders and supplier records"], ["NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns", "IN", "shop", "fraudulent shopping campaigns via new registrations, automated checkout, merchant IDs"], ["USENIX/2011/measuring-and-analyzing-search-redirection-attacks-in-the-illicit-online-prescri", "IN", "shop", "unlicensed-pharmacy storefronts reached by search redirection; conversion estimate"], ["USENIX/2011/show-me-the-money-characterizing-spam-advertised-revenue", "IN", "shop", "spam-advertised pharmacy/software storefront revenue from sequential order numbers and purchases"], ["USENIX/2012/pharmaleaks-understanding-the-business-of-online-pharmaceutical-affiliate-progra", "IN", "shop", "leaked affiliate-program databases: customers, revenue, payment"], ["IMC/2019/what-you-see-is-not-what-you-get-discovering-and-tracking-social-engineering-att", "IN", "socialeng", "social-engineering ad campaigns clustered by screenshot; tech-support, scareware, survey scams"], ["IMC/2020/when-push-comes-to-ads-measuring-the-rise-of-malicious-push-advertising", "IN", "socialeng", "malicious web-push ad campaigns; clustering; many are scams"], ["USENIX/2023/trident-towards-detecting-and-mitigating-web-based-social-engineering-attacks", "IN", "socialeng", "social-engineering ad detector on crawled sites; tech-support, scareware"], ["WWW/2021/where-are-you-taking-me-understanding-abusive-traffic-distribution-systems", "IN", "socialeng", "abusive traffic distribution systems; multi-profile, multi-vantage crawl; scam destinations"], ["IEEE-SP/2018/surveylance-automatically-detecting-online-survey-scams", "IN", "survey", "survey scam detector over search-derived crawl"], ["NDSS/2017/dial-one-for-scam-a-large-scale-analysis-of-technical-support-scams", "IN", "techsupport", "technical support scams: malvertising crawl, infrastructure, calls to scammers"], ["WWW/2018/exposing-search-and-advertisement-abuse-tactics-and-infrastructure-of-technical", "IN", "techsupport", "tech-support scams via search results and search ads"], ["CCS/2012/knowing-your-enemy-understanding-and-detecting-malicious-web-advertising", "CONTEXT", "ads", "malvertising chains; fake-AV scams are one category"], ["IMC/2017/exploring-the-dynamics-of-search-advertiser-fraud", "CONTEXT", "ads", "fraudulent advertisers seen from inside Bing; scam ads are one monetisation"], ["PETS/2025/more-and-scammier-ads-the-perils-of-youtubes-ad-privacy-settings", "CONTEXT", "ads", "predatory/scam ad rates on YouTube"], ["PETS/2025/sheeps-clothing-wolfish-intent-automated-detection-and-evaluation-of-problematic", "CONTEXT", "ads", "problematic acceptable ads incl. scams"], ["USENIX/2023/problematic-advertising-and-its-disparate-exposure-on-facebook", "CONTEXT", "ads", "Facebook ads coded incl. scams (privacy:advertising)"], ["CCS/2015/drops-for-stuff-an-analysis-of-reshipping-mule-scams", "CONTEXT", "ecosystem", "reshipping-scam back-office databases; mules recruited by job scams"], ["CCS/2023/cybercrime-bitcoin-revenue-estimations-quantifying-the-impact-of-methodology-and", "CONTEXT", "ecosystem", "how method and coverage change cybercrime revenue estimates from addresses (harm method)"], ["IEEE-SP/2022/analyzing-ground-truth-data-of-mobile-gambling-scams", "CONTEXT", "mobile-app", "mobile gambling scam apps and their payment channels; apps, not websites"], ["NDSS/2014/scambaiter-understanding-targeted-nigerian-scams-on-craigslist", "CONTEXT", "ecosystem", "Craigslist advance-fee scams via email; scam-baiting method and ethics"], ["NDSS/2019/cybercriminal-minds-an-investigative-study-of-cryptocurrency-abuses-in-the-dark-web", "CONTEXT", "ecosystem", "dark-web Bitcoin addresses including investment-scam sites"], ["PETS/2026/ai-in-the-loop-privacy-preserving-real-time-scam-detection-and-conversational-sc", "CONTEXT", "llm", "LLM scam detection / scam-baiting on conversations"], ["USENIX/2026/love-lies-and-language-models-investigating-ais-role-in-romance-baiting-scams", "CONTEXT", "llm", "LLM agents in romance-baiting scams (WhatsApp)"], ["CCS/2022/understanding-security-issues-in-the-nft-ecosystem", "CONTEXT", "onchain", "NFT counterfeit/fraudulent trading"], ["CCS/2024/characterizing-ethereum-address-poisoning-attack", "CONTEXT", "onchain", "address poisoning"], ["CCS/2024/tokenscout-early-detection-of-ethereum-scam-tokens-via-temporal-graph-learning", "CONTEXT", "onchain", "scam tokens"], ["IMC/2022/challenges-in-decentralized-name-management-the-case-of-ens", "CONTEXT", "onchain", "ENS records incl. scam sites"], ["USENIX/2019/the-anatomy-of-a-cryptocurrency-pump-and-dump-scheme", "CONTEXT", "onchain", "pump-and-dump via Telegram"], ["USENIX/2019/the-art-of-the-scam-demystifying-honeypots-in-ethereum-smart-contracts", "CONTEXT", "onchain", "honeypot contracts"], ["USENIX/2023/token-spammers-rug-pulls-and-sniper-bots-an-analysis-of-the-ecosystem-of-tokens", "CONTEXT", "onchain", "rug pulls"], ["USENIX/2025/blockchain-address-poisoning", "CONTEXT", "onchain", "address poisoning"], ["USENIX/2026/a-midsummer-memes-dream-investigating-market-manipulations-in-the-meme-coin-ecos", "CONTEXT", "onchain", "meme-coin manipulation"], ["WWW/2018/detecting-ponzi-schemes-on-ethereum-towards-healthier-blockchain-technology", "CONTEXT", "onchain", "Ponzi contracts"], ["WWW/2024/interface-illusions-uncovering-the-rise-of-visual-scams-in-cryptocurrency-wallet", "CONTEXT", "onchain", "wallet visual scams (token metadata)"], ["WWW/2025/serial-scammers-and-attack-of-the-clones-how-scammers-coordinate-multiple-rug-pu", "CONTEXT", "onchain", "rug pulls"], ["CCS/2023/txphishscope-towards-detecting-and-understanding-transaction-based-phishing-on-e", "CONTEXT", "phishing-boundary", "transaction-based phishing (signing), security:phishing counts it"], ["IMC/2025/unmasking-the-shadow-economy-a-deep-dive-into-drainer-as-a-service-phishing-on-e", "CONTEXT", "phishing-boundary", "drainer-as-a-service"], ["NDSS/2024/drainclog-detecting-rogue-accounts-with-illegally-obtained-nfts-using-classifiers-learned-on-graphs", "CONTEXT", "phishing-boundary", "NFT drainer accounts"], ["NDSS/2025/dissecting-payload-based-transaction-phishing-on-ethereum", "CONTEXT", "phishing-boundary", "payload transaction phishing"], ["NDSS/2026/ctphishcapture-uncovering-credential-theft-based-phishing-scams-targeting-cryptocurrency-wallets", "CONTEXT", "phishing-boundary", "credential-theft wallet phishing (security:phishing)"], ["CCS/2010/spam-the-underground-on-140-characters-or-less", "CONTEXT", "platform", "Twitter spam; URLs crawled but the object is the account/tweet"], ["CCS/2025/poster-longitudinal-analysis-of-romance-scam-infrastructure-evolution-evidence-o", "CONTEXT", "platform", "romance-scam profiles (dating sites)"], ["IMC/2011/suspended-accounts-in-retrospect-an-analysis-of-twitter-spam", "CONTEXT", "platform", "Twitter spam accounts; affiliate programs named"], ["IMC/2025/exploration-of-the-dynamics-of-buy-and-sale-of-social-media-accounts", "CONTEXT", "platform", "account markets; scams among buyers/sellers"], ["NDSS/2026/tbtrackerx-fantastic-trigger-bots-and-where-to-find-malicious-campaigns-on-x", "CONTEXT", "platform", "trigger bots on X"], ["USENIX/2012/efficient-and-scalable-socware-detection-in-online-social-networks", "CONTEXT", "platform", "Facebook socware (scam posts)"], ["USENIX/2024/the-imitation-game-exploring-brand-impersonation-attacks-on-social-media-platfor", "CONTEXT", "platform", "brand impersonation accounts"], ["USENIX/2025/please-dont-send-that-bot-anything-a-mixed-methods-study-of-personal-impersonati", "CONTEXT", "platform", "payment impersonation on X/Bluesky"], ["USENIX/2026/chameleon-channels-measuring-youtube-accounts-repurposed-for-deception-and-profi", "CONTEXT", "platform", "repurposed YouTube channels"], ["WWW/2017/pinning-down-abuse-on-google-maps", "CONTEXT", "platform", "fake Google Maps listings"], ["WWW/2025/pirates-of-charity-exploring-donation-based-abuses-in-social-media-platforms", "CONTEXT", "platform", "donation scams on social platforms"], ["CCS/2011/cloak-and-dagger-dynamics-of-web-search-cloaking", "CONTEXT", "search-abuse", "search cloaking method (user vs crawler identities); destinations mixed"], ["CCS/2011/fashion-crimes-trending-term-exploitation-on-the-web", "CONTEXT", "search-abuse", "trending-term exploitation; fake AV is one of several monetisations"], ["CCS/2011/surf-detecting-and-measuring-search-poisoning", "CONTEXT", "search-abuse", "search poisoning detector; destinations not primarily scams"], ["IEEE-SP/2012/evilseed-a-guided-approach-to-finding-malicious-web-pages", "CONTEXT", "search-abuse", "guided discovery of malicious pages from seeds"], ["IEEE-SP/2016/seeking-nonsense-looking-for-trouble-efficient-promotional-infection-detection-t", "CONTEXT", "search-abuse", "promotional infections"], ["IEEE-SP/2017/how-to-learn-klingon-without-a-dictionary-detection-and-measurement-of-black-key", "CONTEXT", "search-abuse", "black keywords (underground promotion)"], ["IEEE-SP/2019/measuring-and-analyzing-search-engine-poisoning-of-linguistic-collisions", "CONTEXT", "search-abuse", "misspelling-query poisoning"], ["NDSS/2013/juice-a-longitudinal-study-of-an-seo-botnet", "CONTEXT", "search-abuse", "SEO botnet; redirection chains"], ["NDSS/2022/auto-draft-261", "CONTEXT", "search-abuse", "local-business search poisoning for illicit drugs"], ["USENIX/2011/deseo-combating-search-result-poisoning", "CONTEXT", "search-abuse", "SEO poisoning infrastructure; fake AV destinations among others"], ["USENIX/2016/towards-measuring-and-mitigating-social-engineering-software-download-attacks", "CONTEXT", "search-abuse", "social-engineering download ads (fake updates, scareware); payload is software not a payment"], ["WWW/2016/characterizing-long-tail-seo-spam-on-cloud-web-hosting-services", "CONTEXT", "search-abuse", "SEO spam on cloud hosting"], ["IEEE-SP/2018/a-machine-learning-approach-to-prevent-malicious-calls-over-telephony-networks", "CONTEXT", "telephony", "malicious calls"], ["IEEE-SP/2025/blind-users-really-do-heed-aural-telephone-scam-warnings", "CONTEXT", "telephony", "aural scam warnings"], ["IEEE-SP/2025/characterizing-robocalls-with-multiple-vantage-points", "CONTEXT", "telephony", "robocalls"], ["IMC/2024/poster-a-comprehensive-categorization-of-sms-scams", "CONTEXT", "telephony", "SMS scams"], ["IMC/2025/fishing-for-smishing-understanding-sms-phishing-infrastructure-and-strategies-by", "CONTEXT", "telephony", "smishing from public user reports"], ["NDSS/2015/phoneypot-data-driven-understanding-of-telephony-threats", "CONTEXT", "telephony", "telephony honeypot"], ["USENIX/2019/users-really-do-answer-telephone-scams", "CONTEXT", "telephony", "telephone-scam experiment"], ["USENIX/2020/whos-calling-characterizing-robocalls-through-audio-and-metadata-analysis", "CONTEXT", "telephony", "robocalls"], ["USENIX/2023/diving-into-robocall-content-with-snorcall", "CONTEXT", "telephony", "robocall content"], ["USENIX/2025/hey-mum-i-dropped-my-phone-down-the-toilet-investigating-hi-mum-and-dad-sms-scam", "CONTEXT", "telephony", "SMS impersonation scams"], ["CCS/2025/is-this-a-scam-the-nature-and-quality-of-reddit-discussion-about-scams", "CONTEXT", "user-study", "Reddit scam discussion (user reports as a source)"], ["IEEE-SP/2026/international-students-and-scams-at-risk-abroad", "CONTEXT", "user-study", "students' scam exposure"], ["NDSS/2025/the-kids-are-all-right-investigating-the-susceptibility-of-teens-and-adults-to-youtube-giveaway-scams", "CONTEXT", "user-study", "giveaway-scam susceptibility experiment"], ["USENIX/2024/i-experienced-more-than-10-defi-scams-on-defi-users-perception-of-security-breac", "CONTEXT", "user-study", "DeFi users' scam experiences"], ["USENIX/2025/scanned-and-scammed-insecurity-by-obsqrity-measuring-user-susceptibility-and-awa", "CONTEXT", "user-study", "QR-code scam susceptibility"], ["IEEE-SP/2015/every-second-counts-quantifying-the-negative-externalities-of-cybercrime-via-typ", "OUT", "domain-abuse", "typosquatting externalities"], ["NDSS/2015/parking-sensors-analyzing-and-detecting-parked-domains", "OUT", "domain-abuse", "parked domains"], ["NDSS/2015/seven-months-worth-of-mistakes-a-longitudinal-study-of-typosquatting-abuse", "OUT", "domain-abuse", "typosquatting"], ["WWW/2013/bitsquatting-exploiting-bit-flips-for-fun-or-profit", "OUT", "domain-abuse", "bitsquatting"], ["CCS/2021/warmonger-inflicting-denial-of-service-via-serverless-functions-in-the-cloud", "OUT", "infra", "\"Scamalytics\" IP checker homonym"], ["IMC/2010/an-empirical-study-of-orphan-dns-servers-in-the-internet", "OUT", "infra", "DNS hygiene; \"scam\" hits are feed names"], ["IMC/2011/monitoring-the-initial-dns-behavior-of-malicious-domains", "OUT", "infra", "DNS behaviour of malicious domains"], ["IMC/2016/yarrping-the-internet-randomized-high-speed-active-topology-discovery", "OUT", "infra", "\"Scamper\" homonym"], ["IMC/2018/clusters-in-the-expanse-understanding-and-unbiasing-ipv6-hitlists", "OUT", "infra", "\"Scamper\" homonym"], ["IMC/2020/flashroute-efficient-traceroute-on-a-massive-scale", "OUT", "infra", "\"Scamper\" homonym"], ["IMC/2025/remaproute-local-remapping-of-internet-path-changes", "OUT", "infra", "\"Scamper\" homonym"], ["USENIX/2023/union-under-duress-understanding-hazards-of-duplicate-resource-mismediation-in-a", "OUT", "infra", "Android libraries"], ["USENIX/2024/6sense-internet-wide-ipv6-scanning-and-its-security-applications", "OUT", "infra", "\"Scamper\" homonym"], ["CCS/2012/manufacturing-compromise-the-emergence-of-exploit-as-a-service", "OUT", "malware", "exploit-as-a-service"], ["NDSS/2014/botcoin-monetizing-stolen-cycles", "OUT", "malware", "mining malware"], ["NDSS/2025/careful-about-what-app-promotion-ads-recommend-detecting-and-explaining-malware-promotion-via-app-promotion-graph", "OUT", "malware", "mobile malware promotion"], ["USENIX/2015/post-mortem-of-a-zombie-conficker-cleanup-after-six-years", "OUT", "malware", "Conficker"], ["USENIX/2015/trends-and-lessons-from-three-years-fighting-malicious-extensions", "OUT", "malware", "malicious extensions"], ["USENIX/2015/webwitness-investigating-categorizing-and-mitigating-malware-download-paths", "OUT", "malware", "malware download paths"], ["USENIX/2016/investigating-commercial-pay-per-install-and-the-distribution-of-unwanted-softwa", "OUT", "malware", "PPI"], ["USENIX/2016/measuring-pup-prevalence-and-pup-distribution-through-pay-per-install-services", "OUT", "malware", "PUP"], ["WWW/2016/automatic-extraction-of-indicators-of-compromise-for-web-applications", "OUT", "malware", "web-app IOCs"], ["CCS/2023/evaluating-the-security-posture-of-real-world-fido2-deployments", "OUT", "mention", "FIDO2"], ["IEEE-SP/2011/design-and-evaluation-of-a-real-time-url-spam-filtering-service", "OUT", "mention", "URL spam filter"], ["IEEE-SP/2024/mawseo-adversarial-wiki-search-poisoning-for-illicit-online-promotion", "OUT", "mention", "wiki poisoning attack"], ["IMC/2018/needle-in-a-haystack-tracking-down-elite-phishing-domains-in-the-wild", "OUT", "mention", "phishing (security:phishing)"], ["NDSS/2024/understanding-and-analyzing-appraisal-systems-in-the-underground-marketplaces", "OUT", "mention", "underground marketplaces"], ["NDSS/2026/hey-there-you-are-using-whatsapp-enumerating-three-billion-accounts-for-security-and-privacy", "OUT", "mention", "WhatsApp enumeration"], ["PETS/2024/a-black-box-privacy-analysis-of-messaging-service-providers-chat-message-process", "OUT", "mention", "messaging privacy"], ["PETS/2025/sok-web-authentication-and-recovery-in-the-age-of-end-to-end-encryption", "OUT", "mention", "web authentication"], ["USENIX/2015/measuring-the-longitudinal-evolution-of-the-online-anonymous-marketplace-ecosyst", "OUT", "mention", "anonymous marketplaces; \"scam\" is exit scams"], ["USENIX/2018/plug-and-prey-measuring-the-commoditization-of-cybercrime-via-online-anonymous-m", "OUT", "mention", "anonymous-market cybercrime listings"], ["USENIX/2019/platforms-in-everything-analyzing-ground-truth-data-on-the-anatomy-and-economics", "OUT", "mention", "bullet-proof hosting"], ["USENIX/2021/phishprint-evading-phishing-detection-crawlers-by-prior-profiling", "OUT", "mention", "phishing-crawler profiling (security:phishing)"], ["USENIX/2024/malla-demystifying-real-world-large-language-model-integrated-malicious-services", "OUT", "mention", "malicious LLM services"], ["USENIX/2025/darkgram-a-large-scale-analysis-of-cybercriminal-activity-channels-on-telegram", "OUT", "mention", "Telegram cybercrime channels"], ["USENIX/2025/when-llms-go-online-the-emerging-threat-of-web-enabled-llms", "OUT", "mention", "web-enabled LLM misuse"], ["WWW/2021/websocket-adoption-and-the-landscape-of-the-real-time-web", "OUT", "mention", "scam push notifications incidental"], ["WWW/2022/conspiracy-brokers-understanding-the-monetization-of-youtube-conspiracy-theories", "OUT", "mention", "YouTube monetisation"], ["WWW/2025/harmful-terms-and-where-to-find-them-measuring-and-modeling-unfavorable-financia", "OUT", "mention", "unfavourable T&Cs on legitimate shops"], ["WWW/2026/doxing-as-a-service-demystifying-the-chinese-online-doxing-ecosystem", "OUT", "mention", "doxing"], ["CCS/2022/watch-your-back-identifying-cybercrime-financial-relationships-in-bitcoin-throug", "OUT", "onchain-other", "malware-family Bitcoin relationships"], ["IMC/2021/selfish-opaque-transaction-ordering-in-the-bitcoin-blockchain-the-case-for-chain", "OUT", "onchain-other", "Bitcoin transaction ordering"], ["NDSS/2024/abusing-the-ethereum-smart-contract-verification-services-for-fun-and-profit", "OUT", "onchain-other", "verification services"], ["NDSS/2026/houston-real-time-anomaly-detection-of-attacks-against-ethereum-defi-protocols", "OUT", "onchain-other", "DeFi attacks"], ["NDSS/2026/phishing-in-wonderland-evaluating-learning-based-ethereum-phishing-transaction-detection-and-pitfalls", "OUT", "onchain-other", "Ethereum phishing-transaction detection"], ["USENIX/2019/tracing-transactions-across-cryptocurrency-ledgers", "OUT", "onchain-other", "cross-ledger tracing"], ["USENIX/2021/evil-under-the-sun-understanding-and-discovering-attacks-on-ethereum-decentraliz", "OUT", "onchain-other", "DApp attacks"], ["USENIX/2023/mixed-signals-analyzing-ground-truth-data-on-the-users-and-economics-of-a-bitcoi", "OUT", "onchain-other", "Bitcoin mixer"], ["USENIX/2026/lost-in-blockchain-address-misuse-hidden-cross-platform-risks-and-their-security", "OUT", "onchain-other", "address misuse"], ["WWW/2024/denseflow-spotting-cryptocurrency-money-laundering-in-ethereum-transaction-graph", "OUT", "onchain-other", "money laundering"], ["WWW/2025/catalog-exploiting-joint-temporal-dependencies-for-enhanced-phishing-detection-o", "OUT", "onchain-other", "phishing-account detection on Ethereum"], ["WWW/2025/hunting-in-the-dark-forest-a-pre-trained-model-for-on-chain-attack-transaction-d", "OUT", "onchain-other", "attack transactions"], ["WWW/2025/safeguarding-blockchain-ecosystem-understanding-and-detecting-attack-transaction", "OUT", "onchain-other", "bridge attacks"], ["CCS/2012/operating-system-framed-in-case-of-mistaken-identity-measuring-the-success-of-we", "OUT", "other-fraud", "spoofed OS password dialogs (credential)"], ["CCS/2014/dialing-back-abuse-on-phone-verified-accounts", "OUT", "other-fraud", "phone-verified account markets"], ["IEEE-SP/2015/ad-injection-at-scale-assessing-deceptive-advertisement-modifications", "OUT", "other-fraud", "ad injection"], ["IMC/2014/handcrafted-fraud-and-extortion-manual-account-hijacking-in-the-wild", "OUT", "other-fraud", "account hijacking"], ["IMC/2015/affiliate-crookies-characterizing-affiliate-marketing-abuse", "OUT", "other-fraud", "affiliate fraud against merchants"], ["IMC/2015/measurement-and-analysis-of-traffic-exchange-services", "OUT", "other-fraud", "traffic exchanges"], ["IMC/2015/the-doppelganger-bot-attack-exploring-identity-impersonation-in-online-social-ne", "OUT", "other-fraud", "impersonation accounts"], ["NDSS/2013/compa-detecting-compromised-accounts-on-social-networks", "OUT", "other-fraud", "compromised accounts"], ["NDSS/2021/understanding-and-detecting-international-revenue-share-fraud", "OUT", "other-fraud", "telephony fraud against operators"], ["NDSS/2023/preventing-sim-box-fraud-using-device-model-fingerprinting", "OUT", "other-fraud", "SIM box"], ["NDSS/2025/all-your-database-are-belong-to-us-characterizing-database-ransomware-attacks", "OUT", "other-fraud", "database ransom"], ["USENIX/2021/having-your-cake-and-eating-it-an-analysis-of-concession-abuse-as-a-service", "OUT", "other-fraud", "concession abuse (fraud against merchants)"], ["USENIX/2021/socialheisting-understanding-stolen-facebook-accounts", "OUT", "other-fraud", "stolen accounts"], ["WWW/2024/identifying-risky-vendors-in-cryptocurrency-p2p-marketplaces", "OUT", "other-fraud", "P2P vendor risk"], ["CCS/2025/postmortem-voice-cloning-individuals-perspectives-of-ownership-and-deceptive-har", "OUT", "user-general", "voice cloning perceptions"], ["CCS/2025/security-and-privacy-perceptions-of-pakistani-facebook-matrimony-group-users", "OUT", "user-general", "matrimony groups"], ["IEEE-SP/2018/computer-security-and-privacy-for-refugees-in-the-united-states", "OUT", "user-general", "qualitative; scams mentioned"], ["IEEE-SP/2019/how-well-do-my-results-generalize-comparing-security-and-privacy-survey-results", "OUT", "user-general", "survey-sample comparison"], ["IEEE-SP/2023/skilled-or-gullible-gender-stereotypes-related-to-computer-security-and-privacy", "OUT", "user-general", "stereotypes survey"], ["IEEE-SP/2024/security-privacy-and-data-sharing-trade-offs-when-moving-to-the-united-states-in", "OUT", "user-general", "immigrants"], ["IEEE-SP/2025/security-and-privacy-experiences-of-first-and-second-generation-pakistani-immigr", "OUT", "user-general", "immigrants"], ["IEEE-SP/2026/searching-for-a-farang-collective-security-among-women-in-pattaya-thailand", "OUT", "user-general", "ethnography"], ["NDSS/2012/insights-into-user-behavior-in-dealing-with-internet-attacks", "OUT", "user-general", "user experiment on attack scenarios"], ["PETS/2017/social-engineering-attacks-on-government-opponents-target-perspectives", "OUT", "user-general", "targeted attacks on activists"], ["PETS/2021/warn-them-or-just-block-them-investigating-privacy-concerns-among-older-and-work", "OUT", "user-general", "older adults privacy"], ["PETS/2024/cross-contextual-examination-of-older-adults-privacy-concerns-behaviors-and-vuln", "OUT", "user-general", "older adults interviews"], ["PETS/2026/how-we-define-privacy-literacy-teaching-experiences-challenges-of-community-enga", "OUT", "user-general", "educators"], ["PETS/2026/linguistic-hooks-investigating-the-role-of-language-triggers-in-phishing-emails", "OUT", "user-general", "phishing emails"], ["USENIX/2022/they-look-at-vulnerability-and-use-that-to-abuse-you-participatory-threat-modell", "OUT", "user-general", "migrant domestic workers"], ["USENIX/2023/anatomy-of-a-high-profile-data-breach-dissecting-the-aftermath-of-a-crypto-walle", "OUT", "user-general", "Ledger breach victims survey"], ["USENIX/2023/millions-of-people-are-watching-you-understanding-the-digital-safety-needs-and-p", "OUT", "user-general", "creators safety"], ["USENIX/2025/digital-security-perceptions-and-practices-around-the-world-a-weird-versus-non-w", "OUT", "user-general", "survey"], ["USENIX/2025/security-and-privacy-advice-for-upi-users-in-india", "OUT", "user-general", "UPI advice"], ["USENIX/2025/url-inspection-tasks-helping-users-detect-phishing-links-in-emails", "OUT", "user-general", "phishing-email user study"], ["USENIX/2026/digital-risks-and-coping-practices-among-roblox-game-creators", "OUT", "user-general", "Roblox creators"], ]; // --------------------------------------------------------------------------- // Hand codes for the IN papers, read from each paper's full text // (notes/scams_papers_{A,B,C}.md hold the verbatim quotes behind every code). // One row per IN paper; report_online_scams.mjs throws if the key sets differ. export const HAND_FIELDS = { channel: { 'search': 'organic search results for chosen queries', 'search-ads': 'sponsored search results', 'ad-networks': 'malvertising / low-tier ad networks / push ads / typosquat parking / URL shorteners', 'social': 'posts, comments or livestreams on a platform', 'ct-logs': 'Certificate Transparency stream', 'new-registrations': 'newly registered domain feed', 'feeds': 'blocklists, VirusTotal or commercial scam feeds', 'user-reports': 'victim or vigilante reports (forums, Reddit, complaints)', 'spam-feeds': 'email spam feeds', 'telemetry': 'security-vendor or platform telemetry / internal index', 'leaked-backend': 'leaked or seized operator databases', 'prior-dataset': "another paper's scam list", }, cluster: { payment: 'bank accounts, acquiring banks, merchant IDs or wallet addresses', phone: 'phone numbers or messenger contacts', whois: 'WHOIS registrant fields', content: 'HTML / text similarity or store templates', screenshot: 'perceptual hash of screenshots', hosting: 'IP / DNS / hosting infrastructure', 'analytics-id': 'shared analytics or tag-manager IDs', 'redirect-graph': 'connected components of redirect chains', 'affiliate-id': 'affiliate identifiers embedded in the page', none: 'no site-to-operator grouping', }, gt: { manual: 'researchers looked at the sites', classifier: 'trained classifier (validated on a labelled sample)', heuristic: 'keyword / structural rules', 'list-as-label': 'a blocklist, VirusTotal or trust score used as the label', source: 'the source itself is the label (leaked back end, curated list, reports)', llm: 'LLM as the classifier', }, cloak: { measured: 'cloaking or victim-only serving measured (two vantages or identities compared)', addressed: 'crawler set up to look like a victim, cloaking not measured', none: 'not stated', }, harm: { 'order-numbers': 'sequential order numbers (purchase pairs)', 'test-purchase': 'completed purchases', 'leaked-revenue': 'leaked or partner revenue / transaction records', 'traffic-stats': 'visitor counts from exposed server stats or traffic panels, times a conversion rate', 'wallet-flows': 'inflows to scam wallet addresses', 'telemetry-exposure': 'users exposed / reaching checkout in vendor telemetry', 'victim-reports': 'losses stated in victim complaints or police records', none: 'not measured', }, interact: { purchases: 'bought from the scam', calls: 'phoned the scammers', chats: 'messaged the scammers', 'checkout-no-pay': 'went to the payment step and stopped', accounts: 'created victim accounts on the scam site', 'ad-clicks': 'clicked ads (cost to advertisers)', 'form-filling': 'completed survey forms automatically', none: 'no interaction', }, }; // [key, channel[], cluster[], gt[], blocklistCoverageMeasured, cloak, harm[], interact[]] export const HAND = [ ['CCS/2010/dissecting-one-click-frauds', ['user-reports'], ['payment', 'phone', 'whois'], ['source', 'manual'], true, 'none', ['victim-reports'], ['none']], ['IEEE-SP/2011/click-trajectories-end-to-end-analysis-of-the-spam-value-chain', ['spam-feeds'], ['content', 'payment'], ['heuristic', 'manual'], false, 'addressed', ['test-purchase'], ['purchases']], ['USENIX/2011/measuring-and-analyzing-search-redirection-attacks-in-the-illicit-online-prescri', ['search'], ['redirect-graph'], ['heuristic', 'source'], true, 'addressed', ['traffic-stats'], ['checkout-no-pay']], ['USENIX/2011/show-me-the-money-characterizing-spam-advertised-revenue', ['spam-feeds'], ['content'], ['heuristic'], false, 'none', ['order-numbers', 'test-purchase', 'leaked-revenue', 'traffic-stats'], ['purchases']], ['CCS/2012/priceless-the-role-of-payments-in-abuse-advertised-goods', ['user-reports', 'prior-dataset'], ['payment', 'content'], ['manual', 'heuristic'], false, 'addressed', ['test-purchase'], ['purchases', 'calls']], ['IMC/2012/tasters-choice-a-comparative-analysis-of-spam-feeds', ['spam-feeds', 'feeds'], ['affiliate-id'], ['heuristic'], true, 'none', ['leaked-revenue'], ['none']], ['USENIX/2012/pharmaleaks-understanding-the-business-of-online-pharmaceutical-affiliate-progra', ['leaked-backend'], ['none'], ['source'], false, 'none', ['leaked-revenue'], ['none']], // its orders were placed in earlier studies and reused to authenticate the leak (generic review F9) ['CCS/2014/a-nearly-four-year-longitudinal-study-of-search-engine-poisoning', ['search'], ['redirect-graph'], ['source', 'heuristic'], false, 'measured', ['none'], ['none']], ['IMC/2014/search-seizure-the-effectiveness-of-interventions-on-seo-campaigns', ['search'], ['content'], ['heuristic', 'classifier', 'manual'], false, 'measured', ['order-numbers', 'test-purchase', 'leaked-revenue', 'traffic-stats'], ['purchases', 'checkout-no-pay']], ['NDSS/2017/dial-one-for-scam-a-large-scale-analysis-of-technical-support-scams', ['ad-networks'], ['phone', 'whois'], ['heuristic', 'manual'], true, 'measured', ['traffic-stats'], ['calls']], ['IEEE-SP/2018/surveylance-automatically-detecting-online-survey-scams', ['search', 'ad-networks'], ['whois', 'screenshot'], ['classifier', 'manual'], false, 'addressed', ['none'], ['form-filling']], ['WWW/2018/betrayed-by-your-dashboard-discovering-malicious-campaigns-via-web-analytics', ['feeds', 'prior-dataset'], ['analytics-id'], ['list-as-label'], true, 'none', ['none'], ['none']], ['WWW/2018/exposing-search-and-advertisement-abuse-tactics-and-infrastructure-of-technical', ['search', 'search-ads'], ['hosting', 'content'], ['classifier'], true, 'addressed', ['none'], ['none']], ['IMC/2019/what-you-see-is-not-what-you-get-discovering-and-tracking-social-engineering-att', ['ad-networks'], ['screenshot'], ['manual', 'list-as-label'], true, 'measured', ['none'], ['ad-clicks']], ['IMC/2020/when-push-comes-to-ads-measuring-the-rise-of-malicious-push-advertising', ['ad-networks'], ['content'], ['list-as-label', 'manual'], true, 'addressed', ['none'], ['ad-clicks']], ['WWW/2021/where-are-you-taking-me-understanding-abusive-traffic-distribution-systems', ['ad-networks', 'search'], ['none'], ['manual', 'classifier'], true, 'measured', ['none'], ['none']], ['IMC/2023/evolving-bots-the-new-generation-of-comment-bots-and-their-underlying-scam-campa', ['social'], ['none'], ['manual', 'list-as-label'], false, 'none', ['none'], ['none']], ['NDSS/2023/double-and-nothing-understanding-and-detecting-cryptocurrency-giveaway-scams', ['ct-logs'], ['whois', 'payment', 'screenshot'], ['heuristic', 'manual'], true, 'addressed', ['wallet-flows'], ['none']], ['USENIX/2023/trident-towards-detecting-and-mitigating-web-based-social-engineering-attacks', ['ad-networks'], ['none'], ['manual', 'list-as-label'], false, 'none', ['none'], ['ad-clicks']], ['IEEE-SP/2023/beyond-phish-toward-detecting-fraudulent-e-commerce-websites-at-scale', ['user-reports'], ['none'], ['classifier', 'manual'], true, 'none', ['none'], ['none']], ['NDSS/2024/like-comment-get-scammed-characterizing-comment-scams-on-media-platforms', ['social'], ['phone'], ['heuristic', 'manual'], true, 'none', ['wallet-flows'], ['chats']], ['IMC/2024/give-and-take-an-end-to-end-investigation-of-giveaway-scam-conversion-rates', ['social', 'prior-dataset'], ['none'], ['heuristic', 'manual'], false, 'measured', ['wallet-flows'], ['none']], ['NDSS/2025/ctrlaltdeceive-quantifying-user-exposure-to-online-scams', ['feeds', 'telemetry'], ['none'], ['list-as-label', 'classifier', 'manual'], true, 'addressed', ['telemetry-exposure'], ['none']], ['NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns', ['new-registrations'], ['payment', 'whois'], ['classifier', 'manual'], false, 'none', ['leaked-revenue'], ['checkout-no-pay']], ['WWW/2025/the-poorest-man-in-babylon-a-longitudinal-study-of-cryptocurrency-investment-sca', ['ct-logs'], ['hosting', 'screenshot', 'phone', 'payment'], ['llm', 'manual'], true, 'none', ['wallet-flows'], ['accounts']], ['USENIX/2025/nokescam-understanding-and-rectifying-non-sense-keywords-spear-scam-in-search-en', ['telemetry', 'user-reports'], ['content', 'whois'], ['classifier', 'manual'], false, 'measured', ['victim-reports', 'traffic-stats'], ['none']], ['NDSS/2026/loki-proactively-discovering-online-scams-by-mining-toxic-search-queries', ['search', 'user-reports'], ['none'], ['classifier'], false, 'none', ['none'], ['none']], ]; // Cited on the page for a method, but outside the candidate set (no scam-destination // focus, so no probe reaches them). Listed so the citation is not mistaken for a verdict. export const EXTRA_CONTEXT = [ ['IEEE-SP/2016/cloak-of-visibility-detecting-when-machines-browse-a-different-web', 'cloaking measurement method (search and ads)'], ['NDSS/2020/into-the-deep-web-understanding-e-commerce-fraud-from-autonomous-chat-with-cybercriminals', 'automated chat with fraud crews; ethics of interacting'], ];
Reading-note quote check, unedited
node scripts/scams_notes_quotecheck.mjs 2>/dev/null > scripts/scams_notes_quotecheck-output.txt. A miss means the quote was not found verbatim in any rendering — most are column splices or elisions across a footnote, not inventions, but they were not individually checked; see Hand verdicts and hand codes.
- scams_notes_quotecheck-output.txt
notes/scams_papers_A.md: 209 of 283 quotes located notes/scams_papers_B.md: 146 of 215 quotes located notes/scams_papers_C.md: 231 of 279 quotes located ALL: 586 of 777 quotes located (75.4%) NOT LOCATED (191): CCS/2010/dissecting-one-click-frauds "Parsing the entire database yields a graph with 1,341 nodes and 5,296 edges... the whole graph G contains 105 connected subgraphs ('clusters')... Including singletons... 112 groups." CCS/2010/dissecting-one-click-frauds "We define an undirected graph G = (V,E) as follows. We create a vertex v ∈ V for each domain name, bank account number, or phone number our database contains. Then, we connect vertices belonging to th" CCS/2010/dissecting-one-click-frauds "we notice that a number of registrant entries shared similarities across seemingly different incidents... By grouping this entries together, we can link seven connected subgraphs." CCS/2010/dissecting-one-click-frauds "We check the 275 IP addresses against a set of eight blacklisting services... l2.apews.org Spam-friendly 90 (32.73%)... None of the URLs we find in our database is listed in the Google Safe Browsing d" CCS/2010/dissecting-one-click-frauds "The average, over 1,268 records in our database, stands at f = 60,462 JPY" CCS/2010/dissecting-one-click-frauds "we make a copy of a dump of our database available (in gzip'ed form) at http://arima.ini.cmu.edu/public/oneclick.sql.gz." CCS/2010/dissecting-one-click-frauds "as long as, for each scam operated, more than four people fall for the scam within a year, the miscreant turns a profit." USENIX/2011/show-me-the-money-characterizing-spam-advertised-revenue " that analyze site content. QUOTE: " USENIX/2011/show-me-the-money-characterizing-spam-advertised-revenue " effort is validating its order-volume/revenue inference against leaked accounting data. QUOTE: " USENIX/2011/show-me-the-money-characterizing-spam-advertised-revenue "we find that between May 31 and June 26, 2010, Rx-Promotion's turnover via electronic payments was $609K... this is consistent with an average revenue per order of $58, very similar to our basket-weig" USENIX/2011/show-me-the-money-characterizing-spam-advertised-revenue "we performed all of our purchases using prepaid Visa credit cards... We used a distinct card for each purchase and went to considerable lengths to emulate real customers. We used valid names and assoc" USENIX/2011/show-me-the-money-characterizing-spam-advertised-revenue "Rx-Promotion's turnover via electronic payments was $609K... this is consistent with an average revenue per order of $58, very similar to our basket-weighted average order price estimate of $57." USENIX/2011/measuring-and-analyzing-search-redirection-attacks-in-the-illicit-online-prescri "We construct a directed graph G = (V, E) as follows. We gather all URIs in our database that are part of a redirection chain ... and assign each second-level domain to a node ... We then create edges " USENIX/2011/measuring-and-analyzing-search-redirection-attacks-in-the-illicit-online-prescri "We use the spin-glass model proposed by Reichardt and Bornhold [43] (with q = 500, γ = 1) because ... it works on directed graphs." USENIX/2011/measuring-and-analyzing-search-redirection-attacks-in-the-illicit-online-prescri "Requests originating from search-engine crawlers ... return a mix of the compromised site's original content plus numerous links to websites promoted by the attacker ... This technique, 'link stuffing" USENIX/2011/measuring-and-analyzing-search-redirection-attacks-in-the-illicit-online-prescri "as a result of this 'cloaking' mechanism, some of the victim sites remain infected for a long time." USENIX/2011/measuring-and-analyzing-search-redirection-attacks-in-the-illicit-online-prescri "Conversion ≈ 0.75 × 855 000 / 20 000 000 = 3.2%." USENIX/2011/measuring-and-analyzing-search-redirection-attacks-in-the-illicit-online-prescri " only up to the start of checkout (items added to cart, payment not completed): QUOTE above (" USENIX/2011/measuring-and-analyzing-search-redirection-attacks-in-the-illicit-online-prescri "Conversion ≈ 0.75 × 855 000 / 20 000 000 = 3.2%." IMC/2012/tasters-choice-a-comparative-analysis-of-spam-feeds " (unlicensed) software — via e-mail spam advertising domains. " IMC/2012/tasters-choice-a-comparative-analysis-of-spam-feeds " spanning five collection types — botnet malware execution, MX honeypots, seeded honey accounts, human-identified (flagged by users at a large webmail provider), and domain blacklists (dbl, uribl) — c" IMC/2012/tasters-choice-a-comparative-analysis-of-spam-feeds "Live Domains Total 564,946" IMC/2012/tasters-choice-a-comparative-analysis-of-spam-feeds "Tagged Domains Total 64,116." IMC/2012/tasters-choice-a-comparative-analysis-of-spam-feeds "Such domains constituted 11-33% of domains in high-purity feeds" IMC/2012/tasters-choice-a-comparative-analysis-of-spam-feeds "Provider agreements preclude us from naming the remaining feeds (Ac1, mx2, Ac2, mx3, Hyb, Hu)" USENIX/2012/pharmaleaks-understanding-the-business-of-online-pharmaceutical-affiliate-progra " storefronts), including scheduled/controlled-substance sales at RX-Promotion. Quote: " USENIX/2012/pharmaleaks-understanding-the-business-of-online-pharmaceutical-affiliate-progra "Customers Billed orders Revenue / 584,199 699,428 $73M / 535,365 704,164 $85M / 59,769 - 69,446 71,294 $12M" USENIX/2012/pharmaleaks-understanding-the-business-of-online-pharmaceutical-affiliate-progra "For all affiliates with over $200 in revenue we link those who share an email address, ICQ number or " USENIX/2012/pharmaleaks-understanding-the-business-of-online-pharmaceutical-affiliate-progra " for 2010). Test purchases as validation, not systematic harm measurement: " USENIX/2012/pharmaleaks-understanding-the-business-of-online-pharmaceutical-affiliate-progra " describes an institutional human-subjects review process and harm-minimization controls, not direct scammer contact beyond the researchers' own historical test purchases. QUOTE: " USENIX/2012/pharmaleaks-understanding-the-business-of-online-pharmaceutical-affiliate-progra " Affiliates/employees are referenced only by non-identifying handles: " USENIX/2012/pharmaleaks-understanding-the-business-of-online-pharmaceutical-affiliate-progra " Acknowledgments cite legal/ethics oversight: " USENIX/2012/pharmaleaks-understanding-the-business-of-online-pharmaceutical-affiliate-progra " technique) is biased relative to ground truth: QUOTE: " USENIX/2012/pharmaleaks-understanding-the-business-of-online-pharmaceutical-affiliate-progra " It also flags its own data-cleaning judgment call: it discarded the first 14 months of the GlavMed dump as unreliable — QUOTE: " CCS/2014/a-nearly-four-year-longitudinal-study-of-search-engine-poisoning "antivirus, software (in general), pirated software, e-books, online gambling, and luxury items (specifically, watches)" CCS/2014/a-nearly-four-year-longitudinal-study-of-search-engine-poisoning " (drug corpus Q, from prior work) plus a " CCS/2014/a-nearly-four-year-longitudinal-study-of-search-engine-poisoning " across six other categories, collected " CCS/2014/a-nearly-four-year-longitudinal-study-of-search-engine-poisoning " via four combined datasets. " CCS/2014/a-nearly-four-year-longitudinal-study-of-search-engine-poisoning "Unique URLs in results 122 382" CCS/2014/a-nearly-four-year-longitudinal-study-of-search-engine-poisoning "Unique domains in results 30 881" CCS/2014/a-nearly-four-year-longitudinal-study-of-search-engine-poisoning "In this section we always look at traffic brokers and pharmacies at the fully-qualified domain name level," CCS/2014/a-nearly-four-year-longitudinal-study-of-search-engine-poisoning " and node degree (in+out) as the grouping unit. Overlap across product types measured via " CCS/2014/a-nearly-four-year-longitudinal-study-of-search-engine-poisoning "Because attackers are known to perform cloaking, that is, to make malicious results look benign when suspecting a visit from an automated agent as opposed to a customer, we periodically spot-checked t" CCS/2014/a-nearly-four-year-longitudinal-study-of-search-engine-poisoning "escape IP-based detection." CCS/2012/priceless-the-role-of-payments-in-abuse-advertised-goods "Taken together, we were able to identify 40 different programs (25 focused on pharmaceutical sales and 15 on the sale of counterfeit \" CCS/2012/priceless-the-role-of-payments-in-abuse-advertised-goods "we also identified a number of new programs by monitoring underground forums [12], since new programs must advertise to acquire new affiliates [11]." CCS/2012/priceless-the-role-of-payments-in-abuse-advertised-goods ") and merchant accounts/descriptors, not individual URLs. QUOTE: " CCS/2012/priceless-the-role-of-payments-in-abuse-advertised-goods "While there is still no \" CCS/2012/priceless-the-role-of-payments-in-abuse-advertised-goods " is undercover human purchasing, not automated crawling, and merchants actively try to detect/refuse undercover buyers (an analog of victim-only serving). QUOTE (purchase setup): " CCS/2012/priceless-the-role-of-payments-in-abuse-advertised-goods " QUOTE (merchant-side discrimination): " CCS/2012/priceless-the-role-of-payments-in-abuse-advertised-goods " banks after discounting trial-only banks — denominator: banks seen processing any of the 429 successful orders. QUOTE: " CCS/2012/priceless-the-role-of-payments-in-abuse-advertised-goods "refusal = processing failure" CCS/2012/priceless-the-role-of-payments-in-abuse-advertised-goods "it is clear that a subset of the programs have become far better at counter-intelligence on such undercover purchases and thus some subset of our refusals may not be due to true payment processing pro" IMC/2014/search-seizure-the-effectiveness-of-interventions-on-seo-campaigns ") doorway/storefront campaigns — " IMC/2014/search-seizure-the-effectiveness-of-interventions-on-seo-campaigns " for the same Nov 13, 2013–Jul 15, 2014 window, e.g. " IMC/2014/search-seizure-the-effectiveness-of-interventions-on-seo-campaigns ".] Seed terms came from (a) manually finding 10 doorways of the " IMC/2014/search-seizure-the-effectiveness-of-interventions-on-seo-campaigns " botnet per vertical and extracting URL-path keywords via site: queries, and (b) Google Suggest autocomplete recursion plus adjective+brand concatenation for non-KEY verticals: " IMC/2014/search-seizure-the-effectiveness-of-interventions-on-seo-campaigns " substrings) validated manually: " IMC/2014/search-seizure-the-effectiveness-of-interventions-on-seo-campaigns "we created such a data set by identifying the SEO campaigns behind a small subset of 491 Web pages... The average accuracy on held-out data was 86.8% for multiway classification of Web pages into 52 d" IMC/2014/search-seizure-the-effectiveness-of-interventions-on-seo-campaigns "We take the orders all the way to the payment processing page, which requires credit card details, before finally leaving the site. The order and customer information we provide are semantically consi" IMC/2014/search-seizure-the-effectiveness-of-interventions-on-seo-campaigns "third party legal counsel or with companies who specialize in brand protection, such as MarkMonitor, OpSec Security and Safenames." IMC/2014/search-seizure-the-effectiveness-of-interventions-on-seo-campaigns " label coverage — denominator: PSRs crawled. " IMC/2014/search-seizure-the-effectiveness-of-interventions-on-seo-campaigns " And root-only labeling gap: " IMC/2014/search-seizure-the-effectiveness-of-interventions-on-seo-campaigns " order-volume metric is only an upper bound, not actual sales: " IMC/2014/search-seizure-the-effectiveness-of-interventions-on-seo-campaigns " It also flags an internal inconsistency risk for the crawl duration itself — Section 4.1 says " IMC/2014/search-seizure-the-effectiveness-of-interventions-on-seo-campaigns " while the abstract, Section 5.1, Table 3 caption, and the conclusion all describe the identical date range as " IMC/2014/search-seizure-the-effectiveness-of-interventions-on-seo-campaigns " — a student re-deriving crawl duration from the dates (Nov 13, 2013 to Jul 15, 2014 ≈ 8 months) should trust " IEEE-SP/2011/click-trajectories-end-to-end-analysis-of-the-spam-value-chain "We obtained seven distinct URL feeds from third-party partners (including multiple commercial anti-spam providers), and harvested URLs from our own botfarm environment." IEEE-SP/2011/click-trajectories-end-to-end-analysis-of-the-spam-value-chain "Note that the 'bot' feeds tend to be focused spam sources, while the other feeds are spam sinks comprised of a blend of spam from a variety of sources. Further, individual feeds, particularly those ga" IEEE-SP/2011/click-trajectories-end-to-end-analysis-of-the-spam-value-chain "Received URLs 968,918,303 / Distinct URLs 93,185,779 (9.6%) / Distinct domains 17,813,952 / Distinct domains crawled 3,495,627 / URLs covered 950,716,776 (98.1%)" IEEE-SP/2011/click-trajectories-end-to-end-analysis-of-the-spam-value-chain "URLs 346,993,046 [pharmacy] / 3,071,828 [software] / 15,330,404 [replicas] ... Domains ... 69,002 ... Web clusters ... 1,039 ... Programs 30 / 5 / 10." IEEE-SP/2011/click-trajectories-end-to-end-analysis-of-the-spam-value-chain "we also purchased from 'canonical' instances of their sites advertised on their online support forums. We verified that they use the same bank, order number format, and email template as the spam-adve" IEEE-SP/2011/click-trajectories-end-to-end-analysis-of-the-spam-value-chain "we restricted our pharmaceutical purchasing to non-prescription goods such as herbal and over-the-counter products, and we restricted our software purchases to items for which we already possessed a s" IEEE-SP/2011/click-trajectories-end-to-end-analysis-of-the-spam-value-chain "Conversely, the 13M distinct domains produced by the Rustock bot are artifacts of a 'blacklist-poisoning' campaign undertaken by the bot operators that comprised millions of 'garbage' domains ... one " NDSS/2017/dial-one-for-scam-a-large-scale-analysis-of-technical-support-scams "Of the 5 million domains resolved by ROBOVIC, 22K URLs were detected as technical support scam pages, belonging to 8,698 unique domains." NDSS/2017/dial-one-for-scam-a-large-scale-analysis-of-technical-support-scams "out of 1,524 scam domains, only 108 (7%) were blacklisted... The rest were blacklisted, on average, 38 days after ROBOVIC's detection." NDSS/2017/dial-one-for-scam-a-large-scale-analysis-of-technical-support-scams "974 of 1524 TLD+1 domains, i.e., approx. 64%, were detectable by, on average, 3.25 AV engines." NDSS/2017/dial-one-for-scam-a-large-scale-analysis-of-technical-support-scams "our browser extension ensured the modification of the browser's user-agent properties to match a typical user browsing the web using a Microsoft Windows OS." NDSS/2017/dial-one-for-scam-a-large-scale-analysis-of-technical-support-scams "33,768 of the 1,688,412 users have paid for unnecessary technical support. With the average price of a technical support scam package being $290 ... just for the 142 mod-status-monitored domains, scam" NDSS/2017/dial-one-for-scam-a-large-scale-analysis-of-technical-support-scams "Of the 5 million domains resolved by ROBOVIC, 22K URLs were detected as technical support scam pages, belonging to 8,698 unique domains." NDSS/2017/dial-one-for-scam-a-large-scale-analysis-of-technical-support-scams "Since VirusTotal does not show the date of first discovery of a malicious domain, we cannot calculate the exact fraction of domains which ROBOVIC discovered before AVs." WWW/2018/betrayed-by-your-dashboard-discovering-malicious-campaigns-via-web-analytics " firewall-telemetry feed. " WWW/2018/betrayed-by-your-dashboard-discovering-malicious-campaigns-via-web-analytics "the data provided to us by Miramirkhani et al. [28] is by definition skewed towards social-engineering attacks, particularly of the fake technical support kind." WWW/2018/betrayed-by-your-dashboard-discovering-malicious-campaigns-via-web-analytics "identifiers enable the clustering of seemingly unrelated websites as part of a common third-party analytics account (i.e. websites whose analytics are managed by a single person or team)." WWW/2018/betrayed-by-your-dashboard-discovering-malicious-campaigns-via-web-analytics "using the VT-sourced URLs, we were able to find an average campaign size of 7.6 domains, with the largest campaign including 480 domains... for the scam-related dataset, the average campaign size was " WWW/2018/betrayed-by-your-dashboard-discovering-malicious-campaigns-via-web-analytics "836 were still active at the time of this writing with 95% of them being flagged as malicious by VirusTotal scanners" WWW/2018/betrayed-by-your-dashboard-discovering-malicious-campaigns-via-web-analytics "To faithfully mimic a user who lands on a malicious domain, our crawler is based on the headless Chrome Browser. Our crawler is capable of intercepting JavaScript alerts, simulate clicks, and extract " WWW/2018/betrayed-by-your-dashboard-discovering-malicious-campaigns-via-web-analytics "a fake-survey campaign that was live for at least 764 days (UA-11040674 seen on 4 captured domains) and a separate fake Flash Player update / fake technical support campaign that was live for at least" WWW/2018/betrayed-by-your-dashboard-discovering-malicious-campaigns-via-web-analytics "we were able to deanonymize 59 malicious actors behind VirusTotal URLs by finding public WHOIS records for other domains sharing the same analytics IDs" WWW/2018/betrayed-by-your-dashboard-discovering-malicious-campaigns-via-web-analytics "100,379 unique APKs were flagged as malware" WWW/2018/betrayed-by-your-dashboard-discovering-malicious-campaigns-via-web-analytics "9,775 were classified by Palo Alto Networks' systems as not malicious" WWW/2018/exposing-search-and-advertisement-abuse-tactics-and-infrastructure-of-technical " (scare-and-block popups) and a newly identified " WWW/2018/exposing-search-and-advertisement-abuse-tactics-and-infrastructure-of-technical " TSS type invisible to malvertising studies. QUOTE: " WWW/2018/exposing-search-and-advertisement-abuse-tactics-and-infrastructure-of-technical " Also seed-phrase choice biases toward TSS-branded language, and query popularity anti-correlates with " WWW/2018/exposing-search-and-advertisement-abuse-tactics-and-infrastructure-of-technical "the total number of unique FQDNs hosting TSS content, |Ff−tss| = 9,221... mapped to 8,104 TLD+1 domains," IMC/2019/what-you-see-is-not-what-you-get-discovering-and-tracking-social-engineering-att "we manually scouted websites and forums that discuss techniques for increasing ad revenue... we compiled an initial list of 11 different ad networks." IMC/2019/what-you-see-is-not-what-you-get-discovering-and-tracking-social-engineering-att " etc. (Table 2), though the authors explicitly designed the system to be publisher-category-agnostic. QUOTE: " IMC/2019/what-you-see-is-not-what-you-get-discovering-and-tracking-social-engineering-att "we compute a perceptual hash, specifically a 128 bit difference hash (dhash)... we define the distance function between such pairs to be the Hamming distance between the dhash values... we use DBSCAN," IMC/2019/what-you-see-is-not-what-you-get-discovering-and-tracking-social-engineering-att "we are releasing all browser logs and screenshots related to the SE attacks that we collected during our experiments. We are also making available the source code for all components of our system... T" IMC/2019/what-you-see-is-not-what-you-get-discovering-and-tracking-social-engineering-att "5th column shows the percentage of SE attack domains that have been blacklisted by Google Safe Browsing (GSB)" IMC/2020/when-push-comes-to-ads-measuring-the-rise-of-malicious-push-advertising "we manually discovered 15 popular ad networks that provide push advertisement services... we obtained a total of 87,622 HTTPS URLs of potential WPN ad publishing web pages... hosted on 82,566 distinct" IMC/2020/when-push-comes-to-ads-measuring-the-rise-of-malicious-push-advertising "WPN messages sent to mobile devices tended to be somewhat different... more tailored to mobile users... malicious mobile WPN messages included fake missed call notifications, fake amber alerts... 'spo" IMC/2020/when-push-comes-to-ads-measuring-the-rise-of-malicious-push-advertising " if it spans more than one second-level source domain. A second " IMC/2020/when-push-comes-to-ads-measuring-the-rise-of-malicious-push-advertising " step builds a bipartite graph linking WPN clusters to shared landing-page domains to catch related campaigns split by the conservative first-pass clustering, plus a " IMC/2020/when-push-comes-to-ads-measuring-the-rise-of-malicious-push-advertising "), then manually verified. Blacklist blindness measured directly: " IMC/2020/when-push-comes-to-ads-measuring-the-rise-of-malicious-push-advertising " Manual verification of blacklist hits: " IMC/2020/when-push-comes-to-ads-measuring-the-rise-of-malicious-push-advertising " (denominator: the 1,388 VT-flagged URLs). Clustering-driven relabeling raised confirmed-malicious WPN ads " IMC/2020/when-push-comes-to-ads-measuring-the-rise-of-malicious-push-advertising " and meta-clustering further raised total labeled WPN ads " IMC/2020/when-push-comes-to-ads-measuring-the-rise-of-malicious-push-advertising "To our surprise, the landing URL was neither blacklisted by Google Safe Browsing nor detected as malicious by any of the web page scanners on Virus Total" IMC/2020/when-push-comes-to-ads-measuring-the-rise-of-malicious-push-advertising "some websites... first create a dynamic JavaScript-based prompt that mimics a browser permission request... Our current version of PushAdMiner is not able to identify 'fake' JavaScript-driven permissi" WWW/2021/where-are-you-taking-me-understanding-abusive-traffic-distribution-systems "ODIN performed 874,494 scrapes over two months (June 19, 2019-August 24, 2019), posing as six different types of users... accumulating over 2TB of data." WWW/2021/where-are-you-taking-me-understanding-abusive-traffic-distribution-systems "Illicit 84,503 (17.1%)... Suspicious 17,085 (3.47%)... Malicious 11,032 (2.24%)." WWW/2021/where-are-you-taking-me-understanding-abusive-traffic-distribution-systems "grouping pages by matching text or perceptual hash, and clustering using the k-nearest neighbor algorithm (KNN). The KNN clustering uses the last layer of DenseNet 201 model trained on the ImageNet da" WWW/2021/where-are-you-taking-me-understanding-abusive-traffic-distribution-systems " (vanilla desktop, referrer-spoofed desktop, no-proxy desktop, Googlebot UA, emulated Android phone, real Nexus 6P phone), proxied through university IPs and a research-friendly VPS /24 subnet (explic" WWW/2021/where-are-you-taking-me-understanding-abusive-traffic-distribution-systems "), with anti-fingerprinting (window size, extensions, locale), self-rate-limiting scheduling, and a dedicated IP-Cloaking sub-experiment comparing a single IP vs. a 240-IP pool. QUOTE: " WWW/2021/where-are-you-taking-me-understanding-abusive-traffic-distribution-systems " IP-based cloaking measured directly: " WWW/2021/where-are-you-taking-me-understanding-abusive-traffic-distribution-systems " No evidence found of proxy-detection or mobile-emulation-detection cloaking: " WWW/2021/where-are-you-taking-me-understanding-abusive-traffic-distribution-systems "we find that 72% of all unique downloaded files are malicious, according to Virus Total. This number is 97.5% in the case of files downloaded from URL shortening services." USENIX/2023/trident-towards-detecting-and-mitigating-web-based-social-engineering-attacks "We deployed the crawlers in 20 docker containers simulating users' interactions with websites from October 2021 to January 2022 to collect training data and in October 2022 to collect data for examini" USENIX/2023/trident-towards-detecting-and-mitigating-web-based-social-engineering-attacks "Landing page screenshots clustering. During crawling, when a new tab is open, or cross-origin navigation occurs, the crawler will take a screenshot of it. Following the methodology in the study [Vadre" USENIX/2023/trident-towards-detecting-and-mitigating-web-based-social-engineering-attacks "a categorical BlockList on Github... Google Safe Browsing (GSB)... and VirusTotal (VT)," USENIX/2023/trident-towards-detecting-and-mitigating-web-based-social-engineering-attacks "Trident detects SE-ads related navigation with 92.63% accuracy, 90.63% precision, 96.28% recall, and 93.37% F-1 score" USENIX/2023/trident-towards-detecting-and-mitigating-web-based-social-engineering-attacks "), but the paper does not describe residential IPs, geography, or explicit anti-cloaking crawler design (contrast with ODIN or PushAdMiner in this set) — vantage/IP diversity is NOT STATED beyond " USENIX/2023/trident-towards-detecting-and-mitigating-web-based-social-engineering-attacks "our crawler made two clicks on the ads for each advertiser on average. Considering the average CPC (cost per click) being USD $0.75, the cost to each advertiser would be USD $1.5 on average." USENIX/2023/trident-towards-detecting-and-mitigating-web-based-social-engineering-attacks "We will release Trident source code at https://github.com/ian7yang/trident." USENIX/2023/trident-towards-detecting-and-mitigating-web-based-social-engineering-attacks "Some low-tier ad networks (e.g., 'PopAds') are known to distribute SE-ads. However, we did not find positive samples from the training dataset for these ad networks. After investigation, we found that" IEEE-SP/2018/surveylance-automatically-detecting-online-survey-scams "we collected English search terms for a period of 14 days... 10,000 most popular search items... After querying the search items using the Microsoft Web Search API, we collected 23,124 URLs... SURVEYL" IEEE-SP/2018/surveylance-automatically-detecting-online-survey-scams "SURVEYLANCE reported 8,623 survey gateways by crawling 2,301,733 URLs" IEEE-SP/2018/surveylance-automatically-detecting-online-survey-scams " unique survey-publisher domains, of which " IEEE-SP/2018/surveylance-automatically-detecting-online-survey-scams " feature, (b) cosine-similarity word-cluster analysis for false-negative triage, (c) structural-similarity screenshot clustering to categorize post-survey landing pages, and (d) WHOIS-field Levenshtei" IEEE-SP/2018/surveylance-automatically-detecting-online-survey-scams "Random Forest... TPR 95.8/97.7%, FPR 0.6/0.9%" IEEE-SP/2018/surveylance-automatically-detecting-online-survey-scams " in an ambiguous similarity band, manually finding " IEEE-SP/2018/surveylance-automatically-detecting-online-survey-scams " experiment additionally crawled with three different browser vendors (Chrome, Firefox, Internet Explorer) explicitly " IEEE-SP/2018/surveylance-automatically-detecting-online-survey-scams "we collected 2,612 unique binaries (unique MD5s) by visiting 22,057 URLs that delivered a binary, yielding 954 distinct polymorphic files. Of the distinct files, 521 samples were not previously submit" IEEE-SP/2018/surveylance-automatically-detecting-online-survey-scams "PUPs 42.2%... Adult 25%... Scams 12.4%... Malware 4.4%... Mal. Docs 3.4%" IEEE-SP/2023/beyond-phish-toward-detecting-fraudulent-e-commerce-websites-at-scale "Beyond Phish: Toward Detecting Fraudulent e-Commerce Websites at Scale" IEEE-SP/2023/beyond-phish-toward-detecting-fraudulent-e-commerce-websites-at-scale "We continuously crawled the /r/Scams subreddit from December 2019 to October 2021... We collected 16,072 submissions of which 6,233 contained live URLs... we also retroactively crawled the prior year " IEEE-SP/2023/beyond-phish-toward-detecting-fraudulent-e-commerce-websites-at-scale "includes 12,330 legitimate and 6,127 fraudulent URLs" IEEE-SP/2023/beyond-phish-toward-detecting-fraudulent-e-commerce-websites-at-scale "a 1.98% false positive rate (FPR) and 1.63% false negative rate (FNR), accounting for 1.86% label noise" IEEE-SP/2023/beyond-phish-toward-detecting-fraudulent-e-commerce-websites-at-scale "13,917 responses from 8,174 users on 2,223 submissions... BeyondPhish predicted the correct label 98.38% of the time." IEEE-SP/2023/beyond-phish-toward-detecting-fraudulent-e-commerce-websites-at-scale "attackers need to spend at least $1,000 to buy an aged domain" IEEE-SP/2023/beyond-phish-toward-detecting-fraudulent-e-commerce-websites-at-scale "we release our collected datasets (though not the proprietary dataset provided by Palo Alto Networks), source code, and the BeyondPhish model" IEEE-SP/2023/beyond-phish-toward-detecting-fraudulent-e-commerce-websites-at-scale " price signal used by prior fake-shop detectors (Carpineto & Romano, Wadleigh et al.) has eroded over time as scammers adapted, which the paper flags as a reason older detection features no longer tra" IMC/2023/evolving-bots-the-new-generation-of-comment-bots-and-their-underlying-scam-campa ", 4.88% of videos), e-commerce/counterfeit-discount scams (3 campaigns), malvertising/malware-download phishing (1 campaign), plus miscellaneous and " IMC/2023/evolving-bots-the-new-generation-of-comment-bots-and-their-underlying-scam-campa " (URL-shortener-suspended) categories. Quote: " IMC/2023/evolving-bots-the-new-generation-of-comment-bots-and-their-underlying-scam-campa ") profile/channel pages those bots controlled — not from an external blocklist/CT/ad-network feed. QUOTE: " IMC/2023/evolving-bots-the-new-generation-of-comment-bots-and-their-underlying-scam-campa " total commenters (Table 1). After DBSCAN clustering + filtering: " IMC/2023/evolving-bots-the-new-generation-of-comment-bots-and-their-underlying-scam-campa "; (2) domain-level clustering by shared second-level domain (SLD) extracted from bot channel pages to group SSBs into " IMC/2023/evolving-bots-the-new-generation-of-comment-bots-and-their-underlying-scam-campa " metric (video views × engagement-rate²) used as a harm proxy, not real financial loss: " IMC/2023/evolving-bots-the-new-generation-of-comment-bots-and-their-underlying-scam-campa " section on GDPR-aware crawling and PII minimization. QUOTE: " IMC/2023/evolving-bots-the-new-generation-of-comment-bots-and-their-underlying-scam-campa " finding that SSBs reply to their own campaign's other bot comments to game ranking: " IMC/2023/evolving-bots-the-new-generation-of-comment-bots-and-their-underlying-scam-campa " were found to be externally sourced — QUOTE: " IMC/2024/give-and-take-an-end-to-end-investigation-of-giveaway-scam-conversion-rates "); YouTube — 343 domains, 1,632 accounts (channels), 2,069 artifacts (livestreams). " IMC/2024/give-and-take-an-end-to-end-investigation-of-giveaway-scam-conversion-rates " (from the 14-day pilot, a subset of the full run). No further clustering of domains into " IMC/2024/give-and-take-an-end-to-end-investigation-of-giveaway-scam-conversion-rates " via HTML/kits/etc. The paper instead clusters cryptocurrency *addresses* via Chainalysis: " IMC/2024/give-and-take-an-end-to-end-investigation-of-giveaway-scam-conversion-rates "the conversion rate of viewers to victims was 0.0039%-or roughly 4 in 100,000 views netting a victim" NDSS/2023/double-and-nothing-understanding-and-detecting-cryptocurrency-giveaway-scams "we use deterministic identifiers to cluster scams into campaigns. To identify connected components which are representative of scam campaigns, we merge the domains whose WHOIS information listed the s" NDSS/2023/double-and-nothing-understanding-and-detecting-cryptocurrency-giveaway-scams "there are additional evasion techniques (e.g. waiting for users to move their mouse before revealing their content, or accepting web notifications) which CryptoScamTracker does not currently handle." NDSS/2023/double-and-nothing-understanding-and-detecting-cryptocurrency-giveaway-scams "Bitcoin (BTC) 860 [unique wallets]... Total Cryptocurrency Amount 940.07... Total USD Value (Min.-Max) $17.8M-$44.9M." NDSS/2023/double-and-nothing-understanding-and-detecting-cryptocurrency-giveaway-scams "belong to online exchanges (such as Coinbase), where multiple victims can share the same outgoing address" NDSS/2024/like-comment-get-scammed-characterizing-comment-scams-on-media-platforms "We define a domain as malicious if at least 3 of the 90 tools that are integrated into VirusTotal labeled it as either 'Suspicious' or 'Malicious.' In total, out of the 24 URLs we submitted, only 1 UR" NDSS/2024/like-comment-get-scammed-characterizing-comment-scams-on-media-platforms "we designed a time-based filter... we captured periodic 'snapshots' of the comments every hour." NDSS/2024/like-comment-get-scammed-characterizing-comment-scams-on-media-platforms "This result stands in contrast to other cryptocurrency-related scams [20], [3], [21], which were characterized by short-lived infrastructure. Our findings suggest this difference is largely due to the" WWW/2025/the-poorest-man-in-babylon-a-longitudinal-study-of-cryptocurrency-investment-sca " tool, and lists as a future-work limitation: " NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns "ScamMagnifier: Piercing the Veil of Fraudulent Shopping Website Campaigns" NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns "ScamMagnifier collected 1,155,237 shopping domains from May 2023 to June 2024." NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns "the checkout process for 41,863 domains... ScamMagnifier was able to extract merchant IDs for three different payment processors for 5,278 total. Of these, 4,484 (the vast majority) were for financial" NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns " is merchant-ID based, not feature/template-based: sites sharing a merchant ID are treated as one operator, and Org A additionally links merchant IDs by shared WHOIS registrant identity: " NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns " Separately, a simple visual clustering was done on screenshots: " NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns "Our security experts analyzed a subset of 1,000 websites identified as a scam by Beyond Phish. This method had 23 false positives which is insignificant." NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns "Note that the axis are deliberately unspecified — financial Org A considers transaction volume to be sensitive, therefore this figure only shows the distribution." NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns "ScamMagnifier was able to extract merchant IDs for three different payment processors for 5,278 total. Of these, 4,484... were for financial Org A" NDSS/2026/loki-proactively-discovering-online-scams-by-mining-toxic-search-queries "The collected websites are passed through the oracle classifier, which labels 52,493 (19.3%) of them as scams. We refer to this set of newly identified fraudulent websites as the 'discovered scams.'" NDSS/2026/loki-proactively-discovering-online-scams-by-mining-toxic-search-queries " is used loosely to mean scam category/vertical, not a clustered group of related sites, e.g. " NDSS/2026/loki-proactively-discovering-online-scams-by-mining-toxic-search-queries " classifier (Gradient Boosting over 103 features) decides scam vs. benign, not manual review of the wild-discovered set: " NDSS/2026/loki-proactively-discovering-online-scams-by-mining-toxic-search-queries " For the 52,493 discovered scams there is no ground truth, so third-party corroboration is used instead: " NDSS/2026/loki-proactively-discovering-online-scams-by-mining-toxic-search-queries "over 90% of scam-linked Instagram accounts have fewer than 5,000 followers, and approximately 50% fall below the 500 follower mark" NDSS/2026/loki-proactively-discovering-online-scams-by-mining-toxic-search-queries "This work is not considered human subjects research by our institution, since we do not interact with humans and do not collect any private information." NDSS/2026/loki-proactively-discovering-online-scams-by-mining-toxic-search-queries "the query ranking outputs can be misused by adversaries as part of curating their own 'content blacklist' to avoid using these keywords in their page source... to prevent these webpages from being ind" NDSS/2026/loki-proactively-discovering-online-scams-by-mining-toxic-search-queries "LOKI (which we make publicly available) will be an effective tool..." NDSS/2026/loki-proactively-discovering-online-scams-by-mining-toxic-search-queries "supplementary data files (raw_data_keywords.csv, query_ner_output.json, keywords_categories.json)" NDSS/2026/loki-proactively-discovering-online-scams-by-mining-toxic-search-queries "we rigorously validate LOKI across 10 major scam categories and demonstrate a 20.58 times improvement in discovery over both heuristic and data-driven baselines across all categories." USENIX/2025/nokescam-understanding-and-rectifying-non-sense-keywords-spear-scam-in-search-en "), pornography scams, and gambling scams — all delivered via search-engine " USENIX/2025/nokescam-understanding-and-rectifying-non-sense-keywords-spear-scam-in-search-en " using non-sense keywords (NSKeywords). " USENIX/2025/nokescam-understanding-and-rectifying-non-sense-keywords-spear-scam-in-search-en " — newly observed/indexed webpages — dataset), sourced from Baidu's live crawl/indexing pipeline, plus a seed ground-truth set drawn from Baidu user complaints. " USENIX/2025/nokescam-understanding-and-rectifying-non-sense-keywords-spear-scam-in-search-en "We designed a clustering methodology for NOKEScam based on two assumptions: 1) websites with similar title patterns, and 2) websites using the same domain registration information, belong to the same " USENIX/2025/nokescam-understanding-and-rectifying-non-sense-keywords-spear-scam-in-search-en "all domains were screened using VirusTotal, resulting in a final dataset of 35,000 legitimate websites." USENIX/2025/nokescam-understanding-and-rectifying-non-sense-keywords-spear-scam-in-search-en "By examining the last 100 complaint reports, we find 20 complaints mentioned the amount of money defrauded, ranging from 100 to 33,443 dollars. The average fraud amount per NOKEScam victim was $2,896." USENIX/2025/nokescam-understanding-and-rectifying-non-sense-keywords-spear-scam-in-search-en "Our analysis strictly followed ethical guidelines, namely the Belmont Report and the Menlo Report... before analysis, Baidu manually reviewed the data and anonymized any personal information by comput" USENIX/2025/nokescam-understanding-and-rectifying-non-sense-keywords-spear-scam-in-search-en "We released our detection script and detected data on http://nokescam.com. However, as our detection data source (Baidu's actual indexing data) cannot be disclosed, thus a direct evaluation of the det" USENIX/2025/nokescam-understanding-and-rectifying-non-sense-keywords-spear-scam-in-search-en "we identified 153,975 NSKeywords across 68,863 domain names, indicating the scale of NOKEScam"
Recall probe, unedited
The title-plus-summary and ad-library probes described under Probes and the candidate set (node scripts/scams_recall_probe.mjs > scripts/scams_recall_probe-output.txt).
- scams_recall_probe-output.txt
title+summary probe outside candidates (empirical, web): 36 2011 IMC Understanding fraudulent activities in online ad exchanges. 2014 CCS Characterizing Large-Scale Click Fraud in ZeroAccess. 2014 USENIX Understanding the Dark Side of Domain Parking 2015 NDSS Liar Buyer Fraud, and How to Curb It 2015 WWW Understanding Malvertising Through Ad-Injecting Browser Extensions. 2017 NDSS Fake Co-visitation Injection Attacks to Recommender Systems 2017 WWW Neural Underpinnings of Website Legitimacy and Familiarity Detection: An fNIRS Study. 2019 IEEE-SP RIDL: Rogue In-Flight Data Load. 2019 IEEE-SP Stealthy Porn: Understanding Real-World Adversarial Images for Illicit Online Promotion. 2019 USENIX All Your Clicks Belong to Me: Investigating Click Interception on the Web 2019 WWW Revisiting Mobile Advertising Threats with MAdLife. 2019 WWW Think Outside the Dataset: Finding Fraudulent Reviews using Cross-Dataset Analysis. 2019 WWW What happened? The Spread of Fake News Publisher Content During the 2016 U.S. Presidential 2020 NDSS Deceptive Previews: A Study of the Link Preview Trustworthiness in Social Platforms 2020 NDSS Into the Deep Web: Understanding E-commerce Fraud from Autonomous Chat with Cybercriminals 2020 PETS Multiple Purposes, Multiple Problems: A User Study of Consent Dialogs after GDPR 2020 USENIX Cached and Confused: Web Cache Deception in the Wild 2020 USENIX Sunrise to Sunset: Analyzing the End-to-end Life Cycle and Effectiveness of Phishing Attac 2021 USENIX AdCube: WebVR Ad Fraud and Practical Confinement of Third-Party Ads 2021 WWW Deepfake Videos in the Wild: Analysis and Detection. 2022 USENIX Web Cache Deception Escalates! 2022 WWW A View into YouTube View Fraud. 2023 IEEE-SP Deepfake Text Detection: Limitations and Opportunities. 2024 IMC Browser Polygraph: Efficient Deployment of Coarse-Grained Browser Fingerprints for Web-Sca 2024 CCS Breaching Security Keys without Root: FIDO2 Deception Attacks via Overlays exploiting Limi 2024 CCS ProFake: Detecting Deepfakes in the Wild against Quality Degradation with Progressive Qual 2023 WWW Who Funds Misinformation? A Systematic Analysis of the Ad-related Profit Routines of Fake 2024 USENIX Moderating Illicit Online Image Promotion for Unsafe User Generated Content Games Using La 2024 USENIX FakeBehalf: Imperceptible Email Spoofing Attacks against the Delegation Mechanism in Email 2025 NDSS Characterizing the Impact of Audio Deepfakes in the Presence of Cochlear Implant 2025 WWW Welcome to the Dark Side: Analyzing the Revenue Flows of Fraud in the Online Ad Ecosystem. 2025 WWW 50 Shades of Deceptive Patterns: A Unified Taxonomy, Multimodal Detection, and Security Im 2025 CCS Automatically Detecting Online Deceptive Patterns. 2025 CCS Phishing Susceptibility and the (In-)Effectiveness of Common Anti-Phishing Interventions i 2025 USENIX Phishing Attacks against Password Manager Browser Extensions 2025 USENIX Characterizing the MrDeepFakes Sexual Deepfake Marketplace ad library/archive AND scam|fraud in title/summary/tools: 0
Bibliography
Entries appended to bibliography in this sitting, generated by bibgen.mjs from the corpus index (publisher DOIs where the index has them). The two USENIX entries had no authors in the index: NOKEScam's 13 authors are from its USENIX landing page; Leontiadis, Moore and Christin from the USENIX Security 2011 technical-sessions page.
- bib_additions_online_scams.bib
@inproceedings{christin2010_dissecting, author = {Christin, Nicolas and Yanagihara, Sally S. and Kamataki, Keisuke}, title = {Dissecting one click frauds}, booktitle = {Proceedings of the ACM SIGSAC Conference on Computer and Communications Security}, year = {2010}, series = {CCS 2010}, doi = {10.1145/1866307.1866310}, } @inproceedings{mccoy2012_priceless, author = {McCoy, Damon and Dharmdasani, Hitesh and Kreibich, Christian and Voelker, Geoffrey M. and Savage, Stefan}, title = {Priceless: the role of payments in abuse-advertised goods}, booktitle = {Proceedings of the ACM SIGSAC Conference on Computer and Communications Security}, year = {2012}, series = {CCS 2012}, doi = {10.1145/2382196.2382285}, } @inproceedings{wang2014_search, author = {Wang, David Y. and Der, Matthew F. and Karami, Mohammad and Saul, Lawrence K. and McCoy, Damon and Savage, Stefan and Voelker, Geoffrey M.}, title = {Search + Seizure: The Effectiveness of Interventions on SEO Campaigns}, booktitle = {Proceedings of the ACM Internet Measurement Conference}, year = {2014}, series = {IMC 2014}, doi = {10.1145/2663716.2663738}, } @inproceedings{starov2018_betrayed, author = {Starov, Oleksii and Zhou, Yuchen and Zhang, Xiao and Miramirkhani, Najmeh and Nikiforakis, Nick}, title = {Betrayed by Your Dashboard: Discovering Malicious Campaigns via Web Analytics}, booktitle = {Proceedings of the ACM Web Conference}, year = {2018}, series = {TheWebConf 2018}, doi = {10.1145/3178876.3186089}, } @inproceedings{srinivasan2018_exposing, author = {Srinivasan, Bharat and Kountouras, Athanasios and Miramirkhani, Najmeh and Alam, Monjur and Nikiforakis, Nick and Antonakakis, Manos and Ahamad, Mustaque}, title = {Exposing Search and Advertisement Abuse Tactics and Infrastructure of Technical Support Scammers}, booktitle = {Proceedings of the ACM Web Conference}, year = {2018}, series = {TheWebConf 2018}, doi = {10.1145/3178876.3186098}, } @inproceedings{kharraz2018_surveylance, author = {Kharraz, Amin and Robertson, William K. and Kirda, Engin}, title = {Surveylance: Automatically Detecting Online Survey Scams}, booktitle = {Proceedings of the IEEE Symposium on Security and Privacy}, year = {2018}, series = {IEEE S&P 2018}, doi = {10.1109/sp.2018.00044}, } @inproceedings{na2023_evolving, author = {Na, Seung Ho and Cho, Sumin and Shin, Seungwon}, title = {Evolving Bots: The New Generation of Comment Bots and their Underlying Scam Campaigns in YouTube}, booktitle = {Proceedings of the ACM Internet Measurement Conference}, year = {2023}, series = {IMC 2023}, doi = {10.1145/3618257.3624822}, } @inproceedings{li2023_double, author = {Li, Xigao and Yepuri, Anurag and Nikiforakis, Nick}, title = {Double and Nothing: Understanding and Detecting Cryptocurrency Giveaway Scams}, booktitle = {Proceedings of the Network and Distributed System Security Symposium}, year = {2023}, series = {NDSS 2023}, url = {https://www.ndss-symposium.org/ndss-paper/double-and-nothing-understanding-and-detecting-cryptocurrency-giveaway-scams/}, } @inproceedings{li2024_like, author = {Li, Xigao and Rahmati, Amir and Nikiforakis, Nick}, title = {Like, Comment, Get Scammed: Characterizing Comment Scams on Media Platforms}, booktitle = {Proceedings of the Network and Distributed System Security Symposium}, year = {2024}, series = {NDSS 2024}, url = {https://www.ndss-symposium.org/ndss-paper/like-comment-get-scammed-characterizing-comment-scams-on-media-platforms/}, } @inproceedings{liu2024_give, author = {Liu, Enze and Kappos, George and Mugnier, Eric and Invernizzi, Luca and Savage, Stefan and Tao, David and Thomas, Kurt and Voelker, Geoffrey M. and Meiklejohn, Sarah}, title = {Give and Take: An End-To-End Investigation of Giveaway Scam Conversion Rates}, booktitle = {Proceedings of the ACM Internet Measurement Conference}, year = {2024}, series = {IMC 2024}, doi = {10.1145/3646547.3689005}, } @inproceedings{bitaab2023_beyond, author = {Bitaab, Marzieh and Cho, Haehyun and Oest, Adam and Lyu, Zhuoer and Wang, Wei and Abraham, Jorij and Wang, Ruoyu and Bao, Tiffany and Shoshitaishvili, Yan and Doupé, Adam}, title = {Beyond Phish: Toward Detecting Fraudulent e-Commerce Websites at Scale}, booktitle = {Proceedings of the IEEE Symposium on Security and Privacy}, year = {2023}, series = {IEEE S&P 2023}, doi = {10.1109/sp46215.2023.10179461}, } @inproceedings{kotzias2025_ctrl, author = {Kotzias, Platon and Pachilakis, Michalis and Iuit, Javier Aldana and Caballero, Juan and Sanchez-Rola, Iskander and Bilge, Leyla}, title = {Ctrl+Alt+Deceive: Quantifying User Exposure to Online Scams}, booktitle = {Proceedings of the Network and Distributed System Security Symposium}, year = {2025}, series = {NDSS 2025}, url = {https://www.ndss-symposium.org/ndss-paper/ctrlaltdeceive-quantifying-user-exposure-to-online-scams/}, } @inproceedings{muzammil2025_poorest, author = {Muzammil, Muhammad and Pitumpe, Abisheka and Li, Xigao and Rahmati, Amir and Nikiforakis, Nick}, title = {The Poorest Man in Babylon: A Longitudinal Study of Cryptocurrency Investment Scams}, booktitle = {Proceedings of the ACM Web Conference}, year = {2025}, series = {TheWebConf 2025}, doi = {10.1145/3696410.3714588}, } @inproceedings{paudel2026_loki, author = {Paudel, Pujan and Stringhini, Gianluca}, title = {LOKI: Proactively Discovering Online Scam Websites by Mining Toxic Search Queries}, booktitle = {Proceedings of the Network and Distributed System Security Symposium}, year = {2026}, series = {NDSS 2026}, url = {https://www.ndss-symposium.org/ndss-paper/loki-proactively-discovering-online-scams-by-mining-toxic-search-queries/}, } @inproceedings{liu2025_nokescam, author = {Liu, Mingxuan and Zhang, Yunyi and Wu, Lijie and Liu, Baojun and Hong, Geng and Zhang, Yiming and Jiang, Hui and Zhang, Jia and Duan, Haixin and Zhang, Min and Guan, Wei and Shi, Fan and Yang, Min}, title = {NOKEScam: Understanding and Rectifying Non-Sense Keywords Spear Scam in Search Engines}, booktitle = {Proceedings of the USENIX Security Symposium}, year = {2025}, series = {USENIX Security 2025}, url = {https://www.usenix.org/conference/usenixsecurity25/presentation/liu-mingxuan}, } @inproceedings{gomez2023_cybercrime, author = {Gómez, Gibran and Liebergen, Kevin van and Caballero, Juan}, title = {Cybercrime Bitcoin Revenue Estimations: Quantifying the Impact of Methodology and Coverage}, booktitle = {Proceedings of the ACM SIGSAC Conference on Computer and Communications Security}, year = {2023}, series = {CCS 2023}, doi = {10.1145/3576915.3623094}, } @inproceedings{wang2020_into, author = {Wang, Peng and Liao, Xiaojing and Qin, Yue and Wang, XiaoFeng}, title = {Into the Deep Web: Understanding E-commerce Fraud from Autonomous Chat with Cybercriminals}, booktitle = {Proceedings of the Network and Distributed System Security Symposium}, year = {2020}, series = {NDSS 2020}, url = {https://www.ndss-symposium.org/ndss-paper/into-the-deep-web-understanding-e-commerce-fraud-from-autonomous-chat-with-cybercriminals/}, } @inproceedings{he2023_txphishscope, author = {He, Bowen and Chen, Yuan and Chen, Zhuo and Hu, Xiaohui and Hu, Yufeng and Wu, Lei and Chang, Rui and Wang, Haoyu and Zhou, Yajin}, title = {TxPhishScope: Towards Detecting and Understanding Transaction-based Phishing on Ethereum}, booktitle = {Proceedings of the ACM SIGSAC Conference on Computer and Communications Security}, year = {2023}, series = {CCS 2023}, doi = {10.1145/3576915.3623210}, } @inproceedings{leontiadis2011_measuring, author = {Leontiadis, Nektarios and Moore, Tyler and Christin, Nicolas}, title = {Measuring and Analyzing Search-Redirection Attacks in the Illicit Online Prescription Drug Trade}, booktitle = {Proceedings of the USENIX Security Symposium}, year = {2011}, series = {USENIX Security 2011}, url = {https://www.usenix.org/legacy/event/sec11/tech/}, }
References
- [1]
- Park, Youngsam; Jones, Jackie; McCoy, Damon; Shi, Elaine; Jakobsson, Markus (2014): "Scambaiter: Understanding Targeted Nigerian Scams on Craigslist", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
- [2]
- Gómez, Gibran; Liebergen, Kevin van; Caballero, Juan (2023): "Cybercrime Bitcoin Revenue Estimations: Quantifying the Impact of Methodology and Coverage", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
- [3]
- He, Bowen; Chen, Yuan; Chen, Zhuo; Hu, Xiaohui; Hu, Yufeng; Wu, Lei; Chang, Rui; Wang, Haoyu; Zhou, Yajin (2023): "TxPhishScope: Towards Detecting and Understanding Transaction-based Phishing on Ethereum", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
