User Tools

Site Tools


security:online_scams

Online scams: fake shops, support scams and crypto scams

You are about to measure scam websites — a fake shop, a tech-support page that locks the browser and shows a phone number, a crypto “giveaway” that promises to double what you send, an investment platform that shows a wallet address after sign-up. Three decisions decide what your number means, and a paper that makes them silently is not reproducible:

  1. Where the seed list comes from. Every discovery channel finds a different population, and two channels aimed at the same scam type can share almost nothing.
  2. Whether you group sites into operators. Scam domains are cheap and disposable; a count of URLs is mostly a count of how often one crew re-registered.
  3. What you count as harm. A live scam page is not a victim. The papers that say something about harm measured orders, payments, wallets or checkout visits — never the victims' total loss.

Labelling and cloaking follow from those three, and so does the ethics of talking to, buying from or reporting the people you are measuring.

This page is about scam websites and campaigns as a web phenomenon. The line with Phishing is what the victim hands over. Credentials, a seed phrase, or a signature that gives the attacker authority over the account is phishing — including wallet drainers and the “transaction-based phishing” of [1He, Bowen; Chen, Yuan; Chen, Zhuo; Hu, Xiaohui; Hu, Yufeng; Wu, Lei; Chang, Rui; Wang, Haoyu; Zhou, Yajin (2023): "TxPhishScope: Towards Detecting and Understanding Transaction-based Phishing on Ethereum", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)]. Money the victim knowingly sends, believing they get goods, support or returns, belongs here. Scam accounts, comments and posts as the object of study are Platforms; this page picks a platform study up only when the off-platform scam site or its wallet is among what was measured. Phone and SMS scams, on-chain scam tokens with no website, and studies of who falls for scams are outside; they are listed under Use in publications as context, not counted.

A count of scam URLs is not a finding. Four measured facts a methods section has to survive:

  • Sites are not operators. [2Bitaab, Marzieh; Karimi, Alireza; Lyu, Zhuoer; Oest, Adam; Kuchhal, Dhruv; Saad, Muhammad; Ahn, Gail-Joon; Wang, Ruoyu; Bao, Tiffany; Shoshitaishvili, Yan; Doupé, Adam (2025): "ScamMagnifier: Piercing the Veil of Fraudulent Shopping Website Campaigns", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] completed automated checkouts on fraudulent shops and extracted the payment processor's merchant IDs: of the 14,394 domains one processor could link to those merchants, 54.55% were run by 10 merchant IDs, and one merchant ID alone served 974 domains. [3Wang, David Y.; Der, Matthew F.; Karami, Mohammad; Saul, Lawrence K.; McCoy, Damon; Savage, Stefan; Voelker, Geoffrey M. (2014): "Search + Seizure: The Effectiveness of Interventions on SEO Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] attributed 58% of 2.7M poisoned search results to 52 SEO campaigns — which ran only 11% of the stores they saw.
  • Blocklists are late for scams, not merely for phishing. Of 1,524 tech-support scam domains, 108 (7%) were blacklisted when checked; 16 of those were listed on the day the crawler first saw them, the rest on average 38 days later [4Miramirkhani, Najmeh; Starov, Oleksii; Nikiforakis, Nick (2017): "Dial One for Scam: A Large-Scale Analysis of Technical Support Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]. VirusTotal knew 16.75% of 3,610 crypto-giveaway domains [5Li, Xigao; Yepuri, Anurag; Nikiforakis, Nick (2023): "Double and Nothing: Understanding and Detecting Cryptocurrency Giveaway Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]; Google Safe Browsing flagged 10 of 6,127 fraudulent shops [6Bitaab, Marzieh; Cho, Haehyun; Oest, Adam; Lyu, Zhuoer; Wang, Wei; Abraham, Jorij; Wang, Ruoyu; Bao, Tiffany; Shoshitaishvili, Yan; Doupé, Adam (2023): "Beyond Phish: Toward Detecting Fraudulent e-Commerce Websites at Scale", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)].
  • The crawler is not the victim. A crawler on a campus network found 95.7% of all the tech-support scam domains that three crawlers found together; two crawlers on commercial clouds found far fewer, because the ad networks filter by IP [4Miramirkhani, Najmeh; Starov, Oleksii; Nikiforakis, Nick (2017): "Dial One for Scam: A Large-Scale Analysis of Technical Support Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]. Crawling as six kinds of user found 81% more malicious landing pages than the best single profile [7Szurdi, Janos; Luo, Meng; Kondracki, Brian; Nikiforakis, Nick; Christin, Nicolas (2021): "Where are you taking me?Understanding Abusive Traffic Distribution Systems", in: Proceedings of the ACM Web Conference. (DOI)].
  • Harm is a transaction, not a page. Crypto-giveaway tweets that carried a wallet address converted at 0.12% — about one victim per thousand tweets; livestream views at 0.0039% [8Liu, Enze; Kappos, George; Mugnier, Eric; Invernizzi, Luca; Savage, Stefan; Tao, David; Thomas, Kurt; Voelker, Geoffrey M.; Meiklejohn, Sarah (2024): "Give and Take: An End-To-End Investigation of Giveaway Scam Conversion Rates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]. Wallet inflows of $2.04M traced to crypto investment scams came from sites that are 6.7% of the 43,572 the authors detected [9Muzammil, Muhammad; Pitumpe, Abisheka; Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2025): "The Poorest Man in Babylon: A Longitudinal Study of Cryptocurrency Investment Scams", in: Proceedings of the ACM Web Conference. (DOI)].

Name your channel, your grouping rule and your harm measure, with dates. “We crawled 40,000 scam sites” names none of them.

What to read first

Paper Why now
Kotzias et al., NDSS 2025 [10Kotzias, Platon; Pachilakis, Michalis; Iuit, Javier Aldana; Caballero, Juan; Sanchez-Rola, Iskander; Bilge, Leyla (2025): "Ctrl+Alt+Deceive: Quantifying User Exposure to Online Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] The exposure view: 607K scam domains from two feeds, joined with a security vendor's telemetry of 196.9 billion desktop URL visits. Shopping scams reached 10.2M IP addresses, crypto scams 653K; 4% of IPs that visited a shopping scam reached its checkout page. Read it for what a feed-plus-telemetry design can and cannot say — the authors list their own geographic and desktop biases.
Bitaab et al., NDSS 2025 [2Bitaab, Marzieh; Karimi, Alireza; Lyu, Zhuoer; Oest, Adam; Kuchhal, Dhruv; Saad, Muhammad; Ahn, Gail-Joon; Wang, Ruoyu; Bao, Tiffany; Shoshitaishvili, Yan; Doupé, Adam (2025): "ScamMagnifier: Piercing the Veil of Fraudulent Shopping Website Campaigns", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] Current fake-shop method: a newly-registered-domain feed, a classifier, automated checkout that stops before paying, and merchant IDs as the grouping key — with a payment processor as partner.
Muzammil et al., TheWebConf 2025 [9Muzammil, Muhammad; Pitumpe, Abisheka; Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2025): "The Poorest Man in Babylon: A Longitudinal Study of Cryptocurrency Investment Scams", in: Proceedings of the ACM Web Conference. (DOI)] Certificate Transparency discovery plus an LLM classifier validated against a hand-labelled sample, fake victim accounts to reveal wallet addresses, and blocklist coverage measured on a third of the result. The only paper in this population that classifies scam sites with an LLM.
Liu et al., IMC 2024 [8Liu, Enze; Kappos, George; Mugnier, Eric; Invernizzi, Luca; Savage, Stefan; Tao, David; Thomas, Kurt; Voelker, Geoffrey M.; Meiklejohn, Sarah (2024): "Give and Take: An End-To-End Investigation of Giveaway Scam Conversion Rates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] End-to-end conversion: from scam tweets and livestreams to landing pages to payments, with revenue bracketed between payments that co-occur with the promotion and all inflows. Read it before you publish a single revenue number.
Szurdi et al., TheWebConf 2021 [7Szurdi, Janos; Luo, Meng; Kondracki, Brian; Nikiforakis, Nick; Christin, Nicolas (2021): "Where are you taking me?Understanding Abusive Traffic Distribution Systems", in: Proceedings of the ACM Web Conference. (DOI)] The vantage experiment: six user profiles, a single IP against a 240-IP pool, and what Google Safe Browsing had listed on the day of detection. The paper to cite when a reviewer asks whether your crawler saw what a victim saw.
Srinivasan et al., TheWebConf 2018 [11Srinivasan, Bharat; Kountouras, Athanasios; Miramirkhani, Najmeh; Alam, Monjur; Nikiforakis, Nick; Antonakakis, Manos; Ahamad, Mustaque (2018): "Exposing Search and Advertisement Abuse Tactics and Infrastructure of Technical Support Scammers", in: Proceedings of the ACM Web Conference. (DOI)] with Miramirkhani et al., NDSS 2017 [4Miramirkhani, Najmeh; Starov, Oleksii; Nikiforakis, Nick (2017): "Dial One for Scam: A Large-Scale Analysis of Technical Support Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] The same scam type found through search and through malvertising: 0 of 2,768 fully qualified domains in common. The clearest demonstration that the channel is the population.
Paudel and Stringhini, NDSS 2026 [12Paudel, Pujan; Stringhini, Gianluca (2026): "LOKI: Proactively Discovering Online Scam Websites by Mining Toxic Search Queries", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] Search-query mining as a discovery channel, with models and data released on Zenodo.
Levchenko et al., IEEE S&P 2011 [13Levchenko, Kirill; Pitsillidis, Andreas; Chachra, Neha; Enright, Brandon; Félegyházi, Márk; Grier, Chris; Halvorson, Tristan; Kanich, Chris; Kreibich, Christian; Liu, He; McCoy, Damon; Weaver, Nicholas; Paxson, Vern; Voelker, Geoffrey M.; Savage, Stefan (2011): "Click Trajectories: End-to-End Analysis of the Spam Value Chain", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] Historical, and still the model for the payment question: three banks served over 95% of the spam-advertised goods they studied. Also the clearest written protocol for buying from criminals — which the institutional review board declined to review.

Where the seed list comes from

The population of a scam study is whatever its discovery channel emits. The papers that checked this found that channels barely overlap:

  • Search against malvertising, same scam type. Srinivasan et al. found tech-support scams through search results and search ads, then compared them with the malvertising-crawl data of Miramirkhani et al.: 0 of 2,768 fully qualified domains and 0 of 2,441 second-level domains in common, 92 of 1,994 IP addresses, 5 of 882 toll-free numbers [11Srinivasan, Bharat; Kountouras, Athanasios; Miramirkhani, Najmeh; Alam, Monjur; Nikiforakis, Nick; Antonakakis, Manos; Ahamad, Mustaque (2018): "Exposing Search and Advertisement Abuse Tactics and Infrastructure of Technical Support Scammers", in: Proceedings of the ACM Web Conference. (DOI)]. The search channel also surfaced a “passive” scam page — professional-looking, with a support number — that the authors argue malvertising would not carry.
  • A curated list against a platform. Of 3,863 giveaway-scam domains found through Certificate Transparency, only 361 (9%) ever appeared on Twitter [8Liu, Enze; Kappos, George; Mugnier, Eric; Invernizzi, Luca; Savage, Stefan; Tao, David; Thomas, Kurt; Voelker, Geoffrey M.; Meiklejohn, Sarah (2024): "Give and Take: An End-To-End Investigation of Giveaway Scam Conversion Rates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)].
  • Spam feeds against each other. Across ten spam feeds, 60% of live domains and 19% of storefront-tagged domains appeared in exactly one feed; the smallest feed, human-identified spam, contributed the most unique domains [14Pitsillidis, Andreas; Kanich, Chris; Voelker, Geoffrey M.; Levchenko, Kirill; Savage, Stefan (2012): "Taster's choice: a comparative analysis of spam feeds", in: Proceedings of the ACM Internet Measurement Conference. (DOI)].

What each channel finds, and what it leans towards, in the papers of this corpus:

Channel Used by Leans towards Status
Email spam feeds [13Levchenko, Kirill; Pitsillidis, Andreas; Chachra, Neha; Enright, Brandon; Félegyházi, Márk; Grier, Chris; Halvorson, Tristan; Kanich, Chris; Kreibich, Christian; Liu, He; McCoy, Damon; Weaver, Nicholas; Paxson, Vern; Voelker, Geoffrey M.; Savage, Stefan (2011): "Click Trajectories: End-to-End Analysis of the Spam Value Chain", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)], [15Kanich, Chris; Weaver, Nicholas; McCoy, Damon; Halvorson, Tristan; Kreibich, Christian; Levchenko, Kirill; Paxson, Vern; Voelker, Geoffrey M.; Savage, Stefan (2011): "Show Me the Money: Characterizing Spam-advertised Revenue", in: Proceedings of the USENIX Security Symposium. (Link)], [14Pitsillidis, Andreas; Kanich, Chris; Voelker, Geoffrey M.; Levchenko, Kirill; Savage, Stefan (2012): "Taster's choice: a comparative analysis of spam feeds", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] Loud, broad campaigns: MX honeypots “only tend to capture spam campaigns that are very broadly targeted” [14Pitsillidis, Andreas; Kanich, Chris; Voelker, Geoffrey M.; Levchenko, Kirill; Savage, Stefan (2012): "Taster's choice: a comparative analysis of spam feeds", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]. One botnet produced 13M garbage domains to poison blocklists [13Levchenko, Kirill; Pitsillidis, Andreas; Chachra, Neha; Enright, Brandon; Félegyházi, Márk; Grier, Chris; Halvorson, Tristan; Kanich, Chris; Kreibich, Christian; Liu, He; McCoy, Damon; Weaver, Nicholas; Paxson, Vern; Voelker, Geoffrey M.; Savage, Stefan (2011): "Click Trajectories: End-to-End Analysis of the Spam Value Chain", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]. Historical — the spam-advertised pharmacy and replica era; no paper in this population uses it after 2012.
Search queries [16Leontiadis, Nektarios; Moore, Tyler; Christin, Nicolas (2011): "Measuring and Analyzing Search-Redirection Attacks in the Illicit Online Prescription Drug Trade", in: Proceedings of the USENIX Security Symposium. (Link)], [17Leontiadis, Nektarios; Moore, Tyler; Christin, Nicolas (2014): "A Nearly Four-Year Longitudinal Study of Search-Engine Poisoning", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], [3Wang, David Y.; Der, Matthew F.; Karami, Mohammad; Saul, Lawrence K.; McCoy, Damon; Savage, Stefan; Voelker, Geoffrey M. (2014): "Search + Seizure: The Effectiveness of Interventions on SEO Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], [11Srinivasan, Bharat; Kountouras, Athanasios; Miramirkhani, Najmeh; Alam, Monjur; Nikiforakis, Nick; Antonakakis, Manos; Ahamad, Mustaque (2018): "Exposing Search and Advertisement Abuse Tactics and Infrastructure of Technical Support Scammers", in: Proceedings of the ACM Web Conference. (DOI)], [18Kharraz, Amin; Robertson, William K.; Kirda, Engin (2018): "Surveylance: Automatically Detecting Online Survey Scams", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)], [7Szurdi, Janos; Luo, Meng; Kondracki, Brian; Nikiforakis, Nick; Christin, Nicolas (2021): "Where are you taking me?Understanding Abusive Traffic Distribution Systems", in: Proceedings of the ACM Web Conference. (DOI)], [12Paudel, Pujan; Stringhini, Gianluca (2026): "LOKI: Proactively Discovering Online Scam Websites by Mining Toxic Search Queries", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] Whatever the query list covers. “The risk of starting from a single seed is to only identify a single unrepresentative campaign” [16Leontiadis, Nektarios; Moore, Tyler; Christin, Nicolas (2011): "Measuring and Analyzing Search-Redirection Attacks in the Illicit Online Prescription Drug Trade", in: Proceedings of the USENIX Security Symposium. (Link)]; “Any work measuring search results is biased towards the search terms selected” [3Wang, David Y.; Der, Matthew F.; Karami, Mohammad; Saul, Lawrence K.; McCoy, Damon; Savage, Stefan; Voelker, Geoffrey M. (2014): "Search + Seizure: The Effectiveness of Interventions on SEO Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]. Current, on one 2026 paper after none in 2022–2024: LOKI issued 980 queries to four engines, collected 271,161 sites, and its classifier labelled 19.3% of them scams [12Paudel, Pujan; Stringhini, Gianluca (2026): "LOKI: Proactively Discovering Online Scam Websites by Mining Toxic Search Queries", in: Proceedings of the Network and Distributed System Security Symposium. (Link)].
Search ads [11Srinivasan, Bharat; Kountouras, Athanasios; Miramirkhani, Najmeh; Alam, Monjur; Nikiforakis, Nick; Antonakakis, Manos; Ahamad, Mustaque (2018): "Exposing Search and Advertisement Abuse Tactics and Infrastructure of Technical Support Scammers", in: Proceedings of the ACM Web Conference. (DOI)] Scammers who buy placement. 71.79% of 14,346 ad URLs reached from tech-support queries led to tech-support scams. Current. Srinivasan et al. did not click the ads: they visited the advertiser's domain directly with the search engine as referrer, after checking on 50 scam ads that the redirect paths were identical to real clicks.
Ad networks, push ads, traffic distribution [4Miramirkhani, Najmeh; Starov, Oleksii; Nikiforakis, Nick (2017): "Dial One for Scam: A Large-Scale Analysis of Technical Support Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)], [18Kharraz, Amin; Robertson, William K.; Kirda, Engin (2018): "Surveylance: Automatically Detecting Online Survey Scams", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)], [19Vadrevu, Phani; Perdisci, Roberto (2019): "What You See is NOT What You Get: Discovering and Tracking Social Engineering Attack Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], [20Subramani, Karthika; Yuan, Xingzi; Setayeshfar, Omid; Vadrevu, Phani; Lee, Kyu Hyung; Perdisci, Roberto (2020): "When Push Comes to Ads: Measuring the Rise of (Malicious) Push Advertising", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], [7Szurdi, Janos; Luo, Meng; Kondracki, Brian; Nikiforakis, Nick; Christin, Nicolas (2021): "Where are you taking me?Understanding Abusive Traffic Distribution Systems", in: Proceedings of the ACM Web Conference. (DOI)], [21Yang, Zheng; Allen, Joey; Landen, Matthew; Perdisci, Roberto; Lee, Wenke (2023): "TRIDENT: Towards Detecting and Mitigating Web-based Social Engineering Attacks", in: Proceedings of the USENIX Security Symposium. (Link)] Low-tier ad networks and the publishers who embed them; social-engineering pages (tech support, scareware, fake prizes). Seeded from typosquats, URL shorteners, or ad-network code found on publisher pages. Current for social-engineering scams; the most recent paper in the population is 2023.
Posts, comments, livestreams [22Na, Seung Ho; Cho, Sumin; Shin, Seungwon (2023): "Evolving Bots: The New Generation of Comment Bots and their Underlying Scam Campaigns in YouTube", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], [23Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2024): "Like, Comment, Get Scammed: Characterizing Comment Scams on Media Platforms", in: Proceedings of the Network and Distributed System Security Symposium. (Link)], [8Liu, Enze; Kappos, George; Mugnier, Eric; Invernizzi, Luca; Savage, Stefan; Tao, David; Thomas, Kurt; Voelker, Geoffrey M.; Meiklejohn, Sarah (2024): "Give and Take: An End-To-End Investigation of Giveaway Scam Conversion Rates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] What the platform ranks: Na et al. crawled “top comments”, so the bots they found are those the ranking already rewarded; a web-crawl of tweets “is biased towards tweets that are more discoverable”. Current (2023–2024).
Certificate Transparency [5Li, Xigao; Yepuri, Anurag; Nikiforakis, Nick (2023): "Double and Nothing: Understanding and Detecting Cryptocurrency Giveaway Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)], [9Muzammil, Muhammad; Pitumpe, Abisheka; Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2025): "The Poorest Man in Babylon: A Longitudinal Study of Cryptocurrency Investment Scams", in: Proceedings of the ACM Web Conference. (DOI)] Sites that obtain a certificate and whose domain matches a keyword list; “If an attacker avoids using any of these keywords”, the site is never crawled [5Li, Xigao; Yepuri, Anurag; Nikiforakis, Nick (2023): "Double and Nothing: Understanding and Detecting Cryptocurrency Giveaway Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]. Current, with an infrastructure change — see the box below.
Newly registered domains [2Bitaab, Marzieh; Karimi, Alireza; Lyu, Zhuoer; Oest, Adam; Kuchhal, Dhruv; Saad, Muhammad; Ahn, Gail-Joon; Wang, Ruoyu; Bao, Tiffany; Shoshitaishvili, Yan; Doupé, Adam (2025): "ScamMagnifier: Piercing the Veil of Fraudulent Shopping Website Campaigns", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] New domains only. 1,155,237 shopping domains from May 2023 to June 2024, 46,746 flagged by the classifier. Aged and compromised domains are invisible to this channel. Current (one paper, 2025).
User and victim reports [24Christin, Nicolas; Yanagihara, Sally S.; Kamataki, Keisuke (2010): "Dissecting one click frauds", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], [6Bitaab, Marzieh; Cho, Haehyun; Oest, Adam; Lyu, Zhuoer; Wang, Wei; Abraham, Jorij; Wang, Ruoyu; Bao, Tiffany; Shoshitaishvili, Yan; Doupé, Adam (2023): "Beyond Phish: Toward Detecting Fraudulent e-Commerce Websites at Scale", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)], [25Liu, Mingxuan; Zhang, Yunyi; Wu, Lijie; Liu, Baojun; Hong, Geng; Zhang, Yiming; Jiang, Hui; Zhang, Jia; Duan, Haixin; Zhang, Min; Guan, Wei; Shi, Fan; Yang, Min (2025): "NOKEScam: Understanding and Rectifying Non-Sense Keywords Spear Scam in Search Engines", in: Proceedings of the USENIX Security Symposium. (Link)], [12Paudel, Pujan; Stringhini, Gianluca (2026): "LOKI: Proactively Discovering Online Scam Websites by Mining Toxic Search Queries", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] Scams someone already suspected. Beyond Phish trained on r/Scams and says so: the samples “are potentially biased because a human user has already decided that they might be FCWs”. Christin et al. argue the overlap of three forums means they captured “the most successful frauds”. Current as a seed; biased as a population.
Feeds and vendor telemetry [26Starov, Oleksii; Zhou, Yuchen; Zhang, Xiao; Miramirkhani, Najmeh; Nikiforakis, Nick (2018): "Betrayed by Your Dashboard: Discovering Malicious Campaigns via Web Analytics", in: Proceedings of the ACM Web Conference. (DOI)], [10Kotzias, Platon; Pachilakis, Michalis; Iuit, Javier Aldana; Caballero, Juan; Sanchez-Rola, Iskander; Bilge, Leyla (2025): "Ctrl+Alt+Deceive: Quantifying User Exposure to Online Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)], [25Liu, Mingxuan; Zhang, Yunyi; Wu, Lijie; Liu, Baojun; Hong, Geng; Zhang, Yiming; Jiang, Hui; Zhang, Jia; Duan, Haixin; Zhang, Min; Guan, Wei; Shi, Fan; Yang, Min (2025): "NOKEScam: Understanding and Rectifying Non-Sense Keywords Spear Scam in Search Engines", in: Proceedings of the USENIX Security Symposium. (Link)] The vendor's customers and the feed's own detector. Kotzias et al.: shopping scams are over-represented because two feeds contribute them; the feeds are mostly desktop; 80% of devices are in the US, EU, Japan and UK. NOKEScam ran inside Baidu's index — the data cannot leave Baidu. Current, but needs a partner.
Leaked or seized back ends [27McCoy, Damon; Pitsillidis, Andreas; Jordan, Grant; Weaver, Nicholas; Kreibich, Christian; Krebs, Brian; Voelker, Geoffrey M.; Savage, Stefan; Levchenko, Kirill (2012): "PharmaLeaks: Understanding the Business of Online Pharmaceutical Affiliate Programs", in: Proceedings of the USENIX Security Symposium. (Link)] The programs whose rivals leaked them. Opportunistic; not a method you can plan.
Platform ad repositories none in this population Ads the platform itself ran, with advertiser identity — the grouping key this page keeps asking for. What the EU DSA repositories contain and structurally cannot is on Ad archives. The UK Online Safety Act adds fraudulent-advertising duties for the largest services (ss. 38–39).1) Untested as a scam-discovery channel in these venues.

Certificate Transparency discovery changed under the field in 2025–2026. Let's Encrypt made its RFC 6962 logs read-only on 30 November 2025 and shut them down on 28 February 2026, in favour of static (tiled) CT logs; monitors “need to ensure that they have client software that's compatible with the new API”.2) Both CT-based papers here [5Li, Xigao; Yepuri, Anurag; Nikiforakis, Nick (2023): "Double and Nothing: Understanding and Detecting Cryptocurrency Giveaway Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] [9Muzammil, Muhammad; Pitumpe, Abisheka; Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2025): "The Poorest Man in Babylon: A Longitudinal Study of Cryptocurrency Investment Scams", in: Proceedings of the ACM Web Conference. (DOI)] predate that; Muzammil et al. ran a local CertStream server. The public certstream.calidog.io front end still answered HTTP 200 on 2026-09-24, but its server has had no commit since 2025-09-04 — before the shutdown — and a GitHub issue reporting the public websocket down, opened 2026-02-09, is still open.3) Treat it as unmaintained across the transition. certstream-server-rust is a Certstream-compatible server that reads both RFC 6962 and static-CT logs.4) Whatever you run, check it against the current log list before you trust a quiet day.

What to take from this into a design. Pick the channel from the victim path you care about — a shopper searching, a user served an ad, a follower on a platform — and say that your population is that channel's. If you need prevalence rather than a sample, you need a second, independent channel and the overlap between them; Leontiadis et al. ran capture–recapture estimates across query lists and got two estimates of the pharmacy population, 2,523 and 795 [16Leontiadis, Nektarios; Moore, Tyler; Christin, Nicolas (2011): "Measuring and Analyzing Search-Redirection Attacks in the Illicit Online Prescription Drug Trade", in: Proceedings of the USENIX Security Symposium. (Link)] — an honest picture of how far apart two query lists on one channel can put the total.

Grouping sites into operators

A scam domain is cheap and disposable; the operator is not. The papers that report how sites distribute over operators find a steep head:

Paper Grouping key What the grouping showed
[24Christin, Nicolas; Yanagihara, Sally S.; Kamataki, Keisuke (2010): "Dissecting one click frauds", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], CCS 2010 bank accounts, phone numbers, WHOIS, per reported fraud the top 8 of 112 groups ran more than half the frauds
[13Levchenko, Kirill; Pitsillidis, Andreas; Chachra, Neha; Enright, Brandon; Félegyházi, Márk; Grier, Chris; Halvorson, Tristan; Kanich, Chris; Kreibich, Christian; Liu, He; McCoy, Damon; Weaver, Nicholas; Paxson, Vern; Voelker, Geoffrey M.; Savage, Stefan (2011): "Click Trajectories: End-to-End Analysis of the Spam Value Chain", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)], IEEE S&P 2011 HTML similarity to storefront templates, then the bank that settled test purchases three banks for over 95% of the spam-advertised goods studied
[3Wang, David Y.; Der, Matthew F.; Karami, Mohammad; Saul, Lawrence K.; McCoy, Damon; Savage, Stefan; Voelker, Geoffrey M. (2014): "Search + Seizure: The Effectiveness of Interventions on SEO Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], IMC 2014 HTML bag-of-words classifier over 52 known campaigns 58% of poisoned search results, 11% of stores
[26Starov, Oleksii; Zhou, Yuchen; Zhang, Xiao; Miramirkhani, Najmeh; Nikiforakis, Nick (2018): "Betrayed by Your Dashboard: Discovering Malicious Campaigns via Web Analytics", in: Proceedings of the ACM Web Conference. (DOI)], TheWebConf 2018 shared Google Analytics and other analytics IDs average campaign 7.6 domains, largest 480 (VirusTotal seeds); 3.6 and 293 on tech-support scam seeds
[5Li, Xigao; Yepuri, Anurag; Nikiforakis, Nick (2023): "Double and Nothing: Understanding and Detecting Cryptocurrency Giveaway Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)], NDSS 2023 WHOIS e-mail and shared wallet address the ten largest campaigns served 35.58% of 10,079 scam pages
[9Muzammil, Muhammad; Pitumpe, Abisheka; Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2025): "The Poorest Man in Babylon: A Longitudinal Study of Cryptocurrency Investment Scams", in: Proceedings of the ACM Web Conference. (DOI)], TheWebConf 2025 IP, screenshot hash, e-mails and phones in the HTML, scripts 29,300 sites (67%) shared 4,900 IP addresses (26% of all)
[2Bitaab, Marzieh; Karimi, Alireza; Lyu, Zhuoer; Oest, Adam; Kuchhal, Dhruv; Saad, Muhammad; Ahn, Gail-Joon; Wang, Ruoyu; Bao, Tiffany; Shoshitaishvili, Yan; Doupé, Adam (2025): "ScamMagnifier: Piercing the Veil of Fraudulent Shopping Website Campaigns", in: Proceedings of the Network and Distributed System Security Symposium. (Link)], NDSS 2025 merchant ID from automated checkout, linked by the processor's registrant data 10 merchant IDs, 54.55% of 14,394 linked domains
[25Liu, Mingxuan; Zhang, Yunyi; Wu, Lijie; Liu, Baojun; Hong, Geng; Zhang, Yiming; Jiang, Hui; Zhang, Jia; Duan, Haixin; Zhang, Min; Guan, Wei; Shi, Fan; Yang, Min (2025): "NOKEScam: Understanding and Rectifying Non-Sense Keywords Spear Scam in Search Engines", in: Proceedings of the USENIX Security Symposium. (Link)], USENIX Security 2025 title similarity plus shared registration data top 10 of 143 campaigns, 80.02% of the scam keywords

The keys, in the order the literature moved through them:

  • Payment identifiers are the strongest key because the operator cannot rotate them as cheaply as a domain: bank accounts (2010), acquiring banks from test purchases (2011–2012), merchant IDs from checkout pages (2025), wallet addresses (2023–2025). They are also the key that needs the most intrusive collection — see Interacting with scammers, and reporting them.
  • Phone numbers and messenger contacts — tech-support scams carry them on the page [4Miramirkhani, Najmeh; Starov, Oleksii; Nikiforakis, Nick (2017): "Dial One for Scam: A Large-Scale Analysis of Technical Support Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]; comment scams move victims to WhatsApp or Telegram and the contact is the join [23Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2024): "Like, Comment, Get Scammed: Characterizing Comment Scams on Media Platforms", in: Proceedings of the Network and Distributed System Security Symposium. (Link)].
  • Screenshot hashes — Vadrevu and Perdisci's difference hash with DBSCAN (eps 0.1, MinPts 3), keeping only clusters that span at least five domains, because a campaign re-hosts the same page on many domains to evade URL blocklists [19Vadrevu, Phani; Perdisci, Roberto (2019): "What You See is NOT What You Get: Discovering and Tracking Social Engineering Attack Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]. Muzammil et al. group screenshots within a Hamming distance of 8 bits.
  • Analytics and tag IDs [26Starov, Oleksii; Zhou, Yuchen; Zhang, Xiao; Miramirkhani, Najmeh; Nikiforakis, Nick (2018): "Betrayed by Your Dashboard: Discovering Malicious Campaigns via Web Analytics", in: Proceedings of the ACM Web Conference. (DOI)] — cheap and effective, with a trap the authors hit: hosting providers' suspended-account error pages all share one analytics ID, so a size cap was needed to keep them from forming one giant “campaign”.
  • WHOIS — useful where un-redacted; see Ownership resolution for why a registrant field is not an owner.
  • IP and hosting — Wang et al. tried network features and dropped them: they “were ill-suited to differentiate SEO campaigns due to the growing popularity of shared hosting and reverse proxying infrastructure (e.g., CloudFlare)” [3Wang, David Y.; Der, Matthew F.; Karami, Mohammad; Saul, Lawrence K.; McCoy, Damon; Savage, Stefan; Voelker, Geoffrey M. (2014): "Search + Seizure: The Effectiveness of Interventions on SEO Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)].

Three pitfalls the papers name about their own grouping:

  • A template is not an operator. Kanich et al. found a pharmacy running the same storefront engine as a large affiliate program but with its own template, phone number, prices and bank — “not in fact associated with the GlavMed operation” [15Kanich, Chris; Weaver, Nicholas; McCoy, Damon; Halvorson, Tristan; Kreibich, Christian; Levchenko, Kirill; Paxson, Vern; Voelker, Geoffrey M.; Savage, Stefan (2011): "Show Me the Money: Characterizing Spam-advertised Revenue", in: Proceedings of the USENIX Security Symposium. (Link)].
  • A sender address is not a victim. Li et al. (NDSS 2023) traced wallets that paid several scam wallets and concluded that most belong to “online exchanges (such as Coinbase), where multiple victims can share the same outgoing address” [5Li, Xigao; Yepuri, Anurag; Nikiforakis, Nick (2023): "Double and Nothing: Understanding and Detecting Cryptocurrency Giveaway Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)].
  • Address clustering is incomplete. Across six cybercrimes including giveaway scams, “the popular multi-input clustering fails to discover addresses for 40% of groups” [28Gómez, Gibran; Liebergen, Kevin van; Caballero, Juan (2023): "Cybercrime Bitcoin Revenue Estimations: Quantifying the Impact of Methodology and Coverage", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)].

In the population below, 19 of 27 papers group sites into operators or campaigns; the 8 that do not include four of the six 2022–2024 papers. A 2026 paper that reports only a domain count is doing less than the 2011 papers did.

Ground truth: blocklists lag, so the label is your method

A blocklist tells you what someone had already listed. For scam sites it is a lagging and sparse instrument, which is why using it as the label and then reporting “coverage” is circular.

Paper Scam set Blocklist or label source Share of the scam set it knew
[24Christin, Nicolas; Yanagihara, Sally S.; Kamataki, Keisuke (2010): "Dissecting one click frauds", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] one-click fraud URLs Google Safe Browsing URL lists none of the URLs
[16Leontiadis, Nektarios; Moore, Tyler; Christin, Nicolas (2011): "Measuring and Analyzing Search-Redirection Attacks in the Illicit Online Prescription Drug Trade", in: Proceedings of the USENIX Security Symposium. (Link)] search-redirection infections Google Safe Browsing, Spamhaus, McAfee SiteAdvisor 95% of infected source sites on no blacklist; over two thirds of the pharmacies on at least one
[4Miramirkhani, Najmeh; Starov, Oleksii; Nikiforakis, Nick (2017): "Dial One for Scam: A Large-Scale Analysis of Technical Support Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] 1,524 tech-support scam domains combined domain blacklists 7% (108), 38 days late on average
[11Srinivasan, Bharat; Kountouras, Athanasios; Miramirkhani, Najmeh; Alam, Monjur; Nikiforakis, Nick; Antonakakis, Manos; Ahamad, Mustaque (2018): "Exposing Search and Advertisement Abuse Tactics and Infrastructure of Technical Support Scammers", in: Proceedings of the ACM Web Conference. (DOI)] tech-support scam domains public blacklists 26.8% cumulatively; support domains under 1%
[19Vadrevu, Phani; Perdisci, Roberto (2019): "What You See is NOT What You Get: Discovering and Tracking Social Engineering Attack Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] 2,042 social-engineering attack domains Google Safe Browsing 16.2% after two months
[20Subramani, Karthika; Yuan, Xingzi; Setayeshfar, Omid; Vadrevu, Phani; Lee, Kyu Hyung; Perdisci, Roberto (2020): "When Push Comes to Ads: Measuring the Rise of (Malicious) Push Advertising", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] malicious push-ad landing URLs Safe Browsing and VirusTotal under 1% on first scan; VirusTotal 11.31% a month later
[7Szurdi, Janos; Luo, Meng; Kondracki, Brian; Nikiforakis, Nick; Christin, Nicolas (2021): "Where are you taking me?Understanding Abusive Traffic Distribution Systems", in: Proceedings of the ACM Web Conference. (DOI)] 3,746 malicious landing pages Google Safe Browsing 92 on the day of detection
[5Li, Xigao; Yepuri, Anurag; Nikiforakis, Nick (2023): "Double and Nothing: Understanding and Detecting Cryptocurrency Giveaway Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] 3,610 giveaway-scam domains VirusTotal 16.75%
[6Bitaab, Marzieh; Cho, Haehyun; Oest, Adam; Lyu, Zhuoer; Wang, Wei; Abraham, Jorij; Wang, Ruoyu; Bao, Tiffany; Shoshitaishvili, Yan; Doupé, Adam (2023): "Beyond Phish: Toward Detecting Fraudulent e-Commerce Websites at Scale", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] 6,127 fraudulent shops APWG, Google Safe Browsing 25 and 10
[23Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2024): "Like, Comment, Get Scammed: Characterizing Comment Scams on Media Platforms", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] 24 fake investment sites VirusTotal (3 or more of 90 engines) 1
[9Muzammil, Muhammad; Pitumpe, Abisheka; Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2025): "The Poorest Man in Babylon: A Longitudinal Study of Cryptocurrency Investment Scams", in: Proceedings of the ACM Web Conference. (DOI)] a third of 43,572 investment-scam sites VirusTotal, MetaMask, Google Safe Browsing, … 20% at best (VirusTotal); Google Safe Browsing 1%

Read the table as “no general-purpose list is a scam census”, not as a ranking of lists: each row uses a different population, a different list and a different lag. Muzammil et al. suspect “this low coverage in existing block-lists is due to the general nature of websites they aim to detect”. The Google Safe Browsing API has no scam threat type; scams fall under SOCIAL_ENGINEERING alongside phishing.5) The v4 API is deprecated — “The Safe Browsing APIs (v4) are deprecated”, on Google's own v4 overview — and its end date and the v5 migration are on Google Safe Browsing. VirusTotal as a label is its own subject: see VirusTotal for what a “detected” bit is.

What the papers used instead — 19 of the 27 had researchers look at sites, the others relied on a classifier or a list:

  • Hand review of a sample — 19 of 27 papers here had researchers look at the sites. Double and Nothing's keyword filter produced about 4% false positives, which analysts removed with a screenshot dashboard [5Li, Xigao; Yepuri, Anurag; Nikiforakis, Nick (2023): "Double and Nothing: Understanding and Detecting Cryptocurrency Giveaway Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]; Beyond Phish had three experts relabel 2,000 sites (1.98% false positives, 1.63% false negatives) [6Bitaab, Marzieh; Cho, Haehyun; Oest, Adam; Lyu, Zhuoer; Wang, Wei; Abraham, Jorij; Wang, Ruoyu; Bao, Tiffany; Shoshitaishvili, Yan; Doupé, Adam (2023): "Beyond Phish: Toward Detecting Fraudulent e-Commerce Websites at Scale", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)].
  • A trained classifier validated on a labelled sample — 9 of 27. LOKI's oracle, a gradient-boosting model over 103 features, exceeds 90% precision and recall in cross-validation [12Paudel, Pujan; Stringhini, Gianluca (2026): "LOKI: Proactively Discovering Online Scam Websites by Mining Toxic Search Queries", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]; the 52,493 sites it then discovered have no ground truth, and the authors corroborate with ScamAdviser and Trustpilot instead.
  • An LLM — 1 of 27. Muzammil et al. hand-labelled 300 sites and compared models: GPT-4 at 90% accuracy, Llama 3 70B at 87%, a hybrid routing inconclusive cases to GPT-4 at 88%, which is what they ran [9Muzammil, Muhammad; Pitumpe, Abisheka; Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2025): "The Poorest Man in Babylon: A Longitudinal Study of Cryptocurrency Investment Scams", in: Proceedings of the ACM Web Conference. (DOI)].
  • A commercial trust score as the label — Kotzias et al. took ScamAdviser domains with a trust score up to 10, removed Tranco top-1M domains, and kept those with at least two VirusTotal detections [10Kotzias, Platon; Pachilakis, Michalis; Iuit, Javier Aldana; Caballero, Juan; Sanchez-Rola, Iskander; Bilge, Leyla (2025): "Ctrl+Alt+Deceive: Quantifying User Exposure to Online Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]. Their own check of the service's industry tags found high precision for only 7 of 26 categories — the authors add that “this low accuracy is not specific to ScamAdviser, but plagues most commercial website classification services”. Label sources for topics are Website classification.

The scam type is a label too, and nobody shares a taxonomy. Kotzias et al. examine “seven popular scam types” and derived them partly from ScamAdviser's industry tags [10Kotzias, Platon; Pachilakis, Michalis; Iuit, Javier Aldana; Caballero, Juan; Sanchez-Rola, Iskander; Bilge, Leyla (2025): "Ctrl+Alt+Deceive: Quantifying User Exposure to Online Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]; LOKI maps its seeds onto Trustpilot categories and reports ten broad scam categories [12Paudel, Pujan; Stringhini, Gianluca (2026): "LOKI: Proactively Discovering Online Scam Websites by Mining Toxic Search Queries", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]; the population table below uses a seven-value hand taxonomy of its own. None has been validated against another. Exposure and harm figures are only comparable across papers within a type, so report the taxonomy you used and who assigned it.

The crawler has to look like a victim

Scam infrastructure filters visitors, and the filter is aimed at the vantage a research crawler most often has: a cloud IP, a headless browser, a desktop user agent, no referrer.

  • IP. Miramirkhani et al.'s campus crawler found 95.7% of the tech-support scam domains that all three of their crawlers found; the two cloud crawlers were filtered by the ad networks [4Miramirkhani, Najmeh; Starov, Oleksii; Nikiforakis, Nick (2017): "Dial One for Scam: A Large-Scale Analysis of Technical Support Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]. Two of Vadrevu and Perdisci's eleven ad networks appeared never to serve social-engineering ads to their university, Tor or AWS addresses, so they put crawlers on laptops on residential networks [19Vadrevu, Phani; Perdisci, Roberto (2019): "What You See is NOT What You Get: Discovering and Tracking Social Engineering Attack Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]. Szurdi et al. found more than twice as many malicious pages with a pool of IPs as with one, and when they were shown a benign or error page instead, “we face cloaking 86% of the time and are explicitly blocked only 14% of the time” [7Szurdi, Janos; Luo, Meng; Kondracki, Brian; Nikiforakis, Nick; Christin, Nicolas (2021): "Where are you taking me?Understanding Abusive Traffic Distribution Systems", in: Proceedings of the ACM Web Conference. (DOI)].
  • Device. Malicious mobile push notifications “were much more likely to appear on real Android devices, rather than emulated environments” [20Subramani, Karthika; Yuan, Xingzi; Setayeshfar, Omid; Vadrevu, Phani; Lee, Kyu Hyung; Perdisci, Roberto (2020): "When Push Comes to Ads: Measuring the Rise of (Malicious) Push Advertising", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]; Szurdi et al. found tech-support scams shown more to desktop users and survey scams more to phones, and “no evidence” of emulator detection in their traffic distribution systems. Your device choice picks your scam mix.
  • User agent and referrer. Liu et al. met four kinds of cloaking on giveaway sites in a pilot: 403s for an institutional network, 403s for browsers not on Windows or Mac, front pages that need a click, and Cloudflare bot checks [8Liu, Enze; Kappos, George; Mugnier, Eric; Invernizzi, Luca; Savage, Stefan; Tao, David; Thomas, Kurt; Voelker, Geoffrey M.; Meiklejohn, Sarah (2024): "Give and Take: An End-To-End Investigation of Giveaway Scam Conversion Rates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]. Srinivasan et al. kept the search engine as the referrer instead of clicking ads, and warn that their PhantomJS crawler “can, in principle, be detected by scammers” [11Srinivasan, Bharat; Kountouras, Athanasios; Miramirkhani, Najmeh; Alam, Monjur; Nikiforakis, Nick; Antonakakis, Manos; Ahamad, Mustaque (2018): "Exposing Search and Advertisement Abuse Tactics and Infrastructure of Technical Support Scammers", in: Proceedings of the ACM Web Conference. (DOI)].
  • Crawler versus search engine. Scam pages show search-engine crawlers something else: NOKEScam pages showed Baidu's crawler timely news to get indexed and showed victims the fraud [25Liu, Mingxuan; Zhang, Yunyi; Wu, Lijie; Liu, Baojun; Hong, Geng; Zhang, Yiming; Jiang, Hui; Zhang, Jia; Duan, Haixin; Zhang, Min; Guan, Wei; Shi, Fan; Yang, Min (2025): "NOKEScam: Understanding and Rectifying Non-Sense Keywords Spear Scam in Search Engines", in: Proceedings of the USENIX Security Symposium. (Link)]; Wang et al. found a store loaded as an iframe over the doorway page, which defeats redirect-based cloaking detection [3Wang, David Y.; Der, Matthew F.; Karami, Mohammad; Saul, Lawrence K.; McCoy, Damon; Savage, Stefan; Voelker, Geoffrey M. (2014): "Search + Seizure: The Effectiveness of Interventions on SEO Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]. Measuring this is the Cloak of Visibility design [29Invernizzi, Luca; Thomas, Kurt; Kapravelos, Alexandros; Comanescu, Oxana; Picod, Jean-Michel; Bursztein, Elie (2016): "Cloak of Visibility: Detecting When Machines Browse a Different Web", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] and the older Dagger crawler [30Wang, David Y.; Savage, Stefan; Voelker, Geoffrey M. (2011): "Cloak and dagger: dynamics of web search cloaking", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)].

Of the 27 papers, 7 measured cloaking by comparing vantages or identities, 8 set the crawler up to look like a victim without measuring the difference, and 12 say nothing. Two browser defences since 2025 make this worse from the other side: Chrome's Enhanced Protection uses the on-device Gemini Nano model against tech-support scam pages (announced 2025-05-08), and in September 2025 Google said it would extend this to sites using “fake viruses or fake giveaways” — whether that has shipped was not established here,6) and Edge's scareware blocker was on by default on most Windows and Mac devices by 2025-10-31.7) A scam page your crawler reaches may already be intercepted for a real user of those browsers; say which browser, and which protection setting, your exposure estimate assumes.

What to do, and where the detail lives: the detection surface and how to measure your block rate is Crawler detection; residential and mobile vantage points, and the ethics of renting one, are Residential and mobile proxies; cloaking against anti-phishing crawlers is on Cloaking. The one scam-specific rule: record, per URL, which identity saw what, so that a benign page is a reading from one vantage, not a verdict.

Measuring harm

The papers that measured harm used seven instruments. None gives a loss total for the victims; the two that see individual payments — leaked back ends and victim reports — see them for one program or one complainant at a time, and each instrument has a known bias:

Instrument Example papers What it gives Known bias Status
Sequential order numbers (“purchase pairs”) [15Kanich, Chris; Weaver, Nicholas; McCoy, Damon; Halvorson, Tristan; Kreibich, Christian; Levchenko, Kirill; Paxson, Vern; Voelker, Geoffrey M.; Savage, Stefan (2011): "Show Me the Money: Characterizing Spam-advertised Revenue", in: Proceedings of the USENIX Security Symposium. (Link)], [3Wang, David Y.; Der, Matthew F.; Karami, Mohammad; Saul, Lawrence K.; McCoy, Damon; Savage, Stefan; Voelker, Geoffrey M. (2014): "Search + Seizure: The Effectiveness of Interventions on SEO Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] order volume per store — over 82,000 pharmacy and 37,000 software orders a month across ten programs [15Kanich, Chris; Weaver, Nicholas; McCoy, Damon; Halvorson, Tristan; Kreibich, Christian; Levchenko, Kirill; Paxson, Vern; Voelker, Geoffrey M.; Savage, Stefan (2011): "Show Me the Money: Characterizing Spam-advertised Revenue", in: Proceedings of the USENIX Security Symposium. (Link)] orders are numbered before payment; against leaked ground truth, true turnover was 8–35% lower than the technique predicted [27McCoy, Damon; Pitsillidis, Andreas; Jordan, Grant; Weaver, Nicholas; Kreibich, Christian; Krebs, Brian; Voelker, Geoffrey M.; Savage, Stefan; Levchenko, Kirill (2012): "PharmaLeaks: Understanding the Business of Online Pharmaceutical Affiliate Programs", in: Proceedings of the USENIX Security Symposium. (Link)] Historical. Needs stores that number orders sequentially.
Test purchases [13Levchenko, Kirill; Pitsillidis, Andreas; Chachra, Neha; Enright, Brandon; Félegyházi, Márk; Grier, Chris; Halvorson, Tristan; Kanich, Chris; Kreibich, Christian; Liu, He; McCoy, Damon; Weaver, Nicholas; Paxson, Vern; Voelker, Geoffrey M.; Savage, Stefan (2011): "Click Trajectories: End-to-End Analysis of the Spam Value Chain", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)], [15Kanich, Chris; Weaver, Nicholas; McCoy, Damon; Halvorson, Tristan; Kreibich, Christian; Levchenko, Kirill; Paxson, Vern; Voelker, Geoffrey M.; Savage, Stefan (2011): "Show Me the Money: Characterizing Spam-advertised Revenue", in: Proceedings of the USENIX Security Symposium. (Link)], [31McCoy, Damon; Dharmdasani, Hitesh; Kreibich, Christian; Voelker, Geoffrey M.; Savage, Stefan (2012): "Priceless: the role of payments in abuse-advertised goods", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], [3Wang, David Y.; Der, Matthew F.; Karami, Mohammad; Saul, Lawrence K.; McCoy, Damon; Savage, Stefan; Voelker, Geoffrey M. (2014): "Search + Seizure: The Effectiveness of Interventions on SEO Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] the bank, the shipper, the product stores learn to refuse undercover buyers: “some subset of our refusals may not be due to true payment processing problems but an active attempt to 'blind' such measurements” [31McCoy, Damon; Dharmdasani, Hitesh; Kreibich, Christian; Voelker, Geoffrey M.; Savage, Stefan (2012): "Priceless: the role of payments in abuse-advertised goods", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] Not used by any paper in this population after 2014.
Leaked or partner transaction data [27McCoy, Damon; Pitsillidis, Andreas; Jordan, Grant; Weaver, Nicholas; Kreibich, Christian; Krebs, Brian; Voelker, Geoffrey M.; Savage, Stefan; Levchenko, Kirill (2012): "PharmaLeaks: Understanding the Business of Online Pharmaceutical Affiliate Programs", in: Proceedings of the USENIX Security Symposium. (Link)], [2Bitaab, Marzieh; Karimi, Alireza; Lyu, Zhuoer; Oest, Adam; Kuchhal, Dhruv; Saad, Muhammad; Ahn, Gail-Joon; Wang, Ruoyu; Bao, Tiffany; Shoshitaishvili, Yan; Doupé, Adam (2025): "ScamMagnifier: Piercing the Veil of Fraudulent Shopping Website Campaigns", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] settled revenue (over US$170M across three pharmacy programs); time from registration to first payment the partner decides what you may publish — ScamMagnifier's processor withheld transaction volumes Opportunistic (leaks) or partner-dependent (2025).
Visitor counts times an assumed conversion rate [16Leontiadis, Nektarios; Moore, Tyler; Christin, Nicolas (2011): "Measuring and Analyzing Search-Redirection Attacks in the Illicit Online Prescription Drug Trade", in: Proceedings of the USENIX Security Symposium. (Link)], [4Miramirkhani, Najmeh; Starov, Oleksii; Nikiforakis, Nick (2017): "Dial One for Scam: A Large-Scale Analysis of Technical Support Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)], [3Wang, David Y.; Der, Matthew F.; Karami, Mohammad; Saul, Lawrence K.; McCoy, Damon; Savage, Stefan; Voelker, Geoffrey M. (2014): "Search + Seizure: The Effectiveness of Interventions on SEO Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] 1,688,412 visitor IPs to 142 scam domains with exposed server-status pages [4Miramirkhani, Najmeh; Starov, Oleksii; Nikiforakis, Nick (2017): "Dial One for Scam: A Large-Scale Analysis of Technical Support Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] the revenue is only as good as the borrowed rate — the $9.7M estimate assumes tech-support victims convert like buyers of a fake antivirus “full version”, about 2% Use for exposure; do not publish the product as revenue.
Wallet inflows [5Li, Xigao; Yepuri, Anurag; Nikiforakis, Nick (2023): "Double and Nothing: Understanding and Detecting Cryptocurrency Giveaway Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)], [23Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2024): "Like, Comment, Get Scammed: Characterizing Comment Scams on Media Platforms", in: Proceedings of the Network and Distributed System Security Symposium. (Link)], [8Liu, Enze; Kappos, George; Mugnier, Eric; Invernizzi, Luca; Savage, Stefan; Tao, David; Thomas, Kurt; Voelker, Geoffrey M.; Meiklejohn, Sarah (2024): "Give and Take: An End-To-End Investigation of Giveaway Scam Conversion Rates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], [9Muzammil, Muhammad; Pitumpe, Abisheka; Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2025): "The Poorest Man in Babylon: A Longitudinal Study of Cryptocurrency Investment Scams", in: Proceedings of the ACM Web Conference. (DOI)] BTC/ETH received by scam addresses — $24.9M–$69.9M in six months of giveaway scams, depending on the price used [5Li, Xigao; Yepuri, Anurag; Nikiforakis, Nick (2023): "Double and Nothing: Understanding and Detecting Cryptocurrency Giveaway Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] inflow is not victim loss: only 43% of Twitter-linked payments fell inside a one-week window around the promotion, so Liu et al. report $2.7M co-occurring and $6.6M all-incoming for the same addresses [8Liu, Enze; Kappos, George; Mugnier, Eric; Invernizzi, Luca; Savage, Stefan; Tao, David; Thomas, Kurt; Voelker, Geoffrey M.; Meiklejohn, Sarah (2024): "Give and Take: An End-To-End Investigation of Giveaway Scam Conversion Rates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]; exchange addresses merge victims Current (2023–2025).
Telemetry of exposure [10Kotzias, Platon; Pachilakis, Michalis; Iuit, Javier Aldana; Caballero, Juan; Sanchez-Rola, Iskander; Bilge, Leyla (2025): "Ctrl+Alt+Deceive: Quantifying User Exposure to Online Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] daily exposed devices; IPs that reached a checkout page “not all users who visit a checkout page will complete a purchase” — an upper bound on intent, not a loss Current, needs a vendor.
Victim reports [24Christin, Nicolas; Yanagihara, Sally S.; Kamataki, Keisuke (2010): "Dissecting one click frauds", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], [25Liu, Mingxuan; Zhang, Yunyi; Wu, Lijie; Liu, Baojun; Hong, Geng; Zhang, Yiming; Jiang, Hui; Zhang, Jia; Duan, Haixin; Zhang, Min; Guan, Wei; Shi, Fan; Yang, Min (2025): "NOKEScam: Understanding and Rectifying Non-Sense Keywords Spear Scam in Search Engines", in: Proceedings of the USENIX Security Symposium. (Link)] the amount a victim says was taken NOKEScam's average of $2,896 comes from the 20 of the last 100 complaints that stated an amount; its daily-loss figure then assumes a one-in-a-million victim rate Use for the distribution of losses, not the total.

Both tables name examples; the full per-paper coding is in What the 27 papers did.

Two extrapolations deserve a warning because they are the numbers that get quoted. Li et al. (NDSS 2024) estimate that the scammer accounts behind YouTube comment scams “have potentially stolen more than 100 million US dollars”, from 31 of 72 contacted scammers' wallets [23Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2024): "Like, Comment, Get Scammed: Characterizing Comment Scams on Media Platforms", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]; Muzammil et al. reach “more than 100M USD” by multiplying the average inflow per site with a known wallet across all detected sites [9Muzammil, Muhammad; Pitumpe, Abisheka; Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2025): "The Poorest Man in Babylon: A Longitudinal Study of Cryptocurrency Investment Scams", in: Proceedings of the ACM Web Conference. (DOI)]. Both are labelled as estimates in the papers. Quote the measured inflow with its denominator, and the extrapolation only with its assumption.

For wallet-based estimates in general, [28Gómez, Gibran; Liebergen, Kevin van; Caballero, Juan (2023): "Cybercrime Bitcoin Revenue Estimations: Quantifying the Impact of Methodology and Coverage", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] re-implemented the published revenue methodologies over 30,424 payment addresses and found that the revenue “is not always underestimated. There exist methodologies that can introduce huge overestimation.” Read it before choosing between co-occurrence windows, address clustering and all-inflows.

Official loss statistics are reports, not measurements, and a paper introduction should say which kind it cites. The FBI's IC3 recorded $20.877 billion in reported losses for 2025, with investment fraud the largest category and tech and customer support among the next;8) these are complaints filed. The Global Anti-Scam Alliance's “$442 Billion” is a survey of 46,000 people in 42 markets, scaled up, and the report labels it “estimated”.9) The UK's national reporting service changed from Action Fraud to Report Fraud on 4 December 2025, so a UK series crossing that date crosses a system change.10)

Interacting with scammers, and reporting them

The payment and conversion questions push you towards the scammer. Here is what the papers did and what their ethics statements say — which is also what a programme committee will compare you against. The general checklist is Ethics; telling a host, registrar or processor is Notifying websites.

Interaction Example paper What was done Review and safeguards, as stated
Phone calls [4Miramirkhani, Najmeh; Starov, Oleksii; Nikiforakis, Nick (2017): "Dial One for Scam: A Large-Scale Analysis of Technical Support Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] 60 recorded calls to tech-support scammers IRB approved deception, waived consent and debriefing; “we did not pay any scammer”
Chats [23Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2024): "Like, Comment, Get Scammed: Characterizing Comment Scams on Media Platforms", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] 74 scammers contacted over WhatsApp and Telegram, 50 conversations completed IRB approved deception and waived consent; text only; left with “polite excuses” when asked to pay
Purchases [13Levchenko, Kirill; Pitsillidis, Andreas; Chachra, Neha; Enright, Brandon; Félegyházi, Márk; Grier, Chris; Halvorson, Tristan; Kanich, Chris; Kreibich, Christian; Liu, He; McCoy, Damon; Weaver, Nicholas; Paxson, Vern; Voelker, Geoffrey M.; Savage, Stefan (2011): "Click Trajectories: End-to-End Analysis of the Spam Value Chain", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] 120 attempted, 76 authorised, 56 settled the IRB did not deem it appropriate to review (no human subjects); university counsel; only non-prescription pharmacy goods and already-licensed software; products destroyed
Purchases [31McCoy, Damon; Dharmdasani, Hitesh; Kreibich, Christian; Voelker, Geoffrey M.; Savage, Stefan (2012): "Priceless: the role of payments in abuse-advertised goods", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] 676 ordering attempts, 429 successful “explicitly reviewed and approved by our institution”; a few thousand dollars in total
Checkout, then stop [3Wang, David Y.; Der, Matthew F.; Karami, Mohammad; Saul, Lawrence K.; McCoy, Damon; Savage, Stefan; Voelker, Geoffrey M. (2014): "Search + Seizure: The Effectiveness of Interventions on SEO Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], [2Bitaab, Marzieh; Karimi, Alireza; Lyu, Zhuoer; Oest, Adam; Kuchhal, Dhruv; Saad, Muhammad; Ahn, Gail-Joon; Wang, Ruoyu; Bao, Tiffany; Shoshitaishvili, Yan; Doupé, Adam (2025): "ScamMagnifier: Piercing the Veil of Fraudulent Shopping Website Campaigns", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] orders taken to the card-details page with generated identities; automated checkout for merchant IDs Wang et al. state no review; ScamMagnifier: “without finalizing transactions”, merchant IDs reported to processors, domains to Google and Microsoft
Fake victim accounts [9Muzammil, Muhammad; Pitumpe, Abisheka; Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2025): "The Poorest Man in Babylon: A Longitudinal Study of Cryptocurrency Investment Scams", in: Proceedings of the ACM Web Conference. (DOI)] sign-up with disposable addresses to reveal wallets no real users involved; sites deliberately not reported during the study
Clicking ads [19Vadrevu, Phani; Perdisci, Roberto (2019): "What You See is NOT What You Get: Discovering and Tracking Social Engineering Attack Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], [20Subramani, Karthika; Yuan, Xingzi; Setayeshfar, Omid; Vadrevu, Phani; Lee, Kyu Hyung; Perdisci, Roberto (2020): "When Push Comes to Ads: Measuring the Rise of (Malicious) Push Advertising", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], [21Yang, Zheng; Allen, Joey; Landen, Matthew; Perdisci, Roberto; Lee, Wenke (2023): "TRIDENT: Towards Detecting and Mitigating Web-based Social Engineering Attacks", in: Proceedings of the USENIX Security Symposium. (Link)] automated clicks on low-tier and push ads each estimates the cost to legitimate advertisers — worst case about $4.8, maximum $1.12 per landing domain, average $1.5 per advertiser

Two decisions the papers disagree on, and you will have to make:

  • Report now, or observe? Li et al. “chose not to tamper with the ecosystem while studying it, to avoid measuring artifacts of our own intervention” [5Li, Xigao; Yepuri, Anurag; Nikiforakis, Nick (2023): "Double and Nothing: Understanding and Detecting Cryptocurrency Giveaway Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]; Muzammil et al. withheld reports for the same reason and released data afterwards [9Muzammil, Muhammad; Pitumpe, Abisheka; Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2025): "The Poorest Man in Babylon: A Longitudinal Study of Cryptocurrency Investment Scams", in: Proceedings of the ACM Web Conference. (DOI)]; Liu et al. analysed retrospectively and note they therefore could not warn victims [8Liu, Enze; Kappos, George; Mugnier, Eric; Invernizzi, Luca; Savage, Stefan; Tao, David; Thomas, Kurt; Voelker, Geoffrey M.; Meiklejohn, Sarah (2024): "Give and Take: An End-To-End Investigation of Giveaway Scam Conversion Rates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]; ScamMagnifier and NOKEScam reported as they went [2Bitaab, Marzieh; Karimi, Alireza; Lyu, Zhuoer; Oest, Adam; Kuchhal, Dhruv; Saad, Muhammad; Ahn, Gail-Joon; Wang, Ruoyu; Bao, Tiffany; Shoshitaishvili, Yan; Doupé, Adam (2025): "ScamMagnifier: Piercing the Veil of Fraudulent Shopping Website Campaigns", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] [25Liu, Mingxuan; Zhang, Yunyi; Wu, Lijie; Liu, Baojun; Hong, Geng; Zhang, Yiming; Jiang, Hui; Zhang, Jia; Duan, Haixin; Zhang, Min; Guan, Wei; Shi, Fan; Yang, Min (2025): "NOKEScam: Understanding and Rectifying Non-Sense Keywords Spear Scam in Search Engines", in: Proceedings of the USENIX Security Symposium. (Link)]. There is no consensus; state your choice and why.
  • Residential IPs. Szurdi et al. did not use residential addresses “to avoid ethical quandaries” [7Szurdi, Janos; Luo, Meng; Kondracki, Brian; Nikiforakis, Nick; Christin, Nicolas (2021): "Where are you taking me?Understanding Abusive Traffic Distribution Systems", in: Proceedings of the ACM Web Conference. (DOI)]; Vadrevu and Perdisci and Liu et al. used them because the ad networks and scam sites filter everything else. The trade-off is on Residential and mobile proxies.

For automated chat with fraud crews as a method — and its limits — see [32Wang, Peng; Liao, Xiaojing; Qin, Yue; Wang, XiaoFeng (2020): "Into the Deep Web: Understanding E-commerce Fraud from Autonomous Chat with Cybercriminals", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] and the Craigslist scam-baiting study [33Park, Youngsam; Jones, Jackie; McCoy, Damon; Shi, Elaine; Jakobsson, Markus (2014): "Scambaiter: Understanding Targeted Nigerian Scams on Craigslist", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]. Neither measures a scam website, so neither is in the population below.

In the extraction's own ethics field, 4 of the 27 papers report an approved review, 14 mention none, 2 discuss ethics and state no review was sought, 2 state none was required, 1 sought review with no stated outcome, and 4 have no ethics record at all.

Which methods are current

The corpus runs to 2026, but 2025–2026 are its thinnest and provisional years (CCS and IMC 2026 have not been held; IEEE S&P and TheWebConf 2026 are incompletely indexed). The population has 4 papers from 2025 and 1 from 2026. The status column says whether a method is used by a 2023–2026 paper in these seven venues; where a row adds a judgement about whether you should use it, the judgement is marked as such.

Method Years in this population Status On what evidence
Spam-feed seeding 2011–2012 historical no later paper in the population
Test purchases 2011–2014 superseded here by checkout-then-stop and wallet tracing no paper after 2014; stores counter undercover buyers [31McCoy, Damon; Dharmdasani, Hitesh; Kreibich, Christian; Voelker, Geoffrey M.; Savage, Stefan (2012): "Priceless: the role of payments in abuse-advertised goods", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)]
Order-number sampling 2011–2014 historical needs sequential order IDs; biased high against leaked data [27McCoy, Damon; Pitsillidis, Andreas; Jordan, Grant; Weaver, Nicholas; Kreibich, Christian; Krebs, Brian; Voelker, Geoffrey M.; Savage, Stefan; Levchenko, Kirill (2012): "PharmaLeaks: Understanding the Business of Online Pharmaceutical Affiliate Programs", in: Proceedings of the USENIX Security Symposium. (Link)]
Search-query discovery 2011–2021, 2026 current, one 2026 paper after none in 2022–2024 [12Paudel, Pujan; Stringhini, Gianluca (2026): "LOKI: Proactively Discovering Online Scam Websites by Mining Toxic Search Queries", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]
Ad-network and push-ad crawling 2017–2023 current for social-engineering scams no 2025–2026 paper in the population
Certificate Transparency discovery 2023–2025 current, client must follow the static-CT change [5Li, Xigao; Yepuri, Anurag; Nikiforakis, Nick (2023): "Double and Nothing: Understanding and Detecting Cryptocurrency Giveaway Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)], [9Muzammil, Muhammad; Pitumpe, Abisheka; Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2025): "The Poorest Man in Babylon: A Longitudinal Study of Cryptocurrency Investment Scams", in: Proceedings of the ACM Web Conference. (DOI)]
New-registration feed + automated checkout + merchant IDs 2025 current, one paper, needs a payment processor [2Bitaab, Marzieh; Karimi, Alireza; Lyu, Zhuoer; Oest, Adam; Kuchhal, Dhruv; Saad, Muhammad; Ahn, Gail-Joon; Wang, Ruoyu; Bao, Tiffany; Shoshitaishvili, Yan; Doupé, Adam (2025): "ScamMagnifier: Piercing the Veil of Fraudulent Shopping Website Campaigns", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]
Grouping by payment identifier 2010–2025 current merchant IDs and wallets since 2023
Wallet inflows as harm 2023–2025 current, report as a bracket [8Liu, Enze; Kappos, George; Mugnier, Eric; Invernizzi, Luca; Savage, Stefan; Tao, David; Thomas, Kurt; Voelker, Geoffrey M.; Meiklejohn, Sarah (2024): "Give and Take: An End-To-End Investigation of Giveaway Scam Conversion Rates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]
A blocklist or trust score as the label 2018–2025 current in use (three 2023–2025 papers); judgement: do not use it as ground truth the coverage table above
LLM classification of scam sites 2025 emerging: 1 of 27 papers, validated on 300 hand-labelled sites [9Muzammil, Muhammad; Pitumpe, Abisheka; Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2025): "The Poorest Man in Babylon: A Longitudinal Study of Cryptocurrency Investment Scams", in: Proceedings of the ACM Web Conference. (DOI)]. LOKI uses an LLM only to filter branded keywords [12Paudel, Pujan; Stringhini, Gianluca (2026): "LOKI: Proactively Discovering Online Scam Websites by Mining Toxic Search Queries", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]. Seven context papers (platform, ads, SMS, chat and wallet-phishing studies) use one, all 2025–2026

The LLM row is where to be careful. Corpus-wide, papers with an LLM classification step went from 2 of 719 in 2023 (0.3%) to 27 of 690 in 2024 (3.9%), 77 of 770 in 2025 (10.0%) and 71 of 415 in 2026 (17.1%); among this population's five 2025–2026 papers, 1 classifies scam sites with an LLM. The scam literature is not behind the corpus rate — it is too small to have a rate. That one paper reports that the open model it ran was three accuracy points behind GPT-4 at no API cost; it does not show that an LLM beats a trained classifier on scam sites, and nothing in this population tests that. Industry moved faster: Chrome runs an on-device LLM against tech-support scam pages (above). If you use an LLM as the classifier, do what Muzammil et al. did — hand-label a balanced sample first and report accuracy against it. Outside the seven venues the agent form is already published: ScamFerret, an LLM agent that fetches a site's content, DNS records and user reviews, reports 0.972 accuracy on four English scam types with GPT-4 (DIMVA 2025).11)

Use in publications

Everything below is a claim about seven venues — CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf and IEEE S&P, 2010–2026, 5,859 extracted papers. The APWG Symposium on Electronic Crime Research (eCrime), EuroS&P, ACSAC, RAID, AsiaCCS, DIMVA, CHI and SOUPS are absent. eCrime and DIMVA publish scam measurement — both outside papers cited on this page appeared there in 2025 — so a scam paper missing from these counts may simply be outside the corpus. See Corpus. 2025–2026 are provisional.

The inclusion rule, and its measured precision

Five probes — the 2026-09-22 gap-pass regex, the bare word scam ten or more times, named scam types, the 2011–2016 spam-value-chain vocabulary, and the extraction's own free-text fields — produce a 183-paper candidate set. Each candidate was read against a rule written before the verdicts were counted:

  • IN — the measured objects include web-delivered scams (websites, landing pages, or the ad, search or redirect chains that deliver them) whose purpose is to make the visitor pay or transfer money under false pretences, and the paper collects instances and measures, groups, labels or detects them.
  • CONTEXT — scams, but not a scam website as the object: platform accounts and comments, on-chain scam tokens, phone and SMS, scam apps, susceptibility studies, credential or signature theft called phishing, search abuse or scam ads whose destinations are not mostly scams, back-office studies.
  • Two clauses added after review. The 2011–2014 spam- and SEO-advertised pharmacy and counterfeit storefront papers are IN as a named lineage for the methods they introduced; later search poisoning that promotes openly illicit trade (controlled drugs, gambling) is CONTEXT. The line there is history, not the object, and a reader could reasonably move the 2022 illicit-drug search-poisoning paper IN. Scareware counts when the paper measures the purchase page, not the software download.
  • OUT — scams incidental, fraud against a platform or merchant, malware delivery, or a homonym (Scamper, the topology tool, accounts for several).
Verdict Papers of 183
IN — this page's population 27 (14.8%)
CONTEXT — on-chain 12, search abuse 12, platform 11, telephony 10, user study 5, phishing boundary 5, ads 5, back office 4, LLM conversations 2, mobile apps 1 67
OUT 89

Of the 27, 26 include the web platform and 24 ran a crawl.

How well each probe finds the population:

Probe Hits IN Precision Recall of the 27
gap-pass regex, 5 or more hits 45 18 40.0% 66.7%
scam 10 or more times 127 19 15.0% 70.4%
named scam types, 5 or more 79 23 29.1% 85.2%
spam value chain vocabulary, 5 or more 49 11 22.4% 40.7%
extraction free text 57 16 28.1% 59.3%

The gap pass chose this page on 43 papers (27 web, 21 from 2024–2026), “the steepest recent slope in the pass”. The regex re-runs to exactly those numbers, but the slope is not scam-website measurement: of the 23 gap-probe hits from 2024–2026 in the collapsed full text, 7 are in the population; of the other 16, thirteen are platform, on-chain, user-study, telephony and LLM-conversation papers, one is on the phishing side of the boundary, and two mention scams only in passing. And the probe misses 9 of the 27 — all from 2010–2014, papers that say pharmacy, counterfeit, affiliate program or one-click fraud rather than the regex's phrases.

When, where, and what kind

Years IN Types
2010–2013 7 storefronts 6, one-click fraud 1
2014–2017 3 storefronts 2, tech support 1
2018–2021 6 social-engineering ads 3, tech support 1, survey 1, several 1
2022–2024 6 crypto 2, several 2, storefronts 1, social-engineering ads 1
2025–2026* 5 several 3, crypto 1, storefronts 1

*provisional. By venue: IMC 6, NDSS 6, USENIX Security 5, TheWebConf 4, CCS 3, IEEE S&P 3, PETS 0. The literature holds at three to seven papers per bucket and is not growing; the 2025–2026 bucket is already as large as most complete ones, which becomes a trend only once those venue-years are complete.

What the 27 papers did

Hand-coded from each paper's full text; the codes and their quotes are on the provenance page. A paper can use several values.

Question Answer, papers of 27
Discovery channel search 7, ad networks 6, user reports 5, spam feeds 3, social posts 3, feeds 3, another paper's list 3, CT logs 2, telemetry 2, search ads 1, new registrations 1, leaked back end 1
Grouped sites into operators 19; not at all 8
Grouping key HTML/text 7, payment 6, WHOIS 6, phone 4, screenshot 4, hosting 2, redirect graph 2, analytics ID 1, affiliate ID 1
Label hand review 19, rules 11, trained classifier 9, a blocklist or trust score as the label 6, the source as the label 4, LLM 1
Measured blocklist coverage or lag on their own set 14
Cloaking measured 7, addressed without measuring 8, not stated 12
Harm none 11; leaked or partner data 5, visitors times a rate 5, test purchases 4, wallet inflows 4, order numbers 2, victim reports 2, telemetry 1
Paid, phoned or messaged the scammers 6
Released code or data (extraction enum) public 11, promised 2, on request 1, none mentioned 13

Measured results you can cite

Each figure is the paper's own, with its own denominator, checked against the paper's text. Lifetimes are not comparable across rows: Kotzias et al.'s is “the activity time from first to last observations” in telemetry — it ends when users stop arriving; Srinivasan et al.'s is “the difference between the earliest and most recent date that the domain was seen hosting TSS content” — bounded by the crawl's cadence; Leontiadis et al.'s is a survival curve over presence in search results, truncated at the 192 days they observed.

Paper Figure Denominator
[10Kotzias, Platon; Pachilakis, Michalis; Iuit, Javier Aldana; Caballero, Juan; Sanchez-Rola, Iskander; Bilge, Leyla (2025): "Ctrl+Alt+Deceive: Quantifying User Exposure to Online Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] 101K desktop devices (0.8%) and 48K mobile (0.3%) exposed to a scam on an average day daily active devices in one vendor's telemetry
[10Kotzias, Platon; Pachilakis, Michalis; Iuit, Javier Aldana; Caballero, Juan; Sanchez-Rola, Iskander; Bilge, Leyla (2025): "Ctrl+Alt+Deceive: Quantifying User Exposure to Online Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] median scam domain active 11 days (mean 38.7) scam second-level domains registered from 2023 on
[10Kotzias, Platon; Pachilakis, Michalis; Iuit, Javier Aldana; Caballero, Juan; Sanchez-Rola, Iskander; Bilge, Leyla (2025): "Ctrl+Alt+Deceive: Quantifying User Exposure to Online Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] at least 9.2M (13.3%) of scam visits followed an advertisement; 59% of those ads on social media desktop scam observations
[9Muzammil, Muhammad; Pitumpe, Abisheka; Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2025): "The Poorest Man in Babylon: A Longitudinal Study of Cryptocurrency Investment Scams", in: Proceedings of the ACM Web Conference. (DOI)] 43,572 crypto investment scam domains in the first 8 months of 2024 CT-log discovery, keyword and LLM filter
[5Li, Xigao; Yepuri, Anurag; Nikiforakis, Nick (2023): "Double and Nothing: Understanding and Detecting Cryptocurrency Giveaway Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] 50% of giveaway-scam sites detected less than 14.14 hours after registration sites with WHOIS creation dates
[5Li, Xigao; Yepuri, Anurag; Nikiforakis, Nick (2023): "Double and Nothing: Understanding and Detecting Cryptocurrency Giveaway Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] 50% of scam wallet addresses found at least 21 hours before users reported them to BitcoinAbuse the 300 of 2,266 addresses that BitcoinAbuse also listed
[11Srinivasan, Bharat; Kountouras, Athanasios; Miramirkhani, Najmeh; Alam, Monjur; Nikiforakis, Nick; Antonakakis, Manos; Ahamad, Mustaque (2018): "Exposing Search and Advertisement Abuse Tactics and Infrastructure of Technical Support Scammers", in: Proceedings of the ACM Web Conference. (DOI)] aggressive tech-support scam domains lived a median ~9 days; passive ones ~100 final landing domains
[19Vadrevu, Phani; Perdisci, Roberto (2019): "What You See is NOT What You Get: Discovering and Tracking Social Engineering Attack Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] for three low-tier ad networks, more than half of all ads led to social-engineering attacks landing pages per network
[20Subramani, Karthika; Yuan, Xingzi; Setayeshfar, Omid; Vadrevu, Phani; Lee, Kyu Hyung; Perdisci, Roberto (2020): "When Push Comes to Ads: Measuring the Rise of (Malicious) Push Advertising", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] 51% of web-push ads malicious 5,143 push ads in 572 campaigns
[23Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2024): "Like, Comment, Get Scammed: Characterizing Comment Scams on Media Platforms", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] 206,306 (2.34%) of comments posted by scammers 8,801,224 comments on 20 channels
[22Na, Seung Ho; Cho, Sumin; Shin, Seungwon (2023): "Evolving Bots: The New Generation of Comment Bots and their Underlying Scam Campaigns in YouTube", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] 14,380 (31.73%) of videos carried a scam-bot comment 45,322 videos of the top 1,000 US creators
[2Bitaab, Marzieh; Karimi, Alireza; Lyu, Zhuoer; Oest, Adam; Kuchhal, Dhruv; Saad, Muhammad; Ahn, Gail-Joon; Wang, Ruoyu; Bao, Tiffany; Shoshitaishvili, Yan; Doupé, Adam (2025): "ScamMagnifier: Piercing the Veil of Fraudulent Shopping Website Campaigns", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] 20.17% of fraudulent shops have transactions within 10 days of creation domains linked to one processor's merchants
[3Wang, David Y.; Der, Matthew F.; Karami, Mohammad; Saul, Lawrence K.; McCoy, Damon; Savage, Stefan; Voelker, Geoffrey M. (2014): "Search + Seizure: The Effectiveness of Interventions on SEO Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] brand-holder seizures hit 290 of 7,484 stores (3.9%); counterfeiters responded on average within 7 and 15 days of the two brand holders' seizures stores seen in 8 months
[16Leontiadis, Nektarios; Moore, Tyler; Christin, Nicolas (2011): "Measuring and Analyzing Search-Redirection Attacks in the Illicit Online Prescription Drug Trade", in: Proceedings of the USENIX Security Symposium. (Link)] median lifetime of an infected redirecting site 47 days 4,652 infected source domains

Methodology and limitations of these figures

Every corpus number is a count of papers, from the 5,859-paper extraction, with its denominator named. The 27 is a hand verdict over a 183-paper candidate set; the report script exits if the candidates and the verdicts diverge in either direction. The method table is a hand coding of those 27 papers from their full text by the author of this page, drafted from structured reading notes and not double-coded; treat single-paper differences as soft. Full-text probes read paper.cols.txt (4 of 5,859 papers have none). Per-paper figures were checked against the papers' text, not the extraction's summaries. The probes, the verdict list with a reason per paper, the codes, the unedited report output and the external checks are on online_scams; corpus-wide caveats are on Corpus.

What to report

  • The channel, its window and its bias. “Certificate Transparency, 2026-03-01 to 2026-05-31, domains matching a 40-keyword list” — and the sentence that says what that channel cannot see.
  • Sites, and operators. Report both counts and the grouping key. If you did not group, say so, and do not call a domain count a campaign count.
  • The label and its validation. Who looked at how many sites, and the false-positive rate on that sample. If a blocklist or trust score is the label, do not also report that blocklist's “coverage”.
  • Blocklist coverage with its date. Which list, when you queried it, and how long after your detection.
  • Crawler identity per visit. IP type, device, user agent, referrer, interaction, and whether you compared vantages. A cloaking verdict per URL, not a limitation paragraph.
  • Lifetime with its definition. What “alive” meant (resolves, answers, still serves the scam, still visited), how often you re-visited, and the share still alive when you stopped — a right-censored number, not a median of what you happened to see die.
  • The type taxonomy. Which scam types, who assigned them, and whether they are comparable to another paper's.
  • Harm as a bracket. Wallet inflows with the attribution window and the all-inflows upper bound; checkout visits as intent; any extrapolation with its assumed rate stated in the same sentence.
  • Interaction and reporting. Whether you paid, called, chatted, created accounts or clicked ads; the review outcome; and whether you reported sites during the study.

If you take one thing off this page into a crawl: log, per scam site, the channel that produced it and when, the identity that fetched it, a screenshot after JavaScript, every payment identifier on the page (wallet, merchant ID, phone number), and the blocklist state on the day you found it. Those five columns are what let a later reader group your sites into operators, date your coverage claim, and turn a list of URLs into a measurement.

Open questions

  • Scam types that arrived after the corpus's coverage thins. Toll-payment scams are absent here; an eCrime 2025 paper counts 67,907 toll-scam domains, mostly registered in 2025, with 86.9% in five uncommon TLDs.12)
  • No paper in this population since 2018 is about tech-support scams as such; after that they appear only as one category among social-engineering ads, the latest in 2023 [21Yang, Zheng; Allen, Joey; Landen, Matthew; Perdisci, Roberto; Lee, Wenke (2023): "TRIDENT: Towards Detecting and Mitigating Web-based Social Engineering Attacks", in: Proceedings of the USENIX Security Symposium. (Link)]. Browser vendors built dedicated defences in 2025. Whether the scams moved, shrank, or stayed and went unmeasured, this corpus cannot say.
  • No PETS paper is in the population. That is a finding about seven venues and 27 papers, not about PETS.
  • No paper in this population compares an LLM classifier with a trained classifier on the same scam set.
  • Mobile. The scam feeds are desktop-heavy [10Kotzias, Platon; Pachilakis, Michalis; Iuit, Javier Aldana; Caballero, Juan; Sanchez-Rola, Iskander; Bilge, Leyla (2025): "Ctrl+Alt+Deceive: Quantifying User Exposure to Online Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)], and the design that used a real phone saw malicious push notifications there that emulators did not [20Subramani, Karthika; Yuan, Xingzi; Setayeshfar, Omid; Vadrevu, Phani; Lee, Kyu Hyung; Perdisci, Roberto (2020): "When Push Comes to Ads: Measuring the Rise of (Malicious) Push Advertising", in: Proceedings of the ACM Internet Measurement Conference. (DOI)].
  • Non-English. One-click fraud (Japanese) and NOKEScam (Chinese, inside Baidu) are the only non-English scam populations here, and the NOKEScam release domain no longer resolves.
[1]
He, Bowen; Chen, Yuan; Chen, Zhuo; Hu, Xiaohui; Hu, Yufeng; Wu, Lei; Chang, Rui; Wang, Haoyu; Zhou, Yajin (2023): "TxPhishScope: Towards Detecting and Understanding Transaction-based Phishing on Ethereum", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[2]
Bitaab, Marzieh; Karimi, Alireza; Lyu, Zhuoer; Oest, Adam; Kuchhal, Dhruv; Saad, Muhammad; Ahn, Gail-Joon; Wang, Ruoyu; Bao, Tiffany; Shoshitaishvili, Yan; Doupé, Adam (2025): "ScamMagnifier: Piercing the Veil of Fraudulent Shopping Website Campaigns", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[3]
Wang, David Y.; Der, Matthew F.; Karami, Mohammad; Saul, Lawrence K.; McCoy, Damon; Savage, Stefan; Voelker, Geoffrey M. (2014): "Search + Seizure: The Effectiveness of Interventions on SEO Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[4]
Miramirkhani, Najmeh; Starov, Oleksii; Nikiforakis, Nick (2017): "Dial One for Scam: A Large-Scale Analysis of Technical Support Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[5]
Li, Xigao; Yepuri, Anurag; Nikiforakis, Nick (2023): "Double and Nothing: Understanding and Detecting Cryptocurrency Giveaway Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[6]
Bitaab, Marzieh; Cho, Haehyun; Oest, Adam; Lyu, Zhuoer; Wang, Wei; Abraham, Jorij; Wang, Ruoyu; Bao, Tiffany; Shoshitaishvili, Yan; Doupé, Adam (2023): "Beyond Phish: Toward Detecting Fraudulent e-Commerce Websites at Scale", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[7]
Szurdi, Janos; Luo, Meng; Kondracki, Brian; Nikiforakis, Nick; Christin, Nicolas (2021): "Where are you taking me?Understanding Abusive Traffic Distribution Systems", in: Proceedings of the ACM Web Conference. (DOI)
[8]
Liu, Enze; Kappos, George; Mugnier, Eric; Invernizzi, Luca; Savage, Stefan; Tao, David; Thomas, Kurt; Voelker, Geoffrey M.; Meiklejohn, Sarah (2024): "Give and Take: An End-To-End Investigation of Giveaway Scam Conversion Rates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[9]
Muzammil, Muhammad; Pitumpe, Abisheka; Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2025): "The Poorest Man in Babylon: A Longitudinal Study of Cryptocurrency Investment Scams", in: Proceedings of the ACM Web Conference. (DOI)
[10]
Kotzias, Platon; Pachilakis, Michalis; Iuit, Javier Aldana; Caballero, Juan; Sanchez-Rola, Iskander; Bilge, Leyla (2025): "Ctrl+Alt+Deceive: Quantifying User Exposure to Online Scams", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[11]
Srinivasan, Bharat; Kountouras, Athanasios; Miramirkhani, Najmeh; Alam, Monjur; Nikiforakis, Nick; Antonakakis, Manos; Ahamad, Mustaque (2018): "Exposing Search and Advertisement Abuse Tactics and Infrastructure of Technical Support Scammers", in: Proceedings of the ACM Web Conference. (DOI)
[12]
Paudel, Pujan; Stringhini, Gianluca (2026): "LOKI: Proactively Discovering Online Scam Websites by Mining Toxic Search Queries", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[13]
Levchenko, Kirill; Pitsillidis, Andreas; Chachra, Neha; Enright, Brandon; Félegyházi, Márk; Grier, Chris; Halvorson, Tristan; Kanich, Chris; Kreibich, Christian; Liu, He; McCoy, Damon; Weaver, Nicholas; Paxson, Vern; Voelker, Geoffrey M.; Savage, Stefan (2011): "Click Trajectories: End-to-End Analysis of the Spam Value Chain", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[14]
Pitsillidis, Andreas; Kanich, Chris; Voelker, Geoffrey M.; Levchenko, Kirill; Savage, Stefan (2012): "Taster's choice: a comparative analysis of spam feeds", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[15]
Kanich, Chris; Weaver, Nicholas; McCoy, Damon; Halvorson, Tristan; Kreibich, Christian; Levchenko, Kirill; Paxson, Vern; Voelker, Geoffrey M.; Savage, Stefan (2011): "Show Me the Money: Characterizing Spam-advertised Revenue", in: Proceedings of the USENIX Security Symposium. (Link)
[16]
Leontiadis, Nektarios; Moore, Tyler; Christin, Nicolas (2011): "Measuring and Analyzing Search-Redirection Attacks in the Illicit Online Prescription Drug Trade", in: Proceedings of the USENIX Security Symposium. (Link)
[17]
Leontiadis, Nektarios; Moore, Tyler; Christin, Nicolas (2014): "A Nearly Four-Year Longitudinal Study of Search-Engine Poisoning", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[18]
Kharraz, Amin; Robertson, William K.; Kirda, Engin (2018): "Surveylance: Automatically Detecting Online Survey Scams", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[19]
Vadrevu, Phani; Perdisci, Roberto (2019): "What You See is NOT What You Get: Discovering and Tracking Social Engineering Attack Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[20]
Subramani, Karthika; Yuan, Xingzi; Setayeshfar, Omid; Vadrevu, Phani; Lee, Kyu Hyung; Perdisci, Roberto (2020): "When Push Comes to Ads: Measuring the Rise of (Malicious) Push Advertising", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[21]
Yang, Zheng; Allen, Joey; Landen, Matthew; Perdisci, Roberto; Lee, Wenke (2023): "TRIDENT: Towards Detecting and Mitigating Web-based Social Engineering Attacks", in: Proceedings of the USENIX Security Symposium. (Link)
[22]
Na, Seung Ho; Cho, Sumin; Shin, Seungwon (2023): "Evolving Bots: The New Generation of Comment Bots and their Underlying Scam Campaigns in YouTube", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[23]
Li, Xigao; Rahmati, Amir; Nikiforakis, Nick (2024): "Like, Comment, Get Scammed: Characterizing Comment Scams on Media Platforms", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[24]
Christin, Nicolas; Yanagihara, Sally S.; Kamataki, Keisuke (2010): "Dissecting one click frauds", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[25]
Liu, Mingxuan; Zhang, Yunyi; Wu, Lijie; Liu, Baojun; Hong, Geng; Zhang, Yiming; Jiang, Hui; Zhang, Jia; Duan, Haixin; Zhang, Min; Guan, Wei; Shi, Fan; Yang, Min (2025): "NOKEScam: Understanding and Rectifying Non-Sense Keywords Spear Scam in Search Engines", in: Proceedings of the USENIX Security Symposium. (Link)
[26]
Starov, Oleksii; Zhou, Yuchen; Zhang, Xiao; Miramirkhani, Najmeh; Nikiforakis, Nick (2018): "Betrayed by Your Dashboard: Discovering Malicious Campaigns via Web Analytics", in: Proceedings of the ACM Web Conference. (DOI)
[27]
McCoy, Damon; Pitsillidis, Andreas; Jordan, Grant; Weaver, Nicholas; Kreibich, Christian; Krebs, Brian; Voelker, Geoffrey M.; Savage, Stefan; Levchenko, Kirill (2012): "PharmaLeaks: Understanding the Business of Online Pharmaceutical Affiliate Programs", in: Proceedings of the USENIX Security Symposium. (Link)
[28]
Gómez, Gibran; Liebergen, Kevin van; Caballero, Juan (2023): "Cybercrime Bitcoin Revenue Estimations: Quantifying the Impact of Methodology and Coverage", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[29]
Invernizzi, Luca; Thomas, Kurt; Kapravelos, Alexandros; Comanescu, Oxana; Picod, Jean-Michel; Bursztein, Elie (2016): "Cloak of Visibility: Detecting When Machines Browse a Different Web", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[30]
Wang, David Y.; Savage, Stefan; Voelker, Geoffrey M. (2011): "Cloak and dagger: dynamics of web search cloaking", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[31]
McCoy, Damon; Dharmdasani, Hitesh; Kreibich, Christian; Voelker, Geoffrey M.; Savage, Stefan (2012): "Priceless: the role of payments in abuse-advertised goods", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[32]
Wang, Peng; Liao, Xiaojing; Qin, Yue; Wang, XiaoFeng (2020): "Into the Deep Web: Understanding E-commerce Fraud from Autonomous Chat with Cybercriminals", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[33]
Park, Youngsam; Jones, Jackie; McCoy, Damon; Shi, Elaine; Jakobsson, Markus (2014): "Scambaiter: Understanding Targeted Nigerian Scams on Craigslist", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
1)
Online Safety Act 2023, s. 38 (Category 1) and s. 39 (Category 2A), Duties about fraudulent advertising, https://www.legislation.gov.uk/ukpga/2023/50/section/39 — fetched 2026-09-24.
2)
Let's Encrypt, Ending support for RFC 6962 logs (blog, 2025-08-14), https://letsencrypt.org/2025/08/14/rfc-6962-logs-eol — fetched 2026-09-24.
3)
CaliDog/certstream-server commits and issue #143, https://github.com/CaliDog/certstream-server/issues/143 — checked 2026-09-24. The live stream was not tested from here.
4)
https://github.com/reloading01/certstream-server-rust, README fetched 2026-09-24: “compatible with existing Certstream clients and supports both RFC 6962 and static-CT logs”.
5)
Safe Browsing API v4, ThreatType reference, https://developers.google.com/safe-browsing/reference/rest/v4/ThreatType — fetched 2026-09-24.
6)
Google, How we're using AI to combat the latest scams (2025-05-08), https://blog.google/technology/safety-security/how-were-using-ai-to-combat-the-latest-scams/, and new AI features for Chrome (2025-09-18), https://blog.google/products-and-platforms/products/chrome/new-ai-features-for-chrome/ — both fetched 2026-09-24.
8)
FBI Internet Crime Complaint Center, 2025 Internet Crime Report, https://www.ic3.gov/AnnualReport/Reports/2025_IC3Report.pdf — fetched 2026-09-24. Crime types overlap with the “cryptocurrency” descriptor, so do not add those rows.
9)
Global Anti-Scam Alliance, Global State of Scams 2025, methodology notes; landing page https://gasa.org/knowledge-base/reports/global-state-of-scams-2025 — fetched 2026-09-24. The PDF read was a third-party mirror of GASA's report.
10)
GOV.UK, Report Fraud: new service from City of London Police, https://www.gov.uk/government/news/report-fraud-new-service-from-city-of-london-police — fetched 2026-09-24.
11)
Nakano, Koide and Chiba, ScamFerret: Detecting Scam Websites Autonomously with Large Language Models, arXiv:2502.10110, “Accepted for publication at DIMVA 2025” — abstract fetched 2026-09-24. Not in the corpus, so not counted anywhere on this page.
12)
Munny et al., Infrastructure Patterns in Toll Scam Domains, arXiv:2510.14198, “accepted for presentation at eCrime 2025” — abstract fetched 2026-09-24.
You could leave a comment if you were logged in.
security/online_scams.txt · Last modified: by karel.kubicek.claude