User Tools

Site Tools


provenance:security:phishing

Provenance: security:phishing

Working log behind phishing. Corpus-wide caveats are on corpus. Citations use the shared bibliography; this page adds no keys of its own.

Run: 2026-08-27. Corpus: 5,859 extracted papers, 7 venues, 2010–2026, data/extract/run1. Item: drain wiki-measuretheweb / “security:phishing (new)”, claimed by key as cursor-drain-phishing, store id 223, run 61. Author of the page and this log: Cursor, not Claude Code. Reviews: three focused passes then one generic, all on GPT 5.6 Luna medium (gpt-5.6-luna-medium), the sitting's requested substitute for the spec's sonnet/fable split.

Creating, not extending. ?do=export_raw on security:phishing returned the HTML error page (not an existence test by byte count). dw.mjs pages does not list provenance:. Neighbour security already linked the red child Phishing; live start already has the same link. Overlap judgement: write the promised child rather than widen virustotal or website_classification — those pages own the vendor panel and the topic label; this one owns feeds and live phishing websites.

No ~~DISCUSSION~~ on this provenance page. Comments belong on the content page.

Scope decisions

Decision Why What a reasonable person might have done instead
PAGE_N = 52, not the 139-paper schema union and not the 1,098-paper keyword hit The union mixes live-site crawls, detectors, feed-as-GT for a different question, user studies, email/SMS, on-chain “phishing”, and mobile UI. Merging them is a meaningless N. The keyword hit is a fact about these being security venues. Publish 139 as “phishing papers”. Rejected.
Single primary ROLE per paper A detector that also crawls still has one object of study. Double-counting would make the role table sum past 139. Multi-label. Defensible for cloaking∩live-site; not this sitting.
GSB on this page, not a sixth security child Already decided on security. GSB is a feed. A GSB-API tutorial page. Rejected: MDN/Google docs exist; the student question is what a hit proves.
Do not visit feed URLs The community feed is a list of live phish. Fetching the list is the instrument; fetching the sites is a crawl we are not running. “Verify a sample is still up.” That is a different paper.
Cite GSB v4 sunset 2027-03-31 only with the Brave/email source The date is in Google's notification mail (Brave issue 56023), not on the migration docs this sitting fetched. Put 2027-03-31 in the lead as a Google-docs fact. That would be a lie about the primary source.
Keep the poster in PAGE_N, show 52 and 51 CCS 2014 PhishTrack is a poster. Dropping it is a sensitivity check, not the headline. Quietly drop posters.
Do not add uncited detector keys to the bibliography The page cites the measurement papers. Phishpedia/KnowPhish/PhishLang appear as names in a sentence, not as [key]. Unused keys still break every page if they collide. Dump every PAGE_N paper into the bib.

Report script

scripts/report_phishing.mjs plus scripts/phish_fold.mjs (generated by scripts/_gen_phish_fold.py; do not hand-edit the ROLE Map). Deterministic. Re-run:

node scripts/report_phishing.mjs > scripts/report_phishing-output.txt

Exits 1 if the schema union and ROLE diverge in either direction, if a ROLE value is unknown, if a ROLE quote is shorter than 20 characters, if missing paper.cols.txt is not 4, or if PhishTank is missing from the feed ranking. A printed FAILURE with exit 0 is forbidden.

python3 pages/openphish_community_feed.py is the published script. Its output is quoted in a <code> block on the content page. It does not visit the listed sites. scripts/phishing_probe.sh re-fetches the feeds and the docs; it prints FAILED per broken check and exits 1.

Queries

# Query Population / denominator Result On the page?
Q1 Extraction records all 5,859 outer frame
Q2 Full-text /\bphish(?:ing|tank|ers)?\b/i 5,855 with .cols (4 missing) 1,098 yes, labelled not-N
Q3 Full-text Google Safe Browsing 5,855 110 provenance + namespace; content page uses the API facts
Q4 Full-text PhishTank / OpenPhish / APWG 5,855 219 provenance
Q5 Schema union: detection phenomenon/technique, classification resource/targetDetail, population.sourceList, or slug matching /phish/ 5,859 139 (92 web, 64 crawled) yes, candidate set
Q6 Hand ROLE over Q5 139 live-site 29, detector 13, cloaking 8, kit 2, feed-gt 19, user-study 29, email-sms 9, onchain 6, mobile-ui 6, mention 18 yes
Q7 PAGE_N = ROLE in {live-site, detector, cloaking, kit} 139 52 (37.4%) yes, the page's N
Q8 Q7 web / crawled 52 49 (94.2%) / 40 (76.9%) yes
Q9 Q7 posters 52 1 (CCS 2014 PhishTrack); without it 51 yes
Q10 Q7 by venue 52 USENIX 17, WWW 11, IMC 9, CCS 8, NDSS 4, IEEE-SP 3, PETS 0 yes
Q11 Q7 by year-bucket 52 1 / 7 / 13 / 18 / 13* yes
Q12 Feed family fold on sourceList + resourceName + tools.name 139 VT 32, PhishTank 30, GSB 27, APWG 13, OpenPhish 10, CertStream 9, … yes
Q13 Same fold on PAGE_N 52 VT 16, GSB 13, APWG 12, PhishTank 11, CertStream 9, OpenPhish 8 yes
Q14 sourceList/resourceName matching /phishtank/i 139 28 provenance (explains the namespace 21)
Q15 Namespace sitting's topNames fold of /phish/ sourceList+resourceName 139 21 yes, as the number the namespace published
Q16 Full-text probes on PAGE_N 52 cloak 34, takedown 31, lifespan 26, phishtank 32, openphish 20, gsb 35, captcha 21 yes, labelled upper bounds
Q17 ROLE quotes vs paper.cols.txt 139 exact 72, partial 0, absent 66, missing-file 0, title-only 1 provenance

Folding and residue

ROLE is a hand map, not a regex. 139 keys, generated from scripts/_gen_phish_fold.py. The report exits 1 if a union paper has no ROLE or a ROLE key is outside the union. Single-label: the ten rows sum to 139.

Feed families (FEED_FAMILIES in phish_fold.mjs): ordered regexes, most-specific first, over sourceList + resourceName + tools.name. A paper is counted once per family. Residue of phish-ish strings that matched no family: 118 strings, printed in full in the report below. They are custom kit names, “phishing emails”, GPPF, PhishFarm-the-experiment, combined vendor lists — not a missing commercial feed. Do not promote residue into a family without reading the string.

Namespace 21 vs this sitting's 30. The namespace page counted a topNames fold of sourceList/resourceName that already matched /phish/, then took the family whose spellings include “phishtank”: 21. Any sourceList/resourceName matching /phishtank/i (including “Google Safe Browsing API, Malware Patrol, PhishTank, …”) is 28. Adding tools[].name and the family fold is 30. The content page uses 30 and explains the 21. Do not mix them.

Quotes spot-checked

ROLE deciding sentences vs paper.cols.txt: 72 exact, 0 partial, 66 absent, 0 missing-file, 1 title-only. The checker compared the stored ROLE quote string against .cols. A focused reviewer spot-checked five “absent” keys (SearchAudit, Compa, Cui 2017, PhishTime, PhishDecloaker) and found the deciding sentence in .cols in each case; at least two were verbatim modulo a column break. That is not a census of the 66 and does not prove the other 61 are clean — it is enough to reject “the roles are fabricated”, and not enough to claim every absent quote is mere ellipsis. The 66 keys are listed in the report.

Title-only / corpus defect: IEEE S&P 2024 from-chatbots-to-phishbots-…. Slug and bibliographic title are Roy et al. on LLM phishing. Stored PDF and paper.cols.txt are Nanayakkara et al. on differential privacy, DOI 10.1109/SP54263.2024.00182 on the PDF. ROLE=mention, quote taken from the title. Do not cite the stored PDF for that slug. Not patched in the extraction; recorded here.

Load-bearing paper figures on the content page were checked against paper.cols.txt (and against detection.prevalence where the extractor captured them):

Paper Needle In .cols?
Lee et al. WWW 2025 54.04 hours / 5.46 hours / 286,237 / 67.23% / 75.84% / 99.80% yes
Zhang et al. IEEE S&P 2021 35,067 of 112,005 (31.31%); 23.32% / 33.70%; 150 artificial; 21 (42%) Click Through in Edge yes
Peng et al. IMC 2019 15 of 68 / 36 simple phishing sites; 4 vendors week three PayPal yes
Oest et al. IEEE S&P 2019 23.0% vs 49.4%; 126 → 238 minutes; mobile no warnings mid-2017–late 2018 yes
Oest et al. USENIX 2020 21 hours; 4.8 million; 7.42%; 71.51% / 43.55% yes
Zhang et al. CCS 2022 96.52% (2,831/2,933); 82.28% of 160,728; 99.31% “bot”; 88.98% (7,540/8,474) yes
Teoh et al. USENIX 2024 0/100; 7.6% (66/869); 16h vs 11h yes
Acharya et al. USENIX 2021 18 of 20 still up after one month (2 of 20 blocked) yes
Subramani et al. IMC 2022 56,027 → 51,859; 23,446 (45%) required input yes
Han et al. CCS 2016 643 kits; eight days; 98% / 12 days; 62% after 75% of victims yes
Bijmans et al. USENIX 2021 mean 45 hours, median 24 hours yes
Tian et al. IMC 2018 1,175 confirmed; 91.5% undetected a month later yes
Moura et al. CCS 2024 28,754 domains; after 24h 80% of .nl, roughly 70% of .ie yes
Lin et al. USENIX 2022 73.98% / 90.08% / 91.36% fingerprint collection on phishing-with-JS yes
Lin et al. USENIX 2021 “this gave us 350K phishing URLs” yes

External sources

Fetched 2026-08-27, re-checkable with scripts/phishing_probe.sh:

  • PhishTank home / API / developer / FAQ / register — all HTTP 200. Operator string “PhishTank is operated by Cisco Talos Intelligence Group”. Unauthenticated dump http://data.phishtank.com/data/online-valid.csv.gz: 73,660 rows, all online=yes and verified=yes, target Other 65,457 (88.9%), gzip 2,549,032 B, submissions 2011-02-18 → 2026-08-27. The documented public dump path is on the content page because that is the file papers train on. Signed or expiring CDN query strings stay off the wiki. stats.php timed out at 20s from this host — documented, not treated as a probe failure.
  • Wikipedia “registration closed” was not confirmed on the live FAQ; register.php returned 200. Not cited.
  • OpenPhish feed.txt: 300 URLs, 269 hosts, 201 https / 99 http, 15,020 bytes. phishing_feeds.html Community row: 12 hours, Limited, Text File, Free. Premium: 5 minutes. Academic programme: 60 days + 30-day archive, institutional email to support@openphish.com.
  • GSB v4 overview: deprecated, non-commercial, Web Risk for commercial. Unauthenticated v4/threatLists and v5/hashLists: HTTP 403. Test page still says “Should show a phishing warning”. Migration guide: “as of 2021, 60% of sites that deliver attacks live less than 10 minutes”; “around 25–30% of missing phishing protection is due to such data staleness.”
  • GSB v4 end date 31 March 2027: Brave brave-browser#56023 quoting the notification mail. Not on the docs page this sitting fetched.
  • APWG eCX: apwg.org/ecx HTTP 200. No public dump.

Rejected: SEO listicles of “best phishing feeds”; any instruction to crawl the URLs in feed.txt; treating 2027-03-31 as a Google-docs fact.

What could not be established

  • A live PhishTank submissions-per-day series (stats.php timed out).
  • Whether PhishTank still accepts new user registrations — register.php is 200 and mostly JS; the FAQ no longer has the 2020 “closed” sentence this sitting could find. Page says Wikipedia was not confirmed, and stops.
  • APWG eCX membership contents without being a member.
  • GSB URL lists — the API does not return them. That is the finding.
  • A complete quote-verbatim map of all 139 ROLE sentences (66 absent from .cols because of extraction ellipsis).
  • PETS “does not do phishing websites” as a venue finding — N=2 in the union, both user-study. Named as a small-N observation.

Published script

pages/openphish_community_feed.py. Stdlib. Advertised invocation: python3 openphish_community_feed.py. Empty feed, non-http(s) rows, or a URL with no hostname: exit 1. HTTP != 200: exit 1. Direct schemes['https'] / schemes['http'] (Counter returns 0 for a missing scheme, which is the right count). The <file> block on the content page is byte-identical to this file (checked before freeze).

Bibliography keys

Collision-checked against a fresh ?do=export_raw of live literature:bibliography on 2026-08-27 (558 keys; local pages/literature_bibliography.txt is stale and was not used).

Already live, not re-added: zhang2021_crawlphish, peng2019_opening, zhang2022_spartacus, alam2026_philter, sanchezrola2023_rods.

Added (cited on the content page): lee2025_7days, moura2024_cctld, lin2022_sheep, oest2019_phishfarm, oest2020_phishtime, oest2020_sunrise, han2016_phisheye, teoh2024_phishdecloaker, subramani2022_phishinpatterns, bijmans2021_catching, tian2018_needle, lin2021_phishpedia, acharya2021_phishprint.

USENIX entries use url= not DOI (index has none). Affiliation tokens (PayPal, Trustwave, EURECOM, NCS Cyber Special Ops) were stripped from author fields after fetch_authors.py glued them on — known USENIX-author-list bug, same as virustotal.

Generated but not added because the content page does not [cite] them: roy2023_freewaters, lim2024_legit, cui2017_tracking, maroofi2020_human, choi2024_phishinwebview, lim2025_phishers, roy2026_phishlang, liu2023_knowledge, li2024_knowphish, liu2022_inferring.

Review log

Drafts frozen in out/freeze_phishing/ before the three focused passes. All four reviewers: GPT 5.6 Luna medium (gpt-5.6-luna-medium), substituting for the spec's sonnet/fable split. Pages were not edited while the three focused reviewers ran.

Focused pass 1 — figures vs script

# Finding Decision
1 PhishEye “5%–95% victim-connection quantiles” not in the report's prevalence string Rejected. The paper defines first/last victim as the 5% and 95% tails (.cols lines 736–744); eight days is 10 − 2. Added as PAPER_LITERAL phisheye_lifetime_cut so the guard sees the digits. The report's prevalence field was incomplete, the page was not.
2 Published script does not reject http://foo bar Rejected. Hostname-None already exits 1. Gold-plating URL syntax is not the script's job; the community feed is a URL list.

Focused pass 2 — citations and quotes

# Finding Decision
1 Cloaking paragraph attributed “invisible to VT/GSB/SmartScreen for a week” to the 66/869 field study Accepted. That week-long 0/100 is the controlled experiment (Table 1, 100 URLs per type). The field study says the 66 were “discovered solely by PhishDecloaker” and gives median blacklist 16h vs 11h. Split the two claims. WRAP important already described the controlled experiment; tightened to “100 URLs per cloaking type”.
2 All 17 citekeys resolve; no bib collisions; author table matches Noted, no change.
3 Corpus defect (Roy slug / Nanayakkara PDF) correctly disclosed Noted, no change.
4 66 “absent” ROLE quotes: at least two are verbatim modulo line wrap, so “all ellipsis” overstates Accepted on provenance only. Softened: absent is truncation/line-wrap/path, not fabrication; not all 66 are ellipsis.

Focused pass 3 — external currency

# Finding Decision
1 PhishTank / OpenPhish / GSB / APWG claims match live primary sources. Feed counts moved 73,660 → 73,668 (expected). OpenPhish 300/269 still matched on re-fetch. 2027-03-31 still not on Google docs. Noted, no change to dated snapshots.
2 phishing_probe.sh dump and feed.txt fetches used curl -sS without -L; both hosts now redirect, so the probe false-failed Accepted. Both fetches now curl -sSL. The probe already exited 1 on FAILED (not the exit-0 defect).

Generic pass ran on the applied text, with no checklist.

Generic pass — no checklist

# Finding Decision
1 Provenance header/log read as if the generic pass had already happened Accepted. This table is that pass.
2 Content publishes the public dump path; provenance said dump URLs stay off the wiki Accepted as a provenance error. The documented public path is the instrument. Policy is now: public dump path on the page; signed CDN query strings off.
3 Lead “the feed is a delayed vote” overgeneralises Accepted. Vote is PhishTank; other feeds are dated provider snapshots.
4 Blocklist table treated online as years-old Accepted. Distinguished current online=yes filter from old submission_time.
5 “Cloud ranges are the first filter” / “what a victim saw” / “IP blocking is the wrong lever” overbroad Accepted. Scoped to Spartacus's proxy figure and Lee et al.'s CDN dataset.
6 WRAP todo claimed five papers “all have that tuple” Accepted. Now the page's recommended instrumentation.
7 Five spot-checks cannot clear all 66 absent quotes Accepted. Provenance now says five is not a census.
8 Detector papers in PAGE_N dilute the crawl/feed question Accepted in part. One sentence on why detectors are in the 52; not moved to provenance (a student evaluating a feed is often about to train).
9 Update security red-link list and 1,097 → 1,098 Accepted in part. WRAP now lists phishing as written. Cell figures stay the namespace snapshot (21, 1,097); the child re-derived 1,098 and PAGE_N 52 and owns those.

No further generic re-run. The applied fixes are local wording, not new figures.

Report output (unedited)

report_phishing-output.txt
# security:phishing — every figure with its denominator
 
Corpus: 5859 extracted papers from CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf and IEEE S&P, 2010–2026.
Crawled population: 1120 papers (19.1%), defined by crawlConfig != null or studyTypes includes automated-web-crawl.
Web platform: 1622 papers (27.7%).
2025–2026 venue-years are provisional: CCS/IMC 2026 not held; IEEE S&P/WWW 2026 incompletely selected.
Run date: figures below are recomputed on every run; external facts are in scripts/phishing_probe.sh and section Z.
 
## A. Why a keyword search is not a population
 
missing paper.cols.txt: 4 (corpus known gap is 4)
PUBLISHED_FT_PHISH 1098
PUBLISHED_FT_GSB 110
PUBLISHED_FT_FEEDS 219
Full-text /\bphish(?:ing|tank|ers)?\b/i: 1098 of 5855 papers with cols.
That count is a fact about these being security venues. It is NOT the page N.
Full-text Google Safe Browsing: 110.
Full-text PhishTank|OpenPhish|APWG: 219 (upper bound).
 
## B. Schema union (candidate set) and the hand map
 
PUBLISHED_UNION 139
PUBLISHED_UNION_WEB 92
PUBLISHED_UNION_CRAWLED 64
PUBLISHED_PAGE_N 52
Schema union (detection phenomenon/technique, classification resource/targetDetail, population.sourceList, or slug matching /phish/): 139 papers.
  of which platforms includes web: 92 (66.2%)
  of which crawled: 64 (46.0%)
PAGE_N = ROLE in {live-site, detector, cloaking, kit}: 52 papers (37.4% of the union).
Excluded from PAGE_N (user-study, email-sms, onchain, mobile-ui, mention, feed-gt): 87.
 
Hand roles over the 139-paper union (one primary role per paper):
Role        Papers  Share of 139  In PAGE_N?
----------  ------  ------------  ----------
live-site   29      20.9%         yes
detector    13      9.4%          yes
cloaking    8       5.8%          yes
kit         2       1.4%          yes
feed-gt     19      13.7%         no
user-study  29      20.9%         no
email-sms   9       6.5%          no
onchain     6       4.3%          no
mobile-ui   6       4.3%          no
mention     18      12.9%         no
 
Sum of role counts: 139 (must equal 139; single-label).
 
## C. PAGE_N shape
 
PAGE_N web: 49 / 52 (94.2%)
PAGE_N crawled: 40 / 52 (76.9%)
PAGE_N posters (slug poster- or title Poster:): 1
 
PAGE_N by venue:
Venue    PAGE_N  Share of 52  Union  Corpus
-------  ------  -----------  -----  ------
USENIX   17      32.7%        51     1410
CCS      8       15.4%        19     990
WWW      11      21.2%        21     843
IMC      9       17.3%        16     638
IEEE-SP  3       5.8%         14     767
NDSS     4       7.7%         16     701
PETS     0       0.0%         2      510
 
PAGE_N by year-bucket (2025–2026 starred = provisional):
Window      PAGE_N  Share of 52  Union
----------  ------  -----------  -----
2010–2013   1       1.9%         8
2014–2017   7       13.5%        19
2018–2021   13      25.0%        36
2022–2024   18      34.6%        40
2025–2026*  13      25.0%        36
 
PAGE_N per calendar year (do not read 2026 as a complete year):
Year   PAGE_N  Union
-----  ------  -----
2010   1       4
2011   0       0
2012   0       1
2013   0       3
2014   2       8
2015   0       4
2016   2       3
2017   3       4
2018   2       7
2019   4       12
2020   3       7
2021   4       10
2022   6       11
2023   5       9
2024   7       20
2025*  8       25
2026*  5       11
 
## D. Feeds named in the schema (union, then PAGE_N)
 
Feeds / oracles folded from sourceList + resourceName + tools.name, papers of the union:
Family                           Papers of union  Share of 139
-------------------------------  ---------------  ------------
VirusTotal (as phishing oracle)  32               23.0%
PhishTank                        30               21.6%
Google Safe Browsing             27               19.4%
APWG / eCX                       13               9.4%
OpenPhish                        10               7.2%
CertStream (discovery)           9                6.5%
Microsoft SmartScreen            3                2.2%
phishunt.io                      3                2.2%
SURBL                            2                1.4%
URLhaus                          2                1.4%
Netcraft                         1                0.7%
PhishStats                       1                0.7%
 
Same fold, papers of PAGE_N:
Family                           Papers of PAGE_N  Share of 52
-------------------------------  ----------------  -----------
VirusTotal (as phishing oracle)  16                30.8%
Google Safe Browsing             13                25.0%
APWG / eCX                       12                23.1%
PhishTank                        11                21.2%
CertStream (discovery)           9                 17.3%
OpenPhish                        8                 15.4%
Microsoft SmartScreen            3                 5.8%
phishunt.io                      3                 5.8%
Netcraft                         1                 1.9%
PhishStats                       1                 1.9%
 
Feed-fold residue (phish-ish string, no family) in the union: 118 strings.
PUBLISHED_FEED_RESIDUE 118
RESIDUE (all):
  IMC/2010/an-empirical-study-of-orphan-dns-servers-in-the-internet	five live feeds of phishing and malware-hosting sites
  IMC/2010/an-empirical-study-of-orphan-dns-servers-in-the-internet	phishing and malware feeds
  CCS/2014/poster-proactive-blacklist-update-for-anti-phishing	PhishTrack (custom)
  CCS/2014/poster-proactive-blacklist-update-for-anti-phishing	PhishNet
  CCS/2014/poster-proactive-blacklist-update-for-anti-phishing	PhishTrack
  IMC/2014/handcrafted-fraud-and-extortion-manual-account-hijacking-in-the-wild	user-reported phishing emails
  IMC/2014/handcrafted-fraud-and-extortion-manual-account-hijacking-in-the-wild	Google Forms taken down for phishing
  IMC/2014/handcrafted-fraud-and-extortion-manual-account-hijacking-in-the-wild	phishing pages targeting Google
  NDSS/2015/nophish-app-evaluation-lab-and-retention-study	NoPhish
  IMC/2014/the-dark-alleys-of-madison-avenue-understanding-malicious-advertisements	49 antivirus, spam and phishing blacklists
  IMC/2014/the-dark-alleys-of-madison-avenue-understanding-malicious-advertisements	Malware and phishing blacklists
  WWW/2016/cracking-classifiers-for-evasion-a-case-study-on-the-googles-phishing-pages-filt	Google's phishing pages filter (GPPF)
  CCS/2017/data-breaches-phishing-or-malware-understanding-the-risks-of-stolen-credentials	custom phishing-kit template rules
  CCS/2017/hiding-in-plain-sight-a-longitudinal-study-of-combosquatting-abuse	custom phishing detector
  WWW/2018/betrayed-by-your-dashboard-discovering-malicious-campaigns-via-web-analytics	custom phishing-target labels
  IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni	custom-generated phishing sites
  IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni	custom-generated phishing sites
  IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni	PhishFarm
  USENIX/2019/cognitive-triaging-of-phishing-attacks	phishing email database provided by Org
  WWW/2019/doppelgangers-on-the-dark-web-a-large-scale-assessment-on-phishing-hidden-web-se	manual phishing verification using external clues
  USENIX/2020/phishtime-continuous-longitudinal-measurement-of-the-effectiveness-of-anti-phish	PhishTime
  USENIX/2020/phishtime-continuous-longitudinal-measurement-of-the-effectiveness-of-anti-phish	PhishFarm
  USENIX/2020/sunrise-to-sunset-analyzing-the-end-to-end-life-cycle-and-effectiveness-of-phish	user-forwarded phishing emails
  USENIX/2020/sunrise-to-sunset-analyzing-the-end-to-end-life-cycle-and-effectiveness-of-phish	previously-proposed phishing URL classification scheme
  USENIX/2021/catching-phishers-by-their-bait-investigating-the-dutch-phishing-landscape-throu	custom phishing-kit fingerprints
  USENIX/2021/phishpedia-a-hybrid-deep-learning-based-approach-to-visually-identify-phishing-w	Phishpedia hybrid deep learning model (custom)
  USENIX/2021/phishpedia-a-hybrid-deep-learning-based-approach-to-visually-identify-phishing-w	PhishCatcher
  USENIX/2021/phishprint-evading-phishing-detection-crawlers-by-prior-profiling	custom phishing-site deployment
  USENIX/2021/phishprint-evading-phishing-detection-crawlers-by-prior-profiling	PhishPrint
  CCS/2022/im-spartacus-no-im-spartacus-proactively-protecting-users-from-phishing-by-inten	Cisco phishing kits
  IEEE-SP/2022/phishing-in-organizations-findings-from-a-large-scale-and-long-term-study	commercial anti-phishing appliance
  IEEE-SP/2022/phishing-in-organizations-findings-from-a-large-scale-and-long-term-study	commercial anti-phishing appliance
  IMC/2022/phishinpatterns-measuring-elicited-user-interactions-at-scale-on-phishing-websit	VisualPhishNet
  IMC/2022/phishinpatterns-measuring-elicited-user-interactions-at-scale-on-phishing-websit	custom intelligent phishing website crawler
  IMC/2022/phishinpatterns-measuring-elicited-user-interactions-at-scale-on-phishing-websit	VisualPhishNet
  IMC/2022/phishinpatterns-measuring-elicited-user-interactions-at-scale-on-phishing-websit	PhishInPattern
  IMC/2022/phishweb-a-progressive-multi-layered-system-for-phishing-websites-detection	PhishStorm
  IMC/2022/phishweb-a-progressive-multi-layered-system-for-phishing-websites-detection	Phishing-Benign
  USENIX/2022/inferring-phishing-intention-via-webpage-appearance-and-dynamics-a-deep-vision-b	Phishpedia dataset
  USENIX/2022/inferring-phishing-intention-via-webpage-appearance-and-dynamics-a-deep-vision-b	Phishpedia dataset
  USENIX/2022/inferring-phishing-intention-via-webpage-appearance-and-dynamics-a-deep-vision-b	Phishpedia dataset
  USENIX/2022/inferring-phishing-intention-via-webpage-appearance-and-dynamics-a-deep-vision-b	PhishIntention deep-learning models
  USENIX/2022/inferring-phishing-intention-via-webpage-appearance-and-dynamics-a-deep-vision-b	PhishIntention
  USENIX/2022/phish-in-sheeps-clothing-exploring-the-authentication-pitfalls-of-browser-finger	Phish-A
  USENIX/2022/phish-in-sheeps-clothing-exploring-the-authentication-pitfalls-of-browser-finger	Phish-B
  IMC/2023/phishing-in-the-free-waters-a-study-of-phishing-attacks-created-using-free-websi	FreePhish Twitter and Facebook streams
  IMC/2023/phishing-in-the-free-waters-a-study-of-phishing-attacks-created-using-free-websi	FreePhish
  IMC/2023/phishing-in-the-free-waters-a-study-of-phishing-attacks-created-using-free-websi	VisualPhishNet
  IMC/2023/phishing-in-the-free-waters-a-study-of-phishing-attacks-created-using-free-websi	PhishIntention
  CCS/2023/txphishscope-towards-detecting-and-understanding-transaction-based-phishing-on-e	Chainabuse Ethereum Phishing Scam Reports
  CCS/2023/txphishscope-towards-detecting-and-understanding-transaction-based-phishing-on-e	TxPhishScope detection results
  CCS/2023/txphishscope-towards-detecting-and-understanding-transaction-based-phishing-on-e	11 large-scale TxPhish events and Web3 security-company reports
  USENIX/2023/knowledge-expansion-and-counterfactual-interaction-for-reference-based-phishing	DynaPhish
  USENIX/2023/knowledge-expansion-and-counterfactual-interaction-for-reference-based-phishing	Phishpedia
  USENIX/2023/knowledge-expansion-and-counterfactual-interaction-for-reference-based-phishing	PhishIntention
  USENIX/2023/rods-with-laser-beams-understanding-browser-fingerprinting-on-phishing-pages	random sample of phishing websites
  USENIX/2023/rods-with-laser-beams-understanding-browser-fingerprinting-on-phishing-pages	phishing websites belonging to two selected signatures
  USENIX/2024/guardians-of-the-galaxy-content-moderation-in-the-interplanetary-file-system	Web2 Anti-Phishing Services (APS)
  USENIX/2024/guardians-of-the-galaxy-content-moderation-in-the-interplanetary-file-system	combined badbits and phishing denylist
  USENIX/2024/knowphish-large-language-models-meet-multimodal-knowledge-graphs-for-enhancing-r	KnowPhish Detector (custom)
  USENIX/2024/knowphish-large-language-models-meet-multimodal-knowledge-graphs-for-enhancing-r	Phishpedia
  USENIX/2024/knowphish-large-language-models-meet-multimodal-knowledge-graphs-for-enhancing-r	PhishIntention
  USENIX/2024/knowphish-large-language-models-meet-multimodal-knowledge-graphs-for-enhancing-r	DynaPhish
  USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish	PhishIntention's logo dataset
  USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish	Phishpedia
  USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish	Phishpedia
  USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish	PhishIntention's protected brand list
  USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish	PhishIntention
  USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish	Phishpedia
  USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish	VisualPhishNet
  USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish	PhishIntention
  USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish	Phishpedia
  USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish	VisualPhishNet
  USENIX/2024/less-defined-knowledge-and-more-true-alarms-reference-based-phishing-detection-w	custom PhishLLM decision rules
  USENIX/2024/less-defined-knowledge-and-more-true-alarms-reference-based-phishing-detection-w	PhishLLM
  USENIX/2024/less-defined-knowledge-and-more-true-alarms-reference-based-phishing-detection-w	Phishpedia
  USENIX/2024/less-defined-knowledge-and-more-true-alarms-reference-based-phishing-detection-w	PhishIntention
  USENIX/2024/less-defined-knowledge-and-more-true-alarms-reference-based-phishing-detection-w	DynaPhish
  USENIX/2024/phishdecloaker-detecting-captcha-cloaked-phishing-websites-via-hybrid-vision-bas	PhishDecloaker
  WWW/2024/are-adversarial-phishing-webpages-a-threat-in-reality-understanding-the-users-pe	Chiew et al. phishing dataset
  WWW/2024/are-adversarial-phishing-webpages-a-threat-in-reality-understanding-the-users-pe	SaTML'23 adversarial phishing webpage dataset
  WWW/2024/phishing-vs-legit-comparative-analysis-of-client-side-resources-of-phishing-and	custom collected phishing kits
  WWW/2024/phishing-vs-legit-comparative-analysis-of-client-side-resources-of-phishing-and	GoPhish
  CCS/2025/systematic-assessment-of-tabular-data-synthesis	Adult, Shoppers, Phishing, Magic, Faults, Bean, Obesity, Robot, Abalone, News, Insurance, and Wine
  NDSS/2025/dissecting-payload-based-transaction-phishing-on-ethereum	public phishing complaints and security-community phishing blogs
  NDSS/2025/dissecting-payload-based-transaction-phishing-on-ethereum	ice-phishing victims identified from phishing transactions
  NDSS/2025/dissecting-payload-based-transaction-phishing-on-ethereum	Etherscan Fake_Phishing nametags
  USENIX/2025/unsafe-llm-based-search-quantitative-analysis-and-mitigation-of-safety-risks-in	PhishLLM
  USENIX/2025/unsafe-llm-based-search-quantitative-analysis-and-mitigation-of-safety-risks-in	PhishLLM
  WWW/2025/7-days-later-analyzing-phishing-site-lifespan-after-detected	Phishing Army
  WWW/2025/7-days-later-analyzing-phishing-site-lifespan-after-detected	Phishing Database
  NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns	Beyond Phish
  NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns	Beyond Phish
  NDSS/2026/loki-proactively-discovering-online-scams-by-mining-toxic-search-queries	BeyondPhish
  NDSS/2026/ctphishcapture-uncovering-credential-theft-based-phishing-scams-targeting-cryptocurrency-wallets	CtPhishCapture
  PETS/2026/linguistic-hooks-investigating-the-role-of-language-triggers-in-phishing-emails	NIST Phish Scale
  PETS/2026/linguistic-hooks-investigating-the-role-of-language-triggers-in-phishing-emails	NIST Phish Scale
  WWW/2026/netting-phish-in-the-ipfs-ocean-real-time-monitoring-and-characterization-of-dec	Phishpedia
  WWW/2026/netting-phish-in-the-ipfs-ocean-real-time-monitoring-and-characterization-of-dec	PhishIntention
  IMC/2025/unmasking-the-shadow-economy-a-deep-dive-into-drainer-as-a-service-phishing-on-e	TxPhishScope
  IMC/2025/unmasking-the-shadow-economy-a-deep-dive-into-drainer-as-a-service-phishing-on-e	large-scale phishing incidents
  IMC/2025/unmasking-the-shadow-economy-a-deep-dive-into-drainer-as-a-service-phishing-on-e	TxPhishScope
  USENIX/2025/evaluating-the-effectiveness-and-robustness-of-visual-similarity-based-phishing	PhishIntention
  USENIX/2025/evaluating-the-effectiveness-and-robustness-of-visual-similarity-based-phishing	Phishpedia
  USENIX/2025/evaluating-the-effectiveness-and-robustness-of-visual-similarity-based-phishing	DynaPhish
  USENIX/2025/evaluating-the-effectiveness-and-robustness-of-visual-similarity-based-phishing	PhishZoo
  USENIX/2025/evaluating-the-effectiveness-and-robustness-of-visual-similarity-based-phishing	VisualPhishNet
  USENIX/2025/darkgram-a-large-scale-analysis-of-cybercriminal-activity-channels-on-telegram	PhishIntention
  USENIX/2025/darkgram-a-large-scale-analysis-of-cybercriminal-activity-channels-on-telegram	PhishIntention
  NDSS/2026/phishlang-a-real-time-fully-client-side-phishing-detection-framework-using-mobilebert	PhishPedia
  NDSS/2026/phishlang-a-real-time-fully-client-side-phishing-detection-framework-using-mobilebert	PhishLang
  NDSS/2026/phishlang-a-real-time-fully-client-side-phishing-detection-framework-using-mobilebert	GoPhish
  IEEE-SP/2021/crawlphish-large-scale-analysis-of-client-side-cloaking-techniques-in-phishing	Public Dataset from various phishing URL sources
  IEEE-SP/2021/crawlphish-large-scale-analysis-of-client-side-cloaking-techniques-in-phishing	CrawlPhish visual-similarity detector
  IEEE-SP/2021/crawlphish-large-scale-analysis-of-client-side-cloaking-techniques-in-phishing	CrawlPhish cloaking technique database
  IEEE-SP/2021/crawlphish-large-scale-analysis-of-client-side-cloaking-techniques-in-phishing	CrawlPhish
  IEEE-SP/2024/practical-attacks-against-dns-reputation-systems	StopForumSpam, Prigent DBL, and Phishing Army
  IEEE-SP/2024/practical-attacks-against-dns-reputation-systems	StopForumSpam, Prigent DBL, and Phishing Army
PUBLISHED_PHISHTANK_UNION 30
PUBLISHED_PHISHTANK_SOURCELIST_RESOURCENAME 28  # any sourceList/resourceName matching /phishtank/i
NAMESPACE_SITTING_PHISHTANK 21  # report_security_namespace.mjs topNames fold of /phish/ sourceList+resourceName; not this sitting's 28 or 30
 
## E. Full-text probes on PAGE_N (upper bounds, then read)
 
Each probe is paper-counted over PAGE_N. Hits are an upper bound until read.
PROBE cloak            hits= 34 / 52  missing_cols=0  65.4%
PROBE takedown         hits= 31 / 52  missing_cols=0  59.6%
PROBE lifespan         hits= 26 / 52  missing_cols=0  50.0%
PROBE phishtank-word   hits= 32 / 52  missing_cols=0  61.5%
PROBE openphish-word   hits= 20 / 52  missing_cols=0  38.5%
PROBE gsb-word         hits= 35 / 52  missing_cols=0  67.3%
PROBE captcha-cloak    hits= 21 / 52  missing_cols=0  40.4%
 
## F. Per-paper figures the page quotes (from detection.prevalence on PAGE_N)
 
 
### IEEE-SP/2021/crawlphish-large-scale-analysis-of-client-side-cloaking-techniques-in-phishing  role=cloaking
  phen="client-side JavaScript cloaking"
  metric="share of phishing websites"
  prevalence="35,067 of 112,005 websites (31.31%)"
  quote="Within our dataset of 112,005 phishing websites, CrawlPhish found that 35,067 (31.31%) phishing websites implement client-side cloaking techniques in total"
  phen="client-side cloaking prevalence over time"
  metric="share of phishing websites"
  prevalence="23.32% in 2018 and 33.70% in 2019"
  quote="23.32% (6,024) in 2018 and 33.70% (29,043) in 2019."
  phen="cloaking detection accuracy"
  metric="false-positive and false-negative rates"
  prevalence="1.45% false positives and 1.75% false negatives"
  quote="CrawlPhish correctly detected 1,965 phishing websites as cloaked and 1,971 as uncloaked, with a false-negative rate of 1.75% (35) and a false-positive rate of 1.45% (29)."
  phen="cloaking semantic categorization"
  metric="categorization accuracy"
  prevalence="100% accuracy on 2,000 cloaked phishing websites"
  quote="We found that CrawlPhish correctly categorized the cloaking type with 100% accuracy."
  phen="phishing ecosystem detection bypass"
  metric="sites blacklisted over seven days"
  prevalence="None blacklisted except 21 Click Through sites in Microsoft Edge"
  quote="At the conclusion of these experiments, we found that none of our phishing websites were blacklisted in any browser, with the exception of Click Through websites, 21 (42%) of which were blocked in Microsoft Edge"
  phen="human exposure to cloaked content"
  metric="share seeing hidden text"
  prevalence="100% Mouse, 97.72% Click Through, and 42.55% Notification Window"
  quote="For the Mouse Movement cloaking technique, 100% of the workers saw the \"Hello World\" text... For the Click Through websites, 97.72% saw the text"
 
### IMC/2019/opening-the-blackbox-of-virustotal-analyzing-online-phishing-scan-engines  role=live-site
  phen="phishing-site detection"
  metric="detected sites per vendor"
  prevalence="Only 15 of 68 vendors detected at least one of 36 simple phishing sites"
  quote="Over multiple scans, only 15 vendors (out of 68) have detected at least one of the 36 simple phishing sites."
  phen="phishing-takedown reaction"
  metric="label changes after takedown"
  prevalence="Only four vendors flipped some malicious labels after week three"
  quote="we find 4 vendors that flip some \"malicious\" labels to \"benign\" after the third week (for PayPal sites only)."
  phen="obfuscation evasion"
  metric="average malicious labels per site"
  prevalence="Image and code obfuscation reduced averages from 12.1 to 4.5 and 2"
  quote="the average number of malicious labels drops from 12.1 to 4.5 and 2 respectively."
 
### CCS/2016/phisheye-live-monitoring-of-sandboxed-phishing-kits  role=kit
  phen="Phishing-kit installation and deployment"
  metric="unique phishing kits"
  prevalence="643 unique kits; 474 correctly installed by 471 attackers"
  quote="We collected 643 unique phishing kits, which were uploaded on our honeypot over a period of five months from September 2015 to the end of January 2016. Out of this initial dataset, 474 kits"
  phen="Victim connections"
  metric="unique victims"
  prevalence="2,468 victims connected to 127 phishing kits"
  quote="After our aggressive filtering, we counted a total number of 2,468 victims who have connected to 127 distinct phishing kits."
  phen="Phishing-kit effective lifetime"
  metric="estimated lifetime"
  prevalence="eight days on average"
  quote="The last victims (at the 95% threshold) connect in average to the phishing page after 10 days. This gives us an estimated lifetime of eight days."
  phen="Blacklist detection latency"
  metric="detection latency"
  prevalence="98% blacklisted; average latency of 12 days"
  quote="almost all phishing pages that were installed on the honeypot (98% of them) were correctly detected and blacklisted by the two phishing blacklists, we observed an average detection latency of 12 days."
  phen="Blacklist timing relative to victims"
  metric="share blacklisted after 75% of victims"
  prevalence="62% of kits were blacklisted only after 75% of victims connected"
  quote="62% of the phishing kits were blacklisted only after 75% of victims had already connected to the corresponding phishing page."
  phen="Blacklist-evasion visitor crowd"
  metric="visitor distribution over elapsed time"
  prevalence="Hundreds of connections followed publication of a random URL by PhishTank"
  quote="After the random link was published by PhishTank we received many connections from researchers and security companies, which connected to the reported link instead of the entry page."
 
### IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni  role=cloaking
  phen="phishing blacklist occurrence"
  metric="proportion of sites blacklisted"
  prevalence="In the full tests, 23.0% of crawled cloaked sites versus 49.4% of non-cloaked sites were blacklisted."
  quote="in the full tests, just 23.0% of our sites with cloaking (which were crawled) ended up being blacklisted in at least one browser- far fewer than the 49.4% sites without cloaking"
  phen="blacklisting timeliness"
  metric="mean time before first blacklisting"
  prevalence="Cloaking increased average time to blacklisting from 126 to 238 minutes in full tests."
  quote="Cloaking also slowed the average time to blacklisting from 126 minutes (for sites without cloaking) to 238 minutes."
  phen="cloaking effectiveness"
  metric="reduction in blacklisting likelihood"
  prevalence="Geolocation, device-type, and JavaScript cloaking reduced blacklisting likelihood by over 55% on average."
  quote="simple cloaking techniques representative of real-world attacks- including those based on geolocation, device type, or JavaScript-were effective in reducing the likelihood of blacklisting by over 55% on average."
  phen="mobile blacklist protection"
  metric="presence of blacklist warnings"
  prevalence="Mobile Chrome, Safari, and Firefox showed no blacklist warnings during the tested period."
  quote="mobile Chrome, Safari, and Firefox failed to show any blacklist warnings between mid-2017 and late 2018"
 
### USENIX/2021/phishprint-evading-phishing-detection-crawlers-by-prior-profiling  role=cloaking
  phen="Phishing-site evasion"
  metric="site lifetime"
  prevalence="18 of 20 cloaked sites remained functional for one month"
  quote="In the one-month period in which we did the monitoring, only 2 sites (say, 'A' and 'B') got blocked."
  phen="Cloaking specificity among users"
  metric="share shown phishing content"
  prevalence="PhishPrint-powered logic showed phishing content to 79% of users"
  quote="Overall, the results showed that PhishPrint-powered cloaking logic decided to show phishing content for 79% of the users."
 
### CCS/2022/im-spartacus-no-im-spartacus-proactively-protecting-users-from-phishing-by-inten  role=cloaking
  phen="fingerprinting-based cloaking"
  metric="share of phishing kits"
  prevalence="96.52% (2,831) of 2,933 phishing kits"
  quote="In total, 96.52% (2,831) out of 2,933 phishing kits contain fingerprinting-based cloaking techniques."
  phen="phishing-content evasion"
  metric="share of phishing URLs protected"
  prevalence="132,247 of 160,728 (82.28%)"
  quote="132,247 (82.28%) did not contain malicious content in Spartacus."
  phen="trigger-word cloaking"
  metric="share of cloaked sites evaded"
  prevalence="The word “bot” evaded 99.31% of cloaked phishing websites"
  quote="99.31% of the cloaked phishing web sites can be evaded by appending bot in the User-Agent."
  phen="IP-based cloaking"
  metric="share of phishing websites evaded"
  prevalence="7,540 of 8,474 (88.98%)"
  quote="Among the visited phishing web sites, Spartacus evaded 88.98% (7,540) of them through the proxy server."
  phen="anti-phishing detection latency"
  metric="median detection time"
  prevalence="154 minutes for Spartacus-evaded phishing sites; 22 minutes for non-evaded sites"
  quote="The median detection time is 154 minutes. However, within 24 hours, they only detect/blacklist 76.16% of the cloaked phishing web sites."
 
### USENIX/2024/phishdecloaker-detecting-captcha-cloaked-phishing-websites-via-hybrid-vision-bas  role=cloaking
  phen="CAPTCHA-cloaking evasion of phishing detectors"
  metric="URLs blacklisted / URLs submitted"
  prevalence="0/100 for every CAPTCHA-cloaking type across VirusTotal, Google Safe Browsing, and SmartScreen"
  quote="All baseline sites are blacklisted within 24 hours, whereas all CAPTCHA-cloaked sites remain undetected for 7 days and counting."
  phen="CAPTCHA-cloaked phishing in the wild"
  metric="share of captured phishing websites"
  prevalence="7.6% (66 of 869 phishing websites)"
  quote="Of these, 7.6% were CAPTCHA-cloaked phishing websites, all of which were discovered solely by PhishDecloaker."
  phen="Phishing detection recovery after decloaking"
  metric="phishing detection rate"
  prevalence="recovered to 78%, 40%, 95%, and 84% of baseline performance for reCAPTCHA, hCaptcha, slider, and rotation"
  quote="PhishDecloaker can recover their detection rates back to a percentage of their original performance: 78% on re-CAPTCHA, 40% on hCaptcha, 95% on slider CAPTCHA, and 84% on rotation CAPTCHA."
  phen="CAPTCHA-cloaked phishing lifespan and blacklist delay"
  metric="median hours"
  prevalence="16 hours to blacklist for CAPTCHA-cloaked sites versus 11 hours for ordinary sites"
  quote="CAPTCHA-cloaked phishing sites take a median time of 16 hours to be blacklisted, which is 45.5% longer than ordinary phishing sites (11 hours)."
 
### WWW/2025/7-days-later-analyzing-phishing-site-lifespan-after-detected  role=live-site
  phen="phishing-site lifespan"
  metric="mean and median lifespan"
  prevalence="54.04 hours mean and 5.46 hours median across 286,237 URLs"
  quote="The average lifespan of phishing websites in our dataset is 54.04 hours, but the median is just 5.46 hours"
  phen="phishing takedown causes"
  metric="share of takedown errors"
  prevalence="DNS resolution failure accounted for 67.23% of errors"
  quote="DNS resolution failures are the most common takedown cause (67.23%, 60,913 domains)"
  phen="visual component changes"
  metric="lifespan by visual-change frequency"
  prevalence="Sites with over 100 changes had a 404.52-hour median lifespan"
  quote="Screenshots with a similarity score above 0.95 are clustered together."
  phen="DOM and resource evolution"
  metric="relative feature change"
  prevalence="30,069 of 286,237 sites showed final modifications"
  quote="The analysis covers changes in website resources between discovery and takedown for 286,237 phishing sites, with 10.5% (30,069) showing final modifications."
  phen="DNS configuration changes"
  metric="share of sites changing DNS settings"
  prevalence="75.84% altered DNS settings between discovery and takedown"
  quote="Our analysis reveals that 75.84% of phishing websites alter their DNS settings between discovery and takedown."
  phen="CDN usage"
  metric="share of phishing-related IPs"
  prevalence="99.80% of phishing-related IPs used CDNs"
  quote="99.80% of phishing-related IPs use CDNs, rendering IP-based blocking largely ineffective."
  phen="phishing server transitions"
  metric="number of websites changing server versions"
  prevalence="13 Nginx websites and four Apache websites changed versions"
  quote="Our findings show that 13 websites using Nginx change versions over time"
 
### USENIX/2020/sunrise-to-sunset-analyzing-the-end-to-end-life-cycle-and-effectiveness-of-phish  role=live-site
  phen="phishing attack life-cycle timing"
  metric="average campaign duration and stage delays"
  prevalence="average campaign from first to last victim takes 21 hours"
  quote="We find the average campaign from start to the last victim takes just 21 hours."
  phen="phishing victim traffic"
  metric="victim count"
  prevalence="4.8 million victims over one year, excluding crawler traffic"
  quote="Over a one year period, our network monitor recorded 4.8 million victims who visited phishing pages, excluding crawler traffic."
  phen="credential compromise and fraud"
  metric="share of distinct Known Visitors with fraudulent transactions"
  prevalence="7.42% of distinct Known Visitors subsequently suffered a fraudulent transaction"
  quote="In our dataset, the accounts of 7.42% of distinct Known Visitors subsequently suffered a fraudulent transaction"
  phen="browser phishing-warning effectiveness"
  metric="relative compromised-visitor ratio"
  prevalence="ratio fell to 71.51% after one hour and 43.55% after two hours"
  quote="within one hour after detection, at which point the ratio of Compromised Visitors drops to 71.51%. By the end of the second hour, the ratio drops further to 43.55%"
  phen="phishing-data visibility"
  metric="visibility ratio"
  prevalence="39.1% of hostnames and 40.9% of URLs"
  quote="We found that our approach had visibility into an average of 39.1% of all hostnames and 40.9% of all URLs"
 
### IMC/2022/phishinpatterns-measuring-elicited-user-interactions-at-scale-on-phishing-websit  role=live-site
  phen="phishing campaign clustering"
  metric="campaign count and brand count"
  prevalence="8,472 campaigns targeting over 381 brands"
  quote="we discovered 8,472 distinct phishing campaigns targeting over 381 brands."
 
### USENIX/2021/catching-phishers-by-their-bait-investigating-the-dutch-phishing-landscape-throu  role=live-site
  phen="potential phishing domains"
  metric="number of domains"
  prevalence="7,936 potential domains"
  quote="our domain detector labeled 7,936 domains as potentially malicious, which meant that these domains reached the threshold value"
  phen="phishing kit deployment"
  metric="number of verified phishing FQDNs"
  prevalence="1,363 verified phishing FQDNs"
  quote="Our final dataset contained 1,363 verified phishing fully qualified domain names (FQDN) which have been online for at least one hour."
  phen="phishing kit family prevalence"
  metric="share of identified phishing websites"
  prevalence="Almost 89% used a kit within the uAdmin family."
  quote="Almost 89% of all identified phishing websites were made with a kit within this family"
  phen="phishing domain uptime"
  metric="median uptime"
  prevalence="24 hours"
  quote="On average, a phishing domain in our dataset is online for 45 hours, but we find a median uptime of 24 hours."
  phen="domain cloaking"
  metric="share of detected domains"
  prevalence="946 domains (69%) returned a blank screen and no favicon."
  quote="In fact, 946 (69%) of the detected phishing domains returned a blank screen - and no favicon - to our crawler"
  phen="external resource loading"
  metric="share of dataset"
  prevalence="104 domains (7.6%) loaded resources from benign counterparts."
  quote="only 104 domains (7.6% of the total dataset) load their resources directly from their benign counterparts."
 
### IMC/2018/needle-in-a-haystack-tracking-down-elite-phishing-domains-in-the-wild  role=live-site
  phen="squatting phishing pages"
  metric="share of squatting domains"
  prevalence="1,175 confirmed domains, approximately 0.2% of 657,663 domains."
  quote="As shown in Table 8, after manual examination, we confirmed 1,175 domains are indeed phishing domains."
  phen="layout obfuscation"
  metric="image-hash distance"
  prevalence="Most brands had average distance around 20 or higher."
  quote="most brands have an average distance around 20 or higher, suggesting that layout obfuscation is very common."
  phen="blacklist evasion"
  metric="share undetected for at least a month"
  prevalence="91.5% remained undetected by the tested blacklists."
  quote="Collectively these blacklists only detected 8.4% of the squatting phishing pages, which means 91.5% of the phishing domains remain undetected for at least a month."
  phen="phishing-page lifetime"
  metric="share remaining alive"
  prevalence="About 80% remained alive after at least a month."
  quote="Most pages (about 80%) still remain alive after at least a month."
  phen="IP geolocation"
  metric="number of IP addresses and countries"
  prevalence="1,021 IP addresses across 53 countries."
  quote="In total, we are able to look up the geolocation of 1,021 IP addresses, hosted in 53 different countries."
 
### WWW/2017/tracking-phishing-attacks-over-time  role=live-site
  phen="replicated phishing attacks"
  metric="share of phishing URLs in flagged clusters"
  prevalence="17,451 of 19,066 URLs (91.53%)"
  quote="17,451 (resp. 11,244) of the 19,066 (resp. 12,859) phishing URLs end up in a flagged cluster, which means that 91.53% (resp. 87.44%)"
  phen="false positives on legitimate sites"
  metric="false positive rate"
  prevalence="20 of 24,800 legitimate sites (0.08%)"
  quote="Only 20 of these sites end up in one of the 2,831 phishing clusters, a false positive rate of only 0.08%."
  phen="phishing attack-class lifespan"
  metric="cluster lifespan"
  prevalence="Average lifespan was 25 days; around 80% lasted less than one month"
  quote="The lifespan is defined as the duration between the first and the last observed attack instance. ... the average lifespan of a cluster in our database is 25 days."
 
### IMC/2023/phishing-in-the-free-waters-a-study-of-phishing-attacks-created-using-free-websi  role=live-site
  phen="FWB-hosted phishing attacks"
  metric="number of zero-day phishing URLs"
  prevalence="31,405 URLs from November 2022 to May 2023"
  quote="We ran FreePhish for a period of six months, from November 2022 to May 2023, identifying 31,405 zero-day phishing attacks"
  phen="Historical FWB phishing prevalence"
  metric="share of phishing URLs using FWBs"
  prevalence="25.2K of 34.7K unique phishing URLs used 17 FWBs"
  quote="25.2K URLs ... utilized 17 unique free website-building services."
  phen="Phishing classifier performance"
  metric="F1-score and median runtime"
  prevalence="F1-score 0.96 and median runtime 2.8 seconds"
  quote="our augmented StackModel has an F1-score of 0.96 and a median run-time of only 2.8 secs."
  phen="Anti-phishing blocklist coverage"
  metric="coverage and response time"
  prevalence="GSB covered 18.4% of FWB phishing URLs; median response time 6 hours"
  quote="GSB covered 18.4% of all FWB phishing URLs while having a median response time of 6 hours"
  phen="FWB takedown response"
  metric="removal rate and response time"
  prevalence="29% removed after two weeks; median speed 9 hours 43 minutes"
  quote="only 29% of the websites were removed by the respective FWB service after two weeks since they first appeared in our dataset"
  phen="Evasive phishing variants"
  metric="counts and shares of URLs"
  prevalence="539 Google Sites URLs linked to other phishing pages; 427 Google Sites and 473 Blogspot URLs embedded phishing iframes"
  quote="We randomly sampled 1K URLs from this set and qualitatively analyzed their content ... We developed heuristics to automatically identify these attack vectors"
 
### WWW/2024/phishing-vs-legit-comparative-analysis-of-client-side-resources-of-phishing-and  role=live-site
  phen="phishing sites without JavaScript"
  metric="share of collected phishing webpages without JavaScript"
  prevalence="22.8% (172,348 of 757,421)"
  quote="There are 22.8% (172,348 out of 757,421) of our collected phishing webpages that do not use JavaScript."
  phen="JavaScript library diversity"
  metric="number of distinct libraries"
  prevalence="132 phishing libraries versus 41 legitimate-site libraries"
  quote="A total of 132 distinct JavaScript libraries are identified in our phishing dataset, in contrast to the 41 distinct JavaScript libraries found in their corresponding legitimate target brand websites."
  phen="outdated JavaScript libraries"
  metric="average version-age difference"
  prevalence="646 days, nearly 21.2 months older on phishing websites"
  quote="On average, phishing websites employ JavaScript libraries that are 646 days older, equivalent to nearly 21.2 months, than the versions utilized by legitimate websites."
  phen="code copying"
  metric="share of clusters with exact or high overlap"
  prevalence="2.5% exact overlap; over 21.5% exceeded 85% overlap"
  quote="Our analysis of the similarities between different clusters revealed that there is an exact overlap of 2.5% in terms of code copying. Additionally, more than 21.5% of the clusters show an overlap exceeding 85%."
 
### USENIX/2021/phishpedia-a-hybrid-deep-learning-based-approach-to-visually-identify-phishing-w  role=detector
  phen="Phishing webpage identification"
  metric="identification rate, detection rate, precision, recall"
  prevalence="99.2% identification rate, 98.2% precision, and 87.1% recall on the reported evaluation"
  quote="Phishpedia 99.2% 98.2% 87.1% 0.19"
  phen="Logo presence in phishing webpages"
  metric="share of phishing webpages with logos"
  prevalence="about 98.6%"
  quote="We randomly sampled 5,000 webpages from the phishing webpage dataset, and manually validated that 70 of them have no logos. That is, the ratio of phishing webpages with logos is about 98.6%."
  phen="Phishing target-brand coverage"
  metric="coverage of phishing webpages"
  prevalence="top 100 brands cover 95.8% of phishing webpages"
  quote="our empirical study on around 30K phishing webpages based on an OpenPhish feed (Section 5.1) shows that the top 100 brands cover 95.8% phishing webpages"
  phen="Phishing discovery in the wild"
  metric="number of real and zero-day phishing webpages"
  prevalence="1,704 real phishing webpages within 30 days; 1,133 not reported by VirusTotal"
  quote="Phishpedia discovered 1,704 phishing webpages within 30 days and 1,133 of them are not detected by any engines in VirusTotal [9]."
 
## G. Quote check of ROLE deciding sentences against paper.cols.txt
 
QUOTE_TITLE_ONLY IEEE-SP/2024/from-chatbots-to-phishbots-phishing-scam-generation-in-commercial-large-language  (stored PDF/cols are a different paper; see provenance)
ROLE quotes vs paper.cols.txt: exact 72, partial 0, absent 66, missing-file 0, title-only 1  of 139
PUBLISHED_QUOTE_EXACT 72
PUBLISHED_QUOTE_PARTIAL 0
PUBLISHED_QUOTE_ABSENT 66
ABSENT quotes (must be spliced or PDF-only; listed, not silently dropped):
  USENIX/2010/searching-the-searchers-with-searchaudit
  NDSS/2013/compa-detecting-compromised-accounts-on-social-networks
  USENIX/2014/the-emperor-s-new-password-manager-security-analysis-of-web-based-password-manag
  IMC/2014/handcrafted-fraud-and-extortion-manual-account-hijacking-in-the-wild
  NDSS/2015/nophish-app-evaluation-lab-and-retention-study
  USENIX/2015/towards-discovering-and-understanding-task-hijacking-in-android
  WWW/2017/tracking-phishing-attacks-over-time
  IMC/2018/characterizing-the-internet-host-population-using-deep-learning-a-universal-and
  USENIX/2018/end-to-end-measurements-of-email-spoofing-attacks
  WWW/2018/betrayed-by-your-dashboard-discovering-malicious-campaigns-via-web-analytics
  NDSS/2019/digital-healthcare-associated-infection-a-case-study-on-the-security-of-a-major-multi-campus-hospital-system
  USENIX/2019/cognitive-triaging-of-phishing-attacks
  USENIX/2019/detecting-and-characterizing-lateral-phishing-at-scale
  USENIX/2019/the-webs-identity-crisis-understanding-the-effectiveness-of-website-identity-ind
  USENIX/2019/users-really-do-answer-telephone-scams
  WWW/2019/doppelgangers-on-the-dark-web-a-large-scale-assessment-on-phishing-hidden-web-se
  WWW/2019/hack-for-hire-exploring-the-emerging-market-for-account-hijacking
  NDSS/2020/deceptive-previews-a-study-of-the-link-preview-trustworthiness-in-social-platforms
  USENIX/2020/phishtime-continuous-longitudinal-measurement-of-the-effectiveness-of-anti-phish
  USENIX/2020/sunrise-to-sunset-analyzing-the-end-to-end-life-cycle-and-effectiveness-of-phish
  WWW/2020/dirty-clicks-a-study-of-the-usability-and-security-implications-of-click-related
  IMC/2021/knock-and-talk-investigating-local-network-communications-on-websites
  USENIX/2021/catching-phishers-by-their-bait-investigating-the-dutch-phishing-landscape-throu
  USENIX/2021/phishprint-evading-phishing-detection-crawlers-by-prior-profiling
  WWW/2021/where-are-you-taking-me-understanding-abusive-traffic-distribution-systems
  IEEE-SP/2022/phishing-in-organizations-findings-from-a-large-scale-and-long-term-study
  IMC/2022/phishinpatterns-measuring-elicited-user-interactions-at-scale-on-phishing-websit
  USENIX/2022/identity-confusion-in-webview-based-mobile-app-in-app-ecosystems
  USENIX/2022/inferring-phishing-intention-via-webpage-appearance-and-dynamics-a-deep-vision-b
  USENIX/2022/helping-hands-measuring-the-impact-of-a-large-threat-intelligence-sharing-commun
  IMC/2023/phishing-in-the-free-waters-a-study-of-phishing-attacks-created-using-free-websi
  CCS/2023/understanding-and-detecting-abused-image-hosting-modules-as-malicious-services
  USENIX/2023/knowledge-expansion-and-counterfactual-interaction-for-reference-based-phishing
  USENIX/2023/rods-with-laser-beams-understanding-browser-fingerprinting-on-phishing-pages
  WWW/2023/cashing-in-on-contacts-characterizing-the-onlyfans-ecosystem
  CCS/2024/employees-attitudes-towards-phishing-simulations-its-like-when-a-child-reaches-o
  USENIX/2024/assessing-suspicious-emails-with-banner-warnings-among-blind-and-low-vision-user
  USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish
  USENIX/2024/less-defined-knowledge-and-more-true-alarms-reference-based-phishing-detection-w
  USENIX/2024/malla-demystifying-real-world-large-language-model-integrated-malicious-services
  USENIX/2024/phishdecloaker-detecting-captcha-cloaked-phishing-websites-via-hybrid-vision-bas
  WWW/2024/phishinwebview-analysis-of-anti-phishing-entities-in-mobile-apps-with-webview-ta
  WWW/2024/are-adversarial-phishing-webpages-a-threat-in-reality-understanding-the-users-pe
  WWW/2024/zipzap-efficient-training-of-language-models-for-large-scale-fraud-detection-on
  WWW/2024/phishing-vs-legit-comparative-analysis-of-client-side-resources-of-phishing-and
  IEEE-SP/2016/sending-out-an-sms-characterizing-the-security-of-the-sms-ecosystem-with-public
  CCS/2025/quantifying-security-training-in-organizations-through-the-analysis-of-u-s-sec-1
  IEEE-SP/2025/restricting-the-link-effects-of-focused-attention-and-time-delay-on-phishing-war
  USENIX/2025/doubly-dangerous-evading-phishing-reporting-systems-by-leveraging-email-tracking
  USENIX/2025/scanned-and-scammed-insecurity-by-obsqrity-measuring-user-susceptibility-and-awa
  USENIX/2025/blockchain-address-poisoning
  NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns
  WWW/2025/whats-in-phishers-a-longitudinal-study-of-security-configurations-in-phishing-we
  USENIX/2025/url-inspection-tasks-helping-users-detect-phishing-links-in-emails
  NDSS/2026/ctphishcapture-uncovering-credential-theft-based-phishing-scams-targeting-cryptocurrency-wallets
  PETS/2026/toward-adaptive-privacy-enhancing-training-a-longitudinal-study-of-how-personali
  USENIX/2026/sok-philter-uncovering-security-and-functional-gaps-in-ai-based-phishing-website
  USENIX/2026/a-large-scale-study-of-personalized-phishing-using-large-language-models
  WWW/2026/netting-phish-in-the-ipfs-ocean-real-time-monitoring-and-characterization-of-dec
  CCS/2025/phishing-susceptibility-and-the-in-effectiveness-of-common-anti-phishing-interve
  IEEE-SP/2025/mantis-detection-of-zero-day-malicious-domains-leveraging-low-reputed-hosting-in
  NDSS/2025/misdirection-of-trust-demystifying-the-abuse-of-dedicated-url-shortening-service
  NDSS/2026/phishlang-a-real-time-fully-client-side-phishing-detection-framework-using-mobilebert
  IEEE-SP/2024/practical-attacks-against-dns-reputation-systems
  IEEE-SP/2024/understanding-the-privacy-practices-of-political-campaigns-a-perspective-from-th
  IEEE-SP/2022/siraj-a-unified-framework-for-aggregation-of-malicious-entity-detectors
 
## H. Sensitivity: drop posters from PAGE_N
 
PAGE_N including posters: 52
PAGE_N excluding posters: 51
Posters in PAGE_N: 1  (CCS/2014/poster-proactive-blacklist-update-for-anti-phishing)
 
## Z. EXTERNAL FIGURES (not from the extraction; primary source in phishing_probe.sh)
 
These are dated facts about the instruments, not corpus counts.
OPENPHISH_COMMUNITY_FEED_URLS 300   # GET https://openphish.com/feed.txt 2026-08-27
OPENPHISH_COMMUNITY_HOSTS 269
OPENPHISH_COMMUNITY_HTTPS 201
OPENPHISH_COMMUNITY_HTTP 99
OPENPHISH_COMMUNITY_BYTES 15020
OPENPHISH_COMMUNITY_REFRESH 12 hours  # openphish.com/phishing_feeds.html
PHISHTANK_DUMP_N 73660   # GET http://data.phishtank.com/data/online-valid.csv.gz unauthenticated 2026-08-27
PHISHTANK_DUMP_ONLINE 73660
PHISHTANK_DUMP_VERIFIED 73660
PHISHTANK_DUMP_TARGET_OTHER 65457
PHISHTANK_DUMP_BYTES_GZIP 2549032
PHISHTANK_DUMP_BYTES_UNCOMPRESSED 14105733
PHISHTANK_DUMP_MIN_SUBMISSION 2011-02-18
PHISHTANK_DUMP_MAX_SUBMISSION 2026-08-27
GSB_V4_THREATLISTS_NOKEY HTTP 403
GSB_V5_HASHLISTS_NOKEY HTTP 403
GSB_V4_SUNSET 2027-03-31  # Google notification mail; Brave issue 56023, not the docs page
GSB_V4_SUNSET_BRAVE_ISSUE 56023
GSB_MIGRATION_UNDER_10MIN 60  # "as of 2021, 60% of sites that deliver attacks live less than 10 minutes"
GSB_MIGRATION_STALENESS 25-30  # "around 25–30% of missing phishing protection is due to such data staleness"
GSB_MIGRATION_10_MINUTES 10
OPENPHISH_ACADEMIC_ACCESS_DAYS 60
OPENPHISH_ACADEMIC_ARCHIVE_DAYS 30
OPENPHISH_PREMIUM_REFRESH_MINUTES 5  # phishing_feeds.html Premium row
PHISHTANK_STATS_TIMEOUT 20s  # stats.php did not answer
APWG_ECX HTTP 200  # https://apwg.org/ecx
CORPUS_DEFECT_DOI 10.1109/SP54263.2024.00182  # Nanayakkara DP paper stored under a phishing slug
CORPUS_DEFECT_DOI_PARTS 54263 00182
 
## Z2. PAPER LITERALS on the page that are not in detection.prevalence
 
These were read from paper.cols.txt on 2026-08-27 and printed so check_page_numbers.mjs can see them.
PAPER_LITERAL crawlphish_artificial_sites 150
PAPER_LITERAL crawlphish_clickthrough_edge 21 of 50
PAPER_LITERAL phishpedia_feed_urls 350K  # paper writes "350K phishing URLs"; page says 350k
PAPER_LITERAL phishpedia_feed_urls_n 350000
PAPER_LITERAL phishinpatterns_feed_urls 56027
PAPER_LITERAL phishinpatterns_crawled 51859
PAPER_LITERAL phishinpatterns_multipage 23446 45% of 51859
PAPER_LITERAL lin2022_fingerprint_rates 73.98 90.08 91.36  # phishing sites with JS traces, three datasets
PAPER_LITERAL moura2024_domains 28754  # 28,754 phishing domains across .nl / .ie / .be
PAPER_LITERAL moura2024_nl_24h 80%  # After 24h, 80% of the .nl domains are mitigated
PAPER_LITERAL moura2024_ie_24h 70%  # roughly 70% of the .ie are mitigated
PAPER_LITERAL phisheye_lifetime_cut 5 95  # first victim = 5% of connections, last = 95%; eight-day estimate is 10 days minus 2
PAPER_LITERAL teoh_controlled_per_type 100  # Table 1: 0/100 on VT, GSB, SmartScreen for each of 5 CAPTCHA types
PAPER_LITERAL teoh_field_sole 66 of 869 discovered solely by PhishDecloaker
 
## I. Arithmetic derived on the page
 
PAGE_N / union = 52 / 139 = 37.4%
excluded / union = 87 / 139 = 62.6%
PhishTank papers / union = 21.6%
PAGE_N cloaking role = 8
PAGE_N live-site role = 29
PAGE_N detector role = 13
PAGE_N kit role = 2
user-study in union = 29
feed-gt in union = 19
PhishTank dump Other share = 65457/73660 = 88.9%
 
--- end ---
provenance/security/phishing.txt · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki