Table of Contents
Provenance: security:phishing
Working log behind phishing. Corpus-wide caveats are on corpus. Citations use the shared bibliography; this page adds no keys of its own.
Run: 2026-08-27. Corpus: 5,859 extracted papers, 7 venues, 2010–2026, data/extract/run1. Item: drain wiki-measuretheweb / “security:phishing (new)”, claimed by key as cursor-drain-phishing, store id 223, run 61. Author of the page and this log: Cursor, not Claude Code. Reviews: three focused passes then one generic, all on GPT 5.6 Luna medium (gpt-5.6-luna-medium), the sitting's requested substitute for the spec's sonnet/fable split.
Creating, not extending. ?do=export_raw on security:phishing returned the HTML error page (not an existence test by byte count). dw.mjs pages does not list provenance:. Neighbour security already linked the red child Phishing; live start already has the same link. Overlap judgement: write the promised child rather than widen virustotal or website_classification — those pages own the vendor panel and the topic label; this one owns feeds and live phishing websites.
No ~~DISCUSSION~~ on this provenance page. Comments belong on the content page.
Scope decisions
| Decision | Why | What a reasonable person might have done instead |
|---|---|---|
| PAGE_N = 52, not the 139-paper schema union and not the 1,098-paper keyword hit | The union mixes live-site crawls, detectors, feed-as-GT for a different question, user studies, email/SMS, on-chain “phishing”, and mobile UI. Merging them is a meaningless N. The keyword hit is a fact about these being security venues. | Publish 139 as “phishing papers”. Rejected. |
| Single primary ROLE per paper | A detector that also crawls still has one object of study. Double-counting would make the role table sum past 139. | Multi-label. Defensible for cloaking∩live-site; not this sitting. |
| GSB on this page, not a sixth security child | Already decided on security. GSB is a feed. | A GSB-API tutorial page. Rejected: MDN/Google docs exist; the student question is what a hit proves. |
| Do not visit feed URLs | The community feed is a list of live phish. Fetching the list is the instrument; fetching the sites is a crawl we are not running. | “Verify a sample is still up.” That is a different paper. |
| Cite GSB v4 sunset 2027-03-31 only with the Brave/email source | The date is in Google's notification mail (Brave issue 56023), not on the migration docs this sitting fetched. | Put 2027-03-31 in the lead as a Google-docs fact. That would be a lie about the primary source. |
| Keep the poster in PAGE_N, show 52 and 51 | CCS 2014 PhishTrack is a poster. Dropping it is a sensitivity check, not the headline. | Quietly drop posters. |
| Do not add uncited detector keys to the bibliography | The page cites the measurement papers. Phishpedia/KnowPhish/PhishLang appear as names in a sentence, not as [key]. Unused keys still break every page if they collide. | Dump every PAGE_N paper into the bib. |
Report script
scripts/report_phishing.mjs plus scripts/phish_fold.mjs (generated by scripts/_gen_phish_fold.py; do not hand-edit the ROLE Map). Deterministic. Re-run:
node scripts/report_phishing.mjs > scripts/report_phishing-output.txt
Exits 1 if the schema union and ROLE diverge in either direction, if a ROLE value is unknown, if a ROLE quote is shorter than 20 characters, if missing paper.cols.txt is not 4, or if PhishTank is missing from the feed ranking. A printed FAILURE with exit 0 is forbidden.
python3 pages/openphish_community_feed.py is the published script. Its output is quoted in a <code> block on the content page. It does not visit the listed sites. scripts/phishing_probe.sh re-fetches the feeds and the docs; it prints FAILED per broken check and exits 1.
Queries
| # | Query | Population / denominator | Result | On the page? |
|---|---|---|---|---|
| Q1 | Extraction records | all | 5,859 | outer frame |
| Q2 | Full-text /\bphish(?:ing|tank|ers)?\b/i | 5,855 with .cols (4 missing) | 1,098 | yes, labelled not-N |
| Q3 | Full-text Google Safe Browsing | 5,855 | 110 | provenance + namespace; content page uses the API facts |
| Q4 | Full-text PhishTank / OpenPhish / APWG | 5,855 | 219 | provenance |
| Q5 | Schema union: detection phenomenon/technique, classification resource/targetDetail, population.sourceList, or slug matching /phish/ | 5,859 | 139 (92 web, 64 crawled) | yes, candidate set |
| Q6 | Hand ROLE over Q5 | 139 | live-site 29, detector 13, cloaking 8, kit 2, feed-gt 19, user-study 29, email-sms 9, onchain 6, mobile-ui 6, mention 18 | yes |
| Q7 | PAGE_N = ROLE in {live-site, detector, cloaking, kit} | 139 | 52 (37.4%) | yes, the page's N |
| Q8 | Q7 web / crawled | 52 | 49 (94.2%) / 40 (76.9%) | yes |
| Q9 | Q7 posters | 52 | 1 (CCS 2014 PhishTrack); without it 51 | yes |
| Q10 | Q7 by venue | 52 | USENIX 17, WWW 11, IMC 9, CCS 8, NDSS 4, IEEE-SP 3, PETS 0 | yes |
| Q11 | Q7 by year-bucket | 52 | 1 / 7 / 13 / 18 / 13* | yes |
| Q12 | Feed family fold on sourceList + resourceName + tools.name | 139 | VT 32, PhishTank 30, GSB 27, APWG 13, OpenPhish 10, CertStream 9, … | yes |
| Q13 | Same fold on PAGE_N | 52 | VT 16, GSB 13, APWG 12, PhishTank 11, CertStream 9, OpenPhish 8 | yes |
| Q14 | sourceList/resourceName matching /phishtank/i | 139 | 28 | provenance (explains the namespace 21) |
| Q15 | Namespace sitting's topNames fold of /phish/ sourceList+resourceName | 139 | 21 | yes, as the number the namespace published |
| Q16 | Full-text probes on PAGE_N | 52 | cloak 34, takedown 31, lifespan 26, phishtank 32, openphish 20, gsb 35, captcha 21 | yes, labelled upper bounds |
| Q17 | ROLE quotes vs paper.cols.txt | 139 | exact 72, partial 0, absent 66, missing-file 0, title-only 1 | provenance |
Folding and residue
ROLE is a hand map, not a regex. 139 keys, generated from scripts/_gen_phish_fold.py. The report exits 1 if a union paper has no ROLE or a ROLE key is outside the union. Single-label: the ten rows sum to 139.
Feed families (FEED_FAMILIES in phish_fold.mjs): ordered regexes, most-specific first, over sourceList + resourceName + tools.name. A paper is counted once per family. Residue of phish-ish strings that matched no family: 118 strings, printed in full in the report below. They are custom kit names, “phishing emails”, GPPF, PhishFarm-the-experiment, combined vendor lists — not a missing commercial feed. Do not promote residue into a family without reading the string.
Namespace 21 vs this sitting's 30. The namespace page counted a topNames fold of sourceList/resourceName that already matched /phish/, then took the family whose spellings include “phishtank”: 21. Any sourceList/resourceName matching /phishtank/i (including “Google Safe Browsing API, Malware Patrol, PhishTank, …”) is 28. Adding tools[].name and the family fold is 30. The content page uses 30 and explains the 21. Do not mix them.
Quotes spot-checked
ROLE deciding sentences vs paper.cols.txt: 72 exact, 0 partial, 66 absent, 0 missing-file, 1 title-only. The checker compared the stored ROLE quote string against .cols. A focused reviewer spot-checked five “absent” keys (SearchAudit, Compa, Cui 2017, PhishTime, PhishDecloaker) and found the deciding sentence in .cols in each case; at least two were verbatim modulo a column break. That is not a census of the 66 and does not prove the other 61 are clean — it is enough to reject “the roles are fabricated”, and not enough to claim every absent quote is mere ellipsis. The 66 keys are listed in the report.
Title-only / corpus defect: IEEE S&P 2024 from-chatbots-to-phishbots-…. Slug and bibliographic title are Roy et al. on LLM phishing. Stored PDF and paper.cols.txt are Nanayakkara et al. on differential privacy, DOI 10.1109/SP54263.2024.00182 on the PDF. ROLE=mention, quote taken from the title. Do not cite the stored PDF for that slug. Not patched in the extraction; recorded here.
Load-bearing paper figures on the content page were checked against paper.cols.txt (and against detection.prevalence where the extractor captured them):
| Paper | Needle | In .cols? |
|---|---|---|
| Lee et al. WWW 2025 | 54.04 hours / 5.46 hours / 286,237 / 67.23% / 75.84% / 99.80% | yes |
| Zhang et al. IEEE S&P 2021 | 35,067 of 112,005 (31.31%); 23.32% / 33.70%; 150 artificial; 21 (42%) Click Through in Edge | yes |
| Peng et al. IMC 2019 | 15 of 68 / 36 simple phishing sites; 4 vendors week three PayPal | yes |
| Oest et al. IEEE S&P 2019 | 23.0% vs 49.4%; 126 → 238 minutes; mobile no warnings mid-2017–late 2018 | yes |
| Oest et al. USENIX 2020 | 21 hours; 4.8 million; 7.42%; 71.51% / 43.55% | yes |
| Zhang et al. CCS 2022 | 96.52% (2,831/2,933); 82.28% of 160,728; 99.31% “bot”; 88.98% (7,540/8,474) | yes |
| Teoh et al. USENIX 2024 | 0/100; 7.6% (66/869); 16h vs 11h | yes |
| Acharya et al. USENIX 2021 | 18 of 20 still up after one month (2 of 20 blocked) | yes |
| Subramani et al. IMC 2022 | 56,027 → 51,859; 23,446 (45%) required input | yes |
| Han et al. CCS 2016 | 643 kits; eight days; 98% / 12 days; 62% after 75% of victims | yes |
| Bijmans et al. USENIX 2021 | mean 45 hours, median 24 hours | yes |
| Tian et al. IMC 2018 | 1,175 confirmed; 91.5% undetected a month later | yes |
| Moura et al. CCS 2024 | 28,754 domains; after 24h 80% of .nl, roughly 70% of .ie | yes |
| Lin et al. USENIX 2022 | 73.98% / 90.08% / 91.36% fingerprint collection on phishing-with-JS | yes |
| Lin et al. USENIX 2021 | “this gave us 350K phishing URLs” | yes |
External sources
Fetched 2026-08-27, re-checkable with scripts/phishing_probe.sh:
- PhishTank home / API / developer / FAQ / register — all HTTP 200. Operator string “PhishTank is operated by Cisco Talos Intelligence Group”. Unauthenticated dump
http://data.phishtank.com/data/online-valid.csv.gz: 73,660 rows, all online=yes and verified=yes, target Other 65,457 (88.9%), gzip 2,549,032 B, submissions 2011-02-18 → 2026-08-27. The documented public dump path is on the content page because that is the file papers train on. Signed or expiring CDN query strings stay off the wiki.stats.phptimed out at 20s from this host — documented, not treated as a probe failure. - Wikipedia “registration closed” was not confirmed on the live FAQ; register.php returned 200. Not cited.
- OpenPhish
feed.txt: 300 URLs, 269 hosts, 201 https / 99 http, 15,020 bytes.phishing_feeds.htmlCommunity row: 12 hours, Limited, Text File, Free. Premium: 5 minutes. Academic programme: 60 days + 30-day archive, institutional email to support@openphish.com. - GSB v4 overview: deprecated, non-commercial, Web Risk for commercial. Unauthenticated
v4/threatListsandv5/hashLists: HTTP 403. Test page still says “Should show a phishing warning”. Migration guide: “as of 2021, 60% of sites that deliver attacks live less than 10 minutes”; “around 25–30% of missing phishing protection is due to such data staleness.” - GSB v4 end date 31 March 2027: Brave brave-browser#56023 quoting the notification mail. Not on the docs page this sitting fetched.
- APWG eCX: apwg.org/ecx HTTP 200. No public dump.
Rejected: SEO listicles of “best phishing feeds”; any instruction to crawl the URLs in feed.txt; treating 2027-03-31 as a Google-docs fact.
What could not be established
- A live PhishTank submissions-per-day series (
stats.phptimed out). - Whether PhishTank still accepts new user registrations — register.php is 200 and mostly JS; the FAQ no longer has the 2020 “closed” sentence this sitting could find. Page says Wikipedia was not confirmed, and stops.
- APWG eCX membership contents without being a member.
- GSB URL lists — the API does not return them. That is the finding.
- A complete quote-verbatim map of all 139 ROLE sentences (66 absent from
.colsbecause of extraction ellipsis). - PETS “does not do phishing websites” as a venue finding — N=2 in the union, both user-study. Named as a small-N observation.
Published script
pages/openphish_community_feed.py. Stdlib. Advertised invocation: python3 openphish_community_feed.py. Empty feed, non-http(s) rows, or a URL with no hostname: exit 1. HTTP != 200: exit 1. Direct schemes['https'] / schemes['http'] (Counter returns 0 for a missing scheme, which is the right count). The <file> block on the content page is byte-identical to this file (checked before freeze).
Bibliography keys
Collision-checked against a fresh ?do=export_raw of live literature:bibliography on 2026-08-27 (558 keys; local pages/literature_bibliography.txt is stale and was not used).
Already live, not re-added: zhang2021_crawlphish, peng2019_opening, zhang2022_spartacus, alam2026_philter, sanchezrola2023_rods.
Added (cited on the content page): lee2025_7days, moura2024_cctld, lin2022_sheep, oest2019_phishfarm, oest2020_phishtime, oest2020_sunrise, han2016_phisheye, teoh2024_phishdecloaker, subramani2022_phishinpatterns, bijmans2021_catching, tian2018_needle, lin2021_phishpedia, acharya2021_phishprint.
USENIX entries use url= not DOI (index has none). Affiliation tokens (PayPal, Trustwave, EURECOM, NCS Cyber Special Ops) were stripped from author fields after fetch_authors.py glued them on — known USENIX-author-list bug, same as virustotal.
Generated but not added because the content page does not [cite] them: roy2023_freewaters, lim2024_legit, cui2017_tracking, maroofi2020_human, choi2024_phishinwebview, lim2025_phishers, roy2026_phishlang, liu2023_knowledge, li2024_knowphish, liu2022_inferring.
Review log
Drafts frozen in out/freeze_phishing/ before the three focused passes. All four reviewers: GPT 5.6 Luna medium (gpt-5.6-luna-medium), substituting for the spec's sonnet/fable split. Pages were not edited while the three focused reviewers ran.
Focused pass 1 — figures vs script
| # | Finding | Decision |
|---|---|---|
| 1 | PhishEye “5%–95% victim-connection quantiles” not in the report's prevalence string | Rejected. The paper defines first/last victim as the 5% and 95% tails (.cols lines 736–744); eight days is 10 − 2. Added as PAPER_LITERAL phisheye_lifetime_cut so the guard sees the digits. The report's prevalence field was incomplete, the page was not. |
| 2 | Published script does not reject http://foo bar | Rejected. Hostname-None already exits 1. Gold-plating URL syntax is not the script's job; the community feed is a URL list. |
Focused pass 2 — citations and quotes
| # | Finding | Decision |
|---|---|---|
| 1 | Cloaking paragraph attributed “invisible to VT/GSB/SmartScreen for a week” to the 66/869 field study | Accepted. That week-long 0/100 is the controlled experiment (Table 1, 100 URLs per type). The field study says the 66 were “discovered solely by PhishDecloaker” and gives median blacklist 16h vs 11h. Split the two claims. WRAP important already described the controlled experiment; tightened to “100 URLs per cloaking type”. |
| 2 | All 17 citekeys resolve; no bib collisions; author table matches | Noted, no change. |
| 3 | Corpus defect (Roy slug / Nanayakkara PDF) correctly disclosed | Noted, no change. |
| 4 | 66 “absent” ROLE quotes: at least two are verbatim modulo line wrap, so “all ellipsis” overstates | Accepted on provenance only. Softened: absent is truncation/line-wrap/path, not fabrication; not all 66 are ellipsis. |
Focused pass 3 — external currency
| # | Finding | Decision |
|---|---|---|
| 1 | PhishTank / OpenPhish / GSB / APWG claims match live primary sources. Feed counts moved 73,660 → 73,668 (expected). OpenPhish 300/269 still matched on re-fetch. 2027-03-31 still not on Google docs. | Noted, no change to dated snapshots. |
| 2 | phishing_probe.sh dump and feed.txt fetches used curl -sS without -L; both hosts now redirect, so the probe false-failed | Accepted. Both fetches now curl -sSL. The probe already exited 1 on FAILED (not the exit-0 defect). |
Generic pass ran on the applied text, with no checklist.
Generic pass — no checklist
| # | Finding | Decision |
|---|---|---|
| 1 | Provenance header/log read as if the generic pass had already happened | Accepted. This table is that pass. |
| 2 | Content publishes the public dump path; provenance said dump URLs stay off the wiki | Accepted as a provenance error. The documented public path is the instrument. Policy is now: public dump path on the page; signed CDN query strings off. |
| 3 | Lead “the feed is a delayed vote” overgeneralises | Accepted. Vote is PhishTank; other feeds are dated provider snapshots. |
| 4 | Blocklist table treated online as years-old | Accepted. Distinguished current online=yes filter from old submission_time. |
| 5 | “Cloud ranges are the first filter” / “what a victim saw” / “IP blocking is the wrong lever” overbroad | Accepted. Scoped to Spartacus's proxy figure and Lee et al.'s CDN dataset. |
| 6 | WRAP todo claimed five papers “all have that tuple” | Accepted. Now the page's recommended instrumentation. |
| 7 | Five spot-checks cannot clear all 66 absent quotes | Accepted. Provenance now says five is not a census. |
| 8 | Detector papers in PAGE_N dilute the crawl/feed question | Accepted in part. One sentence on why detectors are in the 52; not moved to provenance (a student evaluating a feed is often about to train). |
| 9 | Update security red-link list and 1,097 → 1,098 | Accepted in part. WRAP now lists phishing as written. Cell figures stay the namespace snapshot (21, 1,097); the child re-derived 1,098 and PAGE_N 52 and owns those. |
No further generic re-run. The applied fixes are local wording, not new figures.
Report output (unedited)
- report_phishing-output.txt
# security:phishing — every figure with its denominator Corpus: 5859 extracted papers from CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf and IEEE S&P, 2010–2026. Crawled population: 1120 papers (19.1%), defined by crawlConfig != null or studyTypes includes automated-web-crawl. Web platform: 1622 papers (27.7%). 2025–2026 venue-years are provisional: CCS/IMC 2026 not held; IEEE S&P/WWW 2026 incompletely selected. Run date: figures below are recomputed on every run; external facts are in scripts/phishing_probe.sh and section Z. ## A. Why a keyword search is not a population missing paper.cols.txt: 4 (corpus known gap is 4) PUBLISHED_FT_PHISH 1098 PUBLISHED_FT_GSB 110 PUBLISHED_FT_FEEDS 219 Full-text /\bphish(?:ing|tank|ers)?\b/i: 1098 of 5855 papers with cols. That count is a fact about these being security venues. It is NOT the page N. Full-text Google Safe Browsing: 110. Full-text PhishTank|OpenPhish|APWG: 219 (upper bound). ## B. Schema union (candidate set) and the hand map PUBLISHED_UNION 139 PUBLISHED_UNION_WEB 92 PUBLISHED_UNION_CRAWLED 64 PUBLISHED_PAGE_N 52 Schema union (detection phenomenon/technique, classification resource/targetDetail, population.sourceList, or slug matching /phish/): 139 papers. of which platforms includes web: 92 (66.2%) of which crawled: 64 (46.0%) PAGE_N = ROLE in {live-site, detector, cloaking, kit}: 52 papers (37.4% of the union). Excluded from PAGE_N (user-study, email-sms, onchain, mobile-ui, mention, feed-gt): 87. Hand roles over the 139-paper union (one primary role per paper): Role Papers Share of 139 In PAGE_N? ---------- ------ ------------ ---------- live-site 29 20.9% yes detector 13 9.4% yes cloaking 8 5.8% yes kit 2 1.4% yes feed-gt 19 13.7% no user-study 29 20.9% no email-sms 9 6.5% no onchain 6 4.3% no mobile-ui 6 4.3% no mention 18 12.9% no Sum of role counts: 139 (must equal 139; single-label). ## C. PAGE_N shape PAGE_N web: 49 / 52 (94.2%) PAGE_N crawled: 40 / 52 (76.9%) PAGE_N posters (slug poster- or title Poster:): 1 PAGE_N by venue: Venue PAGE_N Share of 52 Union Corpus ------- ------ ----------- ----- ------ USENIX 17 32.7% 51 1410 CCS 8 15.4% 19 990 WWW 11 21.2% 21 843 IMC 9 17.3% 16 638 IEEE-SP 3 5.8% 14 767 NDSS 4 7.7% 16 701 PETS 0 0.0% 2 510 PAGE_N by year-bucket (2025–2026 starred = provisional): Window PAGE_N Share of 52 Union ---------- ------ ----------- ----- 2010–2013 1 1.9% 8 2014–2017 7 13.5% 19 2018–2021 13 25.0% 36 2022–2024 18 34.6% 40 2025–2026* 13 25.0% 36 PAGE_N per calendar year (do not read 2026 as a complete year): Year PAGE_N Union ----- ------ ----- 2010 1 4 2011 0 0 2012 0 1 2013 0 3 2014 2 8 2015 0 4 2016 2 3 2017 3 4 2018 2 7 2019 4 12 2020 3 7 2021 4 10 2022 6 11 2023 5 9 2024 7 20 2025* 8 25 2026* 5 11 ## D. Feeds named in the schema (union, then PAGE_N) Feeds / oracles folded from sourceList + resourceName + tools.name, papers of the union: Family Papers of union Share of 139 ------------------------------- --------------- ------------ VirusTotal (as phishing oracle) 32 23.0% PhishTank 30 21.6% Google Safe Browsing 27 19.4% APWG / eCX 13 9.4% OpenPhish 10 7.2% CertStream (discovery) 9 6.5% Microsoft SmartScreen 3 2.2% phishunt.io 3 2.2% SURBL 2 1.4% URLhaus 2 1.4% Netcraft 1 0.7% PhishStats 1 0.7% Same fold, papers of PAGE_N: Family Papers of PAGE_N Share of 52 ------------------------------- ---------------- ----------- VirusTotal (as phishing oracle) 16 30.8% Google Safe Browsing 13 25.0% APWG / eCX 12 23.1% PhishTank 11 21.2% CertStream (discovery) 9 17.3% OpenPhish 8 15.4% Microsoft SmartScreen 3 5.8% phishunt.io 3 5.8% Netcraft 1 1.9% PhishStats 1 1.9% Feed-fold residue (phish-ish string, no family) in the union: 118 strings. PUBLISHED_FEED_RESIDUE 118 RESIDUE (all): IMC/2010/an-empirical-study-of-orphan-dns-servers-in-the-internet five live feeds of phishing and malware-hosting sites IMC/2010/an-empirical-study-of-orphan-dns-servers-in-the-internet phishing and malware feeds CCS/2014/poster-proactive-blacklist-update-for-anti-phishing PhishTrack (custom) CCS/2014/poster-proactive-blacklist-update-for-anti-phishing PhishNet CCS/2014/poster-proactive-blacklist-update-for-anti-phishing PhishTrack IMC/2014/handcrafted-fraud-and-extortion-manual-account-hijacking-in-the-wild user-reported phishing emails IMC/2014/handcrafted-fraud-and-extortion-manual-account-hijacking-in-the-wild Google Forms taken down for phishing IMC/2014/handcrafted-fraud-and-extortion-manual-account-hijacking-in-the-wild phishing pages targeting Google NDSS/2015/nophish-app-evaluation-lab-and-retention-study NoPhish IMC/2014/the-dark-alleys-of-madison-avenue-understanding-malicious-advertisements 49 antivirus, spam and phishing blacklists IMC/2014/the-dark-alleys-of-madison-avenue-understanding-malicious-advertisements Malware and phishing blacklists WWW/2016/cracking-classifiers-for-evasion-a-case-study-on-the-googles-phishing-pages-filt Google's phishing pages filter (GPPF) CCS/2017/data-breaches-phishing-or-malware-understanding-the-risks-of-stolen-credentials custom phishing-kit template rules CCS/2017/hiding-in-plain-sight-a-longitudinal-study-of-combosquatting-abuse custom phishing detector WWW/2018/betrayed-by-your-dashboard-discovering-malicious-campaigns-via-web-analytics custom phishing-target labels IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni custom-generated phishing sites IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni custom-generated phishing sites IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni PhishFarm USENIX/2019/cognitive-triaging-of-phishing-attacks phishing email database provided by Org WWW/2019/doppelgangers-on-the-dark-web-a-large-scale-assessment-on-phishing-hidden-web-se manual phishing verification using external clues USENIX/2020/phishtime-continuous-longitudinal-measurement-of-the-effectiveness-of-anti-phish PhishTime USENIX/2020/phishtime-continuous-longitudinal-measurement-of-the-effectiveness-of-anti-phish PhishFarm USENIX/2020/sunrise-to-sunset-analyzing-the-end-to-end-life-cycle-and-effectiveness-of-phish user-forwarded phishing emails USENIX/2020/sunrise-to-sunset-analyzing-the-end-to-end-life-cycle-and-effectiveness-of-phish previously-proposed phishing URL classification scheme USENIX/2021/catching-phishers-by-their-bait-investigating-the-dutch-phishing-landscape-throu custom phishing-kit fingerprints USENIX/2021/phishpedia-a-hybrid-deep-learning-based-approach-to-visually-identify-phishing-w Phishpedia hybrid deep learning model (custom) USENIX/2021/phishpedia-a-hybrid-deep-learning-based-approach-to-visually-identify-phishing-w PhishCatcher USENIX/2021/phishprint-evading-phishing-detection-crawlers-by-prior-profiling custom phishing-site deployment USENIX/2021/phishprint-evading-phishing-detection-crawlers-by-prior-profiling PhishPrint CCS/2022/im-spartacus-no-im-spartacus-proactively-protecting-users-from-phishing-by-inten Cisco phishing kits IEEE-SP/2022/phishing-in-organizations-findings-from-a-large-scale-and-long-term-study commercial anti-phishing appliance IEEE-SP/2022/phishing-in-organizations-findings-from-a-large-scale-and-long-term-study commercial anti-phishing appliance IMC/2022/phishinpatterns-measuring-elicited-user-interactions-at-scale-on-phishing-websit VisualPhishNet IMC/2022/phishinpatterns-measuring-elicited-user-interactions-at-scale-on-phishing-websit custom intelligent phishing website crawler IMC/2022/phishinpatterns-measuring-elicited-user-interactions-at-scale-on-phishing-websit VisualPhishNet IMC/2022/phishinpatterns-measuring-elicited-user-interactions-at-scale-on-phishing-websit PhishInPattern IMC/2022/phishweb-a-progressive-multi-layered-system-for-phishing-websites-detection PhishStorm IMC/2022/phishweb-a-progressive-multi-layered-system-for-phishing-websites-detection Phishing-Benign USENIX/2022/inferring-phishing-intention-via-webpage-appearance-and-dynamics-a-deep-vision-b Phishpedia dataset USENIX/2022/inferring-phishing-intention-via-webpage-appearance-and-dynamics-a-deep-vision-b Phishpedia dataset USENIX/2022/inferring-phishing-intention-via-webpage-appearance-and-dynamics-a-deep-vision-b Phishpedia dataset USENIX/2022/inferring-phishing-intention-via-webpage-appearance-and-dynamics-a-deep-vision-b PhishIntention deep-learning models USENIX/2022/inferring-phishing-intention-via-webpage-appearance-and-dynamics-a-deep-vision-b PhishIntention USENIX/2022/phish-in-sheeps-clothing-exploring-the-authentication-pitfalls-of-browser-finger Phish-A USENIX/2022/phish-in-sheeps-clothing-exploring-the-authentication-pitfalls-of-browser-finger Phish-B IMC/2023/phishing-in-the-free-waters-a-study-of-phishing-attacks-created-using-free-websi FreePhish Twitter and Facebook streams IMC/2023/phishing-in-the-free-waters-a-study-of-phishing-attacks-created-using-free-websi FreePhish IMC/2023/phishing-in-the-free-waters-a-study-of-phishing-attacks-created-using-free-websi VisualPhishNet IMC/2023/phishing-in-the-free-waters-a-study-of-phishing-attacks-created-using-free-websi PhishIntention CCS/2023/txphishscope-towards-detecting-and-understanding-transaction-based-phishing-on-e Chainabuse Ethereum Phishing Scam Reports CCS/2023/txphishscope-towards-detecting-and-understanding-transaction-based-phishing-on-e TxPhishScope detection results CCS/2023/txphishscope-towards-detecting-and-understanding-transaction-based-phishing-on-e 11 large-scale TxPhish events and Web3 security-company reports USENIX/2023/knowledge-expansion-and-counterfactual-interaction-for-reference-based-phishing DynaPhish USENIX/2023/knowledge-expansion-and-counterfactual-interaction-for-reference-based-phishing Phishpedia USENIX/2023/knowledge-expansion-and-counterfactual-interaction-for-reference-based-phishing PhishIntention USENIX/2023/rods-with-laser-beams-understanding-browser-fingerprinting-on-phishing-pages random sample of phishing websites USENIX/2023/rods-with-laser-beams-understanding-browser-fingerprinting-on-phishing-pages phishing websites belonging to two selected signatures USENIX/2024/guardians-of-the-galaxy-content-moderation-in-the-interplanetary-file-system Web2 Anti-Phishing Services (APS) USENIX/2024/guardians-of-the-galaxy-content-moderation-in-the-interplanetary-file-system combined badbits and phishing denylist USENIX/2024/knowphish-large-language-models-meet-multimodal-knowledge-graphs-for-enhancing-r KnowPhish Detector (custom) USENIX/2024/knowphish-large-language-models-meet-multimodal-knowledge-graphs-for-enhancing-r Phishpedia USENIX/2024/knowphish-large-language-models-meet-multimodal-knowledge-graphs-for-enhancing-r PhishIntention USENIX/2024/knowphish-large-language-models-meet-multimodal-knowledge-graphs-for-enhancing-r DynaPhish USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish PhishIntention's logo dataset USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish Phishpedia USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish Phishpedia USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish PhishIntention's protected brand list USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish PhishIntention USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish Phishpedia USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish VisualPhishNet USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish PhishIntention USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish Phishpedia USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish VisualPhishNet USENIX/2024/less-defined-knowledge-and-more-true-alarms-reference-based-phishing-detection-w custom PhishLLM decision rules USENIX/2024/less-defined-knowledge-and-more-true-alarms-reference-based-phishing-detection-w PhishLLM USENIX/2024/less-defined-knowledge-and-more-true-alarms-reference-based-phishing-detection-w Phishpedia USENIX/2024/less-defined-knowledge-and-more-true-alarms-reference-based-phishing-detection-w PhishIntention USENIX/2024/less-defined-knowledge-and-more-true-alarms-reference-based-phishing-detection-w DynaPhish USENIX/2024/phishdecloaker-detecting-captcha-cloaked-phishing-websites-via-hybrid-vision-bas PhishDecloaker WWW/2024/are-adversarial-phishing-webpages-a-threat-in-reality-understanding-the-users-pe Chiew et al. phishing dataset WWW/2024/are-adversarial-phishing-webpages-a-threat-in-reality-understanding-the-users-pe SaTML'23 adversarial phishing webpage dataset WWW/2024/phishing-vs-legit-comparative-analysis-of-client-side-resources-of-phishing-and custom collected phishing kits WWW/2024/phishing-vs-legit-comparative-analysis-of-client-side-resources-of-phishing-and GoPhish CCS/2025/systematic-assessment-of-tabular-data-synthesis Adult, Shoppers, Phishing, Magic, Faults, Bean, Obesity, Robot, Abalone, News, Insurance, and Wine NDSS/2025/dissecting-payload-based-transaction-phishing-on-ethereum public phishing complaints and security-community phishing blogs NDSS/2025/dissecting-payload-based-transaction-phishing-on-ethereum ice-phishing victims identified from phishing transactions NDSS/2025/dissecting-payload-based-transaction-phishing-on-ethereum Etherscan Fake_Phishing nametags USENIX/2025/unsafe-llm-based-search-quantitative-analysis-and-mitigation-of-safety-risks-in PhishLLM USENIX/2025/unsafe-llm-based-search-quantitative-analysis-and-mitigation-of-safety-risks-in PhishLLM WWW/2025/7-days-later-analyzing-phishing-site-lifespan-after-detected Phishing Army WWW/2025/7-days-later-analyzing-phishing-site-lifespan-after-detected Phishing Database NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns Beyond Phish NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns Beyond Phish NDSS/2026/loki-proactively-discovering-online-scams-by-mining-toxic-search-queries BeyondPhish NDSS/2026/ctphishcapture-uncovering-credential-theft-based-phishing-scams-targeting-cryptocurrency-wallets CtPhishCapture PETS/2026/linguistic-hooks-investigating-the-role-of-language-triggers-in-phishing-emails NIST Phish Scale PETS/2026/linguistic-hooks-investigating-the-role-of-language-triggers-in-phishing-emails NIST Phish Scale WWW/2026/netting-phish-in-the-ipfs-ocean-real-time-monitoring-and-characterization-of-dec Phishpedia WWW/2026/netting-phish-in-the-ipfs-ocean-real-time-monitoring-and-characterization-of-dec PhishIntention IMC/2025/unmasking-the-shadow-economy-a-deep-dive-into-drainer-as-a-service-phishing-on-e TxPhishScope IMC/2025/unmasking-the-shadow-economy-a-deep-dive-into-drainer-as-a-service-phishing-on-e large-scale phishing incidents IMC/2025/unmasking-the-shadow-economy-a-deep-dive-into-drainer-as-a-service-phishing-on-e TxPhishScope USENIX/2025/evaluating-the-effectiveness-and-robustness-of-visual-similarity-based-phishing PhishIntention USENIX/2025/evaluating-the-effectiveness-and-robustness-of-visual-similarity-based-phishing Phishpedia USENIX/2025/evaluating-the-effectiveness-and-robustness-of-visual-similarity-based-phishing DynaPhish USENIX/2025/evaluating-the-effectiveness-and-robustness-of-visual-similarity-based-phishing PhishZoo USENIX/2025/evaluating-the-effectiveness-and-robustness-of-visual-similarity-based-phishing VisualPhishNet USENIX/2025/darkgram-a-large-scale-analysis-of-cybercriminal-activity-channels-on-telegram PhishIntention USENIX/2025/darkgram-a-large-scale-analysis-of-cybercriminal-activity-channels-on-telegram PhishIntention NDSS/2026/phishlang-a-real-time-fully-client-side-phishing-detection-framework-using-mobilebert PhishPedia NDSS/2026/phishlang-a-real-time-fully-client-side-phishing-detection-framework-using-mobilebert PhishLang NDSS/2026/phishlang-a-real-time-fully-client-side-phishing-detection-framework-using-mobilebert GoPhish IEEE-SP/2021/crawlphish-large-scale-analysis-of-client-side-cloaking-techniques-in-phishing Public Dataset from various phishing URL sources IEEE-SP/2021/crawlphish-large-scale-analysis-of-client-side-cloaking-techniques-in-phishing CrawlPhish visual-similarity detector IEEE-SP/2021/crawlphish-large-scale-analysis-of-client-side-cloaking-techniques-in-phishing CrawlPhish cloaking technique database IEEE-SP/2021/crawlphish-large-scale-analysis-of-client-side-cloaking-techniques-in-phishing CrawlPhish IEEE-SP/2024/practical-attacks-against-dns-reputation-systems StopForumSpam, Prigent DBL, and Phishing Army IEEE-SP/2024/practical-attacks-against-dns-reputation-systems StopForumSpam, Prigent DBL, and Phishing Army PUBLISHED_PHISHTANK_UNION 30 PUBLISHED_PHISHTANK_SOURCELIST_RESOURCENAME 28 # any sourceList/resourceName matching /phishtank/i NAMESPACE_SITTING_PHISHTANK 21 # report_security_namespace.mjs topNames fold of /phish/ sourceList+resourceName; not this sitting's 28 or 30 ## E. Full-text probes on PAGE_N (upper bounds, then read) Each probe is paper-counted over PAGE_N. Hits are an upper bound until read. PROBE cloak hits= 34 / 52 missing_cols=0 65.4% PROBE takedown hits= 31 / 52 missing_cols=0 59.6% PROBE lifespan hits= 26 / 52 missing_cols=0 50.0% PROBE phishtank-word hits= 32 / 52 missing_cols=0 61.5% PROBE openphish-word hits= 20 / 52 missing_cols=0 38.5% PROBE gsb-word hits= 35 / 52 missing_cols=0 67.3% PROBE captcha-cloak hits= 21 / 52 missing_cols=0 40.4% ## F. Per-paper figures the page quotes (from detection.prevalence on PAGE_N) ### IEEE-SP/2021/crawlphish-large-scale-analysis-of-client-side-cloaking-techniques-in-phishing role=cloaking phen="client-side JavaScript cloaking" metric="share of phishing websites" prevalence="35,067 of 112,005 websites (31.31%)" quote="Within our dataset of 112,005 phishing websites, CrawlPhish found that 35,067 (31.31%) phishing websites implement client-side cloaking techniques in total" phen="client-side cloaking prevalence over time" metric="share of phishing websites" prevalence="23.32% in 2018 and 33.70% in 2019" quote="23.32% (6,024) in 2018 and 33.70% (29,043) in 2019." phen="cloaking detection accuracy" metric="false-positive and false-negative rates" prevalence="1.45% false positives and 1.75% false negatives" quote="CrawlPhish correctly detected 1,965 phishing websites as cloaked and 1,971 as uncloaked, with a false-negative rate of 1.75% (35) and a false-positive rate of 1.45% (29)." phen="cloaking semantic categorization" metric="categorization accuracy" prevalence="100% accuracy on 2,000 cloaked phishing websites" quote="We found that CrawlPhish correctly categorized the cloaking type with 100% accuracy." phen="phishing ecosystem detection bypass" metric="sites blacklisted over seven days" prevalence="None blacklisted except 21 Click Through sites in Microsoft Edge" quote="At the conclusion of these experiments, we found that none of our phishing websites were blacklisted in any browser, with the exception of Click Through websites, 21 (42%) of which were blocked in Microsoft Edge" phen="human exposure to cloaked content" metric="share seeing hidden text" prevalence="100% Mouse, 97.72% Click Through, and 42.55% Notification Window" quote="For the Mouse Movement cloaking technique, 100% of the workers saw the \"Hello World\" text... For the Click Through websites, 97.72% saw the text" ### IMC/2019/opening-the-blackbox-of-virustotal-analyzing-online-phishing-scan-engines role=live-site phen="phishing-site detection" metric="detected sites per vendor" prevalence="Only 15 of 68 vendors detected at least one of 36 simple phishing sites" quote="Over multiple scans, only 15 vendors (out of 68) have detected at least one of the 36 simple phishing sites." phen="phishing-takedown reaction" metric="label changes after takedown" prevalence="Only four vendors flipped some malicious labels after week three" quote="we find 4 vendors that flip some \"malicious\" labels to \"benign\" after the third week (for PayPal sites only)." phen="obfuscation evasion" metric="average malicious labels per site" prevalence="Image and code obfuscation reduced averages from 12.1 to 4.5 and 2" quote="the average number of malicious labels drops from 12.1 to 4.5 and 2 respectively." ### CCS/2016/phisheye-live-monitoring-of-sandboxed-phishing-kits role=kit phen="Phishing-kit installation and deployment" metric="unique phishing kits" prevalence="643 unique kits; 474 correctly installed by 471 attackers" quote="We collected 643 unique phishing kits, which were uploaded on our honeypot over a period of five months from September 2015 to the end of January 2016. Out of this initial dataset, 474 kits" phen="Victim connections" metric="unique victims" prevalence="2,468 victims connected to 127 phishing kits" quote="After our aggressive filtering, we counted a total number of 2,468 victims who have connected to 127 distinct phishing kits." phen="Phishing-kit effective lifetime" metric="estimated lifetime" prevalence="eight days on average" quote="The last victims (at the 95% threshold) connect in average to the phishing page after 10 days. This gives us an estimated lifetime of eight days." phen="Blacklist detection latency" metric="detection latency" prevalence="98% blacklisted; average latency of 12 days" quote="almost all phishing pages that were installed on the honeypot (98% of them) were correctly detected and blacklisted by the two phishing blacklists, we observed an average detection latency of 12 days." phen="Blacklist timing relative to victims" metric="share blacklisted after 75% of victims" prevalence="62% of kits were blacklisted only after 75% of victims connected" quote="62% of the phishing kits were blacklisted only after 75% of victims had already connected to the corresponding phishing page." phen="Blacklist-evasion visitor crowd" metric="visitor distribution over elapsed time" prevalence="Hundreds of connections followed publication of a random URL by PhishTank" quote="After the random link was published by PhishTank we received many connections from researchers and security companies, which connected to the reported link instead of the entry page." ### IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni role=cloaking phen="phishing blacklist occurrence" metric="proportion of sites blacklisted" prevalence="In the full tests, 23.0% of crawled cloaked sites versus 49.4% of non-cloaked sites were blacklisted." quote="in the full tests, just 23.0% of our sites with cloaking (which were crawled) ended up being blacklisted in at least one browser- far fewer than the 49.4% sites without cloaking" phen="blacklisting timeliness" metric="mean time before first blacklisting" prevalence="Cloaking increased average time to blacklisting from 126 to 238 minutes in full tests." quote="Cloaking also slowed the average time to blacklisting from 126 minutes (for sites without cloaking) to 238 minutes." phen="cloaking effectiveness" metric="reduction in blacklisting likelihood" prevalence="Geolocation, device-type, and JavaScript cloaking reduced blacklisting likelihood by over 55% on average." quote="simple cloaking techniques representative of real-world attacks- including those based on geolocation, device type, or JavaScript-were effective in reducing the likelihood of blacklisting by over 55% on average." phen="mobile blacklist protection" metric="presence of blacklist warnings" prevalence="Mobile Chrome, Safari, and Firefox showed no blacklist warnings during the tested period." quote="mobile Chrome, Safari, and Firefox failed to show any blacklist warnings between mid-2017 and late 2018" ### USENIX/2021/phishprint-evading-phishing-detection-crawlers-by-prior-profiling role=cloaking phen="Phishing-site evasion" metric="site lifetime" prevalence="18 of 20 cloaked sites remained functional for one month" quote="In the one-month period in which we did the monitoring, only 2 sites (say, 'A' and 'B') got blocked." phen="Cloaking specificity among users" metric="share shown phishing content" prevalence="PhishPrint-powered logic showed phishing content to 79% of users" quote="Overall, the results showed that PhishPrint-powered cloaking logic decided to show phishing content for 79% of the users." ### CCS/2022/im-spartacus-no-im-spartacus-proactively-protecting-users-from-phishing-by-inten role=cloaking phen="fingerprinting-based cloaking" metric="share of phishing kits" prevalence="96.52% (2,831) of 2,933 phishing kits" quote="In total, 96.52% (2,831) out of 2,933 phishing kits contain fingerprinting-based cloaking techniques." phen="phishing-content evasion" metric="share of phishing URLs protected" prevalence="132,247 of 160,728 (82.28%)" quote="132,247 (82.28%) did not contain malicious content in Spartacus." phen="trigger-word cloaking" metric="share of cloaked sites evaded" prevalence="The word “bot” evaded 99.31% of cloaked phishing websites" quote="99.31% of the cloaked phishing web sites can be evaded by appending bot in the User-Agent." phen="IP-based cloaking" metric="share of phishing websites evaded" prevalence="7,540 of 8,474 (88.98%)" quote="Among the visited phishing web sites, Spartacus evaded 88.98% (7,540) of them through the proxy server." phen="anti-phishing detection latency" metric="median detection time" prevalence="154 minutes for Spartacus-evaded phishing sites; 22 minutes for non-evaded sites" quote="The median detection time is 154 minutes. However, within 24 hours, they only detect/blacklist 76.16% of the cloaked phishing web sites." ### USENIX/2024/phishdecloaker-detecting-captcha-cloaked-phishing-websites-via-hybrid-vision-bas role=cloaking phen="CAPTCHA-cloaking evasion of phishing detectors" metric="URLs blacklisted / URLs submitted" prevalence="0/100 for every CAPTCHA-cloaking type across VirusTotal, Google Safe Browsing, and SmartScreen" quote="All baseline sites are blacklisted within 24 hours, whereas all CAPTCHA-cloaked sites remain undetected for 7 days and counting." phen="CAPTCHA-cloaked phishing in the wild" metric="share of captured phishing websites" prevalence="7.6% (66 of 869 phishing websites)" quote="Of these, 7.6% were CAPTCHA-cloaked phishing websites, all of which were discovered solely by PhishDecloaker." phen="Phishing detection recovery after decloaking" metric="phishing detection rate" prevalence="recovered to 78%, 40%, 95%, and 84% of baseline performance for reCAPTCHA, hCaptcha, slider, and rotation" quote="PhishDecloaker can recover their detection rates back to a percentage of their original performance: 78% on re-CAPTCHA, 40% on hCaptcha, 95% on slider CAPTCHA, and 84% on rotation CAPTCHA." phen="CAPTCHA-cloaked phishing lifespan and blacklist delay" metric="median hours" prevalence="16 hours to blacklist for CAPTCHA-cloaked sites versus 11 hours for ordinary sites" quote="CAPTCHA-cloaked phishing sites take a median time of 16 hours to be blacklisted, which is 45.5% longer than ordinary phishing sites (11 hours)." ### WWW/2025/7-days-later-analyzing-phishing-site-lifespan-after-detected role=live-site phen="phishing-site lifespan" metric="mean and median lifespan" prevalence="54.04 hours mean and 5.46 hours median across 286,237 URLs" quote="The average lifespan of phishing websites in our dataset is 54.04 hours, but the median is just 5.46 hours" phen="phishing takedown causes" metric="share of takedown errors" prevalence="DNS resolution failure accounted for 67.23% of errors" quote="DNS resolution failures are the most common takedown cause (67.23%, 60,913 domains)" phen="visual component changes" metric="lifespan by visual-change frequency" prevalence="Sites with over 100 changes had a 404.52-hour median lifespan" quote="Screenshots with a similarity score above 0.95 are clustered together." phen="DOM and resource evolution" metric="relative feature change" prevalence="30,069 of 286,237 sites showed final modifications" quote="The analysis covers changes in website resources between discovery and takedown for 286,237 phishing sites, with 10.5% (30,069) showing final modifications." phen="DNS configuration changes" metric="share of sites changing DNS settings" prevalence="75.84% altered DNS settings between discovery and takedown" quote="Our analysis reveals that 75.84% of phishing websites alter their DNS settings between discovery and takedown." phen="CDN usage" metric="share of phishing-related IPs" prevalence="99.80% of phishing-related IPs used CDNs" quote="99.80% of phishing-related IPs use CDNs, rendering IP-based blocking largely ineffective." phen="phishing server transitions" metric="number of websites changing server versions" prevalence="13 Nginx websites and four Apache websites changed versions" quote="Our findings show that 13 websites using Nginx change versions over time" ### USENIX/2020/sunrise-to-sunset-analyzing-the-end-to-end-life-cycle-and-effectiveness-of-phish role=live-site phen="phishing attack life-cycle timing" metric="average campaign duration and stage delays" prevalence="average campaign from first to last victim takes 21 hours" quote="We find the average campaign from start to the last victim takes just 21 hours." phen="phishing victim traffic" metric="victim count" prevalence="4.8 million victims over one year, excluding crawler traffic" quote="Over a one year period, our network monitor recorded 4.8 million victims who visited phishing pages, excluding crawler traffic." phen="credential compromise and fraud" metric="share of distinct Known Visitors with fraudulent transactions" prevalence="7.42% of distinct Known Visitors subsequently suffered a fraudulent transaction" quote="In our dataset, the accounts of 7.42% of distinct Known Visitors subsequently suffered a fraudulent transaction" phen="browser phishing-warning effectiveness" metric="relative compromised-visitor ratio" prevalence="ratio fell to 71.51% after one hour and 43.55% after two hours" quote="within one hour after detection, at which point the ratio of Compromised Visitors drops to 71.51%. By the end of the second hour, the ratio drops further to 43.55%" phen="phishing-data visibility" metric="visibility ratio" prevalence="39.1% of hostnames and 40.9% of URLs" quote="We found that our approach had visibility into an average of 39.1% of all hostnames and 40.9% of all URLs" ### IMC/2022/phishinpatterns-measuring-elicited-user-interactions-at-scale-on-phishing-websit role=live-site phen="phishing campaign clustering" metric="campaign count and brand count" prevalence="8,472 campaigns targeting over 381 brands" quote="we discovered 8,472 distinct phishing campaigns targeting over 381 brands." ### USENIX/2021/catching-phishers-by-their-bait-investigating-the-dutch-phishing-landscape-throu role=live-site phen="potential phishing domains" metric="number of domains" prevalence="7,936 potential domains" quote="our domain detector labeled 7,936 domains as potentially malicious, which meant that these domains reached the threshold value" phen="phishing kit deployment" metric="number of verified phishing FQDNs" prevalence="1,363 verified phishing FQDNs" quote="Our final dataset contained 1,363 verified phishing fully qualified domain names (FQDN) which have been online for at least one hour." phen="phishing kit family prevalence" metric="share of identified phishing websites" prevalence="Almost 89% used a kit within the uAdmin family." quote="Almost 89% of all identified phishing websites were made with a kit within this family" phen="phishing domain uptime" metric="median uptime" prevalence="24 hours" quote="On average, a phishing domain in our dataset is online for 45 hours, but we find a median uptime of 24 hours." phen="domain cloaking" metric="share of detected domains" prevalence="946 domains (69%) returned a blank screen and no favicon." quote="In fact, 946 (69%) of the detected phishing domains returned a blank screen - and no favicon - to our crawler" phen="external resource loading" metric="share of dataset" prevalence="104 domains (7.6%) loaded resources from benign counterparts." quote="only 104 domains (7.6% of the total dataset) load their resources directly from their benign counterparts." ### IMC/2018/needle-in-a-haystack-tracking-down-elite-phishing-domains-in-the-wild role=live-site phen="squatting phishing pages" metric="share of squatting domains" prevalence="1,175 confirmed domains, approximately 0.2% of 657,663 domains." quote="As shown in Table 8, after manual examination, we confirmed 1,175 domains are indeed phishing domains." phen="layout obfuscation" metric="image-hash distance" prevalence="Most brands had average distance around 20 or higher." quote="most brands have an average distance around 20 or higher, suggesting that layout obfuscation is very common." phen="blacklist evasion" metric="share undetected for at least a month" prevalence="91.5% remained undetected by the tested blacklists." quote="Collectively these blacklists only detected 8.4% of the squatting phishing pages, which means 91.5% of the phishing domains remain undetected for at least a month." phen="phishing-page lifetime" metric="share remaining alive" prevalence="About 80% remained alive after at least a month." quote="Most pages (about 80%) still remain alive after at least a month." phen="IP geolocation" metric="number of IP addresses and countries" prevalence="1,021 IP addresses across 53 countries." quote="In total, we are able to look up the geolocation of 1,021 IP addresses, hosted in 53 different countries." ### WWW/2017/tracking-phishing-attacks-over-time role=live-site phen="replicated phishing attacks" metric="share of phishing URLs in flagged clusters" prevalence="17,451 of 19,066 URLs (91.53%)" quote="17,451 (resp. 11,244) of the 19,066 (resp. 12,859) phishing URLs end up in a flagged cluster, which means that 91.53% (resp. 87.44%)" phen="false positives on legitimate sites" metric="false positive rate" prevalence="20 of 24,800 legitimate sites (0.08%)" quote="Only 20 of these sites end up in one of the 2,831 phishing clusters, a false positive rate of only 0.08%." phen="phishing attack-class lifespan" metric="cluster lifespan" prevalence="Average lifespan was 25 days; around 80% lasted less than one month" quote="The lifespan is defined as the duration between the first and the last observed attack instance. ... the average lifespan of a cluster in our database is 25 days." ### IMC/2023/phishing-in-the-free-waters-a-study-of-phishing-attacks-created-using-free-websi role=live-site phen="FWB-hosted phishing attacks" metric="number of zero-day phishing URLs" prevalence="31,405 URLs from November 2022 to May 2023" quote="We ran FreePhish for a period of six months, from November 2022 to May 2023, identifying 31,405 zero-day phishing attacks" phen="Historical FWB phishing prevalence" metric="share of phishing URLs using FWBs" prevalence="25.2K of 34.7K unique phishing URLs used 17 FWBs" quote="25.2K URLs ... utilized 17 unique free website-building services." phen="Phishing classifier performance" metric="F1-score and median runtime" prevalence="F1-score 0.96 and median runtime 2.8 seconds" quote="our augmented StackModel has an F1-score of 0.96 and a median run-time of only 2.8 secs." phen="Anti-phishing blocklist coverage" metric="coverage and response time" prevalence="GSB covered 18.4% of FWB phishing URLs; median response time 6 hours" quote="GSB covered 18.4% of all FWB phishing URLs while having a median response time of 6 hours" phen="FWB takedown response" metric="removal rate and response time" prevalence="29% removed after two weeks; median speed 9 hours 43 minutes" quote="only 29% of the websites were removed by the respective FWB service after two weeks since they first appeared in our dataset" phen="Evasive phishing variants" metric="counts and shares of URLs" prevalence="539 Google Sites URLs linked to other phishing pages; 427 Google Sites and 473 Blogspot URLs embedded phishing iframes" quote="We randomly sampled 1K URLs from this set and qualitatively analyzed their content ... We developed heuristics to automatically identify these attack vectors" ### WWW/2024/phishing-vs-legit-comparative-analysis-of-client-side-resources-of-phishing-and role=live-site phen="phishing sites without JavaScript" metric="share of collected phishing webpages without JavaScript" prevalence="22.8% (172,348 of 757,421)" quote="There are 22.8% (172,348 out of 757,421) of our collected phishing webpages that do not use JavaScript." phen="JavaScript library diversity" metric="number of distinct libraries" prevalence="132 phishing libraries versus 41 legitimate-site libraries" quote="A total of 132 distinct JavaScript libraries are identified in our phishing dataset, in contrast to the 41 distinct JavaScript libraries found in their corresponding legitimate target brand websites." phen="outdated JavaScript libraries" metric="average version-age difference" prevalence="646 days, nearly 21.2 months older on phishing websites" quote="On average, phishing websites employ JavaScript libraries that are 646 days older, equivalent to nearly 21.2 months, than the versions utilized by legitimate websites." phen="code copying" metric="share of clusters with exact or high overlap" prevalence="2.5% exact overlap; over 21.5% exceeded 85% overlap" quote="Our analysis of the similarities between different clusters revealed that there is an exact overlap of 2.5% in terms of code copying. Additionally, more than 21.5% of the clusters show an overlap exceeding 85%." ### USENIX/2021/phishpedia-a-hybrid-deep-learning-based-approach-to-visually-identify-phishing-w role=detector phen="Phishing webpage identification" metric="identification rate, detection rate, precision, recall" prevalence="99.2% identification rate, 98.2% precision, and 87.1% recall on the reported evaluation" quote="Phishpedia 99.2% 98.2% 87.1% 0.19" phen="Logo presence in phishing webpages" metric="share of phishing webpages with logos" prevalence="about 98.6%" quote="We randomly sampled 5,000 webpages from the phishing webpage dataset, and manually validated that 70 of them have no logos. That is, the ratio of phishing webpages with logos is about 98.6%." phen="Phishing target-brand coverage" metric="coverage of phishing webpages" prevalence="top 100 brands cover 95.8% of phishing webpages" quote="our empirical study on around 30K phishing webpages based on an OpenPhish feed (Section 5.1) shows that the top 100 brands cover 95.8% phishing webpages" phen="Phishing discovery in the wild" metric="number of real and zero-day phishing webpages" prevalence="1,704 real phishing webpages within 30 days; 1,133 not reported by VirusTotal" quote="Phishpedia discovered 1,704 phishing webpages within 30 days and 1,133 of them are not detected by any engines in VirusTotal [9]." ## G. Quote check of ROLE deciding sentences against paper.cols.txt QUOTE_TITLE_ONLY IEEE-SP/2024/from-chatbots-to-phishbots-phishing-scam-generation-in-commercial-large-language (stored PDF/cols are a different paper; see provenance) ROLE quotes vs paper.cols.txt: exact 72, partial 0, absent 66, missing-file 0, title-only 1 of 139 PUBLISHED_QUOTE_EXACT 72 PUBLISHED_QUOTE_PARTIAL 0 PUBLISHED_QUOTE_ABSENT 66 ABSENT quotes (must be spliced or PDF-only; listed, not silently dropped): USENIX/2010/searching-the-searchers-with-searchaudit NDSS/2013/compa-detecting-compromised-accounts-on-social-networks USENIX/2014/the-emperor-s-new-password-manager-security-analysis-of-web-based-password-manag IMC/2014/handcrafted-fraud-and-extortion-manual-account-hijacking-in-the-wild NDSS/2015/nophish-app-evaluation-lab-and-retention-study USENIX/2015/towards-discovering-and-understanding-task-hijacking-in-android WWW/2017/tracking-phishing-attacks-over-time IMC/2018/characterizing-the-internet-host-population-using-deep-learning-a-universal-and USENIX/2018/end-to-end-measurements-of-email-spoofing-attacks WWW/2018/betrayed-by-your-dashboard-discovering-malicious-campaigns-via-web-analytics NDSS/2019/digital-healthcare-associated-infection-a-case-study-on-the-security-of-a-major-multi-campus-hospital-system USENIX/2019/cognitive-triaging-of-phishing-attacks USENIX/2019/detecting-and-characterizing-lateral-phishing-at-scale USENIX/2019/the-webs-identity-crisis-understanding-the-effectiveness-of-website-identity-ind USENIX/2019/users-really-do-answer-telephone-scams WWW/2019/doppelgangers-on-the-dark-web-a-large-scale-assessment-on-phishing-hidden-web-se WWW/2019/hack-for-hire-exploring-the-emerging-market-for-account-hijacking NDSS/2020/deceptive-previews-a-study-of-the-link-preview-trustworthiness-in-social-platforms USENIX/2020/phishtime-continuous-longitudinal-measurement-of-the-effectiveness-of-anti-phish USENIX/2020/sunrise-to-sunset-analyzing-the-end-to-end-life-cycle-and-effectiveness-of-phish WWW/2020/dirty-clicks-a-study-of-the-usability-and-security-implications-of-click-related IMC/2021/knock-and-talk-investigating-local-network-communications-on-websites USENIX/2021/catching-phishers-by-their-bait-investigating-the-dutch-phishing-landscape-throu USENIX/2021/phishprint-evading-phishing-detection-crawlers-by-prior-profiling WWW/2021/where-are-you-taking-me-understanding-abusive-traffic-distribution-systems IEEE-SP/2022/phishing-in-organizations-findings-from-a-large-scale-and-long-term-study IMC/2022/phishinpatterns-measuring-elicited-user-interactions-at-scale-on-phishing-websit USENIX/2022/identity-confusion-in-webview-based-mobile-app-in-app-ecosystems USENIX/2022/inferring-phishing-intention-via-webpage-appearance-and-dynamics-a-deep-vision-b USENIX/2022/helping-hands-measuring-the-impact-of-a-large-threat-intelligence-sharing-commun IMC/2023/phishing-in-the-free-waters-a-study-of-phishing-attacks-created-using-free-websi CCS/2023/understanding-and-detecting-abused-image-hosting-modules-as-malicious-services USENIX/2023/knowledge-expansion-and-counterfactual-interaction-for-reference-based-phishing USENIX/2023/rods-with-laser-beams-understanding-browser-fingerprinting-on-phishing-pages WWW/2023/cashing-in-on-contacts-characterizing-the-onlyfans-ecosystem CCS/2024/employees-attitudes-towards-phishing-simulations-its-like-when-a-child-reaches-o USENIX/2024/assessing-suspicious-emails-with-banner-warnings-among-blind-and-low-vision-user USENIX/2024/it-doesnt-look-like-anything-to-me-using-diffusion-model-to-subvert-visual-phish USENIX/2024/less-defined-knowledge-and-more-true-alarms-reference-based-phishing-detection-w USENIX/2024/malla-demystifying-real-world-large-language-model-integrated-malicious-services USENIX/2024/phishdecloaker-detecting-captcha-cloaked-phishing-websites-via-hybrid-vision-bas WWW/2024/phishinwebview-analysis-of-anti-phishing-entities-in-mobile-apps-with-webview-ta WWW/2024/are-adversarial-phishing-webpages-a-threat-in-reality-understanding-the-users-pe WWW/2024/zipzap-efficient-training-of-language-models-for-large-scale-fraud-detection-on WWW/2024/phishing-vs-legit-comparative-analysis-of-client-side-resources-of-phishing-and IEEE-SP/2016/sending-out-an-sms-characterizing-the-security-of-the-sms-ecosystem-with-public CCS/2025/quantifying-security-training-in-organizations-through-the-analysis-of-u-s-sec-1 IEEE-SP/2025/restricting-the-link-effects-of-focused-attention-and-time-delay-on-phishing-war USENIX/2025/doubly-dangerous-evading-phishing-reporting-systems-by-leveraging-email-tracking USENIX/2025/scanned-and-scammed-insecurity-by-obsqrity-measuring-user-susceptibility-and-awa USENIX/2025/blockchain-address-poisoning NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns WWW/2025/whats-in-phishers-a-longitudinal-study-of-security-configurations-in-phishing-we USENIX/2025/url-inspection-tasks-helping-users-detect-phishing-links-in-emails NDSS/2026/ctphishcapture-uncovering-credential-theft-based-phishing-scams-targeting-cryptocurrency-wallets PETS/2026/toward-adaptive-privacy-enhancing-training-a-longitudinal-study-of-how-personali USENIX/2026/sok-philter-uncovering-security-and-functional-gaps-in-ai-based-phishing-website USENIX/2026/a-large-scale-study-of-personalized-phishing-using-large-language-models WWW/2026/netting-phish-in-the-ipfs-ocean-real-time-monitoring-and-characterization-of-dec CCS/2025/phishing-susceptibility-and-the-in-effectiveness-of-common-anti-phishing-interve IEEE-SP/2025/mantis-detection-of-zero-day-malicious-domains-leveraging-low-reputed-hosting-in NDSS/2025/misdirection-of-trust-demystifying-the-abuse-of-dedicated-url-shortening-service NDSS/2026/phishlang-a-real-time-fully-client-side-phishing-detection-framework-using-mobilebert IEEE-SP/2024/practical-attacks-against-dns-reputation-systems IEEE-SP/2024/understanding-the-privacy-practices-of-political-campaigns-a-perspective-from-th IEEE-SP/2022/siraj-a-unified-framework-for-aggregation-of-malicious-entity-detectors ## H. Sensitivity: drop posters from PAGE_N PAGE_N including posters: 52 PAGE_N excluding posters: 51 Posters in PAGE_N: 1 (CCS/2014/poster-proactive-blacklist-update-for-anti-phishing) ## Z. EXTERNAL FIGURES (not from the extraction; primary source in phishing_probe.sh) These are dated facts about the instruments, not corpus counts. OPENPHISH_COMMUNITY_FEED_URLS 300 # GET https://openphish.com/feed.txt 2026-08-27 OPENPHISH_COMMUNITY_HOSTS 269 OPENPHISH_COMMUNITY_HTTPS 201 OPENPHISH_COMMUNITY_HTTP 99 OPENPHISH_COMMUNITY_BYTES 15020 OPENPHISH_COMMUNITY_REFRESH 12 hours # openphish.com/phishing_feeds.html PHISHTANK_DUMP_N 73660 # GET http://data.phishtank.com/data/online-valid.csv.gz unauthenticated 2026-08-27 PHISHTANK_DUMP_ONLINE 73660 PHISHTANK_DUMP_VERIFIED 73660 PHISHTANK_DUMP_TARGET_OTHER 65457 PHISHTANK_DUMP_BYTES_GZIP 2549032 PHISHTANK_DUMP_BYTES_UNCOMPRESSED 14105733 PHISHTANK_DUMP_MIN_SUBMISSION 2011-02-18 PHISHTANK_DUMP_MAX_SUBMISSION 2026-08-27 GSB_V4_THREATLISTS_NOKEY HTTP 403 GSB_V5_HASHLISTS_NOKEY HTTP 403 GSB_V4_SUNSET 2027-03-31 # Google notification mail; Brave issue 56023, not the docs page GSB_V4_SUNSET_BRAVE_ISSUE 56023 GSB_MIGRATION_UNDER_10MIN 60 # "as of 2021, 60% of sites that deliver attacks live less than 10 minutes" GSB_MIGRATION_STALENESS 25-30 # "around 25–30% of missing phishing protection is due to such data staleness" GSB_MIGRATION_10_MINUTES 10 OPENPHISH_ACADEMIC_ACCESS_DAYS 60 OPENPHISH_ACADEMIC_ARCHIVE_DAYS 30 OPENPHISH_PREMIUM_REFRESH_MINUTES 5 # phishing_feeds.html Premium row PHISHTANK_STATS_TIMEOUT 20s # stats.php did not answer APWG_ECX HTTP 200 # https://apwg.org/ecx CORPUS_DEFECT_DOI 10.1109/SP54263.2024.00182 # Nanayakkara DP paper stored under a phishing slug CORPUS_DEFECT_DOI_PARTS 54263 00182 ## Z2. PAPER LITERALS on the page that are not in detection.prevalence These were read from paper.cols.txt on 2026-08-27 and printed so check_page_numbers.mjs can see them. PAPER_LITERAL crawlphish_artificial_sites 150 PAPER_LITERAL crawlphish_clickthrough_edge 21 of 50 PAPER_LITERAL phishpedia_feed_urls 350K # paper writes "350K phishing URLs"; page says 350k PAPER_LITERAL phishpedia_feed_urls_n 350000 PAPER_LITERAL phishinpatterns_feed_urls 56027 PAPER_LITERAL phishinpatterns_crawled 51859 PAPER_LITERAL phishinpatterns_multipage 23446 45% of 51859 PAPER_LITERAL lin2022_fingerprint_rates 73.98 90.08 91.36 # phishing sites with JS traces, three datasets PAPER_LITERAL moura2024_domains 28754 # 28,754 phishing domains across .nl / .ie / .be PAPER_LITERAL moura2024_nl_24h 80% # After 24h, 80% of the .nl domains are mitigated PAPER_LITERAL moura2024_ie_24h 70% # roughly 70% of the .ie are mitigated PAPER_LITERAL phisheye_lifetime_cut 5 95 # first victim = 5% of connections, last = 95%; eight-day estimate is 10 days minus 2 PAPER_LITERAL teoh_controlled_per_type 100 # Table 1: 0/100 on VT, GSB, SmartScreen for each of 5 CAPTCHA types PAPER_LITERAL teoh_field_sole 66 of 869 discovered solely by PhishDecloaker ## I. Arithmetic derived on the page PAGE_N / union = 52 / 139 = 37.4% excluded / union = 87 / 139 = 62.6% PhishTank papers / union = 21.6% PAGE_N cloaking role = 8 PAGE_N live-site role = 29 PAGE_N detector role = 13 PAGE_N kit role = 2 user-study in union = 29 feed-gt in union = 19 PhishTank dump Other share = 65457/73660 = 88.9% --- end ---
