Table of Contents
Provenance: Messaging Groups and Channels
Working notes behind messaging_channels: every probe with its population, the hand audit that produced the 26-paper population, the hand codes behind every “N of 26” on the page, the quote and figure checks, the external sources with fetch dates and the ones rejected, and the decisions taken along the way. Corpus-level caveats — the seven-venue scope, the provisional 2025–2026 slice, extraction stability — are on corpus and are not repeated here.
This page is a log, not prose. It is for somebody checking a number.
The run
| Item | Value |
|---|---|
| Date | 2026-09-27 |
| Corpus | data/extract/run1/extractions.jsonl, 5,859 extracted papers, 5,855 with paper.cols.txt (the 2026-08-11 extension, commit 8a6b843). The ~20% free-text stability and 0.9% unlocatable-quote figures used on corpus were measured on the older 4,322-paper run; this page uses neither. |
| Venues | CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P, 2010–2026 |
| Scope check | sitemap.mjs on 2026-09-27: 195 pages, no page on messaging groups or channels, no red link promising one. platforms (rev 1790286327) had Telegram only as rank 9 of its platform table and no route row for joining groups. Created a new child page; nothing overlapped. |
| Page created | design:platforms:messaging_channels, first save rev 1790510475 (54,672 bytes) |
| Report script | scripts/report_messaging_channels.mjs + scripts/msgch_probes.mjs (probes) + scripts/msgch_fold.mjs (rule, verdicts, hand codes) — all three and the output are below, unedited |
| Figure and quote verifier | scripts/verify_messaging_channels_figures.mjs — output below |
| External checks | scripts/external_checks_messaging_channels.sh — script and output below; exit status is the number of FAILED checks |
| Number guard | check_page_numbers.mjs whole-page against the concatenation of the four outputs below: every figure traces. That guard is a presence check, not a binding check; figures were also read against the output by hand and by the figures reviewer. |
| Bibliography | 27 entries appended to bibliography at rev 1790510451 (1,135 → 1,162 entries); generated by bibgen.mjs from the corpus index, USENIX author lists fetched from the landing pages by fetch_authors.py; bib_dedup_scan.py: 0 definite duplicates, and every candidate pair involving a new key checked by hand as a different paper; no literal @ inside a field; ?purge=true issued before the citing page was saved |
| Agents | One Opus session (queries, triage of 103 candidates, verdicts, hand codes, drafting, verification). Four Sonnet readers read the 44 candidates that looked like group studies or neighbours, into unpublished notes (brief below). One Opus sub-agent did the external-source pass (71 claims); every load-bearing claim was then re-fetched by the external-check script. Four review sub-agents — see Review log. |
Mistakes made in this run
- A citekey was guessed from the title and was wrong. The draft cited Detecting and Understanding the Promotion of Illicit Goods and Services on Twitter as
xu2025_detecting; the first author is Hongyu Wang, andbibgen.mjsproducedwang2025_detecting. Caught by the unresolved-key check before saving. - Three figures in the first draft did not match the hand codes they came from. “Telethon: 6 papers, 5 of them 2024–2026” (all six are); “directories: 7 papers, 4 of them 2024–2026” (six are); “two papers release data on request” (the extraction says five, plus three restricted). All three were written from memory of the codes rather than from the report; the report now prints the year list behind every code the page dates (section E), and the page was corrected from it.
- The draft over-generalised two paper quotes. It credited “two criminology-trained groups” with a consent-waiver argument that only one paper makes in those terms, and wrote “no paper here describes entering a group by deception” without saying that one of the 26 interacted with the operators it studied. Both rewritten against the text.
- The external-check script silently skipped its last section on the first run.
set -uaborted on an unset$1inside the Crossref loop, and the run printed nine FAILED lines above it — all nine were my regexes (tag-stripped text searched for values that live in HTML attributes), not the sources. Fixed; the script now reaches its finalFAILED checks: Nline, and exits with N. - Two hand codes behind headline figures were wrong until review. [1Vu, Anh V.; Collier, Ben; Thomas, Daniel R.; Kristoff, John; Clayton, Richard; Hutchings, Alice (2025): "Assessing the Aftermath: the Effects of a Global Takedown against DDoS-for-hire Services", in: Proceedings of the USENIX Security Symposium. (Link)] also monitored Discord channels (Discord 5 → 6), and [2Gao, Yiran; Xia, Pengcheng; Wang, Liu; Liu, Tianming; Wang, Haoyu (2026): "Doxing-as-a-Service: Demystifying the Chinese Online Doxing Ecosystem", in: Proceedings of the ACM Web Conference. (DOI)]'s only denominator caveat concerns a secondary sample, so “all nine core papers acknowledge the denominator” was eight. Three more codes (two join modes, one ethics review) were recoded; see Review log.
- The external-source sub-agent said TeraGram had no DOI yet. It is in the ICWSM 2026 proceedings; the external-currency reviewer caught it.
- The page missed an application route. It said no platform offers researchers a way in while its own parent page documents that Meta's Content Library covers WhatsApp Channels. The generic reviewer caught it by reading the two pages together.
- A located note quote is not a published quote. The readers' notes carry 606 quotes; 501 (82.7%) were located verbatim. The page quotes only strings that the page's own verifier locates, and the notes stay unpublished.
No credential or token was printed, logged or exposed. GH_TOKEN is referenced in the external script only as ${GH_TOKEN:+x} and inside an Authorization header.
No ~~DISCUSSION~~ block on this page, following the default set on ad_archives: comments belong on the content page.
Judgement calls
- A child page, organised by route. platforms refuses per-company pages and is organised by route in; joining groups is a route it did not cover. The page is about the route across Telegram, WhatsApp and Discord (and other messengers' groups, which the probes also look for), not about the companies.
- Account enumeration is a pointer section, not part of the method. The four contact-discovery papers measure accounts against the numbering plan, not groups against a seed; they share no discovery, joining or denominator logic with the route. They stay on the page as a short section because their rate-limit and disclosure history (2012 → 2021 → 2026) is the clearest record of how a messenger responds to measurement at scale, and platforms already carries the largest of them.
- The inclusion rule counts metadata. Practical Traffic Analysis Attacks on Secure Messaging Applications (NDSS 2020) joined over 1,000 public Telegram channels to record message timing and sizes for a traffic model. The reader coded it OUT (traffic analysis); I coded it IN-section, because the rule says “messages, members, media or metadata” and the rule was written first. A reader who prefers “content only” should subtract one: 25 papers, 16 section.
- Two other reader verdicts were overridden, each recorded in its verdict note in
msgch_fold.mjs: the NDSS 2024 scam-baiting paper (reader IN-section → CONTEXT dm: one group entered on a scammer's invitation is incidental) and the USENIX 2025 WeChat-groups paper (reader CONTEXT user-study → OUT recruitment: the groups were joined to recruit interviewees). One was questioned and kept: the CCS 2020 impersonation-as-a-service paper (reader OUT; it collected a handful of examples from a Telegram channel reached through the marketplace it studied, which the rule could be read to cover; kept OUT as incidental). - Public-channel reading counts as the route. Six papers say they read public Telegram channels without saying whether they joined, and six more do not say how they got in at all. The page keeps them in the population and reports the join mode separately, rather than defining them out — a channel read through Telethon or the web preview is still content from inside the platform's group layer, found by a seed, with the same denominator problem.
- Commercial scraper services count. Two papers bought Telegram channel data from Apify/Telemetrio. The route is the same; the seed is the vendor's and cannot be described, which the page says.
- “Current” verdicts. “Historical in these venues” for WhatsApp group monitoring rests on a corpus fact (no WhatsApp group collection after 2021 in the seven venues) and one outside methods paper proposing donation instead; it is not a claim that nobody does it anywhere. “Historical as a route” for Twitter link harvesting rests on the X API change documented on platforms. “Forbidden” for Discord user-account automation rests on Discord's own guidelines, support article and developer policy, fetched 2026-09-27.
- Hand codes are single-coder. Verdicts and codes were decided by one session from the reader notes and the text; there is no second coder and no agreement figure. The page says so.
Queries, with populations and denominators
Denominators used
| Figure on the page | Denominator | Where |
|---|---|---|
| 40 / 41 / 13 papers naming Telegram / WhatsApp / Discord ≥ 10 times | 5,855 papers with full text | report section A |
| 147 candidates, 26 in the population | 144 probe candidates + 3 recall additions | sections A–B |
| precision / recall of each probe | the probe's own set; recall over the 26 | section C |
| N of 26 for every hand code | the 26 IN papers | section E |
| per 1,000 corpus papers | all corpus papers in the year bucket | section D |
| 17 of 26 state a review outcome vs 33.8% | the 26 (all empirical) vs 5,118 empirical papers — the extraction's reading; the hand codes give 16, the difference being [3Acharya, Bhupendra; Lazzaro, Dario; Cinà, Antonio Emanuele; Holz, Thorsten (2025): "Pirates of Charity: Exploring Donation-based Abuses in Social Media Platforms", in: Proceedings of the ACM Web Conference. (DOI)] | section G |
| 5 / 3 / 9 on request / restricted / public | the 26 vs 5,118 empirical papers | section G |
| 24 vs 21 Telegram papers | platforms_report.mjs --list Telegram vs the 21 Telegram IN papers | section I |
Probe 1: reproducing the gap pass
The task brief's counts — Telegram 40 (23 from 2024–2026), WhatsApp 41 (18), Discord 13 (8), “≥ 10 mentions of the platform name” — reproduce exactly with the case-sensitive patterns \bTelegram\b, \bWhatsApp\b, \bDiscord\b at ≥ 10 hits, on either paper.norm.txt or paper.cols.txt (scripts/_msgch_gapcheck.mjs). Case-insensitive gives 43 / 53 / 15; the extra WhatsApp papers write “Whatsapp”. Discord stays case-sensitive in every probe because lowercase “discord” is an English noun.
Probe 2: the candidate set
Defined in msgch_probes.mjs (below). A paper is a candidate if any of: (a) a platform name ≥ 10 times (case-insensitive for Telegram and WhatsApp); (b) any hit of the route vocabulary — invite-link URL patterns, “invite link”, Telethon, TDLib, Pyrogram, TGStat, Telemetr, WhatsApp Web, whatsapp-web.js, Baileys, yowsup, discord.py, self-bot, DiscordChatExporter, Disboard, “<platform> group/channel/server” — and ≥ 3 platform mentions; © a platform named in the title, a population source list, a detection phenomenon, or a tool the paper used or produced. 144 papers; channel overlap in section A.
Probe 3: recall outside the candidate set
scripts/_msgch_recall.mjs: over the 372 non-candidates that name any of the three at least once, a first-person collection verb (join, collect, scrape, crawl, monitor, subscribe, track, gather, download) within 80 characters of a messenger name — 3 hits. Plus OTHER_RE (other messengers' groups: Signal, Viber, WeChat, LINE, Slack, Kik, GroupMe, Matrix, Element, QQ, KakaoTalk, Snapchat followed by group/channel/chat/server) at ≥ 2 hits. Three papers were added by hand as RECALL_ADDED; none turned out to be in the population. This probe is narrow on purpose (a loose “Telegram channel” probe returns every cybercrime paper that mentions one); its recall is not measured.
Probe 4: reconciliation with the parent page
Section I runs platforms_report.mjs --list Telegram (the parent's four-signal subject rule) and compares titles with the 21 Telegram IN papers: 17 shared, 4 here only, 7 there only, each listed with this page's verdict.
The hand audit
- Inclusion rule, written before any verdict was counted: printed at the top of section B and in
msgch_fold.mjs. - 44 candidates read in full (
READ_IN_FULLin the fold; every IN paper is among them, and the report fails if one is not). Four Sonnet readers, one brief (below), one markdown file each, every factual field with a verbatim quote. Their verdicts were reviewed against the quotes; three were overridden (see Judgement calls). - 103 candidates decided from sentence contexts:
_msgch_ctx.mjsprinted, for every candidate, up to six contexts around a platform name near a collection verb and around every route-vocabulary hit. The dominant OUT codes — passing mention 29, user study of messengers in general 22, recruitment channel 19, app analysis 14, traffic 10 — are visible in the context sentences without reading further. The risk in this step is a group study whose contexts look like a mention; the recall probe and the 144-wide candidate set are the only defences, and neither is a second reading. - Hand codes for the 26 (
HANDin the fold): platforms, object family, discovery seed, join mode, history back-fill, member data, tooling, limits, ethics review, passive-observer statement, minimisation, and whether the paper acknowledges its denominator, plus the found / collected / message counts shown in the page's paper table. “not-stated” means the paper does not say. - The invariant: the report throws if a candidate lacks a verdict, a verdict names a non-candidate, an IN paper lacks hand codes, hand codes exist for a non-IN paper, or a hand-code record lacks a field.
Quote and figure checks
- Page quotes:
verify_messaging_channels_figures.mjspulls every//"…"//span out of the page source (68 at the last run) and requires each to be located in the paper of a citekey on the same line, or to be an EXTERNAL span whose external check printed OK, or to be on a two-entry NOT_A_QUOTE list. One span needed a ligature fallback (the IMC 2020 PDF has an unmappedffglyph in “different” in both renderings), and one a de-hyphenation fallback (“self- constructed” across a line break); both routes are labelled in the output. - Per-paper figures: 83 needles, located in
.colsor, for twelve, only in thepypdfre-extraction. Three mutated needles (250→350, 0.8%→0.9%, 4,709→4,790) must not be found and are not. - Two numbers are verified in the form the PDF prints them: [2Gao, Yiran; Xia, Pengcheng; Wang, Liu; Liu, Tianming; Wang, Haoyu (2026): "Doxing-as-a-Service: Demystifying the Chinese Online Doxing Ecosystem", in: Proceedings of the ACM Web Conference. (DOI)]'s “411, 707” and [4Saha, Punyajoy; Mathew, Binny; Garimella, Kiran; Mukherjee, Animesh (2021): ""Short is the Road that Leads from Fear to Hate": Fear Speech in Indian WhatsApp Groups", in: Proceedings of the ACM Web Conference. (DOI)]'s “8, 000 … 1, 000 … 3, 000” carry a stray space inside the number in every rendering.
- Reader notes:
msgch_notes_quotecheck.py(output below) — 501 of 606 quotes located (82.7%). The misses are dominated by readers stitching two sentences with an ellipsis, against the brief, and by column splices. No page quote depends on an unlocated note quote. - Attribution guard:
check_attributions.mjswas not run — the page names no “X et al.” in the text; every author named in a footnote (Garimella and Tyson, Baumgartner et al., La Morgia et al., Golovin et al., Barbosa and Milan, Garimella and Chauchard) is checked against Crossref or arXiv metadata by the external script.
External sources
Every load-bearing external claim is re-fetched by external_checks_messaging_channels.sh (script and output below); the output prints the matched context for each. Pages that refuse curl were fetched with Playwright's headless shell: WhatsApp's legal pages (HTTP 400 to curl), Discord's support and developer-policy pages (403), Disboard (403). EUR-Lex was not needed.
| Claim on the page | Primary source |
|---|---|
| Telethon GitHub archived, moved to Codeberg; 1.45.0 on PyPI 2026-09-10; no 2.x | GitHub API, raw README, PyPI JSON, Codeberg API |
| Pyrogram and GramJS archived; Kurigram and teleproto are the forks | GitHub API, PyPI, npm |
| TDLib tags stop at v1.8.0; master declares 1.8.x | GitHub API, CMakeLists.txt |
whatsapp-web.js moved to wwebjs; its README's warning; Baileys and whatsmeow maintained; whatsmeow has only pseudo-versions; yowsup's last commit 2021 | GitHub API, raw README, npm, Go module proxy |
| discord.py 2.7.1; DiscordChatExporter's README warning | PyPI, raw README |
| Telegram content-licensing, general ToS and API ToS §1.5 clauses | telegram.org/tos/content-licensing, telegram.org/tos, core.telegram.org/api/terms |
| one api_id per number; unofficial-client logins “under observation” | core.telegram.org/api/obtaining_api_id |
join cap 500 / 1,000; recommendations capped at 10; CHANNELS_TOO_MUCH; FLOOD_WAIT_X | core.telegram.org/api/config, /method/channels.joinChannel, /api/errors |
public-channel web preview documented; t.me/s/ serves posts with ?before= pagination | telegram.org/tour/channels; t.me/s/telegram |
| People Nearby removed 2024-09-06; search clean-up and disclosure change 2024-09-23 | t.me/durov/343, t.me/durov/345 (Durov's own channel) |
| Telegram below the VLOP threshold; Telegram and Discord absent from the designation list | telegram.org/tos/eu-dsa; the Commission's list |
| WhatsApp designated a VLOP on 2026-01-26 because of Channels; private messaging excluded | Commission press release |
| WhatsApp Channels global launch 2023-09-13 | Meta newsroom |
| WhatsApp ToS “automated or other means” and “collect information of or about our users” | whatsapp.com/legal/terms-of-service (Playwright) |
| Discord ToS, guidelines (self-bots) and developer policy (mine or scrape; ML training) | discord.com/terms, discord.com/guidelines, support-dev.discord.com (Playwright) |
| Discord invites 7 days by default; Server Discovery 1,000 members | support.discord.com (Playwright) |
| TGStat “More than 2 864 885 channels and groups”; Disboard is a self-listing directory | tgstat.com; disboard.org (Playwright) |
| Pushshift Telegram, TGDataset, Garimella–Tyson, Barbosa–Milan, Garimella–Chauchard | Crossref records (DOI, title, venue, date, authors printed) |
| TeraGram figures and first author | arXiv abstract page and citation_author meta |
| AoIR 3.0 approval date and absence of any messaging-group passage | aoir.org PDF, full-text search |
| Barbosa and Milan: “avoid by all means covert bypasses” | westminsterpapers.org article page |
Stated from a secondary copy, and labelled so on the page
- Telemetr.io's catalogue size (“11M+” and “7M+” on the same page): the live site is Cloudflare-walled to curl, Playwright and WebFetch; read from a Web Archive capture of 2026-09-26. The page gives both numbers and says they are inconsistent.
Sources rejected
- techcrunch.com, engadget.com, techpolicy.press, Wikipedia — for the WhatsApp Channels launch and Telegram's EU user numbers. Primary sources were reachable. The press claim that the Commission doubts Telegram's self-reported numbers was not used: no Commission statement was found.
- fastsocial.co, peakbot.pro, discordify.net, cybrancee.com, geeksforgeeks.org — SEO how-to pages on Discord Discovery and invites; Discord's own help pages used instead.
- hostafrica.com, thecondia.com, today.com, nepalnews.com, yourstory.com, gsmarena.com, wgg-agency.com — press and blog summaries of WhatsApp Channels.
- researchgate.net copies of papers, awesomepapers.io — not the publisher of record.
- darcmode.org (a volunteer Discord research community) — relevant as a community, but not evidence of any Discord research-access policy.
What could not be established
- When Telegram introduced its content-licensing terms. The page carries no date. The page quotes the terms and says it could not date them.
- When Discord's default invite expiry changed from the one day [5Hoseini, Mohamad; Melo, Philipe; Junior, Manoel; Benevenuto, Fabrício; Chandrasekaran, Balakrishnan; Feldmann, Anja; Zannettou, Savvas (2020): "Demystifying the Messaging Platforms' Ecosystem Through the Lens of Twitter", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] describes to the seven days Discord's help page states. No dated Discord announcement was found.
- Whether WhatsApp runs a DSA Article 40 data-access process for Channels. Designation verified; process not checked.
- Whether Discord or Telegram has any research-access route. None found on their own pages; for Discord, the absence rests on the policy pages plus a web search, which is weaker than an exhaustive check.
- How many groups exist on any platform. No paper in the corpus estimates it; the page proposes capture–recapture as an open question.
- The recall of the candidate set against a truly independent sample. The recall probe is narrow; a group study that never names Telegram, WhatsApp or Discord and never uses the route vocabulary would be missed. At least one such paper (WeChat groups, USENIX 2025) turned up, and it was a recruitment use.
- A second coder's agreement on the verdicts and hand codes. Not done.
Follow-ups filed and changes to other pages
- platforms (rev 1790286327 → 1790510558): route-table row, currency row, a paragraph in Should There Be … Pages?, and a Related Pages entry for the new child.
- design (rev 1790032321 → 1790510524): table row, “18 pages” → 19, “the only child page” → two child pages, “Re-derived 2026-09-27”.
report_namespace_overviews.mjsre-run:design 19 19 equal. - ethics: a Related Pages pointer to the new page; a real section on entering groups is filed as a separate work item (see the review log), because it needs its own population across the whole corpus, not just the 26 papers here.
Review log
Four reviewers, each told that the author's context may not be exhaustive and each handed the page, the provenance page, the report script and its output, the verifier and the external checks. The three focused passes ran in parallel on the frozen snapshot (content rev 1790510757, provenance rev 1790510789); nothing was edited while they ran. Their findings were applied together, then the generic pass ran on the corrected pages. Findings files: notes/msgch_review_*.md in the workdir.
Pass 1 — figures against the script (model: sonnet)
Re-ran the report: byte-identical to the committed output. 5 findings.
| # | Finding | Disposition |
|---|---|---|
| F1 | high — the page says [1Vu, Anh V.; Collier, Ben; Thomas, Daniel R.; Kristoff, John; Clayton, Richard; Hutchings, Alice (2025): "Assessing the Aftermath: the Effects of a Global Takedown against DDoS-for-hire Services", in: Proceedings of the USENIX Security Symposium. (Link)] monitored booter Discord channels, and the paper says so, but the hand code listed Telegram only; the Discord count (5) and the multi-platform count (2) were therefore one short | accepted. Code fixed; Discord 6, multi-platform 3; tip box and currency table corrected |
| F2 | medium — the tooling sentence enumerated every category except the PumpOlymp API (1 paper), so its parts did not add up to the report | accepted, category added |
| F3 | low — the verifier's footer claimed the TGStat Russia tile and the Telemetr.io Web Archive numbers were re-fetched, but no check covered them | accepted: three checks added to the external script (Russia tile, both Telemetr.io numbers from the capture) |
| F4 | low — the page said “at least three mentions of a platform” where the code sums the three platforms (3 borderline candidates, none in the population) | accepted, page reworded to match the code rather than the code changed |
| F5 | low — the needle for [3Acharya, Bhupendra; Lazzaro, Dario; Cinà, Antonio Emanuele; Holz, Thorsten (2025): "Pirates of Charity: Exploring Donation-based Abuses in Social Media Platforms", in: Proceedings of the ACM Web Conference. (DOI)]'s 85,402 did not contain the number | accepted, needle now “(76,111/85,402) from Telegram were filtered” |
Pass 2 — citations and quotes (model: sonnet)
All 37 citekeys resolve; no duplicate keys or DOIs; all 15 URL-only entries (13 USENIX, 2 NDSS) plus wang2025_detecting match the papers' author blocks. 14 findings (4 medium, 10 low).
| # | Finding | Disposition |
|---|---|---|
| C1 | medium — the 196-channel takedown was presented as the outcome of reporting the 339 monitored channels; it is the outcome for 196 new channels found through links on Telegram and Facebook, while the monitored channels' outcome is “only 64 channels (19%) were removed” | accepted; both outcomes now stated, with their denominators, in What to Read First and the results table |
| C2 | medium — join = public-read for [6Xu, Jiahua; Livshits, Benjamin (2019): "The Anatomy of a Cryptocurrency Pump-and-Dump Scheme", in: Proceedings of the USENIX Security Symposium. (Link)] rests on the organiser's public channel, not on what the researchers did | accepted, recoded not-stated |
| C3 | medium — same for [7Vu, Anh V.; Hutchings, Alice; Anderson, Ross J. (2024): "No Easy Way Out: the Effectiveness of Deplatforming an Extremist Forum to Suppress Hate and Harassment", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] (“Both channels permit public access” describes the channels) | accepted, recoded not-stated; the join counts became 11 / 6 / 2 / 1 / 6 and the page now explains the distinction |
| C4 | medium — irb = not-required for [3Acharya, Bhupendra; Lazzaro, Dario; Cinà, Antonio Emanuele; Holz, Thorsten (2025): "Pirates of Charity: Exploring Donation-based Abuses in Social Media Platforms", in: Proceedings of the ACM Web Conference. (DOI)]: the paper says only that it “did not directly involve interaction with any human subjects” | accepted, recoded not-stated; “say nothing about review” 9 → 10 in three places; the page now explains why the extraction's 17 is one higher than the hand codes' 16 |
| C5 | medium — denominator = acknowledged for [2Gao, Yiran; Xia, Pengcheng; Wang, Liu; Liu, Tianming; Wang, Haoyu (2026): "Doxing-as-a-Service: Demystifying the Chinese Online Doxing Ecosystem", in: Proceedings of the ACM Web Conference. (DOI)]: its only caveat is about a secondary 15-group sample; for the main top-100 dataset it calls the corpus “representative” | accepted, recoded not-stated. This broke a headline: “all nine core papers acknowledge” became eight of nine, and the page says which one does not |
| C6 | low — “refugee support group” is the page's gloss | accepted, now “a refugee participants' WhatsApp group” |
| C7 | low — the NDSS 2024 scam-baiting paper literally satisfies the IN rule for one group | rejected, already recorded: kept CONTEXT dm as a documented judgement call; the page says it entered one group on a scammer's invitation |
| C8–C14 | low — seven client-table clauses had no re-fetch in the external script (TDLib tags, Pyrogram README wording, Kurigram fork and currency, GramJS → teleproto on npm, yowsup's Python ceiling, discord.py upload date, DiscordChatExporter 2.48) | accepted: eight checks added; all pass. The page's “re-fetches every row” is now true |
Pass 3 — external currency (model: sonnet)
Independently re-fetched every tool state, term quote, API limit, designation, directory and dataset. 2 findings.
| # | Finding | Disposition |
|---|---|---|
| E1 | medium — TeraGram is no longer only an arXiv preprint: Crossref records it in the ICWSM 2026 proceedings (doi 10.1609/icwsm.v20i1.42783, issued 2026-05-25) | accepted; footnote cites the DOI, and a Crossref check was added. The external-source sub-agent had reported “arXiv; no DOI yet” — its search was wrong, which is exactly what this pass exists to catch |
| E2 | low — Discord's help page conditions never-expiring invites on Community servers | accepted, qualifier added |
Pass 4 — generic (model: fable)
Ran on the corrected pages (content rev 1790511893, provenance rev 1790511895). 16 findings (1 high, 6 medium, 9 low); it also re-derived several of the fixes above and found them sound.
| # | Finding | Disposition |
|---|---|---|
| G1 | high — the page said none of the three platforms has a research route and called WhatsApp Channels “new and untested”, while platforms documents that Meta's Content Library covers WhatsApp Channels | accepted. Re-fetched Meta's page (headless browser): “public content archive from Facebook, Instagram and WhatsApp Channels”. The intro box, the terms section, the currency table and the open questions now say that Channels have an application route, that groups do not, and that no paper here used it; an external check was added |
| G2 | medium — “specific, recent” terms applied to 2019–2024 papers when the page cannot date the clause | accepted: “recent” dropped; “read against today's terms”, and the page says whether the clauses existed at collection time was not established |
| G3 | medium — the artefact sentence gave base rates only for the two rows where group studies exceed the corpus | accepted: public 34.6% vs 47.7% added; the sentence now says they share less openly and shift to gated release |
| G4 | medium — “the only compliant path [on Discord] is a bot an admin invites” ignores a person reading as a member, which the page's own codes record | accepted, rewritten in the terms section and the open questions |
| G5 | medium — “still possible at scale in 2025” generalised one platform's pre-mitigation window to the whole route | accepted: now “capped on most platforms”, with WhatsApp's window, Meta's mitigation statement and the capped platforms named |
| G6 | medium — “what the careful papers converge on” rests on one paper for two bullets, and one is contradicted by [8Saha Roy, Sayak; Pourabbas Vafa, Elham; Khanmohamaddi, Kobra; Nilizadeh, Shirin (2025): "DarkGram: A Large-Scale Analysis of Cybercriminal Activity Channels on Telegram", in: Proceedings of the USENIX Security Symposium. (Link)], which sent 11,800 collected posts to the GPT-4 API | accepted: reframed as what individual papers did; the third-party-service bullet is marked contested, cites both papers, and states the page's position |
| G7 | medium — “what the 2020–2024 Discord papers describe or imply” attributes user-account automation to six papers when only one is coded that way | accepted: names [5Hoseini, Mohamad; Melo, Philipe; Junior, Manoel; Benevenuto, Fabrício; Chandrasekaran, Balakrishnan; Feldmann, Anja; Zannettou, Savvas (2020): "Demystifying the Messaging Platforms' Ecosystem Through the Lens of Twitter", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] only |
| G8 | low — Twitter as a seed surface is still used (2025); only the free API is historical | accepted, currency row split accordingly |
| G9 | low — “known official channels” vs the fourth n/a paper reached through a forum advert | accepted: “known groups or channels” |
| G10 | low — the recruitment row named two platforms where the 19 include Telegram and WeChat; two cells had no citation | accepted: platforms listed; the two uncited cells point to the verdict list |
| G11 | low — the seed table omitted the “not stated” row (2 papers) | accepted, row added |
| G12 | low — “almost all of it is Telegram” (81%) | accepted: “four papers in five” |
| G13 | low — the Discord 100-server cap rests on a 2020 measurement while Telegram's is documented | partly accepted: both measured caps are now dated “in 2020” and the Discord one is marked “not re-checked against Discord's current documentation”; Discord's support page was not re-fetched for a current figure |
| G14 | low — the page explains Discord bots' limits but not Telegram bots' | accepted: a table row on bot accounts, with Telegram's own privacy-mode sentence, and an external check |
| G15 | low — message-level attrition (deletions between back-fill and live polling) not mentioned | accepted: [9Kireev, Klim; Mykhno, Yevhen; Troncoso, Carmela; Overdorf, Rebekah (2025): "Characterizing and Detecting Propaganda-Spreading Accounts on Telegram", in: Proceedings of the USENIX Security Symposium. (Link)] added — export history “does not contain deleted messages”, and moderators removed from below 20% to over 80% of propaganda messages |
| G16 | low — no read-first entry for Discord or for the commoner “one source among several” case | accepted: [10Shen, Xinyue; Chen, Zeyuan; Backes, Michael; Shen, Yun; Zhang, Yang (2024): ""Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] and [1Vu, Anh V.; Collier, Ben; Thomas, Daniel R.; Kristoff, John; Clayton, Richard; Hutchings, Alice (2025): "Assessing the Aftermath: the Effects of a Global Takedown against DDoS-for-hire Services", in: Proceedings of the USENIX Security Symposium. (Link)] named under What to Read First |
Not re-run after these fixes: the generic pass's own findings were applied as stated, every changed figure re-checked by the number guard, and every new quote by the verifier (68 spans, 83 needles, 0 not located) and the external script. A second generic pass was not run.
What the review layer was worth
37 findings over four passes; 35 accepted in full, 1 in part, 1 rejected as an already-recorded judgement call. The two that changed headline figures came from the focused passes (the missing Discord code, and the one core paper that does not acknowledge its denominator). The one that changed what the page tells a reader to do — an application route to WhatsApp Channels the page had missed — came from the generic pass, reading this page against its parent.
The report script
- report_messaging_channels.mjs
// report_messaging_channels.mjs — every corpus figure on design:platforms:messaging_channels // ("Messaging Groups and Channels"), with its denominator. The candidate set comes from // msgch_probes.mjs, the population from the hand verdicts in msgch_fold.mjs; this script throws if // the two diverge in either direction, or if an IN paper has no hand codes. // node scripts/report_messaging_channels.mjs > scripts/report_messaging_channels-output.txt // node scripts/report_messaging_channels.mjs --list (every verdict, one per line) import { loadExtractions, POPULATIONS, pct, table, isSentinel } from './lib.mjs'; import { runProbes, isCandidate, NAME, NAME_THRESH, key, readCollapsed } from './msgch_probes.mjs'; import { VERDICTS, RECALL_ADDED, HAND, INCLUSION_RULE, CODES, HAND_FIELDS, READ_IN_FULL } from './msgch_fold.mjs'; const die = (m) => { console.error('FAILURE: ' + m); process.exit(1); }; const rows = loadExtractions(); const byKey = new Map(rows.map((p) => [key(p), p])); if (rows.length !== 5859) die(`corpus contract: expected 5,859 records, got ${rows.length}`); const venues = new Set(rows.map((p) => p.venue)); if (venues.size !== 7) die(`corpus contract: expected 7 venues, got ${[...venues]}`); console.log('design:platforms:messaging_channels — report'); console.log(`corpus: ${rows.length} extraction records, venues ${[...venues].sort().join(', ')}`); // ---------------------------------------------------------------- A. probes const { hits, missing } = runProbes(rows); console.log(`\n== A. Probes (full text = paper.cols.txt, whitespace-collapsed; ${hits.length} papers with text, ${missing} records without a text file) ==`); const CS = { telegram: /\bTelegram\b/g, whatsapp: /\bWhatsApp\b/g, discord: /\bDiscord\b/g }; // The gap pass's rule: case-sensitive name >= 10. Recomputed here rather than trusted. const gap = {}; for (const r of hits) { const t = readCollapsed(r.p); r.cs = Object.fromEntries(Object.entries(CS).map(([n, re]) => [n, (t.match(re) || []).length])); } const aRows = []; for (const n of Object.keys(CS)) { const csSet = hits.filter((r) => r.cs[n] >= NAME_THRESH); const ciSet = hits.filter((r) => r[n] >= NAME_THRESH); gap[n] = csSet; aRows.push([n, csSet.length, csSet.filter((r) => r.p.year >= 2024).length, ciSet.length, ciSet.filter((r) => r.p.year >= 2024).length, hits.filter((r) => r[n] >= 1).length]); } console.log(table(['platform', 'cs>=10 (gap rule)', 'of which 2024-26', 'probe>=10', 'of which 2024-26', 'any mention'], aRows)); const gapUnion = new Set(Object.values(gap).flat().map((r) => r.key)); console.log(`gap-rule union (any of the three, case-sensitive >= 10): ${gapUnion.size} papers`); const cand = hits.filter(isCandidate); const candKeys = new Set(cand.map((r) => r.key)); const ch = { name: 0, route: 0, schema: 0 }; for (const r of cand) { if (r.telegram >= NAME_THRESH || r.whatsapp >= NAME_THRESH || r.discord >= NAME_THRESH) ch.name++; if (r.route >= 1 && r.telegram + r.whatsapp + r.discord >= 3) ch.route++; if (r.schema.length) ch.schema++; } console.log(`candidate set: ${cand.length} papers (enters by name>=10: ${ch.name}; by route vocabulary + >=3 mentions: ${ch.route}; by schema field: ${ch.schema}; channels overlap)`); for (const k of RECALL_ADDED.map((x) => x[0])) { if (candKeys.has(k)) die(`RECALL_ADDED key is already a candidate: ${k}`); if (!byKey.has(k)) die(`RECALL_ADDED key not in extraction: ${k}`); } console.log(`recall additions (hand-added from the recall probes, each with its reason): ${RECALL_ADDED.length}`); for (const [k, why] of RECALL_ADDED) console.log(` + ${k} — ${why}`); // ---------------------------------------------------------------- B. verdict invariant const V = new Map(); for (const v of VERDICTS) { if (V.has(v[0])) die(`duplicate verdict ${v[0]}`); V.set(v[0], v); } const universe = new Set([...candKeys, ...RECALL_ADDED.map((x) => x[0])]); for (const k of universe) if (!V.has(k)) die(`candidate without a verdict: ${k}`); for (const k of V.keys()) if (!universe.has(k)) die(`verdict for a paper that is not a candidate: ${k}`); for (const [k, verdict, code] of VERDICTS) { if (!['IN', 'CONTEXT', 'OUT'].includes(verdict)) die(`bad verdict ${verdict} for ${k}`); if (!(code in CODES[verdict])) die(`unknown ${verdict} code '${code}' for ${k}`); } const IN = VERDICTS.filter((v) => v[1] === 'IN'); for (const [k] of IN) if (!HAND[k]) die(`IN paper without hand codes: ${k}`); for (const k of Object.keys(HAND)) if (!IN.find((v) => v[0] === k)) die(`hand codes for a paper that is not IN: ${k}`); for (const [k, h] of Object.entries(HAND)) for (const f of HAND_FIELDS) if (!(f in h)) die(`hand codes for ${k} lack field ${f}`); console.log(`\n== B. Verdicts over ${universe.size} candidates (${cand.length} probe + ${RECALL_ADDED.length} recall) ==`); console.log(INCLUSION_RULE); const tally = {}; for (const [, verdict, code] of VERDICTS) tally[`${verdict} ${code}`] = (tally[`${verdict} ${code}`] || 0) + 1; console.log(table(['verdict code', 'papers', 'meaning'], Object.entries(tally).sort((a, b) => a[0].localeCompare(b[0])).map(([vc, n]) => { const [v, c] = vc.split(' '); return [vc, n, CODES[v][c]]; }))); const nIn = IN.length, nCore = IN.filter((v) => v[2] === 'core').length; const nCtx = VERDICTS.filter((v) => v[1] === 'CONTEXT').length, nOut = VERDICTS.filter((v) => v[1] === 'OUT').length; console.log(`IN ${nIn} (core ${nCore}, section ${nIn - nCore}); CONTEXT ${nCtx}; OUT ${nOut}; total ${VERDICTS.length}`); const inKeys = new Set(IN.map((v) => v[0])); for (const k of READ_IN_FULL) if (!V.has(k)) die(`READ_IN_FULL key without a verdict: ${k}`); if (new Set(READ_IN_FULL).size !== READ_IN_FULL.length) die('READ_IN_FULL has duplicates'); for (const k of inKeys) if (!READ_IN_FULL.includes(k)) die(`IN paper not read in full: ${k}`); console.log(`read in full: ${READ_IN_FULL.length}; decided from sentence contexts: ${VERDICTS.length - READ_IN_FULL.length}; every IN paper was read in full`); // ---------------------------------------------------------------- C. precision and recall of the probes console.log('\n== C. Precision and recall of each probe against the hand verdicts =='); const pr = (name, set) => { const s = [...set]; const tp = s.filter((k) => inKeys.has(k)).length; return [name, s.length, tp, pct(tp, s.length), `${tp}/${nIn}`, pct(tp, nIn)]; }; const cRows = []; for (const n of Object.keys(CS)) cRows.push(pr(`gap rule: ${n} cs>=10`, gap[n].map((r) => r.key))); cRows.push(pr('gap rule: union of the three', gapUnion)); cRows.push(pr('this page: candidate set', candKeys)); cRows.push(pr('candidate set + recall additions', universe)); console.log(table(['probe', 'papers', 'IN', 'precision', 'recall', 'recall %'], cRows)); const missedByGap = IN.filter((v) => !gapUnion.has(v[0])).map((v) => v[0]); console.log(`IN papers the gap rule misses (${missedByGap.length}):`); for (const k of missedByGap) { const r = hits.find((h) => h.key === k); console.log(` ${k} (Telegram ${r.cs.telegram}, WhatsApp ${r.cs.whatsapp}, Discord ${r.cs.discord})`); } const gapNotIn = [...gapUnion].filter((k) => !inKeys.has(k)); const gnTally = {}; for (const k of gapNotIn) { const v = V.get(k); const c = `${v[1]} ${v[2]}`; gnTally[c] = (gnTally[c] || 0) + 1; } console.log(`gap-rule papers that are not IN (${gapNotIn.length}), by verdict: ${Object.entries(gnTally).sort((a, b) => b[1] - a[1]).map(([c, n]) => `${c} ${n}`).join('; ')}`); // ---------------------------------------------------------------- D. shape of the IN set console.log(`\n== D. The ${nIn} IN papers: platform, year, venue ==`); const inP = IN.map((v) => byKey.get(v[0])); const plat = {}; for (const [k] of IN) for (const pl of HAND[k].platforms) plat[pl] = (plat[pl] || 0) + 1; console.log('platform (a paper can name several): ' + Object.entries(plat).sort((a, b) => b[1] - a[1]).map(([p, n]) => `${p} ${n}`).join('; ')); const multi = IN.filter(([k]) => HAND[k].platforms.length > 1).length; console.log(`papers with more than one platform: ${multi}`); const BUCKETS = [['2010–2018', 2010, 2018], ['2019–2021', 2019, 2021], ['2022–2024', 2022, 2024], ['2025–2026*', 2025, 2026]]; const corpusBy = (a, b) => rows.filter((p) => p.year >= a && p.year <= b).length; console.log(table(['bucket', 'corpus papers', 'IN', 'IN core', 'IN per 1,000 corpus papers', 'Telegram IN'], BUCKETS.map(([n, a, b]) => { const s = IN.filter(([k]) => { const y = byKey.get(k).year; return y >= a && y <= b; }); return [n, corpusBy(a, b), s.length, s.filter((v) => v[2] === 'core').length, (1000 * s.length / corpusBy(a, b)).toFixed(1), s.filter(([k]) => HAND[k].platforms.includes('Telegram')).length]; }))); console.log('* 2025–2026 is provisional: CCS 2026 and IMC 2026 not yet held; IEEE S&P 2026 and TheWebConf 2026 under-selected by construction.'); const yr = {}; for (const p of inP) yr[p.year] = (yr[p.year] || 0) + 1; console.log('per year: ' + Object.keys(yr).sort().map((y) => `${y}:${yr[y]}`).join(', ')); const vn = {}; for (const p of inP) vn[p.venue] = (vn[p.venue] || 0) + 1; console.log('per venue: ' + Object.entries(vn).sort((a, b) => b[1] - a[1]).map(([v, n]) => `${v} ${n}`).join(', ')); const obj = {}; for (const [k] of IN) obj[HAND[k].object] = (obj[HAND[k].object] || 0) + 1; console.log('object (hand-coded family): ' + Object.entries(obj).sort((a, b) => b[1] - a[1]).map(([o, n]) => `${o} ${n}`).join('; ')); // ---------------------------------------------------------------- E. hand codes over the IN set console.log(`\n== E. Hand codes over the ${nIn} IN papers (from each paper's own text; msgch_fold.mjs HAND) ==`); const multiTally = (field) => { const t = {}; for (const [k] of IN) { const vals = HAND[k][field]; for (const v of (Array.isArray(vals) ? vals : [vals])) t[v] = (t[v] || 0) + 1; } return Object.entries(t).sort((a, b) => b[1] - a[1] || a[0].localeCompare(b[0])); }; for (const f of ['discovery', 'join', 'history', 'members', 'tooling', 'limits', 'irb', 'passive', 'minimise', 'denominator']) { console.log(`-- ${f} (of ${nIn}; multi-valued fields do not sum):`); for (const [v, n] of multiTally(f)) console.log(` ${String(n).padStart(3)} ${pct(n, nIn).padStart(6)} ${v}`); } console.log('\n-- years behind the codes the page dates (a paper counts once per code):'); for (const [f, v] of [['discovery', 'directory'], ['discovery', 'in-app-search'], ['discovery', 'other-platform-links'], ['tooling', 'Telethon'], ['join', 'scraper-service'], ['platforms', 'WhatsApp'], ['platforms', 'Discord'], ['platforms', 'Telegram']]) { const ys = IN.filter(([k]) => [].concat(HAND[k][f]).includes(v)).map(([k]) => byKey.get(k).year).sort(); console.log(` ${f}=${v}: ${ys.length} papers, ${ys.filter((y) => y >= 2024).length} from 2024-2026, ${ys.filter((y) => y >= 2022).length} from 2022-2026; years ${ys.join(', ')}`); } console.log(' core papers by bucket and object:'); for (const [n, a, b] of [['2019-2021', 2019, 2021], ['2022-2024', 2022, 2024], ['2025-2026*', 2025, 2026]]) console.log(` ${n}: ` + IN.filter(([k, , c]) => c === 'core' && byKey.get(k).year >= a && byKey.get(k).year <= b).map(([k]) => HAND[k].object).join('; ')); console.log('\n-- CORE papers only (the ones whose main dataset is group/channel data):'); const coreKeys = IN.filter((v) => v[2] === 'core').map((v) => v[0]); for (const f of ['discovery', 'join', 'tooling', 'irb', 'denominator']) { const t = {}; for (const k of coreKeys) for (const v of [].concat(HAND[k][f])) t[v] = (t[v] || 0) + 1; console.log(` ${f} (of ${coreKeys.length}): ` + Object.entries(t).sort((a, b) => b[1] - a[1]).map(([v, n]) => `${v} ${n}`).join('; ')); } // ---------------------------------------------------------------- F. per-paper table console.log(`\n== F. Per-paper table (IN, ${nIn}) ==`); console.log(table(['year', 'venue', 'role', 'platforms', 'found', 'collected', 'messages', 'discovery', 'join', 'tooling', 'slug'], IN.map(([k, , code]) => { const p = byKey.get(k), h = HAND[k]; return [p.year, p.venue, code, h.platforms.join('+'), h.found, h.collected, h.messages, [].concat(h.discovery).join('+'), h.join, [].concat(h.tooling).join('+'), p.slug.slice(0, 48)]; }) .sort((a, b) => a[0] - b[0] || a[1].localeCompare(b[1])))); // ---------------------------------------------------------------- G. what the extraction says about the same papers console.log(`\n== G. The extraction schema over the IN set vs the corpus baseline ==`); const emp = rows.filter(POPULATIONS.empirical); const inEmp = inP.filter(POPULATIONS.empirical); const eth = (ps) => { const t = {}; for (const p of ps) { const v = p.ethics === null ? '(no ethics object)' : p.ethics.reviewOutcome; t[v] = (t[v] || 0) + 1; } return t; }; const eIn = eth(inEmp), eAll = eth(emp); const eKeys = [...new Set([...Object.keys(eIn), ...Object.keys(eAll)])].sort((a, b) => (eAll[b] || 0) - (eAll[a] || 0)); console.log(`ethics.reviewOutcome — IN papers that are empirical (${inEmp.length} of ${nIn}) vs all empirical papers (${emp.length}):`); console.log(table(['reviewOutcome', `IN (${inEmp.length})`, 'share', `corpus empirical (${emp.length})`, 'share'], eKeys.map((k) => [k, eIn[k] || 0, pct(eIn[k] || 0, inEmp.length), eAll[k] || 0, pct(eAll[k] || 0, emp.length)]))); const stated = (t, n) => n - (t['none-mentioned'] || 0) - (t['(no ethics object)'] || 0); console.log(`states any review outcome: IN ${stated(eIn, inEmp.length)}/${inEmp.length} (${pct(stated(eIn, inEmp.length), inEmp.length)}); corpus ${stated(eAll, emp.length)}/${emp.length} (${pct(stated(eAll, emp.length), emp.length)})`); const art = (ps) => { const t = {}; for (const p of ps) { const v = p.artifacts === null ? '(no artifacts object)' : p.artifacts.availability; t[v] = (t[v] || 0) + 1; } return t; }; const aIn = art(inEmp), aAll = art(emp); console.log(`artifacts.availability — IN empirical (${inEmp.length}) vs corpus empirical (${emp.length}):`); console.log(table(['availability', 'IN', 'share', 'corpus', 'share'], [...new Set([...Object.keys(aIn), ...Object.keys(aAll)])].sort((a, b) => (aAll[b] || 0) - (aAll[a] || 0)).map((k) => [k, aIn[k] || 0, pct(aIn[k] || 0, inEmp.length), aAll[k] || 0, pct(aAll[k] || 0, emp.length)]))); // Schema recall for the client library: does tools[] name the library the paper's own text names? const LIB = /telethon|tdlib|pyrogram|gramjs|whatsapp[- ]web|baileys|whatsmeow|discord\.py|discord\.js|discordchatexporter/i; const handLib = IN.filter(([k]) => [].concat(HAND[k].tooling).some((t) => LIB.test(t))); const schemaLib = handLib.filter(([k]) => byKey.get(k).tools.some((t) => (t.usedOrMentioned === 'used' || t.usedOrMentioned === 'produced') && LIB.test(t.name))); console.log(`client library named in the paper (hand code): ${handLib.length}; of those, named in tools[] used/produced: ${schemaLib.length}`); for (const [k] of handLib) if (!schemaLib.find((x) => x[0] === k)) console.log(` schema misses the library: ${k} (hand: ${[].concat(HAND[k].tooling).join('+')})`); // ---------------------------------------------------------------- H. CONTEXT papers, listed console.log('\n== H. CONTEXT papers, listed with code and note =='); for (const [k, , code, note] of VERDICTS.filter((v) => v[1] === 'CONTEXT').sort((a, b) => a[2].localeCompare(b[2]))) console.log(` ${code.padEnd(12)} ${k}\n ${note}`); // ---------------------------------------------------------------- I. reconciliation with design:platforms // That page's Telegram row is platforms_report.mjs --list Telegram (a four-signal subject rule over the schema). { const { execFileSync } = await import('node:child_process'); const out = execFileSync('node', ['scripts/platforms_report.mjs', '--list', 'Telegram'], { encoding: 'utf8', maxBuffer: 1 << 26, stdio: ['ignore', 'pipe', 'ignore'] }); const head = out.split('\n')[0]; const norm = (t) => t.trim().replace(/\.$/, '').toLowerCase(); const theirs = new Set(out.trim().split('\n').slice(1).map((l) => norm(l.split('\t')[3]))); const mine = IN.filter(([k]) => HAND[k].platforms.includes('Telegram')).map(([k]) => byKey.get(k)); const both = mine.filter((p) => theirs.has(norm(p.title))); console.log('\n== I. Reconciliation with design:platforms (its Telegram row) =='); console.log(`platforms_report.mjs: ${head}`); console.log(`Telegram IN papers here: ${mine.length}; in both: ${both.length}; here only: ${mine.length - both.length}; there only: ${theirs.size - both.length}`); for (const p of mine.filter((q) => !theirs.has(norm(q.title)))) console.log(` here only: ${p.year} ${p.venue} ${p.title}`); const mineTitles = new Set(mine.map((p) => norm(p.title))); for (const t of theirs) if (!mineTitles.has(t)) { const r = rows.find((p) => norm(p.title) === t); const v = V.get(key(r)); console.log(` there only: ${r.year} ${r.venue} ${r.title} [this page: ${v ? v[1] + ' ' + v[2] : 'not a candidate'}]`); } } if (process.argv.includes('--list')) { console.log('\n== --list: every verdict =='); for (const [k, v, c, note] of VERDICTS) console.log(`${v}\t${c}\t${k}\t${note}`); } console.log('\nOK: candidates and verdicts agree in both directions; every IN paper has hand codes.');
The report script's output
- report_messaging_channels-output.txt
design:platforms:messaging_channels — report corpus: 5859 extraction records, venues CCS, IEEE-SP, IMC, NDSS, PETS, USENIX, WWW == A. Probes (full text = paper.cols.txt, whitespace-collapsed; 5855 papers with text, 4 records without a text file) == platform cs>=10 (gap rule) of which 2024-26 probe>=10 of which 2024-26 any mention -------- ----------------- ---------------- --------- ---------------- ----------- telegram 40 23 43 24 168 whatsapp 41 18 52 24 338 discord 13 8 13 8 128 gap-rule union (any of the three, case-sensitive >= 10): 76 papers candidate set: 144 papers (enters by name>=10: 85; by route vocabulary + >=3 mentions: 96; by schema field: 66; channels overlap) recall additions (hand-added from the recall probes, each with its reason): 3 + CCS/2020/impersonation-as-a-service-characterizing-the-emerging-criminal-infrastructure-f — first-person collection sentence near "Telegram channel" (2 mentions, below every threshold) + USENIX/2025/privacy-law-enforcement-under-centralized-governance-a-qualitative-analysis-of-f — "we joined five WeChat groups" (other-messenger groups; no Telegram/WhatsApp/Discord mention) + WWW/2019/semi-supervised-graph-classification-a-hierarchical-graph-perspective — OTHER_RE: 16 mentions of "QQ group" data == B. Verdicts over 147 candidates (144 probe + 3 recall) == INCLUSION RULE (written before any verdict was counted): IN — the paper collects data from INSIDE messaging-platform groups, channels or servers (Telegram, WhatsApp, Discord, or another messenger's group feature) as a measurement source: it finds groups/channels, joins, subscribes, reads their public preview or has a scraper do so, and records messages, members, media or metadata. Sub-coded CORE (the group/channel data is the paper's main dataset) or SECTION (one component among several). CONTEXT — an adjacent route or object, but not collection from inside groups (codes below). OUT — protocol/crypto attacks, traffic analysis, app analysis, a messenger as C2 or as a recruitment channel, user studies of messengers in general, passing mentions. verdict code papers meaning ------------------- ------ -------------------------------------------------------------------------------- CONTEXT bots 1 chatbot ecosystem in a messenger, not group content CONTEXT dm 1 one-to-one conversations, not groups CONTEXT enumeration 4 accounts discovered through contact discovery / phone-number lookups, not groups CONTEXT links-only 3 invite links or group handles harvested elsewhere; groups never entered CONTEXT operator 1 group data supplied by the platform operator CONTEXT reuse 2 a dataset someone else collected from groups CONTEXT user-study 4 people's experience of groups, or a paper that declined to collect from them IN core 9 group/channel data is the main dataset IN section 17 group/channel data is one component among several OUT app-analysis 14 analysis of apps, features or exports OUT attack 3 an attack that involves a messenger OUT c2 1 messenger used as malware command-and-control OUT mention 29 passing or incidental mention OUT protocol 7 protocol or cryptographic analysis OUT recruitment 19 groups or servers used only to recruit study participants OUT traffic 10 traffic analysis / network measurement of messenger apps OUT user-general 22 user study of messengers or privacy in general IN 26 (core 9, section 17); CONTEXT 16; OUT 105; total 147 read in full: 44; decided from sentence contexts: 103; every IN paper was read in full == C. Precision and recall of each probe against the hand verdicts == probe papers IN precision recall recall % -------------------------------- ------ -- --------- ------ -------- gap rule: telegram cs>=10 40 19 47.5% 19/26 73.1% gap rule: whatsapp cs>=10 41 4 9.8% 4/26 15.4% gap rule: discord cs>=10 13 4 30.8% 4/26 15.4% gap rule: union of the three 76 23 30.3% 23/26 88.5% this page: candidate set 144 26 18.1% 26/26 100.0% candidate set + recall additions 147 26 17.7% 26/26 100.0% IN papers the gap rule misses (3): USENIX/2021/having-your-cake-and-eating-it-an-analysis-of-concession-abuse-as-a-service (Telegram 5, WhatsApp 0, Discord 0) WWW/2024/getting-bored-of-cyberwar-exploring-the-role-of-low-level-cybercrime-actors-in-t (Telegram 9, WhatsApp 0, Discord 0) USENIX/2024/dont-listen-to-me-understanding-and-exploring-jailbreak-prompts-of-large-languag (Telegram 0, WhatsApp 0, Discord 1) gap-rule papers that are not IN (53), by verdict: OUT user-general 12; OUT app-analysis 11; OUT traffic 8; CONTEXT enumeration 4; OUT protocol 4; OUT mention 4; CONTEXT user-study 2; CONTEXT links-only 2; OUT attack 2; CONTEXT dm 1; CONTEXT reuse 1; CONTEXT bots 1; OUT c2 1 == D. The 26 IN papers: platform, year, venue == platform (a paper can name several): Telegram 21; Discord 6; WhatsApp 3 papers with more than one platform: 3 bucket corpus papers IN IN core IN per 1,000 corpus papers Telegram IN ---------- ------------- -- ------- -------------------------- ----------- 2010–2018 1534 0 0 0.0 0 2019–2021 1185 9 5 7.6 7 2022–2024 1955 7 0 3.6 4 2025–2026* 1185 10 4 8.4 10 * 2025–2026 is provisional: CCS 2026 and IMC 2026 not yet held; IEEE S&P 2026 and TheWebConf 2026 under-selected by construction. per year: 2019:2, 2020:3, 2021:4, 2023:1, 2024:6, 2025:7, 2026:3 per venue: USENIX 12, WWW 7, IMC 3, IEEE-SP 2, NDSS 1, CCS 1 object (hand-coded family): cybercrime market 9; politics and misinformation 5; scams and impersonation 2; extremism and harassment 2; AI prompts 2; crypto manipulation 1; the messaging ecosystem itself 1; censorship circumvention 1; traffic model for an attack 1; engagement manipulation 1; unsafe content 1 == E. Hand codes over the 26 IN papers (from each paper's own text; msgch_fold.mjs HAND) == -- discovery (of 26; multi-valued fields do not sum): 8 30.8% in-app-search 7 26.9% directory 5 19.2% other-platform-links 4 15.4% snowball 3 11.5% known-official-channel 3 11.5% prior-list 3 11.5% web-search 2 7.7% not-stated -- join (of 26; multi-valued fields do not sum): 11 42.3% joined 6 23.1% not-stated 6 23.1% public-read 2 7.7% scraper-service 1 3.8% vendor -- history (of 26; multi-valued fields do not sum): 18 69.2% not-stated 7 26.9% yes 1 3.8% mixed -- members (of 26; multi-valued fields do not sum): 16 61.5% no 6 23.1% counts only 1 3.8% admin-only lists, not collected 1 3.8% member lists (phone numbers hashed) 1 3.8% member profiles 1 3.8% posters only -- tooling (of 26; multi-valued fields do not sum): 7 26.9% Telegram API (client unnamed) 6 23.1% Telethon 4 15.4% manual 3 11.5% not-stated 2 7.7% commercial scraper (Apify, Telemetrio) 2 7.7% phones + Garimella-Tyson tool 1 3.8% Discord API (user account) 1 3.8% PumpOlymp API 1 3.8% Selenium 1 3.8% vendor crawler 1 3.8% WhatsApp Web client -- limits (of 26; multi-valued fields do not sum): 6 23.1% channels-removed 2 7.7% rate-limit 1 3.8% data-gap 1 3.8% device-capacity 1 3.8% invite-expiry 1 3.8% join-cap -- irb (of 26; multi-valued fields do not sum): 10 38.5% approved 10 38.5% not-stated 2 7.7% no-board-named 2 7.7% not-required 1 3.8% exempt 1 3.8% other-part-only -- passive (of 26; multi-valued fields do not sum): 19 73.1% not-stated 7 26.9% stated -- minimise (of 26; multi-valued fields do not sum): 2 7.7% no de-anonymisation 1 3.8% account mentions removed 1 3.8% aggregate analysis 1 3.8% aggregate release only 1 3.8% aggregate reporting 1 3.8% anonymised 1 3.8% controlled-access release 1 3.8% de-identified 1 3.8% deletion on request 1 3.8% file type and size limits 1 3.8% identifiers pseudonymised 1 3.8% masked before storage 1 3.8% no payload download 1 3.8% phone numbers hashed 1 3.8% PII anonymised in release 1 3.8% PII anonymised, stored separately 1 3.8% quotes paraphrased 1 3.8% usernames masked 1 3.8% users anonymised -- denominator (of 26; multi-valued fields do not sum): 14 53.8% acknowledged 8 30.8% not-stated 4 15.4% n/a (known group or channel) -- years behind the codes the page dates (a paper counts once per code): discovery=directory: 7 papers, 6 from 2024-2026, 6 from 2022-2026; years 2021, 2024, 2024, 2025, 2025, 2025, 2026 discovery=in-app-search: 8 papers, 7 from 2024-2026, 8 from 2022-2026; years 2023, 2024, 2025, 2025, 2025, 2025, 2026, 2026 discovery=other-platform-links: 5 papers, 3 from 2024-2026, 3 from 2022-2026; years 2020, 2021, 2025, 2025, 2026 tooling=Telethon: 6 papers, 6 from 2024-2026, 6 from 2022-2026; years 2024, 2024, 2025, 2025, 2025, 2026 join=scraper-service: 2 papers, 2 from 2024-2026, 2 from 2022-2026; years 2024, 2025 platforms=WhatsApp: 3 papers, 0 from 2024-2026, 0 from 2022-2026; years 2019, 2020, 2021 platforms=Discord: 6 papers, 4 from 2024-2026, 4 from 2022-2026; years 2020, 2021, 2024, 2024, 2024, 2025 platforms=Telegram: 21 papers, 13 from 2024-2026, 14 from 2022-2026; years 2019, 2020, 2020, 2020, 2021, 2021, 2021, 2023, 2024, 2024, 2024, 2025, 2025, 2025, 2025, 2025, 2025, 2025, 2026, 2026, 2026 core papers by bucket and object: 2019-2021: crypto manipulation; the messaging ecosystem itself; politics and misinformation; politics and misinformation; engagement manipulation 2022-2024: 2025-2026*: cybercrime market; cybercrime market; cybercrime market; politics and misinformation -- CORE papers only (the ones whose main dataset is group/channel data): discovery (of 9): directory 4; in-app-search 3; web-search 3; other-platform-links 2; snowball 2; prior-list 1 join (of 9): joined 6; not-stated 2; public-read 1 tooling (of 9): Telegram API (client unnamed) 5; Telethon 2; phones + Garimella-Tyson tool 2; PumpOlymp API 1; WhatsApp Web client 1; Discord API (user account) 1 irb (of 9): approved 3; not-stated 3; not-required 1; no-board-named 1; exempt 1 denominator (of 9): acknowledged 8; not-stated 1 == F. Per-paper table (IN, 26) == year venue role platforms found collected messages discovery join tooling slug ---- ------- ------- ------------------------- ----------- ---------------- -------------------- ------------------------------------------- --------------- ---------------------------------------------------------------------------- ------------------------------------------------ 2019 USENIX core Telegram 300+ 300+ n/s prior-list not-stated Telegram API (client unnamed)+PumpOlymp API the-anatomy-of-a-cryptocurrency-pump-and-dump-sc 2019 WWW core WhatsApp 3,444 141 + 364 121,781 + 789,914 web-search joined phones + Garimella-Tyson tool mis-information-dissemination-in-whatsapp-gather 2020 IMC core WhatsApp+Telegram+Discord 351,535 616 8,255,069 other-platform-links joined WhatsApp Web client+Telegram API (client unnamed)+Discord API (user account) demystifying-the-messaging-platforms-ecosystem-t 2020 NDSS section Telegram n/s 1,000+ n/s not-stated joined Telegram API (client unnamed) practical-traffic-analysis-attacks-on-secure-mes 2020 WWW core Telegram 38,000 URLs 873 n/s web-search+snowball public-read Telegram API (client unnamed) the-pod-people-understanding-manipulation-of-soc 2021 IMC section Telegram+Discord n/s 2,916 (Telegram) n/s prior-list vendor vendor crawler a-large-scale-characterization-of-online-incitem 2021 USENIX section Telegram n/s 50 n/s snowball public-read manual catching-phishers-by-their-bait-investigating-th 2021 USENIX section Telegram 1 1 17,898 other-platform-links joined not-stated having-your-cake-and-eating-it-an-analysis-of-co 2021 WWW core WhatsApp n/s 5,010 1,426,482 directory+web-search joined phones + Garimella-Tyson tool short-is-the-road-that-leads-from-fear-to-hate-f 2023 USENIX section Telegram 6 6 n/s in-app-search+snowball joined manual strategies-and-vulnerabilities-of-participants-i 2024 CCS section Discord 20 6 n/s directory not-stated not-stated do-anything-now-characterizing-and-evaluating-in 2024 IEEE-SP section Telegram 2 2 525k known-official-channel not-stated Telethon no-easy-way-out-the-effectiveness-of-deplatformi 2024 USENIX section Telegram n/s n/s 133,399 in-app-search scraper-service commercial scraper (Apify, Telemetrio) the-imitation-game-exploring-brand-impersonation 2024 USENIX section Discord n/s 2 n/s not-stated not-stated Selenium dont-listen-to-me-understanding-and-exploring-ja 2024 USENIX section Discord n/s n/s 210 images directory joined manual moderating-illicit-online-image-promotion-for-un 2024 WWW section Telegram 1 1 441 + 57,757 replies known-official-channel public-read Telethon getting-bored-of-cyberwar-exploring-the-role-of- 2025 IEEE-SP section Telegram n/s 81 54K directory+other-platform-links public-read Telethon learning-from-censored-experiences-social-media- 2025 IMC section Telegram n/s n/s n/s in-app-search+other-platform-links joined manual unmasking-the-shadow-economy-a-deep-dive-into-dr 2025 USENIX core Telegram 4,709 339 64,801 directory not-stated Telegram API (client unnamed) darkgram-a-large-scale-analysis-of-cybercriminal 2025 USENIX section Telegram+Discord n/s 52 34,438 prior-list public-read Telethon assessing-the-aftermath-the-effects-of-a-global- 2025 USENIX core Telegram n/s 13 17.3M directory+in-app-search joined Telethon characterizing-and-detecting-propaganda-spreadin 2025 WWW section Telegram n/s n/s 85,402 in-app-search scraper-service commercial scraper (Apify, Telemetrio) pirates-of-charity-exploring-donation-based-abus 2025 WWW section Telegram n/s 15,537 4,309,880 in-app-search not-stated Telegram API (client unnamed) exposing-cross-platform-coordinated-inauthentic- 2026 USENIX core Telegram 21k 1,521 ~14M in-app-search+other-platform-links+snowball joined Telethon stayin-alive-how-global-stolen-data-markets-thri 2026 USENIX section Telegram 1 1 n/s known-official-channel public-read not-stated from-mirai-to-gorilla-deep-dive-into-a-long-last 2026 WWW core Telegram 312 100 25,972 directory+in-app-search joined Telegram API (client unnamed) doxing-as-a-service-demystifying-the-chinese-onl == G. The extraction schema over the IN set vs the corpus baseline == ethics.reviewOutcome — IN papers that are empirical (26 of 26) vs all empirical papers (5118): reviewOutcome IN (26) share corpus empirical (5118) share ------------------------------ ------- ----- ----------------------- ----- none-mentioned 8 30.8% 2744 53.6% approved 11 42.3% 994 19.4% (no ethics object) 1 3.8% 646 12.6% explicitly-discussed-no-review 1 3.8% 318 6.2% not-required 3 11.5% 179 3.5% exempt 1 3.8% 157 3.1% sought-outcome-unstated 1 3.8% 80 1.6% states any review outcome: IN 17/26 (65.4%); corpus 1728/5118 (33.8%) artifacts.availability — IN empirical (26) vs corpus empirical (5118): availability IN share corpus share -------------------------- -- ----- ------ ----- public 9 34.6% 2439 47.7% none-mentioned 7 26.9% 1964 38.4% (no artifacts object) 0 0.0% 264 5.2% promised-not-yet-available 2 7.7% 240 4.7% on-request 5 19.2% 86 1.7% restricted 3 11.5% 73 1.4% explicitly-withheld 0 0.0% 52 1.0% client library named in the paper (hand code): 7; of those, named in tools[] used/produced: 7 == H. CONTEXT papers, listed with code and note == bots IMC/2022/exploring-the-security-and-privacy-risks-of-chatbots-in-messaging-services Discord bot ecosystem from top.gg; bots tested in researcher-made guilds dm NDSS/2024/like-comment-get-scammed-characterizing-comment-scams-on-media-platforms 50 one-to-one scam-baiting chats over WhatsApp/Telegram; one Telegram group entered on a scammer's invitation (judgement call: incidental, not the route) enumeration NDSS/2021/all-the-numbers-are-us-large-scale-abuse-of-contact-discovery-in-mobile-messengers contact-discovery crawl of WhatsApp, Signal and Telegram; 10% of US numbers on WhatsApp enumeration NDSS/2026/hey-there-you-are-using-whatsapp-enumerating-three-billion-accounts-for-security-and-privacy 3,546,479,731 WhatsApp accounts enumerated through whatsmeow enumeration NDSS/2026/connecting-the-dots-an-investigative-study-on-linking-private-user-data-across-messaging-apps Korean 010 number space enumerated on Telegram, KakaoTalk, WhatsApp, Signal enumeration NDSS/2012/guess-who-s-texting-you-evaluating-the-security-of-smartphone-messaging-applicat ten million numbers uploaded to WhatsApp contact sync in 2.5 hours (2012) links-only IMC/2025/exploration-of-the-dynamics-of-buy-and-sale-of-social-media-accounts Telegram/WhatsApp/Discord handles appear as contact fields in marketplace listings; never entered links-only USENIX/2023/token-spammers-rug-pulls-and-sniper-bots-an-analysis-of-the-ecosystem-of-tokens 19,096 BSC and 1,334 Ethereum token contracts link a Telegram group; groups not entered links-only WWW/2025/detecting-and-understanding-the-promotion-of-illicit-goods-and-services-on-twitt Telegram/WeChat/QQ/WhatsApp contact URLs extracted from tweets; never entered operator WWW/2019/semi-supervised-graph-classification-a-hierarchical-graph-perspective recall addition: Tencent QQ group graphs supplied by the operator (Tencent authors) reuse WWW/2023/a-prompt-log-analysis-of-text-to-image-generation-systems reuses the Midjourney Discord dataset crawled by others reuse USENIX/2025/bots-can-snoop-uncovering-and-mitigating-privacy-risks-of-bots-in-group-chats reuses DISCO (Discord) and Pushshift (Telegram) datasets user-study PETS/2026/bot-among-us-exploring-user-awareness-and-privacy-concerns-about-chatbots-in-gro survey of 374 users on chatbots in their group chats user-study USENIX/2025/investigating-the-impact-of-online-community-involvement-on-safety-practices-and drug-user interviews; states it limited content collection to public platforms for ethical reasons, excluding Telegram groups user-study USENIX/2021/collective-information-security-in-large-scale-urban-protests-the-case-of-hong-k interviews with Hong Kong protesters about large public vs small private Telegram/WhatsApp groups; no group data collected user-study USENIX/2024/understanding-the-security-and-privacy-implications-of-online-toxic-content-on-r refugee focus group whose invitation-only WhatsApp group was infiltrated after its invite link leaked == I. Reconciliation with design:platforms (its Telegram row) == platforms_report.mjs: # 24 papers matched family "Telegram" with role=subject Telegram IN papers here: 21; in both: 17; here only: 4; there only: 7 here only: 2021 USENIX Having Your Cake and Eating It: An Analysis of Concession-Abuse-as-a-Service here only: 2025 IMC Unmasking the Shadow Economy: A Deep Dive into Drainer-as-a-Service Phishing on Ethereum. here only: 2026 USENIX From Mirai to Gorilla: Deep Dive into a Long-Lasting DDoS-for-Hire Botnet here only: 2021 IMC A large-scale characterization of online incitements to harassment across platforms. there only: 2020 PETS The Road Not Taken: Re-thinking the Feasibility of Voice Calling Over Tor [this page: OUT traffic] there only: 2022 IEEE-SP Four Attacks and a Proof for Telegram. [this page: OUT protocol] there only: 2022 USENIX Inferring Phishing Intention via Webpage Appearance and Dynamics: A Deep Vision Based Approach [this page: OUT mention] there only: 2023 IEEE-SP From 5G Sniffing to Harvesting Leakages of Privacy-Preserving Messengers. [this page: OUT traffic] there only: 2024 NDSS Like, Comment, Get Scammed: Characterizing Comment Scams on Media Platforms [this page: CONTEXT dm] there only: 2025 USENIX Bots can Snoop: Uncovering and Mitigating Privacy Risks of Bots in Group Chats [this page: CONTEXT reuse] there only: 2025 USENIX Investigating the Impact of Online Community Involvement on Safety Practices and Perceived Risks Among People Who Use Drugs [this page: CONTEXT user-study] OK: candidates and verdicts agree in both directions; every IN paper has hand codes.
The probes
- msgch_probes.mjs
// msgch_probes.mjs — candidate probes for design:platforms:messaging_channels // ("Messaging groups and channels as a data source"). Imported by msgch_fold.mjs and // report_messaging_channels.mjs. Nothing here is a population: every probe produces a // CANDIDATE set, and the population is the hand verdict in msgch_fold.mjs. // // Texts are read one at a time from paper.cols.txt with whitespace collapsed (a PDF line // break inside "invite link" otherwise undercounts), and never held all at once. import fs from 'node:fs'; import path from 'node:path'; import { dataRoot } from './lib.mjs'; // The 2026-09-22 gap pass's rule as the task brief states it: ">= 10 mentions of the platform // name" in full text. Discord is case-sensitive on purpose: lowercase "discord" is an English noun. export const NAME = { telegram: /\bTelegram\b/gi, whatsapp: /\bWhats ?App\b/gi, discord: /\bDiscord\b/g, }; export const NAME_THRESH = 10; // The ROUTE vocabulary: the artefacts of joining groups/channels. Any one hit enters the // candidate set, because a paper that joined 5,000 groups may name the platform only a few times. export const ROUTE_RE = new RegExp([ String.raw`chat\.whatsapp\.com`, String.raw`\bt\.me\b`, String.raw`telegram\.me\b`, String.raw`discord\.gg\b`, String.raw`discord(app)?\.com/invite`, String.raw`invite[- ]?links?`, String.raw`invitation links?`, String.raw`group (invite|invitation)s?`, String.raw`\bTelethon\b`, String.raw`\bTDLib\b`, String.raw`\bPyrogram\b`, String.raw`\bMTProto API\b`, String.raw`Telegram (Core |Client )?API`, String.raw`\bTGStat\b`, String.raw`\bTelemetr(io)?\b`, String.raw`\bTgScan\b`, String.raw`\bLyzem\b`, String.raw`WhatsApp Web\b`, String.raw`whatsapp-web\.js`, String.raw`\bBaileys\b`, String.raw`\byowsup\b`, String.raw`WhatsApp Channels?\b`, String.raw`discord\.py\b`, String.raw`discord\.js\b`, String.raw`self-?bots?\b`, String.raw`DiscordChatExporter`, String.raw`\bDisboard\b`, String.raw`discord\.me\b`, String.raw`(public|open) (Telegram|WhatsApp|Discord) (groups?|channels?|servers?|chats?)`, String.raw`(Telegram|WhatsApp|Discord) (groups?|channels?|servers?|chats?|communities)`, ].join('|'), 'gi'); // Other messengers' groups, as a RECALL probe for the boundary (Signal/Slack/LINE are // homographs, so only the "<name> group|channel" form counts, and case-sensitively). export const OTHER_RE = /\b(Signal|Viber|WeChat|LINE|Slack|Kik|GroupMe|Matrix|Element|QQ|KakaoTalk|Snapchat) (groups?|channels?|chats?|chat groups?|servers?|communities)\b/g; // Schema channel: the extraction's own free-text fields. export const SCHEMA_RE = /telegram|whats ?app|discord/i; export function textPath(p) { return path.join(dataRoot(), 'fulltext', String(p.year), p.venue, p.slug, 'paper.cols.txt'); } export function readCollapsed(p) { const f = textPath(p); if (!fs.existsSync(f)) return null; return fs.readFileSync(f, 'latin1').replace(/\s+/g, ' '); } export const count = (t, re) => (t.match(new RegExp(re.source, re.flags.includes('g') ? re.flags : re.flags + 'g')) || []).length; export function schemaHit(p) { const fields = []; for (const x of p.population) if (SCHEMA_RE.test(String(x.sourceList)) || SCHEMA_RE.test(String(x.unit))) fields.push('population'); for (const x of p.detection) if (SCHEMA_RE.test(String(x.phenomenon))) fields.push('detection'); for (const x of p.tools) if ((x.usedOrMentioned === 'used' || x.usedOrMentioned === 'produced') && SCHEMA_RE.test(String(x.name))) fields.push('tools'); if (SCHEMA_RE.test(p.title)) fields.push('title'); return [...new Set(fields)]; } export const key = (p) => `${p.venue}/${p.year}/${p.slug}`; // One pass over the corpus. Returns per-paper probe counts for every paper with text, plus // the number of records with no text file, so no paper drops silently out of a denominator. export function runProbes(rows) { const out = []; let missing = 0; for (const p of rows) { const t = readCollapsed(p); if (t === null) { missing += 1; continue; } const r = { key: key(p), p, route: count(t, ROUTE_RE), other: count(t, OTHER_RE), schema: schemaHit(p) }; for (const [n, re] of Object.entries(NAME)) r[n] = count(t, re); out.push(r); } return { hits: out, missing }; } export const isCandidate = (r) => r.telegram >= NAME_THRESH || r.whatsapp >= NAME_THRESH || r.discord >= NAME_THRESH || r.route >= 1 && (r.telegram + r.whatsapp + r.discord) >= 3 || r.schema.length > 0;
The inclusion rule, the 147 verdicts and the 26 papers' hand codes
- msgch_fold.mjs
// msgch_fold.mjs — inclusion rule, hand verdicts and hand codes behind // design:platforms:messaging_channels ("Messaging Groups and Channels"). Imported by // report_messaging_channels.mjs; nothing here is computed, it is the audit surface. // // The probes (msgch_probes.mjs) produce a CANDIDATE set of 144 papers; three more were added by hand // from the recall probes (RECALL_ADDED). Every one of the 147 has a verdict below, and the report // throws if a candidate lacks a verdict or a verdict names a non-candidate. // // Verdicts for the 44 papers read in full were made from sub-agent reading notes // (notes/msgch_papers_[A-D].md, every quote machine-checked by msgch_notes_quotecheck.py) and // decided by the main session; three reader verdicts were overridden, each marked "judgement call" // or "recall addition" in its note. The other 103 were decided from sentence contexts around every // platform-name and route-vocabulary hit (_msgch_ctx.mjs). export const INCLUSION_RULE = `INCLUSION RULE (written before any verdict was counted): IN — the paper collects data from INSIDE messaging-platform groups, channels or servers (Telegram, WhatsApp, Discord, or another messenger's group feature) as a measurement source: it finds groups/channels, joins, subscribes, reads their public preview or has a scraper do so, and records messages, members, media or metadata. Sub-coded CORE (the group/channel data is the paper's main dataset) or SECTION (one component among several). CONTEXT — an adjacent route or object, but not collection from inside groups (codes below). OUT — protocol/crypto attacks, traffic analysis, app analysis, a messenger as C2 or as a recruitment channel, user studies of messengers in general, passing mentions.`; export const CODES = { IN: { core: 'group/channel data is the main dataset', section: 'group/channel data is one component among several', }, CONTEXT: { enumeration: 'accounts discovered through contact discovery / phone-number lookups, not groups', 'links-only': 'invite links or group handles harvested elsewhere; groups never entered', reuse: 'a dataset someone else collected from groups', bots: 'chatbot ecosystem in a messenger, not group content', dm: 'one-to-one conversations, not groups', 'user-study': "people's experience of groups, or a paper that declined to collect from them", operator: 'group data supplied by the platform operator', }, OUT: { mention: 'passing or incidental mention', 'user-general': 'user study of messengers or privacy in general', recruitment: 'groups or servers used only to recruit study participants', 'app-analysis': 'analysis of apps, features or exports', traffic: 'traffic analysis / network measurement of messenger apps', protocol: 'protocol or cryptographic analysis', attack: 'an attack that involves a messenger', c2: 'messenger used as malware command-and-control', }, }; // Added by hand from the recall probes (_msgch_recall.mjs and OTHER_RE), each with the reason. export const RECALL_ADDED = [ ['CCS/2020/impersonation-as-a-service-characterizing-the-emerging-criminal-infrastructure-f', 'first-person collection sentence near "Telegram channel" (2 mentions, below every threshold)'], ['USENIX/2025/privacy-law-enforcement-under-centralized-governance-a-qualitative-analysis-of-f', '"we joined five WeChat groups" (other-messenger groups; no Telegram/WhatsApp/Discord mention)'], ['WWW/2019/semi-supervised-graph-classification-a-hierarchical-graph-perspective', 'OTHER_RE: 16 mentions of "QQ group" data'], ]; // [key, verdict, code, note] export const VERDICTS = [ ["USENIX/2025/darkgram-a-large-scale-analysis-of-cybercriminal-activity-channels-on-telegram", "IN", "core", "339 cybercrime channels picked from 4,709 Telemetr.io listings; 64,801 posts via the official API"], ["USENIX/2026/stayin-alive-how-global-stolen-data-markets-thrive-on-telegram", "IN", "core", "1,521 stolen-data channels joined (448 private) from 21k snowball candidates; Telethon; 14M messages"], ["WWW/2026/doxing-as-a-service-demystifying-the-chinese-online-doxing-ecosystem", "IN", "core", "top 100 of 312 Chinese doxing channels (TGStat/Telemetr + search bots); 25,972 messages; five query groups"], ["USENIX/2019/the-anatomy-of-a-cryptocurrency-pump-and-dump-scheme", "IN", "core", "message history of 300+ pump channels listed by PumpOlymp; 412 pump events"], ["USENIX/2021/catching-phishers-by-their-bait-investigating-the-dutch-phishing-landscape-throu", "IN", "section", "phishing kits snowballed from 50 public Telegram channels; main data is CT logs"], ["USENIX/2021/having-your-cake-and-eating-it-an-analysis-of-concession-abuse-as-a-service", "IN", "section", "joined one provider supergroup: 1,076 members, 17,898 messages; main data is four forums"], ["IMC/2025/unmasking-the-shadow-economy-a-deep-dive-into-drainer-as-a-service-phishing-on-e", "IN", "section", "joined drainer operators' Telegram groups to obtain toolkits; main data is on-chain"], ["USENIX/2026/from-mirai-to-gorilla-deep-dive-into-a-long-lasting-ddos-for-hire-botnet", "IN", "section", "monitored the botnet's own Telegram marketing channel; one of eight sources"], ["USENIX/2025/assessing-the-aftermath-the-effects-of-a-global-takedown-against-ddos-for-hire-s", "IN", "section", "52 booter chat/news channels via Telethon; one of eight datasets"], ["CCS/2020/impersonation-as-a-service-characterizing-the-emerging-criminal-infrastructure-f", "OUT", "mention", "recall addition: a handful of examples from a Telegram channel reached through the marketplace; incidental to a web-marketplace study"], ["IMC/2025/exploration-of-the-dynamics-of-buy-and-sale-of-social-media-accounts", "CONTEXT", "links-only", "Telegram/WhatsApp/Discord handles appear as contact fields in marketplace listings; never entered"], ["WWW/2025/pirates-of-charity-exploring-donation-based-abuses-in-social-media-platforms", "IN", "section", "donation-scam posts from public Telegram channels via commercial scraper APIs; one of five platforms"], ["USENIX/2024/the-imitation-game-exploring-brand-impersonation-attacks-on-social-media-platfor", "IN", "section", "brand-impersonating Telegram accounts/channels via keyword search and scraper APIs; 133,399 posts; one of four platforms"], ["IMC/2020/demystifying-the-messaging-platforms-ecosystem-through-the-lens-of-twitter", "IN", "core", "351,535 group URLs from Twitter; 616 WhatsApp/Telegram/Discord groups joined"], ["WWW/2019/mis-information-dissemination-in-whatsapp-gathering-analyzing-and-countermeasure", "IN", "core", "3,444 chat.whatsapp.com links from search; 1,828 valid; 141 and 364 groups joined with phones"], ["WWW/2021/short-is-the-road-that-leads-from-fear-to-hate-fear-speech-in-indian-whatsapp-gr", "IN", "core", "over 5,000 public political WhatsApp groups joined; 1.4M posts"], ["USENIX/2025/characterizing-and-detecting-propaganda-spreading-accounts-on-telegram", "IN", "core", "13 political/news channels chosen from TGStat and search; Telethon + export-history; 17.3M messages"], ["WWW/2025/exposing-cross-platform-coordinated-inauthentic-activity-in-the-run-up-to-the-20", "IN", "section", "public Telegram chats by keyword via the Telegram API; one of three platforms"], ["IEEE-SP/2024/no-easy-way-out-the-effectiveness-of-deplatforming-an-extremist-forum-to-suppres", "IN", "section", "the forum's two Telegram channels via Telethon over their whole lifespan; one of several sources"], ["WWW/2024/getting-bored-of-cyberwar-exploring-the-role-of-low-level-cybercrime-actors-in-t", "IN", "section", "the IT Army of Ukraine channel via Telethon from inception; one of four sources"], ["IEEE-SP/2025/learning-from-censored-experiences-social-media-discussions-around-censorship-ci", "IN", "section", "81 Persian/English VPN-sharing channels (Telemetr.io + Twitter); Telethon; with Twitter"], ["USENIX/2023/strategies-and-vulnerabilities-of-participants-in-venezuelan-influence-operation", "IN", "section", "six Telegram groups found by in-app search and joined; member counts; main data is 19 interviews"], ["IMC/2021/a-large-scale-characterization-of-online-incitements-to-harassment-across-platfo", "IN", "section", "2,916 Telegram channels and Discord servers from a vendor's curated list and crawlers; one of five platform types"], ["NDSS/2020/practical-traffic-analysis-attacks-on-secure-messaging-applications", "IN", "section", "joined over 1,000 public Telegram channels to record message timing and sizes for a traffic model (judgement call: the rule covers metadata; the paper is an attack)"], ["WWW/2020/the-pod-people-understanding-manipulation-of-social-media-popularity-via-recipro", "IN", "core", "Instagram engagement pods hosted as public Telegram groups; 38,000 URLs -> 873 active public groups -> 432 pods"], ["CCS/2024/do-anything-now-characterizing-and-evaluating-in-the-wild-jailbreak-prompts-on-l", "IN", "section", "six Discord servers found via Disboard; prompt-collection channels; one of four sources"], ["USENIX/2024/dont-listen-to-me-understanding-and-exploring-jailbreak-prompts-of-large-languag", "IN", "section", "two Discord channels among five jailbreak-prompt sources; Selenium crawler"], ["USENIX/2024/moderating-illicit-online-image-promotion-for-unsafe-user-generated-content-game", "IN", "section", "Roblox game servers on Discord entered via a server-listing site; 210 images; minor in-the-wild set"], ["WWW/2023/a-prompt-log-analysis-of-text-to-image-generation-systems", "CONTEXT", "reuse", "reuses the Midjourney Discord dataset crawled by others"], ["USENIX/2025/bots-can-snoop-uncovering-and-mitigating-privacy-risks-of-bots-in-group-chats", "CONTEXT", "reuse", "reuses DISCO (Discord) and Pushshift (Telegram) datasets"], ["PETS/2026/bot-among-us-exploring-user-awareness-and-privacy-concerns-about-chatbots-in-gro", "CONTEXT", "user-study", "survey of 374 users on chatbots in their group chats"], ["IMC/2022/exploring-the-security-and-privacy-risks-of-chatbots-in-messaging-services", "CONTEXT", "bots", "Discord bot ecosystem from top.gg; bots tested in researcher-made guilds"], ["USENIX/2025/investigating-the-impact-of-online-community-involvement-on-safety-practices-and", "CONTEXT", "user-study", "drug-user interviews; states it limited content collection to public platforms for ethical reasons, excluding Telegram groups"], ["USENIX/2025/privacy-law-enforcement-under-centralized-governance-a-qualitative-analysis-of-f", "OUT", "recruitment", "recall addition: five WeChat groups joined through personal networks to recruit interviewees; group content not analysed"], ["WWW/2019/semi-supervised-graph-classification-a-hierarchical-graph-perspective", "CONTEXT", "operator", "recall addition: Tencent QQ group graphs supplied by the operator (Tencent authors)"], ["USENIX/2023/token-spammers-rug-pulls-and-sniper-bots-an-analysis-of-the-ecosystem-of-tokens", "CONTEXT", "links-only", "19,096 BSC and 1,334 Ethereum token contracts link a Telegram group; groups not entered"], ["WWW/2025/detecting-and-understanding-the-promotion-of-illicit-goods-and-services-on-twitt", "CONTEXT", "links-only", "Telegram/WeChat/QQ/WhatsApp contact URLs extracted from tweets; never entered"], ["CCS/2019/the-art-and-craft-of-fraudulent-app-promotion-in-google-play", "OUT", "recruitment", "WhatsApp/Facebook groups of ASO workers as a recruitment venue"], ["USENIX/2026/cracks-in-the-walled-garden-dissecting-the-gray-market-of-unauthorized-ios-app-d", "OUT", "mention", "QQ/WeChat named only as product or contact-field labels"], ["NDSS/2024/like-comment-get-scammed-characterizing-comment-scams-on-media-platforms", "CONTEXT", "dm", "50 one-to-one scam-baiting chats over WhatsApp/Telegram; one Telegram group entered on a scammer's invitation (judgement call: incidental, not the route)"], ["NDSS/2021/all-the-numbers-are-us-large-scale-abuse-of-contact-discovery-in-mobile-messengers", "CONTEXT", "enumeration", "contact-discovery crawl of WhatsApp, Signal and Telegram; 10% of US numbers on WhatsApp"], ["NDSS/2026/hey-there-you-are-using-whatsapp-enumerating-three-billion-accounts-for-security-and-privacy", "CONTEXT", "enumeration", "3,546,479,731 WhatsApp accounts enumerated through whatsmeow"], ["NDSS/2026/connecting-the-dots-an-investigative-study-on-linking-private-user-data-across-messaging-apps", "CONTEXT", "enumeration", "Korean 010 number space enumerated on Telegram, KakaoTalk, WhatsApp, Signal"], ["NDSS/2012/guess-who-s-texting-you-evaluating-the-security-of-smartphone-messaging-applicat", "CONTEXT", "enumeration", "ten million numbers uploaded to WhatsApp contact sync in 2.5 hours (2012)"], ["IEEE-SP/2025/you-have-to-ignore-the-dangers-user-perceptions-of-the-security-and-privacy-bene", "OUT", "user-general", "interviews on WhatsApp mods; WhatsApp groups only as a recruitment channel"], ["PETS/2025/if-you-want-to-encrypt-it-really-really-hardcore-user-perceptions-of-key-transpa", "OUT", "user-general", "interviews on WhatsApp key transparency"], ["USENIX/2021/collective-information-security-in-large-scale-urban-protests-the-case-of-hong-k", "CONTEXT", "user-study", "interviews with Hong Kong protesters about large public vs small private Telegram/WhatsApp groups; no group data collected"], ["PETS/2024/what-do-privacy-advertisements-communicate-to-consumers", "OUT", "user-general", "WhatsApp privacy ad campaign as a stimulus"], ["IEEE-SP/2022/four-attacks-and-a-proof-for-telegram", "OUT", "protocol", "MTProto cryptanalysis"], ["PETS/2025/can-social-media-privacy-and-safety-features-protect-targets-of-interpersonal-at", "OUT", "app-analysis", "feature audit of apps incl. Telegram/WhatsApp"], ["IEEE-SP/2024/injection-attacks-against-end-to-end-encrypted-applications", "OUT", "protocol", "E2EE backup injection attacks"], ["USENIX/2023/cryptographic-deniability-a-multi-perspective-study-of-user-perceptions-and-expe", "OUT", "user-general", "deniability perceptions; WhatsApp court cases"], ["IMC/2025/protocol-compliance-in-popular-rtc-applications", "OUT", "traffic", "RTC protocol compliance"], ["NDSS/2023/hope-of-delivery-extracting-user-locations-from-mobile-instant-messengers", "OUT", "traffic", "location inference from delivery receipts"], ["USENIX/2020/i-have-too-much-respect-for-my-elders-understanding-south-african-mobile-users-p", "OUT", "user-general", "South African users on Facebook/WhatsApp privacy"], ["PETS/2024/a-black-box-privacy-analysis-of-messaging-service-providers-chat-message-process", "OUT", "app-analysis", "link-preview/server-side processing of chat messages"], ["IMC/2025/hello-genai-dissecting-human-to-generative-ai-calling", "OUT", "traffic", "GenAI calling apps"], ["IEEE-SP/2021/defensive-technology-use-by-political-activists-during-the-sudanese-revolution", "OUT", "user-general", "interviews with Sudanese activists"], ["PETS/2025/real-world-deniability-in-messaging", "OUT", "user-general", "court cases citing WhatsApp messages"], ["CCS/2025/hidden-in-plain-bytes-investigating-interpersonal-account-compromise-with-data-e", "OUT", "app-analysis", "data exports incl. Discord"], ["IEEE-SP/2017/obstacles-to-the-adoption-of-secure-communication-tools", "OUT", "user-general", "secure-messaging adoption interviews"], ["USENIX/2025/on-the-virtues-of-information-security-in-the-uk-climate-movement", "OUT", "user-general", "ethnography of UK climate movement; messengers discussed, no group corpus"], ["NDSS/2021/on-the-insecurity-of-sms-one-time-password-messages-against-local-attackers-in-modern-mobile-devices", "OUT", "app-analysis", "SMS OTP; Telegram clients"], ["USENIX/2023/hiding-in-plain-sight-an-empirical-study-of-web-application-abuse-in-malware", "OUT", "c2", "Telegram/Discord as malware C2"], ["PETS/2026/user-perceptions-and-attitudes-toward-untraceability-in-messaging-platforms", "OUT", "user-general", "untraceability perceptions"], ["IEEE-SP/2023/from-5g-sniffing-to-harvesting-leakages-of-privacy-preserving-messengers", "OUT", "traffic", "5G sniffing of messengers"], ["NDSS/2020/deceptive-previews-a-study-of-the-link-preview-trustworthiness-in-social-platforms", "OUT", "app-analysis", "link previews across platforms"], ["USENIX/2023/account-security-interfaces-important-unintuitive-and-untrustworthy", "OUT", "app-analysis", "account security UIs"], ["IEEE-SP/2017/the-password-reset-mitm-attack", "OUT", "protocol", "password-reset MitM"], ["IEEE-SP/2025/sok-self-generated-nudes-over-private-chats-how-can-technology-contribute-to-a-s", "OUT", "mention", "SoK on sexting features"], ["PETS/2024/the-medium-is-the-message-how-secure-messaging-apps-leak-sensitive-data-to-push", "OUT", "app-analysis", "push notification leakage"], ["USENIX/2026/analyzing-the-webrtc-ecosystem-and-breaking-authentication-in-dtls-srtp", "OUT", "protocol", "DTLS-SRTP"], ["IEEE-SP/2020/the-many-kinds-of-creepware-used-for-interpersonal-attacks", "OUT", "app-analysis", "creepware apps"], ["PETS/2020/a-privacy-focused-systematic-analysis-of-online-status-indicators", "OUT", "app-analysis", "online status indicators"], ["USENIX/2024/rise-of-inspectron-automated-black-box-auditing-of-cross-platform-electron-apps", "OUT", "app-analysis", "Electron apps"], ["USENIX/2021/evaluating-in-workflow-messages-for-improving-mental-models-of-end-to-end-encryp", "OUT", "user-general", "E2EE mental models"], ["NDSS/2022/auto-draft-262", "OUT", "app-analysis", "team-chat app extensions (Slack etc.)"], ["NDSS/2023/tactics-threats-targets-modeling-disinformation-and-its-mitigation", "OUT", "user-general", "interviews with disinformation practitioners; TGStat named as a tool they use"], ["PETS/2026/what-app-app-usage-detection-using-encrypted-lte-5g-traffic", "OUT", "traffic", "app usage from LTE/5G traffic"], ["USENIX/2023/eavesdropping-mobile-app-activity-via-radio-frequency-energy-harvesting", "OUT", "traffic", "RF energy harvesting side channel"], ["PETS/2025/sok-web-authentication-and-recovery-in-the-age-of-end-to-end-encryption", "OUT", "mention", "SoK"], ["USENIX/2022/adversarial-detection-avoidance-attacks-evaluating-the-robustness-of-perceptual", "OUT", "mention", "perceptual hashing CSS"], ["PETS/2023/creative-beyond-tiktoks-investigating-adolescents-social-privacy-management-on-t", "OUT", "user-general", "adolescents on TikTok"], ["IEEE-SP/2025/decentralization-of-ethereums-builder-market", "OUT", "mention", "Telegram trading bots in MEV"], ["PETS/2020/the-road-not-taken-re-thinking-the-feasibility-of-voice-calling-over-tor", "OUT", "traffic", "voice over Tor; Pyrogram as a test client"], ["USENIX/2024/did-they-f-ing-consent-to-that-safer-digital-intimacy-via-proactive-protection-a", "OUT", "user-general", "image-based sexual abuse"], ["NDSS/2025/impact-tracing-identifying-the-culprit-of-misinformation-in-encrypted-messaging-systems", "OUT", "protocol", "misinformation traceback in E2EE"], ["NDSS/2021/awakening-the-webs-sleeper-agents-misusing-service-workers-for-privacy-leakage", "OUT", "attack", "service-worker attack incl. inferring WhatsApp group membership"], ["IEEE-SP/2017/smarper-context-aware-and-automatic-runtime-permissions-for-mobile-devices", "OUT", "mention", "runtime permissions"], ["WWW/2020/the-chameleon-attack-manipulating-content-display-in-online-social-media", "OUT", "attack", "content-display manipulation"], ["NDSS/2025/tweezers-a-framework-for-security-event-detection-via-event-attribution-centric-tweet-embedding", "OUT", "mention", "tweet embeddings"], ["PETS/2019/investigating-sources-of-pii-used-in-facebook-s-targeted-advertising", "OUT", "mention", "Facebook ad PII sources"], ["NDSS/2022/auto-draft-218", "OUT", "traffic", "app fingerprinting on wireless traffic"], ["PETS/2023/on-the-role-and-form-of-personal-information-disclosure-in-cyberbullying-inciden", "OUT", "user-general", "cyberbullying incidents"], ["IEEE-SP/2025/born-with-a-silver-spoon-on-the-in-security-of-native-granted-app-privileges-in", "OUT", "app-analysis", "custom ROM privileges"], ["NDSS/2020/detecting-probe-resistant-proxies", "OUT", "protocol", "MTProto proxies"], ["IEEE-SP/2025/ownership-and-gatekeeping-vs-safeguarding-and-consent-how-migrant-parents-naviga", "OUT", "recruitment", "WhatsApp groups as a recruitment channel"], ["USENIX/2026/love-lies-and-language-models-investigating-ais-role-in-romance-baiting-scams", "OUT", "user-general", "romance-baiting experiment over WhatsApp one-to-one"], ["USENIX/2023/medusa-attack-exploring-security-hazards-of-in-app-qr-code-scanning", "OUT", "attack", "in-app QR scanning"], ["USENIX/2024/enabling-developers-protecting-users-investigating-harassment-and-safety-in-vr", "OUT", "recruitment", "Discord servers as a recruitment channel"], ["NDSS/2026/anchors-of-trust-a-usability-study-on-user-awareness-consent-and-control-in-cross-device-authentication", "OUT", "user-general", "cross-device authentication"], ["NDSS/2021/to-err-is-human-characterizing-the-threat-of-unintended-urls-in-social-media", "OUT", "app-analysis", "unintended URLs"], ["USENIX/2024/understanding-the-security-and-privacy-implications-of-online-toxic-content-on-r", "CONTEXT", "user-study", "refugee focus group whose invitation-only WhatsApp group was infiltrated after its invite link leaked"], ["PETS/2026/i-just-press-allow-understanding-privacy-practices-of-new-internet-users-in-urba", "OUT", "user-general", "new internet users in urban India"], ["USENIX/2018/return-of-bleichenbacher-s-oracle-threat-robot", "OUT", "protocol", "TLS scan"], ["CCS/2022/watch-your-back-identifying-cybercrime-financial-relationships-in-bitcoin-throug", "OUT", "mention", "a Telegram channel cited as evidence for a wallet seed"], ["WWW/2023/misbehavior-and-account-suspension-in-an-online-financial-communication-platform", "OUT", "mention", "StockTwits profile links"], ["WWW/2023/online-advertising-in-ukraine-and-russia-during-the-2022-russian-invasion", "OUT", "mention", "an ad linking to a t.me/s/ preview"], ["USENIX/2024/guardians-of-the-galaxy-content-moderation-in-the-interplanetary-file-system", "OUT", "mention", "t.me citations"], ["CCS/2025/security-and-privacy-measurements-in-cellular-networks-novel-approaches-in-a-glo", "OUT", "traffic", "cellular roaming"], ["IMC/2025/poster-uncovering-lesser-studied-rtc-applications-using-rtp", "OUT", "traffic", "RTP"], ["PETS/2025/what-are-they-gonna-do-with-my-data-privacy-expectations-concerns-and-behaviors", "OUT", "recruitment", "Discord servers as a recruitment channel"], ["USENIX/2026/a-midsummer-memes-dream-investigating-market-manipulations-in-the-meme-coin-ecos", "OUT", "mention", "cites Telegram-focused work and TGDataset; no group data of its own"], ["IEEE-SP/2023/blue-is-the-new-black-market-privacy-leaks-and-re-victimization-from-police-auct", "OUT", "mention", "a phone owner was in a fraud Telegram group"], ["IEEE-SP/2022/scraping-sticky-leftovers-app-user-information-left-on-servers-after-account-del", "OUT", "app-analysis", "leftover account data"], ["CCS/2022/helping-or-hindering-how-browser-extensions-undermine-security", "OUT", "mention", "WhatsApp Web notifier extension"], ["USENIX/2024/i-really-just-leaned-on-my-community-for-support-barriers-challenges-and-coping", "OUT", "user-general", "tech-abuse survivors"], ["NDSS/2025/exploring-user-perceptions-of-security-auditing-in-the-web3-ecosystem", "OUT", "mention", "rug-pull Telegram group deleted (citation)"], ["USENIX/2026/digital-risks-and-coping-practices-among-roblox-game-creators", "OUT", "recruitment", "Discord servers as a recruitment channel"], ["USENIX/2026/chameleon-channels-measuring-youtube-accounts-repurposed-for-deception-and-profi", "OUT", "mention", "YouTube channels advertise WhatsApp/Telegram group links; not entered"], ["NDSS/2020/massbrowser-unblocking-the-censored-web-for-the-masses-by-the-masses", "OUT", "mention", "Telegram traffic whitelist rule"], ["CCS/2023/txphishscope-towards-detecting-and-understanding-transaction-based-phishing-on-e", "OUT", "mention", "Telegram groups selling toolkits observed; t.me references"], ["CCS/2024/modern-problems-require-modern-solutions-community-developed-techniques-for-onli", "OUT", "mention", "a private Telegram group named in a video"], ["IEEE-SP/2025/exploring-parent-child-perceptions-on-safety-in-generative-ai-concerns-mitigatio", "OUT", "user-general", "generative-AI safety interviews"], ["USENIX/2025/security-and-privacy-advice-for-upi-users-in-india", "OUT", "mention", "WhatsApp calls to reach interviewees"], ["IEEE-SP/2022/desperate-times-call-for-desperate-measures-user-concerns-with-mobile-loan-apps", "OUT", "user-general", "mobile loan apps"], ["IEEE-SP/2023/in-eighty-percent-of-the-cases-i-select-the-password-for-them-security-and-priva", "OUT", "mention", "WhatsApp calls to reach interviewees"], ["USENIX/2022/inferring-phishing-intention-via-webpage-appearance-and-dynamics-a-deep-vision-b", "OUT", "mention", "phishing detection"], ["PETS/2023/investigating-privacy-decision-making-processes-among-nigerian-men-and-women", "OUT", "user-general", "sharing to WhatsApp groups as a survey item"], ["CCS/2024/defying-the-odds-solanas-unexpected-resilience-in-spite-of-the-security-challeng", "OUT", "recruitment", "private Discord channels as a recruitment channel"], ["CCS/2024/skipping-the-security-side-quests-a-qualitative-study-on-security-practices-and", "OUT", "recruitment", "Discord servers as a recruitment channel (moderator permission)"], ["WWW/2024/bots-elections-and-controversies-twitter-insights-from-brazils-polarised-electio", "OUT", "mention", "WhatsApp groups in related work"], ["WWW/2024/identifying-risky-vendors-in-cryptocurrency-p2p-marketplaces", "OUT", "mention", "Telegram channel as a complaint venue"], ["WWW/2025/causal-insights-into-parlers-content-moderation-shift-effects-on-toxicity-and-fa", "OUT", "mention", "Telegram channels in related work"], ["IEEE-SP/2018/the-spyware-used-in-intimate-partner-violence", "OUT", "mention", "WhatsApp Web citation"], ["IEEE-SP/2025/security-perceptions-of-users-in-stablecoins-advantages-and-risks-within-the-cry", "OUT", "recruitment", "Discord/Telegram as recruitment"], ["IEEE-SP/2023/we-are-a-startup-to-the-core-a-qualitative-interview-study-on-the-security-and-p", "OUT", "recruitment", "Discord/Slack as recruitment"], ["USENIX/2023/a-mixed-methods-study-of-security-practices-of-smart-contract-developers", "OUT", "recruitment", "Discord channels as recruitment"], ["IEEE-SP/2024/understanding-parents-perceptions-and-practices-toward-childrens-security-and-pr", "OUT", "recruitment", "Discord servers as recruitment (moderator permission)"], ["USENIX/2024/i-experienced-more-than-10-defi-scams-on-defi-users-perception-of-security-breac", "OUT", "recruitment", "Discord/Telegram as recruitment"], ["CCS/2025/a-qualitative-analysis-of-fuzzer-usability-and-challenges", "OUT", "recruitment", "Discord servers as recruitment"], ["IEEE-SP/2023/vulnerability-discovery-for-all-experiences-of-marginalization-in-vulnerability", "OUT", "recruitment", "WhatsApp groups/Slack as recruitment"], ["USENIX/2018/how-do-tor-users-interact-with-onion-services", "OUT", "mention", "tool list"], ["USENIX/2023/lost-at-c-a-user-study-on-the-security-implications-of-large-language-model-code", "OUT", "recruitment", "Discord as recruitment"], ["PETS/2026/privacy-by-voice-designing-usable-privacy-notices-for-the-voice-interface", "OUT", "recruitment", "Discord servers as recruitment"], ["USENIX/2025/im-trying-to-learn-and-im-shooting-myself-in-the-foot-beginners-struggles-when-s", "OUT", "recruitment", "Discord channels as recruitment"], ["PETS/2025/free-wifi-is-not-ultimately-free-privacy-perceptions-of-users-in-the-us-regardin", "OUT", "recruitment", "Discord server as recruitment"], ]; // The 44 candidates read in full by the sub-agent readers (notes/msgch_papers_[A-D].md); the other // 103 verdicts were made from sentence contexts (out/msg/triage_ctx.md, written by _msgch_ctx.mjs). export const READ_IN_FULL = [ "USENIX/2025/darkgram-a-large-scale-analysis-of-cybercriminal-activity-channels-on-telegram", "USENIX/2026/stayin-alive-how-global-stolen-data-markets-thrive-on-telegram", "WWW/2026/doxing-as-a-service-demystifying-the-chinese-online-doxing-ecosystem", "USENIX/2019/the-anatomy-of-a-cryptocurrency-pump-and-dump-scheme", "USENIX/2021/catching-phishers-by-their-bait-investigating-the-dutch-phishing-landscape-throu", "USENIX/2021/having-your-cake-and-eating-it-an-analysis-of-concession-abuse-as-a-service", "IMC/2025/unmasking-the-shadow-economy-a-deep-dive-into-drainer-as-a-service-phishing-on-e", "USENIX/2026/from-mirai-to-gorilla-deep-dive-into-a-long-lasting-ddos-for-hire-botnet", "USENIX/2025/assessing-the-aftermath-the-effects-of-a-global-takedown-against-ddos-for-hire-s", "CCS/2020/impersonation-as-a-service-characterizing-the-emerging-criminal-infrastructure-f", "IMC/2025/exploration-of-the-dynamics-of-buy-and-sale-of-social-media-accounts", "WWW/2025/pirates-of-charity-exploring-donation-based-abuses-in-social-media-platforms", "USENIX/2024/the-imitation-game-exploring-brand-impersonation-attacks-on-social-media-platfor", "IMC/2020/demystifying-the-messaging-platforms-ecosystem-through-the-lens-of-twitter", "WWW/2019/mis-information-dissemination-in-whatsapp-gathering-analyzing-and-countermeasure", "WWW/2021/short-is-the-road-that-leads-from-fear-to-hate-fear-speech-in-indian-whatsapp-gr", "USENIX/2025/characterizing-and-detecting-propaganda-spreading-accounts-on-telegram", "WWW/2025/exposing-cross-platform-coordinated-inauthentic-activity-in-the-run-up-to-the-20", "IEEE-SP/2024/no-easy-way-out-the-effectiveness-of-deplatforming-an-extremist-forum-to-suppres", "WWW/2024/getting-bored-of-cyberwar-exploring-the-role-of-low-level-cybercrime-actors-in-t", "IEEE-SP/2025/learning-from-censored-experiences-social-media-discussions-around-censorship-ci", "USENIX/2023/strategies-and-vulnerabilities-of-participants-in-venezuelan-influence-operation", "IMC/2021/a-large-scale-characterization-of-online-incitements-to-harassment-across-platfo", "NDSS/2020/practical-traffic-analysis-attacks-on-secure-messaging-applications", "WWW/2020/the-pod-people-understanding-manipulation-of-social-media-popularity-via-recipro", "CCS/2024/do-anything-now-characterizing-and-evaluating-in-the-wild-jailbreak-prompts-on-l", "USENIX/2024/dont-listen-to-me-understanding-and-exploring-jailbreak-prompts-of-large-languag", "USENIX/2024/moderating-illicit-online-image-promotion-for-unsafe-user-generated-content-game", "WWW/2023/a-prompt-log-analysis-of-text-to-image-generation-systems", "USENIX/2025/bots-can-snoop-uncovering-and-mitigating-privacy-risks-of-bots-in-group-chats", "PETS/2026/bot-among-us-exploring-user-awareness-and-privacy-concerns-about-chatbots-in-gro", "IMC/2022/exploring-the-security-and-privacy-risks-of-chatbots-in-messaging-services", "USENIX/2025/investigating-the-impact-of-online-community-involvement-on-safety-practices-and", "USENIX/2025/privacy-law-enforcement-under-centralized-governance-a-qualitative-analysis-of-f", "WWW/2019/semi-supervised-graph-classification-a-hierarchical-graph-perspective", "USENIX/2023/token-spammers-rug-pulls-and-sniper-bots-an-analysis-of-the-ecosystem-of-tokens", "WWW/2025/detecting-and-understanding-the-promotion-of-illicit-goods-and-services-on-twitt", "CCS/2019/the-art-and-craft-of-fraudulent-app-promotion-in-google-play", "USENIX/2026/cracks-in-the-walled-garden-dissecting-the-gray-market-of-unauthorized-ios-app-d", "NDSS/2024/like-comment-get-scammed-characterizing-comment-scams-on-media-platforms", "NDSS/2021/all-the-numbers-are-us-large-scale-abuse-of-contact-discovery-in-mobile-messengers", "NDSS/2026/hey-there-you-are-using-whatsapp-enumerating-three-billion-accounts-for-security-and-privacy", "NDSS/2026/connecting-the-dots-an-investigative-study-on-linking-private-user-data-across-messaging-apps", "NDSS/2012/guess-who-s-texting-you-evaluating-the-security-of-smartphone-messaging-applicat", ]; // Hand codes for every IN paper, from the paper's own text (quotes in notes/msgch_papers_[A-C].md). // "n/s" = the paper does not state the number. Multi-valued fields are arrays. // discovery: directory | in-app-search | other-platform-links | web-search | snowball | prior-list // (a third party's list or vendor) | known-official-channel | not-stated // join: joined | public-read (the paper says it read a public channel, not whether it joined) | // scraper-service | vendor | not-stated (the paper does not say how it got in; "Both channels // permit public access" describes the channel, not the researchers — coded not-stated) // history: yes (back-filled messages from before collection began) | mixed | not-stated // irb: approved | exempt | not-required | no-board-named | other-part-only | not-stated // passive: stated = the paper says it did not post, interact or contact members // denominator: acknowledged = the paper says groups found are not groups that exist, or names its seed bias, // for the dataset the page reports (a caveat about a secondary sub-sample does not count) export const HAND_FIELDS = ['platforms', 'object', 'discovery', 'join', 'history', 'members', 'tooling', 'limits', 'irb', 'passive', 'minimise', 'denominator', 'found', 'collected', 'messages']; export const HAND = { "USENIX/2025/darkgram-a-large-scale-analysis-of-cybercriminal-activity-channels-on-telegram": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["directory"], "join": "not-stated", "history": "not-stated", "members": "admin-only lists, not collected", "tooling": ["Telegram API (client unnamed)"], "limits": ["rate-limit", "channels-removed"], "irb": "not-required", "passive": "stated", "minimise": ["no payload download", "PII anonymised in release"], "denominator": "acknowledged", "found": "4,709", "collected": "339", "messages": "64,801"}, "USENIX/2026/stayin-alive-how-global-stolen-data-markets-thrive-on-telegram": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["in-app-search", "other-platform-links", "snowball"], "join": "joined", "history": "yes", "members": "posters only", "tooling": ["Telethon"], "limits": ["rate-limit", "channels-removed"], "irb": "approved", "passive": "stated", "minimise": ["file type and size limits", "no de-anonymisation", "controlled-access release"], "denominator": "acknowledged", "found": "21k", "collected": "1,521", "messages": "~14M"}, "WWW/2026/doxing-as-a-service-demystifying-the-chinese-online-doxing-ecosystem": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["directory", "in-app-search"], "join": "joined", "history": "yes", "members": "no", "tooling": ["Telegram API (client unnamed)"], "limits": ["channels-removed"], "irb": "no-board-named", "passive": "stated", "minimise": ["masked before storage"], "denominator": "not-stated", "found": "312", "collected": "100", "messages": "25,972"}, "USENIX/2019/the-anatomy-of-a-cryptocurrency-pump-and-dump-scheme": {"platforms": ["Telegram"], "object": "crypto manipulation", "discovery": ["prior-list"], "join": "not-stated", "history": "yes", "members": "counts only", "tooling": ["Telegram API (client unnamed)", "PumpOlymp API"], "limits": ["channels-removed"], "irb": "not-stated", "passive": "not-stated", "minimise": [], "denominator": "acknowledged", "found": "300+", "collected": "300+", "messages": "n/s"}, "USENIX/2021/catching-phishers-by-their-bait-investigating-the-dutch-phishing-landscape-throu": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["snowball"], "join": "public-read", "history": "not-stated", "members": "no", "tooling": ["manual"], "limits": [], "irb": "not-stated", "passive": "not-stated", "minimise": [], "denominator": "not-stated", "found": "n/s", "collected": "50", "messages": "n/s"}, "USENIX/2021/having-your-cake-and-eating-it-an-analysis-of-concession-abuse-as-a-service": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["other-platform-links"], "join": "joined", "history": "not-stated", "members": "member profiles", "tooling": ["not-stated"], "limits": [], "irb": "approved", "passive": "not-stated", "minimise": ["anonymised"], "denominator": "n/a (known group or channel)", "found": "1", "collected": "1", "messages": "17,898"}, "IMC/2025/unmasking-the-shadow-economy-a-deep-dive-into-drainer-as-a-service-phishing-on-e": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["in-app-search", "other-platform-links"], "join": "joined", "history": "not-stated", "members": "no", "tooling": ["manual"], "limits": [], "irb": "no-board-named", "passive": "not-stated", "minimise": [], "denominator": "acknowledged", "found": "n/s", "collected": "n/s", "messages": "n/s"}, "USENIX/2026/from-mirai-to-gorilla-deep-dive-into-a-long-lasting-ddos-for-hire-botnet": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["known-official-channel"], "join": "public-read", "history": "not-stated", "members": "counts only", "tooling": ["not-stated"], "limits": ["channels-removed"], "irb": "not-stated", "passive": "not-stated", "minimise": [], "denominator": "n/a (known group or channel)", "found": "1", "collected": "1", "messages": "n/s"}, "USENIX/2025/assessing-the-aftermath-the-effects-of-a-global-takedown-against-ddos-for-hire-s": {"platforms": ["Telegram", "Discord"], "object": "cybercrime market", "discovery": ["prior-list"], "join": "public-read", "history": "not-stated", "members": "no", "tooling": ["Telethon"], "limits": ["channels-removed"], "irb": "approved", "passive": "not-stated", "minimise": ["quotes paraphrased", "aggregate analysis"], "denominator": "acknowledged", "found": "n/s", "collected": "52", "messages": "34,438"}, "WWW/2025/pirates-of-charity-exploring-donation-based-abuses-in-social-media-platforms": {"platforms": ["Telegram"], "object": "scams and impersonation", "discovery": ["in-app-search"], "join": "scraper-service", "history": "not-stated", "members": "no", "tooling": ["commercial scraper (Apify, Telemetrio)"], "limits": [], "irb": "not-stated", "passive": "stated", "minimise": [], "denominator": "not-stated", "found": "n/s", "collected": "n/s", "messages": "85,402"}, "USENIX/2024/the-imitation-game-exploring-brand-impersonation-attacks-on-social-media-platfor": {"platforms": ["Telegram"], "object": "scams and impersonation", "discovery": ["in-app-search"], "join": "scraper-service", "history": "not-stated", "members": "no", "tooling": ["commercial scraper (Apify, Telemetrio)"], "limits": [], "irb": "not-stated", "passive": "not-stated", "minimise": [], "denominator": "acknowledged", "found": "n/s", "collected": "n/s", "messages": "133,399"}, "IMC/2020/demystifying-the-messaging-platforms-ecosystem-through-the-lens-of-twitter": {"platforms": ["WhatsApp", "Telegram", "Discord"], "object": "the messaging ecosystem itself", "discovery": ["other-platform-links"], "join": "joined", "history": "mixed", "members": "member lists (phone numbers hashed)", "tooling": ["WhatsApp Web client", "Telegram API (client unnamed)", "Discord API (user account)"], "limits": ["join-cap", "invite-expiry"], "irb": "approved", "passive": "not-stated", "minimise": ["phone numbers hashed"], "denominator": "acknowledged", "found": "351,535", "collected": "616", "messages": "8,255,069"}, "WWW/2019/mis-information-dissemination-in-whatsapp-gathering-analyzing-and-countermeasure": {"platforms": ["WhatsApp"], "object": "politics and misinformation", "discovery": ["web-search"], "join": "joined", "history": "not-stated", "members": "no", "tooling": ["phones + Garimella-Tyson tool"], "limits": ["device-capacity"], "irb": "not-stated", "passive": "not-stated", "minimise": ["identifiers pseudonymised"], "denominator": "acknowledged", "found": "3,444", "collected": "141 + 364", "messages": "121,781 + 789,914"}, "WWW/2021/short-is-the-road-that-leads-from-fear-to-hate-fear-speech-in-indian-whatsapp-gr": {"platforms": ["WhatsApp"], "object": "politics and misinformation", "discovery": ["directory", "web-search"], "join": "joined", "history": "not-stated", "members": "no", "tooling": ["phones + Garimella-Tyson tool"], "limits": [], "irb": "exempt", "passive": "not-stated", "minimise": ["PII anonymised, stored separately"], "denominator": "acknowledged", "found": "n/s", "collected": "5,010", "messages": "1,426,482"}, "USENIX/2025/characterizing-and-detecting-propaganda-spreading-accounts-on-telegram": {"platforms": ["Telegram"], "object": "politics and misinformation", "discovery": ["directory", "in-app-search"], "join": "joined", "history": "yes", "members": "no", "tooling": ["Telethon"], "limits": [], "irb": "approved", "passive": "stated", "minimise": ["users anonymised", "deletion on request"], "denominator": "acknowledged", "found": "n/s", "collected": "13", "messages": "17.3M"}, "WWW/2025/exposing-cross-platform-coordinated-inauthentic-activity-in-the-run-up-to-the-20": {"platforms": ["Telegram"], "object": "politics and misinformation", "discovery": ["in-app-search"], "join": "not-stated", "history": "not-stated", "members": "no", "tooling": ["Telegram API (client unnamed)"], "limits": [], "irb": "not-stated", "passive": "not-stated", "minimise": [], "denominator": "not-stated", "found": "n/s", "collected": "15,537", "messages": "4,309,880"}, "IEEE-SP/2024/no-easy-way-out-the-effectiveness-of-deplatforming-an-extremist-forum-to-suppres": {"platforms": ["Telegram"], "object": "extremism and harassment", "discovery": ["known-official-channel"], "join": "not-stated", "history": "yes", "members": "no", "tooling": ["Telethon"], "limits": [], "irb": "approved", "passive": "not-stated", "minimise": [], "denominator": "n/a (known group or channel)", "found": "2", "collected": "2", "messages": "525k"}, "WWW/2024/getting-bored-of-cyberwar-exploring-the-role-of-low-level-cybercrime-actors-in-t": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["known-official-channel"], "join": "public-read", "history": "yes", "members": "counts only", "tooling": ["Telethon"], "limits": [], "irb": "approved", "passive": "not-stated", "minimise": ["aggregate reporting"], "denominator": "n/a (known group or channel)", "found": "1", "collected": "1", "messages": "441 + 57,757 replies"}, "IEEE-SP/2025/learning-from-censored-experiences-social-media-discussions-around-censorship-ci": {"platforms": ["Telegram"], "object": "censorship circumvention", "discovery": ["directory", "other-platform-links"], "join": "public-read", "history": "not-stated", "members": "no", "tooling": ["Telethon"], "limits": [], "irb": "not-stated", "passive": "not-stated", "minimise": ["usernames masked"], "denominator": "acknowledged", "found": "n/s", "collected": "81", "messages": "54K"}, "USENIX/2023/strategies-and-vulnerabilities-of-participants-in-venezuelan-influence-operation": {"platforms": ["Telegram"], "object": "politics and misinformation", "discovery": ["in-app-search", "snowball"], "join": "joined", "history": "not-stated", "members": "counts only", "tooling": ["manual"], "limits": [], "irb": "approved", "passive": "stated", "minimise": ["de-identified"], "denominator": "not-stated", "found": "6", "collected": "6", "messages": "n/s"}, "IMC/2021/a-large-scale-characterization-of-online-incitements-to-harassment-across-platfo": {"platforms": ["Telegram", "Discord"], "object": "extremism and harassment", "discovery": ["prior-list"], "join": "vendor", "history": "not-stated", "members": "counts only", "tooling": ["vendor crawler"], "limits": [], "irb": "approved", "passive": "stated", "minimise": ["aggregate release only"], "denominator": "acknowledged", "found": "n/s", "collected": "2,916 (Telegram)", "messages": "n/s"}, "NDSS/2020/practical-traffic-analysis-attacks-on-secure-messaging-applications": {"platforms": ["Telegram"], "object": "traffic model for an attack", "discovery": ["not-stated"], "join": "joined", "history": "not-stated", "members": "no", "tooling": ["Telegram API (client unnamed)"], "limits": [], "irb": "not-stated", "passive": "not-stated", "minimise": [], "denominator": "not-stated", "found": "n/s", "collected": "1,000+", "messages": "n/s"}, "WWW/2020/the-pod-people-understanding-manipulation-of-social-media-popularity-via-recipro": {"platforms": ["Telegram"], "object": "engagement manipulation", "discovery": ["web-search", "snowball"], "join": "public-read", "history": "yes", "members": "counts only", "tooling": ["Telegram API (client unnamed)"], "limits": ["data-gap"], "irb": "not-stated", "passive": "not-stated", "minimise": [], "denominator": "acknowledged", "found": "38,000 URLs", "collected": "873", "messages": "n/s"}, "CCS/2024/do-anything-now-characterizing-and-evaluating-in-the-wild-jailbreak-prompts-on-l": {"platforms": ["Discord"], "object": "AI prompts", "discovery": ["directory"], "join": "not-stated", "history": "not-stated", "members": "no", "tooling": ["not-stated"], "limits": [], "irb": "not-required", "passive": "not-stated", "minimise": ["no de-anonymisation"], "denominator": "not-stated", "found": "20", "collected": "6", "messages": "n/s"}, "USENIX/2024/dont-listen-to-me-understanding-and-exploring-jailbreak-prompts-of-large-languag": {"platforms": ["Discord"], "object": "AI prompts", "discovery": ["not-stated"], "join": "not-stated", "history": "not-stated", "members": "no", "tooling": ["Selenium"], "limits": [], "irb": "other-part-only", "passive": "not-stated", "minimise": [], "denominator": "not-stated", "found": "n/s", "collected": "2", "messages": "n/s"}, "USENIX/2024/moderating-illicit-online-image-promotion-for-unsafe-user-generated-content-game": {"platforms": ["Discord"], "object": "unsafe content", "discovery": ["directory"], "join": "joined", "history": "not-stated", "members": "no", "tooling": ["manual"], "limits": [], "irb": "approved", "passive": "not-stated", "minimise": ["account mentions removed"], "denominator": "acknowledged", "found": "n/s", "collected": "n/s", "messages": "210 images"}, };
The figure and quote verifier's output
- verify_messaging_channels_figures-output.txt
control ok (mutated needle not located) hoseini2020_demystifying "the limit for WhatsApp is between 250 and 350 groups per user" control ok (mutated needle not located) marjanov2026_stayin "identify 0.9% of the total stolen data channels through the" control ok (mutated needle not located) roy2025_darkgram "allowed us to discover 4,790 English-based channels" == A. Every //"…"// span on the page == OK EXTERNAL line 12 Telegram Content Licensing: access beyond ordinary use prohibited "for any purpose other than ordinary, legitimate, and intended use of the Telegram platform as its user" OK paper.cols.txt line 24 hoseini2020_demystifying "For WhatsApp, even without an account, we could collect an impressive number of over 34K phone numbers. Moreover, after joining groups, we obtain another 20K ph" OK paper.cols.txt line 25 resende2019_information "was constrained by the available devices and their resources (memory)" OK paper.cols.txt line 26 marjanov2026_stayin "We do not lie or pretend to be an interested buyer to be admitted into groups" OK paper.cols.txt line 27 roy2025_darkgram "this approach might have introduced potential biases by omitting smaller or newly emerging channels" OK paper.cols.txt line 27 roy2025_darkgram "only 64 channels (19%) were removed" OK paper.cols.txt line 27 roy2025_darkgram "led to the removal of all 196 channels, with a median response time of 4 days" OK paper.cols.txt line 28 kireev2025_characterizing "officially registered as such on the Telegram website" OK paper.cols.txt line 28 kireev2025_characterizing "returned all messages from either the past 36 months or up to a limit" OK paper.cols.txt line 28 kireev2025_characterizing "given the small selection, we cannot make any statement about the pervasiveness of this propaganda activity in Telegram" OK paper.cols.txt line 62 hoseini2020_demystifying "presumably owing to Discord group URLs automatically expiring after a day" OK EXTERNAL line 62 Discord invites default to 7 days "a 7 days access link by default" OK EXTERNAL line 63 Discord Server Discovery needs 1,000 members "at least 1,000 members" OK EXTERNAL line 63 TGStat self-reported catalogue size "More than 2 864 885 channels and groups" OK EXTERNAL line 64 durov/345: problematic content no longer accessible in Search "All the problematic content we identified in Search is no longer accessible" OK EXTERNAL line 74 Telegram bots: privacy mode on by default unless added as admin "Privacy mode is enabled by default for all bots, except bots that were added to a group as admins" OK EXTERNAL line 75 Telegram documents the public-channel web preview "The contents of public channels can be seen on the Web without a Telegram account" OK paper.cols.txt line 82 kireev2025_characterizing "never interacted with the channels" OK paper.cols.txt line 83 hoseini2020_demystifying "translates to a need for hundreds of phones and SIM cards to join all discovered groups" OK EXTERNAL line 83 Telegram: unofficial-client logins put under observation "all accounts that log in using unofficial Telegram API clients are automatically put under observation" OK paper.cols.txt line 89 resende2019_information "we are not aware of an approach that would allow us to assess the representativeness of our data as even the total number of groups available in the country is " OK paper.cols.txt line 90 saha2021_short "Given that WhatsApp does not provide an API or tools to access the data, there is no way of knowing the representativeness of our dataset" OK paper.cols.txt line 90 saha2021_short "a convenience sample" OK paper.cols.txt+lig line 91 hoseini2020_demystifying "The use of Twitter as the only data source for discovering public groups of the different messaging platforms potentially introduces some bias in our sample" OK paper.cols.txt line 97 marjanov2026_stayin "inactive/banned" OK paper.cols.txt line 97 marjanov2026_stayin "the most influential predictor of channel durability" OK paper.cols.txt line 97 marjanov2026_stayin (WARN: located in a citekey other than the nearest) "might miss short-lived channels due to the retroactive nature of data collection" OK paper.cols.txt line 97 kireev2025_characterizing "The data collected through "Export chat history" does not contain deleted messages" OK paper.cols.txt line 99 arunasalam2024_security "that can only be joined via invitation" OK paper.cols.txt line 99 arunasalam2024_security "an unintentional leak of the group's "invite link"" OK NOT-A-QUOTE line 102 template of the sentence NOT to write, in the denominator box "we collected N messages from Telegram" OK EXTERNAL line 114 Telethon README: moved to Codeberg "Moved to https://codeberg.org/Lonami/Telethon" OK EXTERNAL line 114 Telethon README (reST links rendered): be careful not to break ToS "be careful not to break Telegram's ToS or Telegram can ban the account" OK EXTERNAL line 118 whatsapp-web.js README: WhatsApp does not allow unofficial clients "WhatsApp does not allow bots or unofficial clients on their platform, so this shouldn't be considered totally safe" OK EXTERNAL line 121 DiscordChatExporter README: automating user accounts is against Discord TOS "automating user accounts is against Discord TOS and may result in you getting banned" OK EXTERNAL line 125 Meta Content Library covers WhatsApp Channels "provide[s] comprehensive access to the public content archive from Facebook, Instagram and WhatsApp Channels" OK EXTERNAL line 127 Telegram ToS: scraping prohibited for all users and third parties "Telegram additionally prohibits data scraping as part of its Content Licensing and AI Scraping Terms, which apply to all users, businesses, and third-party serv" OK EXTERNAL line 127 Telegram Content Licensing: access beyond ordinary use prohibited "Access to user-generated content for any purpose other than ordinary, legitimate, and intended use of the Telegram platform as its user is prohibited" OK EXTERNAL line 127 Telegram API ToS 1.5: no ML training on API data "prohibited from using, accessing or aggregating data obtained from the Telegram platform to train, fine-tune or otherwise engage in the development" OK EXTERNAL line 127 Telegram self-reports below the VLOP threshold "significantly fewer than 45 million" OK EXTERNAL line 128 WhatsApp ToS: no access or collection through automated means "through automated or other means" OK EXTERNAL line 128 WhatsApp ToS: no collecting information about users in impermissible ways "collect information of or about our users in any impermissible or unauthorized manner" OK EXTERNAL line 128 EC: private messaging stays excluded "private messaging service" OK EXTERNAL line 129 Discord ToS: no scraping without written consent "scraping our services without our written consent" OK EXTERNAL line 129 Discord Community Guidelines: no self-bots "Do not use self-bots or user-bots" OK EXTERNAL line 129 Discord Developer Policy: do not mine or scrape "Do not mine or scrape any data" OK EXTERNAL line 134 FLOOD_WAIT_X: a wait of X seconds is required "A wait of X seconds is required" OK paper.cols.txt line 134 roy2025_darkgram "at 10-minute intervals to capture new posts while adhering to the API rate limits" OK paper.cols.txt line 135 vu2025_assessing "often banned rapidly by Discord" OK pypdf line 142 schrittwieser2012_guess "the WhatsApp server did not prevent us from uploading ten million phone numbers and returned 21095 valid phone numbers" OK paper.cols.txt line 143 hagen2021_numbers "accounts get banned when excessively using the contact discovery service" OK pypdf line 145 gegenhuber2026_there "We encountered no rate limits, our accounts were not banned from the platform" OK paper.cols.txt line 159 kireev2025_characterizing "never interacted with the channels" OK paper.cols.txt line 165 marjanov2026_stayin "We do not lie or pretend to be an interested buyer to be admitted into groups. We also do not attempt to join any groups that require payments or vouching by an" OK paper.cols.txt line 165 he2025_unmasking "joined several related Telegram groups" OK paper.cols.txt line 165 he2025_unmasking "communicated with operators, acquired wallet drainers" OK paper.cols.txt line 166 vu2025_assessing "sending thousands of messages could be regarded as spamming" OK paper.cols.txt line 167 roy2025_darkgram "we did not download or read the payload files" OK pypdf line 168 marjanov2026_stayin "to avoid sending stolen and potentially sensitive data to third-party servers" OK paper.cols.txt line 168 roy2025_darkgram "we utilized the GPT-4 API" OK pypdf+dehyph line 169 li2025_investigating "Although PWUD is also active in other online communities such as Telegram groups and self-constructed forums, for ethical reasons, we limited our online data co" OK paper.cols.txt line 171 albrecht2021_collective "all participants in our study also assumed police monitoring of the public Telegram groups" OK EXTERNAL line 173 Barbosa and Milan 2019: avoid by all means covert bypasses "by all means covert bypasses" OK paper.cols.txt line 262 hoseini2020_demystifying "exposes at least one social media account for 30% of the Discord users we monitored" OK paper.cols.txt line 265 roy2025_darkgram "only 64 channels (19%) were removed" OK paper.cols.txt line 265 roy2025_darkgram "led to the removal of all 196 channels, with a median response time of 4 days" OK paper.cols.txt line 266 gao2026_doxing "over 300,000 unique individuals" OK pypdf line 268 saha2021_short "8% of these fear speech users are also admins in the groups where they post fear speech" spans: 68; by route: EXTERNAL 25, paper.cols.txt 36, paper.cols.txt+lig 1, NOT-A-QUOTE 1, pypdf 4, pypdf+dehyph 1 == B. Per-paper figures == OK paper.cols.txt hoseini2020_demystifying | 351,535 URLs "we discover 351,535 group URLs" OK paper.cols.txt hoseini2020_demystifying | Table 2 totals: 2,234,128 tweets, 616 joined "Total 2,234,128 806,372 351,535 616 8,255,069 761,712" OK pypdf hoseini2020_demystifying | 416 WhatsApp groups "In total, we select and join 416 random public groups" OK paper.cols.txt hoseini2020_demystifying | join caps "the limit for WhatsApp is between 250 and 300 groups per user, while on Discord it is up to 100 servers" OK pypdf hoseini2020_demystifying | 24 of 100 member lists "we obtain, hence, the member list only in 24 groups (out of the 100)" OK paper.cols.txt hoseini2020_demystifying | 49% Discord member coverage "representing 10.8% and 49% for Telegram and Discord, respectively" OK pypdf hoseini2020_demystifying | WhatsApp history from join "WhatsApp gives access to messages shared on the group, after our joining date" OK paper.cols.txt hoseini2020_demystifying | 38-day window "between April 8 and May 15, 2020" OK paper.cols.txt hoseini2020_demystifying | hashed phone numbers "we store only a hash of the phone numbers" OK paper.cols.txt hoseini2020_demystifying | ethics approval "obtained approval prior to collecting any data" OK paper.cols.txt resende2019_information | 3,444 / 1,828 "we found 3,444 distinct links for publicly accessible groups, out of which only 1,828 were valid" OK paper.cols.txt resende2019_information | 141 and 364 groups "with 141 and 364 groups monitored, respectively" OK paper.cols.txt resende2019_information | 11,728 URLs "a total of 11,728 URLs were shared" OK paper.cols.txt resende2019_information | 92,654 URLs "this number reached 92,654 URLs" OK paper.cols.txt resende2019_information | joined with phones "This monitoring involves joining each group using a cell phone" OK paper.cols.txt saha2021_short | 5,000 groups "we joined and monitored over 5,000 political groups discussing politics" OK paper.cols.txt saha2021_short | exempt "approved the data collection as exempt" OK pypdf saha2021_short | ~8,000 messages, ~1,000 groups, ~3,000 users (the text renders thousands as "8, 000") "000 messages spread across∼ 1, 000 groups and spread by∼ 3, 000 users" OK pypdf saha2021_short | ~8,000 fear-speech messages "fear speech to∼ 8, 000 messages" OK paper.cols.txt saha2021_short | 1.43M posts, 5,010 groups "#posts 1,426,482 #groups present 5,010 #users present 109,542" OK paper.cols.txt marjanov2026_stayin | 21k candidates "we identify 21k candidates and examine roughly half of them" OK paper.cols.txt marjanov2026_stayin | 1,500 joined "we join over 1,500 channels that fit the inclusion criteria" OK paper.cols.txt marjanov2026_stayin | 1,521 / 448 private / 14M "approximately 14 million messages across 1,521 channels, of which 1,073 (70%) are public and 448 (30%) are private" OK paper.cols.txt marjanov2026_stayin | 0.8% overlap "identify 0.8% of the total stolen data channels through the" OK paper.cols.txt marjanov2026_stayin | 79% inactive/banned "1,006 (79%) are inactive/banned" OK paper.cols.txt marjanov2026_stayin | 1,282 non-gateway "We have 1,282 non-gateway channels that were classified as containing stolen data" OK paper.cols.txt marjanov2026_stayin | BSC guidance "British Society of Criminology" OK paper.cols.txt marjanov2026_stayin | Telethon "custom-built Telethon data collection tool" OK pypdf marjanov2026_stayin | local LLM "we use local models instead of cloud-based LLMs" OK paper.cols.txt marjanov2026_stayin | collection window "between August 2024 and August 2025" OK paper.cols.txt roy2025_darkgram | 4,709 channels "allowed us to discover 4,709 English-based channels with 10,000 or more followers" OK paper.cols.txt roy2025_darkgram | 339 kept "we identified 339 channels that exhibited one or more of these characteristics" OK paper.cols.txt roy2025_darkgram | Telemetr.io seed "we utilized Telemetr.io" OK paper.cols.txt roy2025_darkgram | IRB not required "deemed not to require an Institutional Review Board (IRB) approval" OK paper.cols.txt roy2025_darkgram | admin-only follower lists "only channel administrators can access follower lists of a channel" OK pypdf kireev2025_characterizing | 17.3M / 13 channels "17.3M labeled messages from 13 political and news-oriented channels" OK paper.cols.txt kireev2025_characterizing | 78.37K / 6,250 "We found 78.37K propaganda messages (1.8% of the dataset), sent by 6,250 propaganda accounts (2.2% of accounts)" OK paper.cols.txt WEAK kireev2025_characterizing | TGStat seed "TGStats catalog" OK paper.cols.txt kireev2025_characterizing | IRB "have been approved by the IRB of our institution" OK paper.cols.txt gao2026_doxing | 312 candidates "retrieved a raw dataset of 312 candidate channels" OK paper.cols.txt gao2026_doxing | top 100 "we selected the top 100 channels by subscriber count" OK paper.cols.txt gao2026_doxing | 411,707 messages, five groups (the text renders the number with a stray space) "we analyzed 411, 707 messages collected from five query groups" OK paper.cols.txt gao2026_doxing | window "from May 30, 2025 to August 25, 2025" OK paper.cols.txt gao2026_doxing | "representative" (the paper's own word for its top 100) "we selected the top 100 channels by subscriber count as our representative corpus" OK paper.cols.txt acharya2025_pirates | the sentence the extraction reads as a review decision "Our research did not directly involve interaction with any human subjects" OK paper.cols.txt roy2025_darkgram | 64 of 339 removed "only 64 channels (19%) were removed" OK paper.cols.txt roy2025_darkgram | 196 new channels found via Telegram and Facebook "takedown 196 new CACs shared on Telegram and Facebook" OK paper.cols.txt vu2025_assessing | Discord channels monitored too "Booters may use Discord; we also monitored these channels" OK paper.cols.txt kireev2025_characterizing | over 80% removed by moderators "more than 80% of propaganda messages removed" OK paper.cols.txt kireev2025_characterizing | below 20% removed "ranging from below 20%" OK paper.cols.txt guo2024_moderating | manual collection "constrained by a manually collected dataset" OK paper.cols.txt roy2025_darkgram | 11,800 posts to the GPT-4 API "across all the 11,800 posts identified by the classifier, we utilized the GPT-4 API" OK paper.cols.txt xu2019_anatomy | 300+ channels "we trace the message history of over 300 Telegram channels" OK paper.cols.txt xu2019_anatomy | 43 deleted "Among those channels, 43 have been deleted from the Telegram sever" OK paper.cols.txt WEAK xu2019_anatomy | PumpOlymp list "PumpOlymp" OK paper.cols.txt weerasinghe2020_people | 38,000 URLs "ultimately collecting a total of 38,000 group URLs" OK paper.cols.txt weerasinghe2020_people | 4,425 -> 873 "we were left with 4,425 URLs. Out of them, 873 belonged to currently active, public Telegram groups" OK paper.cols.txt weerasinghe2020_people | 432 pods "we identified 432 Telegram groups as pods" OK paper.cols.txt vu2025_assessing | 34,438 messages, 52 channels "We collected 34 438 messages, 5 246 replies, and associated metadata including 6 290 emoji reactions in 52 channels" OK paper.cols.txt vu2025_assessing | paraphrased quotes "All quotes were paraphrased to prevent attribution" OK paper.cols.txt vu2024_easy | 525k messages "encompassing 525k messages" OK paper.cols.txt vu2024_getting | 441 posts / 57,757 replies "We collect 441 announcements with 57 757 replies" OK pypdf vafa2025_learning | 81 channels "the 81 channels shared 1,459 unique VPN installation files" OK paper.cols.txt aliapoulios2021_characterization | 2,916 channels "The data collection from Telegram spans 2,916 different channels with 126,432 users" OK pypdf bijmans2021_catching | 50 channels "manually inspecting 50 public Telegram channels" OK pypdf sun2021_having | 17,898 messages "We found 1,076 members posted 17,898 messages within this period" OK paper.cols.txt recabarren2023_strategies | six groups "we have identified and joined six Telegram groups" OK paper.cols.txt shen2024_anything | top 20 servers "we manually inspect the top 20 servers with the most members" OK paper.cols.txt shen2024_anything | six servers "In the end, we discover six Discord servers" OK paper.cols.txt WEAK acharya2024_imitation | 133,399 posts "Telegram (133,399)" OK paper.cols.txt acharya2025_pirates | 85,402 Telegram posts (denominator of the filter) "(76,111/85,402) from Telegram were filtered" OK paper.cols.txt cinus2025_exposing | 15,537 channels "Telegram 15,537 4,309,880" OK paper.cols.txt guo2024_moderating | 210 images "leading to the discovery of 210 instances for image" OK pypdf schrittwieser2012_guess | 21095 "returned 21095 valid phone numbers that are using the WhatsApp application" OK pypdf schrittwieser2012_guess | 2.5 hours "The entire process finished in less than 2.5 hours" OK paper.cols.txt hagen2021_numbers | 10% / 100% "we were able to query 10 % of all US mobile phone numbers for WhatsApp and 100 % for Signal" OK paper.cols.txt hagen2021_numbers | Telegram 100,000 "only 100,000 numbers were checked during that time" OK paper.cols.txt gegenhuber2026_there | 3,546,479,731 "we discovered a total of 3,546,479,731 accounts" OK paper.cols.txt gegenhuber2026_there | 7,000/s "With our query rate of 7,000 phone numbers per second (and session)" OK paper.cols.txt kang2026_connecting | IRB "We obtained IRB approval for our investigative studies" OK paper.cols.txt kang2026_connecting | no profile images "we did not store any profile images" OK paper.cols.txt chou2025_bots | Pushshift = channels only "The Pushshift dataset includes only Telegram "channels,"" OK paper.cols.txt WEAK yu2024_listen | yu2024: Discord named (weak by design; role checked in the hand codes) "Discord" needles: 83; by route: paper.cols.txt 71, pypdf 12; weak (<20 chars): 4 === EXTERNAL FIGURES (non-corpus; each re-fetched by external_checks_messaging_channels.sh — see its labels) === Telegram: channels_limit_default 500, channels_limit_premium 1000, recommended_channels_limit_default 10 (core.telegram.org/api/config) Telegram: one api_id per phone number; People Nearby removed 2024-09-06; search clean-up and disclosure change 2024-09-23 Telethon 1.45.0 (PyPI, 2026-09-10); TDLib tags stop at v1.8.0 (2021); discord.py 2.7.1 (2026-03-03); DiscordChatExporter 2.48 (2026-08-27) yowsup last commit 2021-12; Python <= 3.7 WhatsApp designated VLOP 2026-01-26 (Channels); Telegram below 45 million EU recipients (self-reported, August 2026) Discord: invites 7 days by default; Server Discovery 1,000 members and 8 weeks TGStat: 2 864 885 channels and groups; about 1.69 million channels in Russia (its own country tiles); Telemetr.io 11M+ / 7M+ (Web Archive 2026-09-26) Pushshift Telegram: 27.8K channels, 317M messages (ICWSM 2020); TGDataset 120,979 channels (KDD 2025); TeraGram 5.9 billion messages, 712 thousand channels and groups (arXiv 2605.15956) AoIR 3.0 approved 6 October 2019, 83 pages; Barbosa and Milan 2019; Garimella and Chauchard 2025 === CORPUS FIGURES (report_messaging_channels.mjs) are checked by check_page_numbers.mjs against its output === not located / failed: 0
The external checks
- external_checks_messaging_channels.sh
#!/usr/bin/env bash # external_checks_messaging_channels.sh — re-fetch every external fact design:platforms:messaging_channels # leans on and print the evidence. Each check prints OK or FAILED; the exit status is the number of # FAILED checks (capped at 255), so a green run is a real assertion, not a transcript. # bash scripts/external_checks_messaging_channels.sh > scripts/external_checks_messaging_channels-output.txt # Honours GH_TOKEN (never printed). Needs curl, python3, node + Playwright's headless shell for the # Cloudflare-walled pages (PLAYWRIGHT_BROWSERS_PATH=/workspace/.playwright). set -u cd "$(dirname "$0")/.." UA='Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/130 Safari/537.36' FAILS=0 TMP=$(mktemp -d) AUTH=() if [ -n "${GH_TOKEN:+x}" ]; then AUTH=(-H "Authorization: Bearer $GH_TOKEN"); fi export PLAYWRIGHT_BROWSERS_PATH=/workspace/.playwright echo "run: $(date -u +%Y-%m-%dT%H:%MZ)" # check <label> <file> <python-regex> — whitespace-collapsed, case-sensitive search in a fetched file check() { if python3 - "$2" "$3" <<'PY' import re, sys, html t = open(sys.argv[1], encoding='utf-8', errors='replace').read() t = html.unescape(re.sub(r'<[^>]+>', ' ', t)) t = re.sub(r'\s+', ' ', t).replace('’', "'").replace('‘', "'").replace('“', '"').replace('”', '"') m = re.search(sys.argv[2], t) if not m: sys.exit(1) s = max(0, m.start() - 60) print(' > ' + t[s:m.end() + 60]) PY then echo "OK $1"; else echo "FAILED $1"; FAILS=$((FAILS+1)); fi } # check_raw: same, but on the raw bytes (for values that live in attributes or class names) check_raw() { if python3 - "$2" "$3" <<'PY' import re, sys t = re.sub(r'\s+', ' ', open(sys.argv[1], encoding='utf-8', errors='replace').read()) m = re.search(sys.argv[2], t) if not m: sys.exit(1) print(' > ' + t[max(0, m.start() - 40):m.end() + 40]) PY then echo "OK $1"; else echo "FAILED $1"; FAILS=$((FAILS+1)); fi } get() { curl -sL -m 40 -A "$UA" "$1" -o "$2"; echo " GET $1 -> $(wc -c < "$2") bytes"; } pw() { node scripts/pw_fetch_text.mjs "$1" "$2" > /dev/null 2>&1; echo " PW $1 -> $(wc -c < "$2" 2>/dev/null || echo 0) bytes"; } gh() { curl -sL -m 40 "${AUTH[@]}" -H 'Accept: application/vnd.github+json' "https://api.github.com/repos/$1" -o "$2"; } echo; echo "== 1. Client tooling (GitHub / Codeberg / PyPI / npm / Go proxy) ==" gh LonamiWebs/Telethon $TMP/telethon.json check 'Telethon GitHub repo is archived' $TMP/telethon.json '"archived": true' get https://raw.githubusercontent.com/LonamiWebs/Telethon/v1/README.rst $TMP/telethon_readme.txt check 'Telethon README: moved to Codeberg' $TMP/telethon_readme.txt 'Moved to https://codeberg.org/Lonami/Telethon' check 'Telethon README: Telegram can ban the account' $TMP/telethon_readme.txt 'Telegram can ban the account' python3 -c "import re;t=open('$TMP/telethon_readme.txt').read();t=re.sub(r'\x60([^\x60]+)\x60_',r'\\1',t);open('$TMP/telethon_readme_plain.txt','w').write(t)" check 'Telethon README (reST links rendered): be careful not to break ToS' $TMP/telethon_readme_plain.txt "be careful not to break Telegram's ToS or Telegram can ban the account" get https://pypi.org/pypi/Telethon/json $TMP/telethon_pypi.json python3 -c "import json;d=json.load(open('$TMP/telethon_pypi.json'));v=d['info']['version'];print(' Telethon PyPI latest',v,d['releases'][v][0]['upload_time']);print(' any 2.x on PyPI:',[r for r in d['releases'] if r.startswith('2.')])" check 'Telethon PyPI latest is 1.45.0' $TMP/telethon_pypi.json '"version": ?"1\.45\.0"' get https://codeberg.org/api/v1/repos/Lonami/Telethon $TMP/telethon_cb.json check 'Telethon Codeberg repo not archived' $TMP/telethon_cb.json '"archived": ?false' gh pyrogram/pyrogram $TMP/pyrogram.json check 'Pyrogram GitHub repo is archived' $TMP/pyrogram.json '"archived": true' get https://raw.githubusercontent.com/pyrogram/pyrogram/master/README.md $TMP/pyrogram_readme.txt check 'Pyrogram README: no longer maintained' $TMP/pyrogram_readme.txt 'no longer maintained' get https://pypi.org/pypi/kurigram/json $TMP/kurigram.json python3 -c "import json;d=json.load(open('$TMP/kurigram.json'));v=d['info']['version'];print(' kurigram PyPI latest',v,d['releases'][v][0]['upload_time'])" check_raw 'Kurigram (Pyrogram fork) is on PyPI' $TMP/kurigram.json '"name": ?"[Kk]urigram"' check_raw 'Kurigram released in 2026' $TMP/kurigram.json '"upload_time": ?"2026-' gh kurigram-org/kurigram $TMP/kurigram_gh.json check 'Kurigram repo not archived' $TMP/kurigram_gh.json '"archived": false' check_raw 'Kurigram describes itself as a Pyrogram fork' $TMP/kurigram.json '[Ff]ork of [Pp]yrogram|[Pp]yrogram fork' get https://registry.npmjs.org/telegram $TMP/npm_telegram.json check_raw 'npm telegram (GramJS) deprecated in favour of teleproto' $TMP/npm_telegram.json '"deprecated":"[^"]*teleproto' gh gram-js/gramjs $TMP/gramjs.json check 'GramJS GitHub repo is archived' $TMP/gramjs.json '"archived": true' gh tdlib/td $TMP/tdlib.json check 'TDLib repo not archived' $TMP/tdlib.json '"archived": false' get https://raw.githubusercontent.com/tdlib/td/master/CMakeLists.txt $TMP/tdlib_cmake.txt check 'TDLib master declares a 1.8.x version' $TMP/tdlib_cmake.txt 'project\(TDLib VERSION 1\.8\.\d+' curl -sL -m 40 "${AUTH[@]}" "https://api.github.com/repos/tdlib/td/tags?per_page=100" -o $TMP/tdlib_tags.json python3 - $TMP/tdlib_tags.json <<'PY' && echo "OK TDLib: newest tag by version is v1.8.0" || { echo "FAILED TDLib: newest tag by version is v1.8.0"; } import json, sys, re tags = [t['name'] for t in json.load(open(sys.argv[1]))] vs = sorted((tuple(int(x) for x in re.findall(r'\d+', t)), t) for t in tags if re.match(r'v\d', t)) print(' > tags:', len(tags), 'newest by version:', vs[-1][1]) sys.exit(0 if vs[-1][1] == 'v1.8.0' else 1) PY [ $? -eq 0 ] || FAILS=$((FAILS+1)) gh wwebjs/whatsapp-web.js $TMP/wwebjs.json check 'whatsapp-web.js lives at wwebjs/whatsapp-web.js, not archived' $TMP/wwebjs.json '"full_name": "wwebjs/whatsapp-web.js".*"archived": false' get https://raw.githubusercontent.com/wwebjs/whatsapp-web.js/main/README.md $TMP/wwebjs_readme.txt check 'whatsapp-web.js README: WhatsApp does not allow unofficial clients' $TMP/wwebjs_readme.txt "WhatsApp does not allow bots or unofficial clients on their platform, so this shouldn't be considered totally safe" get https://registry.npmjs.org/whatsapp-web.js/latest $TMP/wwebjs_npm.json check 'whatsapp-web.js npm latest 1.x' $TMP/wwebjs_npm.json '"version":"1\.\d+\.\d+"' gh WhiskeySockets/Baileys $TMP/baileys.json check 'Baileys not archived' $TMP/baileys.json '"archived": false' get https://proxy.golang.org/go.mau.fi/whatsmeow/@latest $TMP/whatsmeow.json check 'whatsmeow resolves only to a pseudo-version' $TMP/whatsmeow.json '"Version":"v0\.0\.0-\d{14}-[0-9a-f]{12}"' gh tgalal/yowsup $TMP/yowsup.json get "https://api.github.com/repos/tgalal/yowsup/commits?per_page=1" $TMP/yowsup_commits.json check 'yowsup: last default-branch commit is from 2021' $TMP/yowsup_commits.json '"date": ?"2021-' get https://raw.githubusercontent.com/tgalal/yowsup/master/README.md $TMP/yowsup_readme.txt check_raw 'yowsup README: requires python <= 3.7' $TMP/yowsup_readme.txt 'python>=2\.7,<=3\.7' get https://pypi.org/pypi/discord.py/json $TMP/dpy.json check 'discord.py PyPI latest 2.x' $TMP/dpy.json '"version": ?"2\.\d+\.\d+"' python3 -c "import json;d=json.load(open('$TMP/dpy.json'));v=d['info']['version'];print(' discord.py',v,d['releases'][v][0]['upload_time'])" check_raw 'discord.py 2.7.1 uploaded March 2026' $TMP/dpy.json '"2\.7\.1": ?\[\{.*?"upload_time": ?"2026-03-' get https://raw.githubusercontent.com/Tyrrrz/DiscordChatExporter/master/Readme.md $TMP/dce.txt [ -s $TMP/dce.txt ] || get https://raw.githubusercontent.com/Tyrrrz/DiscordChatExporter/prime/Readme.md $TMP/dce.txt check 'DiscordChatExporter README: automating user accounts is against Discord TOS' $TMP/dce.txt 'automating user accounts is against Discord TOS and may result in you getting banned' curl -sL -m 40 "${AUTH[@]}" https://api.github.com/repos/Tyrrrz/DiscordChatExporter/releases/latest -o $TMP/dce_rel.json check 'DiscordChatExporter latest release 2.48 (August 2026)' $TMP/dce_rel.json '"tag_name": ?"2\.48".*"published_at": ?"2026-08-' echo; echo "== 2. Terms ==" get https://telegram.org/tos/content-licensing $TMP/tg_cl.html check 'Telegram Content Licensing: access beyond ordinary use prohibited' $TMP/tg_cl.html 'Access to user-generated content for any purpose other than ordinary, legitimate, and intended use of the Telegram platform as its user is prohibited' check 'Telegram Content Licensing: no scraping for AI/ML training' $TMP/tg_cl.html 'firmly prohibits the scraping, indexing, harvesting, aggregation or use of data obtained from its platform to train' get https://telegram.org/tos $TMP/tg_tos.html check 'Telegram ToS: scraping prohibited for all users and third parties' $TMP/tg_tos.html 'Telegram additionally prohibits data scraping as part of its Content Licensing and AI Scraping Terms ?, which apply to all users, businesses, and third-party services accessing the platform' get https://core.telegram.org/api/terms $TMP/tg_apiterms.html check 'Telegram API ToS 1.5: no ML training on API data' $TMP/tg_apiterms.html 'prohibited from using, accessing or aggregating data obtained from the Telegram platform to train, fine-tune or otherwise engage in the development' get https://core.telegram.org/api/obtaining_api_id $TMP/tg_apiid.html check 'Telegram: one api_id per phone number' $TMP/tg_apiid.html 'each number can only have one api_id connected to it' check 'Telegram: unofficial-client logins put under observation' $TMP/tg_apiid.html 'automatically put under observation' get https://telegram.org/tour/channels $TMP/tg_tour.html check 'Telegram documents the public-channel web preview' $TMP/tg_tour.html 'The contents of public channels can be seen on the Web without a Telegram account' get https://t.me/s/telegram $TMP/tme_s.html check_raw 't.me/s/<channel> serves channel posts' $TMP/tme_s.html 'tgme_widget_message_wrap' check_raw 't.me/s/<channel> paginates with ?before=' $TMP/tme_s.html 'href="/s/telegram\?before=\d+"' pw https://www.whatsapp.com/legal/terms-of-service $TMP/wa_tos.txt check 'WhatsApp ToS: no access or collection through automated means' $TMP/wa_tos.txt 'through automated or other means, access, use, copy' check 'WhatsApp ToS: no collecting information about users in impermissible ways' $TMP/wa_tos.txt 'collect information of or about our users in any impermissible or unauthorized manner' get https://discord.com/terms $TMP/dc_terms.html check 'Discord ToS: no scraping without written consent' $TMP/dc_terms.html 'scraping our services without our written consent' get https://discord.com/guidelines $TMP/dc_guide.html check 'Discord Community Guidelines: no self-bots' $TMP/dc_guide.html 'Do not use self-bots or user-bots' pw https://support-dev.discord.com/hc/en-us/articles/8563934450327-Discord-Developer-Policy $TMP/dc_devpol.txt check 'Discord Developer Policy: do not mine or scrape' $TMP/dc_devpol.txt 'Do not mine or scrape any data' check 'Discord Developer Policy: no ML training on message content' $TMP/dc_devpol.txt 'Do not use message content obtained through the APIs to train machine learning' pw https://support.discord.com/hc/en-us/articles/208866998-Invites-101 $TMP/dc_inv.txt check 'Discord invites default to 7 days' $TMP/dc_inv.txt 'displaying a 7 days access link by default' pw https://support.discord.com/hc/en-us/articles/360030843331-Enabling-Server-Discovery $TMP/dc_disc.txt check 'Discord Server Discovery needs 1,000 members' $TMP/dc_disc.txt 'at least 1,000 members' pw https://transparency.meta.com/researchtools/meta-content-library/ $TMP/meta_mcl.html check_raw 'Meta Content Library covers WhatsApp Channels' $TMP/meta_mcl.html 'provide comprehensive access to the public content archive from Facebook, Instagram and WhatsApp Channels' get https://core.telegram.org/bots/features $TMP/tg_botfeat.html check 'Telegram bots: privacy mode on by default unless added as admin' $TMP/tg_botfeat.html 'Privacy mode is enabled by default for all bots, except bots that were added to a group as admins' echo; echo "== 3. Telegram API limits ==" get https://core.telegram.org/api/config $TMP/tg_cfg.html check 'Telegram default join limit 500 channels/supergroups' $TMP/tg_cfg.html '" channels_limit_default ": 500' check 'Telegram Premium join limit 1000' $TMP/tg_cfg.html '" channels_limit_premium ": 1000' check 'Telegram channel recommendations capped at 10 (non-Premium)' $TMP/tg_cfg.html '" recommended_channels_limit_default ": 10' get https://core.telegram.org/method/channels.joinChannel $TMP/tg_join.html check 'channels.joinChannel lists CHANNELS_TOO_MUCH' $TMP/tg_join.html 'CHANNELS_TOO_MUCH' check 'channels.joinChannel lists INVITE_HASH_EXPIRED' $TMP/tg_join.html 'INVITE_HASH_EXPIRED' get https://core.telegram.org/api/errors $TMP/tg_err.html check 'FLOOD_WAIT_X documented' $TMP/tg_err.html 'FLOOD_WAIT_X' check 'FLOOD_WAIT_X: a wait of X seconds is required' $TMP/tg_err.html 'A wait of X seconds is required' echo; echo "== 4. Platform changes ==" get "https://t.me/durov/343?embed=1" $TMP/durov343.html check 'People Nearby removed (Durov, 2024-09-06)' $TMP/durov343.html "removed the People Nearby feature" check_raw 'durov/343 dated 2024-09-06' $TMP/durov343.html '2024-09-06' get "https://t.me/durov/345?embed=1" $TMP/durov345.html check 'Search clean-up + IP/phone disclosure (Durov, 2024-09-23)' $TMP/durov345.html 'IP addresses and phone numbers of those who violate our rules can be disclosed' check 'durov/345: problematic content no longer accessible in Search' $TMP/durov345.html 'All the problematic content we identified in Search is no longer accessible' check_raw 'durov/345 dated 2024-09-23' $TMP/durov345.html '2024-09-23' get https://telegram.org/tos/eu-dsa $TMP/tg_dsa.html check 'Telegram self-reports below the VLOP threshold' $TMP/tg_dsa.html 'significantly fewer than 45 million' get https://digital-strategy.ec.europa.eu/en/news/commission-designates-whatsapp-very-large-online-platform-under-digital-services-act $TMP/ec_wa.html check 'EC designates WhatsApp a VLOP via Channels' $TMP/ec_wa.html 'formally designated WhatsApp as a Very Large Online Platform' check 'EC: private messaging stays excluded' $TMP/ec_wa.html 'private messaging service' get https://digital-strategy.ec.europa.eu/en/policies/list-designated-vlops-and-vloses $TMP/ec_list.html if grep -qi 'telegram' $TMP/ec_list.html; then echo "FAILED Telegram now appears on the VLOP list"; FAILS=$((FAILS+1)); else echo "OK Telegram is not on the VLOP list"; fi if grep -qi 'discord' $TMP/ec_list.html; then echo "FAILED Discord now appears on the VLOP list"; FAILS=$((FAILS+1)); else echo "OK Discord is not on the VLOP list"; fi get https://about.fb.com/news/2023/09/whatsapp-channels-global-launch/ $TMP/wa_ch.html check 'WhatsApp Channels global launch 2023-09-13' $TMP/wa_ch.html '2023-09-13' echo; echo "== 5. Directories ==" get https://tgstat.com/ $TMP/tgstat.html check_raw 'TGStat self-reported catalogue size' $TMP/tgstat.html 'More than [0-9 ]+ channels and groups' check 'TGStat country tile: Russia ~1.69 million channels' $TMP/tgstat.html 'Russia Channels 1 69\d \d{3}' get https://web.archive.org/web/20260926223639/https://telemetr.io/en $TMP/telemetr_wb.html check_raw 'Telemetr.io (Web Archive 2026-09-26): 11M+ channels' $TMP/telemetr_wb.html 'Over 11M\+ channels' check 'Telemetr.io (Web Archive 2026-09-26): 7M+ channels' $TMP/telemetr_wb.html 'has 7M\+ channels' pw https://disboard.org/ $TMP/disboard.txt check 'Disboard is a self-listing directory' $TMP/disboard.txt 'list/find Discord servers' echo; echo "== 6. Datasets and guidance outside the seven venues (Crossref / arXiv) ==" for doi in 10.1609/icwsm.v12i1.14989 10.1609/icwsm.v14i1.7348 10.1145/3690624.3709397 10.16997/wpcc.313 10.1177/20501579251326809; do curl -sL -m 30 "https://api.crossref.org/works/$doi" -o $TMP/cr.json python3 - "$doi" $TMP/cr.json <<'PY' && echo "OK Crossref $doi" || echo "FAILED Crossref $doi" import json, sys try: m = json.load(open(sys.argv[2]))['message'] except Exception: sys.exit(1) a = '; '.join(f"{x.get('family','')}" for x in m.get('author', [])) print(' >', sys.argv[1], '|', m['title'][0], '|', (m.get('container-title') or [''])[0], '|', m['issued']['date-parts'][0], '|', a) PY [ $? -eq 0 ] || FAILS=$((FAILS+1)) done get https://www.westminsterpapers.org/article/id/274/ $TMP/wpcc.html check 'Barbosa and Milan 2019: avoid by all means covert bypasses' $TMP/wpcc.html 'avoid by all means covert bypasses' get https://arxiv.org/abs/2605.15956 $TMP/teragram.html check_raw 'TeraGram arXiv first author Golovin' $TMP/teragram.html 'citation_author" content="Golovin, Anastasia"' check 'TeraGram arXiv: 5.9 billion messages, 712 thousand channels and groups' $TMP/teragram.html 'over 5\.9 billion messages dating from 2015 to 2025, collected from 712 thousand channels and groups' curl -sL -m 30 https://api.crossref.org/works/10.1609/icwsm.v20i1.42783 -o $TMP/tg_cr.json check_raw 'TeraGram published at ICWSM 2026 (Crossref)' $TMP/tg_cr.json 'TeraGram: A Structured Longitudinal Dataset of the Telegram Messenger.*Web and Social Media' get https://aoir.org/reports/ethics3.pdf $TMP/aoir.pdf python3 - $TMP/aoir.pdf <<'PY' && echo "OK AoIR 3.0: approved 6 October 2019, no messaging-group passage" || { echo "FAILED AoIR 3.0"; exit 1; } import sys, re, pypdf, io, contextlib with contextlib.redirect_stderr(io.StringIO()): t = re.sub(r'\s+', ' ', ' '.join(p.extract_text() or '' for p in pypdf.PdfReader(sys.argv[1]).pages)) assert 'approved by the AoIR membership October 6, 2019' in t hits = {w: len(re.findall(w, t, re.I)) for w in ['Telegram', 'Discord', 'closed group', 'private group']} print(' > hits:', hits) assert all(v == 0 for v in hits.values()) PY [ $? -eq 0 ] || FAILS=$((FAILS+1)) rm -rf $TMP echo; echo "FAILED checks: $FAILS" [ $FAILS -gt 255 ] && FAILS=255 exit $FAILS
- external_checks_messaging_channels-output.txt
run: 2026-09-27T12:35Z == 1. Client tooling (GitHub / Codeberg / PyPI / npm / Go proxy) == > scussions": false, "forks_count": 1645, "mirror_url": null, "archived": true, "disabled": false, "open_issues_count": 1, "license": { "k OK Telethon GitHub repo is archived GET https://raw.githubusercontent.com/LonamiWebs/Telethon/v1/README.rst -> 2665 bytes > Moved to https://codeberg.org/Lonami/Telethon. The GitHub repository may be deleted in the future. ---- T OK Telethon README: moved to Codeberg > for Telegram, be careful not to break `Telegram's ToS`_ or `Telegram can ban the account`_. What is this? ------------- Telegram is a popular messag OK Telethon README: Telegram can ban the account > w to migrate. As with any third-party library for Telegram, be careful not to break Telegram's ToS or Telegram can ban the account. What is this? ------------- Telegram is a popular messagin OK Telethon README (reST links rendered): be careful not to break ToS GET https://pypi.org/pypi/Telethon/json -> 354130 bytes Telethon PyPI latest 1.45.0 2026-09-10T14:28:03 any 2.x on PyPI: [] > mmary":"Full-featured Telegram client library for Python 3","version":"1.45.0","yanked":false,"yanked_reason":null},"last_serial":40933366 OK Telethon PyPI latest is 1.45.0 GET https://codeberg.org/api/v1/repos/Lonami/Telethon -> 2545 bytes > pen_pr_counter":2,"release_counter":0,"default_branch":"v1","archived":false,"created_at":"2026-02-21T20:01:22+01:00","updated_at":"2026 OK Telethon Codeberg repo not archived > scussions": false, "forks_count": 1383, "mirror_url": null, "archived": true, "disabled": false, "open_issues_count": 279, "license": { OK Pyrogram GitHub repo is archived GET https://raw.githubusercontent.com/pyrogram/pyrogram/master/README.md -> 2324 bytes > on • Releases • News ## Pyrogram > [!NOTE] > The project is no longer maintained or supported. Thanks for appreciating it. > Elegant, modern OK Pyrogram README: no longer maintained GET https://pypi.org/pypi/kurigram/json -> 54193 bytes kurigram PyPI latest 2.2.26 2026-09-12T17:45:06 > mail":"Danipulok <danipulok@gmail.com>","name":"Kurigram","package_url":"https://pypi.org/project OK Kurigram (Pyrogram fork) is on PyPI > requires_python":">=3.8","size":5457052,"upload_time":"2026-01-28T12:35:37","upload_time_iso_8601":" OK Kurigram released in 2026 > iscussions": false, "forks_count": 220, "mirror_url": null, "archived": false, "disabled": false, "open_issues_count": 23, "license": { " OK Kurigram repo not archived > s\n\nKurigram is an actively maintained pyrogram fork for Python designed as a drop-in replac OK Kurigram describes itself as a Pyrogram fork GET https://registry.npmjs.org/telegram -> 1088391 bytes > e":"kixxauth","email":"kris@kixx.name"},"deprecated":"This package is archived and no longer maintained. Development continues in teleproto (https://npmjs.com/package/teleproto), a largely compatible, actively maintained fork. See the migration guide at https://docs.teleproto.dev/migrating-from-gramjs","_npmVersion OK npm telegram (GramJS) deprecated in favour of teleproto > discussions": true, "forks_count": 236, "mirror_url": null, "archived": true, "disabled": false, "open_issues_count": 323, "license": { OK GramJS GitHub repo is archived > scussions": false, "forks_count": 2181, "mirror_url": null, "archived": false, "disabled": false, "open_issues_count": 78, "license": { " OK TDLib repo not archived GET https://raw.githubusercontent.com/tdlib/td/master/CMakeLists.txt -> 55694 bytes > cmake_minimum_required(VERSION 3.10 FATAL_ERROR) project(TDLib VERSION 1.8.67 LANGUAGES CXX C) if (NOT DEFINED CMAKE_MODULE_PATH) set(CMA OK TDLib master declares a 1.8.x version > tags: 10 newest by version: v1.8.0 OK TDLib: newest tag by version is v1.8.0 > EwOlJlcG9zaXRvcnkxNzEwNzI5Njc=", "name": "whatsapp-web.js", "full_name": "wwebjs/whatsapp-web.js", "private": false, "owner": { "login": "wwebjs", "id": 87630360, "node_id": "MDEyOk9yZ2FuaXphdGlvbjg3NjMwMzYw", "avatar_url": "https://avatars.githubusercontent.com/u/87630360?v=4", "gravatar_id": "", "url": "https://api.github.com/users/wwebjs", "html_url": "https://github.com/wwebjs", "followers_url": "https://api.github.com/users/wwebjs/followers", "following_url": "https://api.github.com/users/wwebjs/following{/other_user}", "gists_url": "https://api.github.com/users/wwebjs/gists{/gist_id}", "starred_url": "https://api.github.com/users/wwebjs/starred{/owner}{/repo}", "subscriptions_url": "https://api.github.com/users/wwebjs/subscriptions", "organizations_url": "https://api.github.com/users/wwebjs/orgs", "repos_url": "https://api.github.com/users/wwebjs/repos", "events_url": "https://api.github.com/users/wwebjs/events{/privacy}", "received_events_url": "https://api.github.com/users/wwebjs/received_events", "type": "Organization", "user_view_type": "public", "site_admin": false }, "html_url": "https://github.com/wwebjs/whatsapp-web.js", "description": "A WhatsApp client library for NodeJS that connects through the WhatsApp Web browser app", "fork": false, "url": "https://api.github.com/repos/wwebjs/whatsapp-web.js", "forks_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/forks", "keys_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/keys{/key_id}", "collaborators_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/collaborators{/collaborator}", "teams_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/teams", "hooks_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/hooks", "issue_events_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/issues/events{/number}", "events_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/events", "assignees_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/assignees{/user}", "branches_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/branches{/branch}", "tags_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/tags", "blobs_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/git/blobs{/sha}", "git_tags_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/git/tags{/sha}", "git_refs_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/git/refs{/sha}", "trees_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/git/trees{/sha}", "statuses_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/statuses/{sha}", "languages_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/languages", "stargazers_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/stargazers", "contributors_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/contributors", "subscribers_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/subscribers", "subscription_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/subscription", "commits_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/commits{/sha}", "git_commits_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/git/commits{/sha}", "comments_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/comments{/number}", "issue_comment_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/issues/comments{/number}", "contents_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/contents/{+path}", "compare_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/compare/{base}...{head}", "merges_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/merges", "archive_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/{archive_format}{/ref}", "downloads_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/downloads", "issues_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/issues{/number}", "pulls_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/pulls{/number}", "milestones_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/milestones{/number}", "notifications_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/notifications{?since,all,participating}", "labels_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/labels{/name}", "releases_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/releases{/id}", "deployments_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/deployments", "created_at": "2019-02-17T02:16:02Z", "updated_at": "2026-09-27T11:54:55Z", "pushed_at": "2026-09-27T12:04:45Z", "git_url": "git://github.com/wwebjs/whatsapp-web.js.git", "ssh_url": "git@github.com:wwebjs/whatsapp-web.js.git", "clone_url": "https://github.com/wwebjs/whatsapp-web.js.git", "svn_url": "https://github.com/wwebjs/whatsapp-web.js", "homepage": "https://wwebjs.dev", "size": 6522, "stargazers_count": 22646, "watchers_count": 22646, "language": "JavaScript", "has_issues": true, "has_projects": false, "has_downloads": false, "has_wiki": false, "has_pages": true, "has_discussions": false, "forks_count": 5169, "mirror_url": null, "archived": false, "disabled": false, "open_issues_count": 113, "license": { OK whatsapp-web.js lives at wwebjs/whatsapp-web.js, not archived GET https://raw.githubusercontent.com/wwebjs/whatsapp-web.js/main/README.md -> 8714 bytes > ot guaranteed you will not be blocked by using this method. WhatsApp does not allow bots or unofficial clients on their platform, so this shouldn't be considered totally safe. ## License Copyright 2019 Pedro S Lopez Licensed under the OK whatsapp-web.js README: WhatsApp does not allow unofficial clients GET https://registry.npmjs.org/whatsapp-web.js/latest -> 2842 bytes > ry-packages-npm-production"},"_id":"whatsapp-web.js@1.34.7","version":"1.34.7"} OK whatsapp-web.js npm latest 1.x > iscussions": true, "forks_count": 3394, "mirror_url": null, "archived": false, "disabled": false, "open_issues_count": 330, "license": { OK Baileys not archived GET https://proxy.golang.org/go.mau.fi/whatsmeow/@latest -> 198 bytes > {"Version":"v0.0.0-20260925162019-b3832c2bd1d1","Time":"2026-09-25T16:20:19Z","Origin":{"VCS":"git","URL":" OK whatsmeow resolves only to a pseudo-version GET https://api.github.com/repos/tgalal/yowsup/commits?per_page=1 -> 4101 bytes > { "name": "Tarek Galal", "email": "tare2.galal@gmail.com", "date": "2021-12-14T08:20:38Z" }, "committer": { "name": "Tarek Galal", "e OK yowsup: last default-branch commit is from 2021 GET https://raw.githubusercontent.com/tgalal/yowsup/master/README.md -> 3099 bytes > 0 yowsup-cli version: 3.2.1 requires: - python>=2.7,<=3.7 - consonance==0.1.5 - python-axolotl==0 OK yowsup README: requires python <= 3.7 GET https://pypi.org/pypi/discord.py/json -> 102972 bytes > n":">=3.8","summary":"A Python wrapper for the Discord API","version":"2.7.1","yanked":false,"yanked_reason":null},"last_serial":34790286 OK discord.py PyPI latest 2.x discord.py 2.7.1 2026-03-03T18:40:44 > ","yanked":false,"yanked_reason":null}],"2.7.1":[{"comment_text":"","core-metadata":{"sha256":"adfb03dd4d98ff00b6b365a615def0d5eb7d00a9397229e1e44828cce4f867e1"},"digests":{"blake2b_256":"f7a717208c3b3f92319e7fad259f1c6d5a5baf8fd0654c54846ced329f83c3eb","md5":"bb7256992891cde615e21efaf87f58e3","sha256":"849dca2c63b171146f3a7f3f8acc04248098e9e6203412ce3cf2745f284f7439"},"downloads":-1,"filename":"discord_py-2.7.1-py3-none-any.whl","has_sig":false,"md5_digest":"bb7256992891cde615e21efaf87f58e3","packagetype":"bdist_wheel","python_version":"py3","requires_python":">=3.8","size":1227550,"upload_time":"2026-03-03T18:40:44","upload_time_iso_8601":"202 OK discord.py 2.7.1 uploaded March 2026 GET https://raw.githubusercontent.com/Tyrrrz/DiscordChatExporter/master/Readme.md -> 6002 bytes > es. > [!WARNING] > While **DiscordChatExporter** allows it, automating user accounts is against Discord TOS and may result in you getting banned. > If possible, use a bot to export chat logs from accessib OK DiscordChatExporter README: automating user accounts is against Discord TOS > ", "site_admin": false }, "node_id": "RE_kwDOBcifOs4Wh8sD", "tag_name": "2.48", "target_commitish": "prime", "name": "2.48", "draft": false, "immutable": false, "prerelease": false, "created_at": "2026-08-27T17:06:34Z", "updated_at": "2026-08-27T17:18:07Z", "published_at": "2026-08-27T17:17:35Z", "assets": [ { "url": "https://api.github.com/ OK DiscordChatExporter latest release 2.48 (August 2026) == 2. Terms == GET https://telegram.org/tos/content-licensing -> 8332 bytes > upload, share and view content in public and private chats. Access to user-generated content for any purpose other than ordinary, legitimate, and intended use of the Telegram platform as its user is prohibited. As a limited exception, Telegram permits access to data re OK Telegram Content Licensing: access beyond ordinary use prohibited > o do so. Large Language Models and AI For clarity, Telegram firmly prohibits the scraping, indexing, harvesting, aggregation or use of data obtained from its platform to train, fine-tune, validate or otherwise engage in the development OK Telegram Content Licensing: no scraping for AI/ML training GET https://telegram.org/tos -> 25192 bytes > ateway API , Telegram API Developers , and Bot Developers . Telegram additionally prohibits data scraping as part of its Content Licensing and AI Scraping Terms , which apply to all users, businesses, and third-party services accessing the platform. Terms of Service for Telegram Stars Telegram users can acq OK Telegram ToS: scraping prohibited for all users and third parties GET https://core.telegram.org/api/terms -> 10294 bytes > ce for Content Licensing and AI Scraping . As such, you are prohibited from using, accessing or aggregating data obtained from the Telegram platform to train, fine-tune or otherwise engage in the development, enhancement or deployment of artificial intelligence, mach OK Telegram API ToS 1.5: no ML training on API data GET https://core.telegram.org/api/obtaining_api_id -> 8330 bytes > parameters required for user authorization. For the moment each number can only have one api_id connected to it. We will be sending important developer notifications to th OK Telegram: one api_id per phone number > ounts that log in using unofficial Telegram API clients are automatically put under observation to avoid violations of the Terms of Service . If you didn't OK Telegram: unofficial-client logins put under observation GET https://telegram.org/tour/channels -> 29886 bytes > vate – but you can edit their profile to make them public . The contents of public channels can be seen on the Web without a Telegram account and are indexed by search engines. For example, try t.me/s/ OK Telegram documents the public-channel web preview GET https://t.me/s/telegram -> 127690 bytes > data-before="441"></a></div><div class="tgme_widget_message_wrap js-widget_message_wrap"><div class="tgm OK t.me/s/<channel> serves channel posts > Very Telegram. Wow. "> <link rel="prev" href="/s/telegram?before=441"> <link rel="canonical" href="/s/telegra OK t.me/s/<channel> paginates with ?before= PW https://www.whatsapp.com/legal/terms-of-service -> 262683 bytes > . You must not (or assist others to) directly, indirectly, through automated or other means, access, use, copy, adapt, modify, prepare derivative works based upon, distri OK WhatsApp ToS: no access or collection through automated means > r our Services through unauthorized or automated means; (f) collect information of or about our users in any impermissible or unauthorized manner; (g) sell, resell, rent, or charge for our Services or data OK WhatsApp ToS: no collecting information about users in impermissible ways GET https://discord.com/terms -> 134546 bytes > to, intentionally overburdening, or attacking our systems; scraping our services without our written consent, including by using any robot, spider, crawler, scraper, or OK Discord ToS: no scraping without written consent GET https://discord.com/guidelines -> 87918 bytes > e our Platform Manipulation Policy Explainer for more.) 14. Do not use self-bots or user-bots. Each account must be associated with a human, not a bot. ( OK Discord Community Guidelines: no self-bots PW https://support-dev.discord.com/hc/en-us/articles/8563934450327-Discord-Developer-Policy -> 157304 bytes > engineer API Data from the form in which you obtain it. 20. Do not mine or scrape any data, content, or information available on or through Discord se OK Discord Developer Policy: do not mine or scrape > Discord services (as defined in our Terms of Service). 21. Do not use message content obtained through the APIs to train machine learning or AI models (including large language models) unless expre OK Discord Developer Policy: no ML training on message content PW https://support.discord.com/hc/en-us/articles/208866998-Invites-101 -> 142204 bytes > sages or copy the invite link. The invite panel will appear displaying a 7 days access link by default if you haven't already adjusted the invite settings for tha OK Discord invites default to 7 days PW https://support.discord.com/hc/en-us/articles/360030843331-Enabling-Server-Discovery -> 146310 bytes > afety requirements for a safe environment. Server must have at least 1,000 members to qualify. Servers need to be at least 8 weeks old to be i OK Discord Server Discovery needs 1,000 members PW https://transparency.meta.com/researchtools/meta-content-library/ -> 685212 bytes > Content Library and Content Library API provide comprehensive access to the public content archive from Facebook, Instagram and WhatsApp Channels. Individuals can apply for access using OK Meta Content Library covers WhatsApp Channels GET https://core.telegram.org/bots/features -> 116993 bytes > chats. All messages from channels where they are a member. Privacy mode is enabled by default for all bots, except bots that were added to a group as admins (bot admins always receive all messages ). It can be disabl OK Telegram bots: privacy mode on by default unless added as admin == 3. Telegram API limits == GET https://core.telegram.org/api/config -> 181819 bytes > drawal_enabled ": true, " channel_wallpaper_level_min ": 9, " channels_limit_default ": 500, " channels_limit_premium ": 1000, " channels_public_limit_ OK Telegram default join limit 500 channels/supergroups > _wallpaper_level_min ": 9, " channels_limit_default ": 500, " channels_limit_premium ": 1000, " channels_public_limit_default ": 10, " channels_public_l OK Telegram Premium join limit 1000 > s_user_max_default ": 1, " reactions_user_max_premium ": 3, " recommended_channels_limit_default ": 10, " recommended_channels_limit_premium ": 100, " restriction OK Telegram channel recommendations capped at 10 (non-Premium) GET https://core.telegram.org/method/channels.joinChannel -> 20799 bytes > n use this method Possible errors Code Type Description 400 CHANNELS_TOO_MUCH You have joined too many channels/supergroups. 400 CHANNEL_ OK channels.joinChannel lists CHANNELS_TOO_MUCH > action. 400 INVITE_HASH_EMPTY The invite hash is empty. 406 INVITE_HASH_EXPIRED The invite link has expired. 400 INVITE_HASH_INVALID The in OK channels.joinChannel lists INVITE_HASH_EXPIRED GET https://core.telegram.org/api/errors -> 20197 bytes > xt messages (SMS) for the same phone number. Error Example: FLOOD_WAIT_X: A wait of X seconds is required (where X is a number) FLOO OK FLOOD_WAIT_X documented > MS) for the same phone number. Error Example: FLOOD_WAIT_X: A wait of X seconds is required (where X is a number) FLOOD_PREMIUM_WAIT_X: A wait of X sec OK FLOOD_WAIT_X: a wait of X seconds is required == 4. Platform changes == GET https://t.me/durov/343?embed=1 -> 18032 bytes > ew features while phasing out a few outdated ones. ⛔️ We've removed the People Nearby feature, which was used by less than 0.1% of Telegram users, but ha OK People Nearby removed (Durov, 2024-09-06) > https://t.me/durov/343"><time datetime="2024-09-06T13:52:50+00:00" class="datetime">Sep 6, OK durov/343 dated 2024-09-06 GET https://t.me/durov/345?embed=1 -> 17335 bytes > e consistent across the world. We've made it clear that the IP addresses and phone numbers of those who violate our rules can be disclosed to relevant authorities in response to valid legal requests OK Search clean-up + IP/phone disclosure (Durov, 2024-09-23) > rators, leveraging AI, has made Telegram Search much safer. All the problematic content we identified in Search is no longer accessible. If you still manage to find something unsafe or illegal in OK durov/345: problematic content no longer accessible in Search > https://t.me/durov/345"><time datetime="2024-09-23T13:09:39+00:00" class="datetime">Sep 23 OK durov/345 dated 2024-09-23 GET https://telegram.org/tos/eu-dsa -> 14425 bytes > forms" under the DSA. As of August 2026, these services had significantly fewer than 45 million average monthly active recipients in the EU over the preced OK Telegram self-reports below the VLOP threshold GET https://digital-strategy.ec.europa.eu/en/news/commission-designates-whatsapp-very-large-online-platform-under-digital-services-act -> 50403 bytes > YTE Publication 26 January 2026 The European Commission has formally designated WhatsApp as a Very Large Online Platform (VLOP) under the Digital Services Act (DSA), as its 'Channe OK EC designates WhatsApp a VLOP via Channels > ns that online platforms in the EU must respect. WhatsApp's private messaging service enabling users to send text messages, voice notes, photos, OK EC: private messaging stays excluded GET https://digital-strategy.ec.europa.eu/en/policies/list-designated-vlops-and-vloses -> 144501 bytes OK Telegram is not on the VLOP list OK Discord is not on the VLOP list GET https://about.fb.com/news/2023/09/whatsapp-channels-global-launch/ -> 355738 bytes > line":"WhatsApp Channels Are Going Global","datePublished":"2023-09-13T20:00:09+00:00","mainEntityOfPage":{"@id":"https://about.fb OK WhatsApp Channels global launch 2023-09-13 == 5. Directories == GET https://tgstat.com/ -> 45922 bytes > ram channels and groups catalog TGStat. More than 2 864 920 channels and groups, classified by countries, languages and OK TGStat self-reported catalogue size > ts Telegram monitoring Telegram channels and groups catalog Russia Channels 1 692 000 Groups 134 600 Total audience 5 673 740 000 Open catalog Uk OK TGStat country tile: Russia ~1.69 million channels GET https://web.archive.org/web/20260926223639/https://telemetr.io/en -> 400241 bytes > and check channels for fake followers. Over 11M+ channels in our database — start for free!"/><me OK Telemetr.io (Web Archive 2026-09-26): 11M+ channels > Find Telegram channels for advertising placement Telemetrio has 7M+ channels categorized into 41 categories Search channels by keywords OK Telemetr.io (Web Archive 2026-09-26): 7M+ channels PW https://disboard.org/ -> 408712 bytes > admin? Add Your Server! DISBOARD is the place where you can list/find Discord servers . Find and join some awesome servers listed here. Or login OK Disboard is a self-listing directory == 6. Datasets and guidance outside the seven venues (Crossref / arXiv) == > 10.1609/icwsm.v12i1.14989 | WhatApp Doc? A First Look at WhatsApp Public Group Data | Proceedings of the International AAAI Conference on Web and Social Media | [2018, 6, 15] | Garimella; Tyson OK Crossref 10.1609/icwsm.v12i1.14989 > 10.1609/icwsm.v14i1.7348 | The Pushshift Telegram Dataset | Proceedings of the International AAAI Conference on Web and Social Media | [2020, 5, 26] | Baumgartner; Zannettou; Squire; Blackburn OK Crossref 10.1609/icwsm.v14i1.7348 > 10.1145/3690624.3709397 | TGDataset: Collecting and Exploring the Largest Telegram Channels Dataset | Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1 | [2025, 7, 20] | La Morgia; Mei; Mongardini OK Crossref 10.1145/3690624.3709397 > 10.16997/wpcc.313 | Do Not Harm in Private Chat Apps: Ethical Issues for Research on and with WhatsApp | Westminster Papers in Communication and Culture | [2019, 8, 14] | Barbosa; Milan OK Crossref 10.16997/wpcc.313 > 10.1177/20501579251326809 | <i>Whatsapp Explorer</i> : A data donation tool to facilitate research on WhatsApp | Mobile Media & Communication | [2025, 3, 28] | Garimella; Chauchard OK Crossref 10.1177/20501579251326809 GET https://www.westminsterpapers.org/article/id/274/ -> 144815 bytes > search and the research ecosystem; embrace transparency and avoid by all means covert bypasses; and guarantee full anonymisation to our research subjects. OK Barbosa and Milan 2019: avoid by all means covert bypasses GET https://arxiv.org/abs/2605.15956 -> 41622 bytes > f the Telegram Messenger" /><meta name="citation_author" content="Golovin, Anastasia" /><meta name="citation_author" content= OK TeraGram arXiv first author Golovin > longitudinal dataset of public Telegram content, comprising over 5.9 billion messages dating from 2015 to 2025, collected from 712 thousand channels and groups, enriched with metadata on forwards, reactions, and polls. OK TeraGram arXiv: 5.9 billion messages, 712 thousand channels and groups > f","is-referenced-by-count":1,"title":["TeraGram: A Structured Longitudinal Dataset of the Telegram Messenger"],"prefix":"10.1609","volume":"20","author":[{"given":"Anastasia","family":"Golovin","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sebastian B.","family":"Mohr","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Arne I.","family":"Gottwald","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ulrik","family":"Hvid","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Srushhti","family":"Trivedi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Joao","family":"Pinheiro Neto","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andreas C.","family":"Schneider","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Viola","family":"Priesemann","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"9382","published-online":{"date-parts":[[2026,5,25]]},"container-title":["Proceedings of the International AAAI Conference on Web and Social Media"],"original-title":[],"link":[{"URL":"h OK TeraGram published at ICWSM 2026 (Crossref) GET https://aoir.org/reports/ethics3.pdf -> 836041 bytes > hits: {'Telegram': 0, 'Discord': 0, 'closed group': 0, 'private group': 0} OK AoIR 3.0: approved 6 October 2019, no messaging-group passage FAILED checks: 0
The reader brief, and the machine check of the readers' notes
- msgch_reader_brief.md
# Reader brief — design:platforms:messaging_channels (hand codes) You are reading research papers for a wiki page about ONE measurement route: collecting data from inside messaging-platform groups, channels or servers (Telegram, WhatsApp, Discord, or another messenger's group feature — WeChat, QQ, Signal...). The page is for a PhD student about to run such a study: how groups are found, what joining gives you, the denominator, client tooling, bans, ethics. ## Where the text is Full text: `/workspace/publications_dataset/data/fulltext/<year>/<venue>/<slug>/paper.cols.txt` (key format below is `<venue>/<year>/<slug>`; note the directory order is year/venue/slug). Grep it whitespace-collapsed: `tr -s '[:space:]' ' ' < paper.cols.txt | grep -oiE ".{0,300}telegram channel.{0,300}"`. Two-column PDFs are sometimes spliced in `.cols`; if a sentence looks broken, also try `python3 /workspace/artifacts/wiki/scripts/pdftext.py <venue>/<year>/<slug> | grep -oiE ".{0,300}PATTERN.{0,300}"`. Read the methodology / data collection / ethics sections fully, not just grep hits. The data is read-only. ## What to produce Write ONE markdown file (path given in your task) with one section per paper, in this exact shape. Every field that makes a factual claim carries a VERBATIM quote from the paper (copy-paste exact characters from the text, no paraphrase inside quotation marks, no ellipsis-joining of two separate sentences). If the paper does not say, write `not stated` — do not infer. Write the file as you go (append after each paper), so partial work survives. ``` ### <key> - verdict: IN-CORE | IN-SECTION | CONTEXT-<code> | OUT-<code> (see rule below) - platforms: Telegram / WhatsApp / Discord / other (name) - object: one line — what community/phenomenon (e.g. stolen-data markets, political misinformation) - discovery: how the groups/channels were found — one or more of: links-on-other-platform(<which>), directory-site(<which, e.g. TGStat, Telemetr, Disboard, top.gg>), in-app-search, snowball(forwards/mentions/links inside channels), prior-dataset-or-list, news/reports/forums, web-search, personal-network, not stated > "verbatim quote" - join: joined-as-member | subscribed/public-channel-read | public-web-preview | third-party-scraper-service | bot-added | operator-supplied | not stated > "verbatim quote" - scale: groups/channels found (candidate) vs joined/collected, messages, users, time window — each number with quote > "verbatim quote" (one per number) - history: did they back-fill message history from before joining? yes / no / not stated > "quote" - members: member lists collected? media downloaded? yes/no/not stated > "quote" - tooling: Telethon / TDLib / Pyrogram / Telegram API (unspecified) / WhatsApp Web automation / phones / Discord API / bot / commercial scraper (name) / manual / not stated > "quote" - bans-limits: bans, rate limits, flood-waits, removal from groups, channels deleted mid-study, invite links expiring — anything reported > "quote" (or not stated) - ethics: IRB/ethics board outcome; passive-only (no posting/interaction) stated?; consent from admins or members?; anonymisation/pseudonymisation; data minimisation (e.g. no media, no phone numbers); disclosure to platform; anything about members' expectation of privacy > "quote" per claim (or not stated) - denominator: does the paper acknowledge that groups found ≠ groups that exist, or discuss the bias of its seed source? > "quote" or not stated - headline-figures: up to 3 measured results with the paper's own denominator > "quote" each - notes: anything a methods-page author should know (surprises, limitations the paper states) ``` ## Verdict rule (fixed; apply it, do not reinterpret) IN-CORE — data collected from inside groups/channels/servers is the paper's main dataset. IN-SECTION — such collection is one component among several (e.g. one data source in a multi-source study). CONTEXT-enumeration — accounts discovered via contact discovery / phone-number lookups, not groups. CONTEXT-links-only — invite links or group URLs harvested from elsewhere, groups never entered. CONTEXT-reuse — uses a dataset someone else collected from groups (say which dataset). CONTEXT-bots — chatbot ecosystem in messengers (bot listings, permissions), not group content. CONTEXT-dm — one-to-one conversations (e.g. with scammers), not groups. CONTEXT-user-study — interviews/surveys about people's group use; or the paper explicitly declined to collect. CONTEXT-operator — group data supplied by the platform operator. OUT-<reason> — protocol attack, traffic analysis, app analysis, messenger used as C2 or as a recruitment channel for a user study, passing mention. For CONTEXT/OUT papers you may shorten: verdict, one deciding quote, and any field that is useful to the page (e.g. an enumeration paper's scale, tooling and ethics are useful — fill those). Your context may not be exhaustive; if a paper seems to be missing text, say so rather than guessing. When done, reply with a one-line summary per paper (key: verdict).
- msgch_notes_quotecheck-output.txt
msgch_papers_A.md: 153/197 quotes located msgch_papers_B.md: 174/198 quotes located msgch_papers_C.md: 83/103 quotes located msgch_papers_D.md: 91/108 quotes located TOTAL: 501/606 located (82.7%); by first rendering: {'paper.cols.txt': 388, 'pypdf': 111, 'paper.norm.txt': 2} MISS: 105 msgch_papers_A.md USENIX/2026/stayin-alive-how-global-stolen-data-markets-thrive-on-telegram "We find candidate channels from cybercriminal forums or marketplaces that specialize in the trade of stolen data. There, the seller often directs potential buyers to their Telegram channel for more information, and the f" msgch_papers_A.md USENIX/2026/stayin-alive-how-global-stolen-data-markets-thrive-on-telegram "We perform recurrent snowballing to expand the set of channels. In each snowballing round, we collect candidate channels appearing in previously collected messages, channel descriptions, and the data sources described ab" msgch_papers_A.md USENIX/2026/stayin-alive-how-global-stolen-data-markets-thrive-on-telegram "We find, based on channel IDs, that the overlap is only six stolen data channels, i.e., we can only identify 0.8% of the total stolen data channels through the (extended) publicly available dataset. Thus, stolen data cha" msgch_papers_A.md WWW/2026/doxing-as-a-service-demystifying-the-chinese-online-doxing-ecosystem "we retrieved a raw dataset of 312 candidate channels ... we selected the top 100 channels by subscriber count as our representative corpus" msgch_papers_A.md WWW/2026/doxing-as-a-service-demystifying-the-chinese-online-doxing-ecosystem "The final dataset comprises 25,972 messages with associated metadata, including 325M cumulative views and 2.27M forwards." msgch_papers_A.md WWW/2026/doxing-as-a-service-demystifying-the-chinese-online-doxing-ecosystem "we analyzed 411,707 messages collected from five query groups from May 30, 2025 to August 25, 2025... personal identity information of over 300,000 unique individuals had been directly exposed within these publicly acces" msgch_papers_A.md WWW/2026/doxing-as-a-service-demystifying-the-chinese-online-doxing-ecosystem "Analyzing 25,972 messages and 13.22 million subscriber links from 100 major channels on Telegram, we demystify its organization, operations, and user engagement." msgch_papers_A.md WWW/2026/doxing-as-a-service-demystifying-the-chinese-online-doxing-ecosystem "personal identity information of over 300,000 unique individuals had been directly exposed within these publicly accessible groups during 3 months...with an average of about 5,100 new successful queries per day from thes" msgch_papers_A.md WWW/2026/doxing-as-a-service-demystifying-the-chinese-online-doxing-ecosystem "A keyword search for \"pian\" (scam) across the group corpus yielded approximately 22,000 mentions (about 10% of total messages), including 3,280 self-reported victimization cases and 17,000 accusations directed at servi" msgch_papers_A.md USENIX/2019/the-anatomy-of-a-cryptocurrency-pump-and-dump-scheme "PumpOlymp discovers those channels by searching pump-related keywords - e.g. \"pump\", \"whales\", \"vip\" and \"coin\" - on Telegram aggregators such as https://tgstat.com/ and https://telegramcryptogroups.com/. Another" msgch_papers_A.md USENIX/2019/the-anatomy-of-a-cryptocurrency-pump-and-dump-scheme "To validate the incoming data from PumpOlymp, we conduct an independent manual search for pump-and-dump channels. We are not able to add new channels to the existing channel list from PumpOlymp, and we are not aware of a" msgch_papers_A.md USENIX/2019/the-anatomy-of-a-cryptocurrency-pump-and-dump-scheme "The largest one being Official McAfee Pump Signals, with a startling 12,333 members." msgch_papers_A.md USENIX/2019/the-anatomy-of-a-cryptocurrency-pump-and-dump-scheme "Official McAfee Pump Signals, with a startling 12,333 members." msgch_papers_A.md USENIX/2019/the-anatomy-of-a-cryptocurrency-pump-and-dump-scheme "43 have been deleted from the Telegram sever, possibly due to inactivity for an extended period of time... This might also imply that the Telegram channels have a \"hit-and-run\" characteristic... channel admins might de" msgch_papers_A.md USENIX/2019/the-anatomy-of-a-cryptocurrency-pump-and-dump-scheme "Pump-and-dump admins, aiming to profit from price manipulation, are certainly unethical. Nevertheless, other pump-and-dump participants are also culpable since their behaviour enables and reinforces the existence of such" msgch_papers_A.md USENIX/2019/the-anatomy-of-a-cryptocurrency-pump-and-dump-scheme "We are not able to add new channels to the existing channel list from PumpOlymp, and we are not aware of any other, more comprehensive pump-and-dump channel list. Therefore, we believe the channel list from PumpOlymp is " msgch_papers_A.md USENIX/2019/the-anatomy-of-a-cryptocurrency-pump-and-dump-scheme "around 100 organized Telegram pump-and-dump channels coordinate on average 2 pumps a day, which generates an aggregate artificial trading volume of 6 million USD a month." msgch_papers_A.md USENIX/2019/the-anatomy-of-a-cryptocurrency-pump-and-dump-scheme "the artificial trading volume generated by those pump-and-dump activities is astonishing: 8,793 BTC (93% from Binance), roughly equivalent to 50 million USD... of trading volume during the pump hours, 9 times as much as " msgch_papers_A.md USENIX/2019/the-anatomy-of-a-cryptocurrency-pump-and-dump-scheme "we estimate that admins made a net profit of 199.52 BTC, equivalent to 1.1 million USD, through 348 pump and dump events during our sample period. The estimated return of insiders averages 18%" msgch_papers_A.md USENIX/2021/catching-phishers-by-their-bait-investigating-the-dutch-phishing-landscape-throu "we manually inspected fraud-related Telegram channels, searched for shared phishing kits and discovered related channels by following shared links in the chat. This snowball approach is a common sampling technique, that " msgch_papers_A.md USENIX/2021/catching-phishers-by-their-bait-investigating-the-dutch-phishing-landscape-throu "First, we collect phishing kits on public Telegram channels employing a so-called 'snowball sampling' approach." msgch_papers_A.md USENIX/2021/catching-phishers-by-their-bait-investigating-the-dutch-phishing-landscape-throu "To gather phishing kits from Telegram, we manually inspected fraud-related Telegram channels, searched for shared phishing kits and discovered related channels by following shared links in the chat." msgch_papers_A.md USENIX/2021/having-your-cake-and-eating-it-an-analysis-of-concession-abuse-as-a-service "We noticed that service providers prefer that scam initiators contact them through external messaging platforms (for privacy reasons), and 17.6% of providers manage groups on external platforms (such as Telegram) in whic" msgch_papers_A.md USENIX/2021/having-your-cake-and-eating-it-an-analysis-of-concession-abuse-as-a-service "Due to the popularity of CAaaS, the provider converted the group to a supergroup (maximum 100,000 members, from the default of 200) on November 16, 2019. We thus collected data from November 16, 2019 to February 28, 2020" msgch_papers_A.md USENIX/2021/having-your-cake-and-eating-it-an-analysis-of-concession-abuse-as-a-service "The provider helped refund the equivalent of $81,159.27 ($41,076.71, €17,130.4, £7,393.29) over three months through scamming merchants in North America and Europe." msgch_papers_A.md IMC/2025/unmasking-the-shadow-economy-a-deep-dive-into-drainer-as-a-service-phishing-on-e "These groups can be discovered by searching for keywords like \"wallet drainer\" on platforms such as Twitter, Github, and Telegram." msgch_papers_A.md IMC/2025/unmasking-the-shadow-economy-a-deep-dive-into-drainer-as-a-service-phishing-on-e "the operator creates Telegram groups where affiliates receive real-time information on phishing websites and updates regarding the wallet drainer." msgch_papers_A.md IMC/2025/unmasking-the-shadow-economy-a-deep-dive-into-drainer-as-a-service-phishing-on-e "All collected materials were securely stored" and "drainer toolkits and other materials collected in this study were analyzed in isolated virtual environments to minimize the risk of malware infection or personal data le" msgch_papers_A.md IMC/2025/unmasking-the-shadow-economy-a-deep-dive-into-drainer-as-a-service-phishing-on-e "drainer toolkits and other materials collected in this study were analyzed in isolated virtual environments to minimize the risk of malware infection or personal data leakage, with all procedures conducted under controll" msgch_papers_A.md IMC/2025/unmasking-the-shadow-economy-a-deep-dive-into-drainer-as-a-service-phishing-on-e "we identified 1,910 profit-sharing contracts, 56 operator accounts, 6,087 affiliate accounts, and 87,077 profit-sharing transactions. In total, the operator accounts earned $23.1 million, while the affiliate accounts ear" msgch_papers_A.md IMC/2025/unmasking-the-shadow-economy-a-deep-dive-into-drainer-as-a-service-phishing-on-e "we collect 867 drainer toolkits from Telegram groups and reported phishing websites." msgch_papers_A.md USENIX/2026/from-mirai-to-gorilla-deep-dive-into-a-long-lasting-ddos-for-hire-botnet "Telegram groups | Chat monitoring | 2025-04-19 ... 2025-07-14 | Monitoring of the Gorilla Telegram group." msgch_papers_A.md USENIX/2025/assessing-the-aftermath-the-effects-of-a-global-takedown-against-ddos-for-hire-s "Many booters operate chat channels to advertise successful attacks, deliver updates, and assist users... We monitor Telegram channels of working booters to track resurrections (if any) by capturing new domains being anno" msgch_papers_A.md USENIX/2025/assessing-the-aftermath-the-effects-of-a-global-takedown-against-ddos-for-hire-s "Our data collection and analysis were approved by our department's research ethics committee. We only scraped public forums and channels, which is lawful... We did not seek the consent of individuals on these forums and " msgch_papers_A.md USENIX/2025/assessing-the-aftermath-the-effects-of-a-global-takedown-against-ddos-for-hire-s "Our analyses were conducted collectively to avoid individuals being identified, which accords with the British Society of Criminology's Statement on Ethics. All quotes were paraphrased to prevent attribution." msgch_papers_A.md USENIX/2025/assessing-the-aftermath-the-effects-of-a-global-takedown-against-ddos-for-hire-s "As academic and industry measurements rely on different approaches and hence generate distinctive views, we use four separate datasets for a more complete and reliable analysis." msgch_papers_A.md USENIX/2025/assessing-the-aftermath-the-effects-of-a-global-takedown-against-ddos-for-hire-s "We found over half of the seized sites in the first wave returned within a median of one day, while all booters seized in the second wave returned within a median of two days." msgch_papers_A.md WWW/2025/pirates-of-charity-exploring-donation-based-abuses-in-social-media-platforms "[47] Daniel Milevski. 2024. Apify Telegram Scraper API... [48] Daniel Milevski. 2024. Telemetrio Telegram Scraper API." (referenced among the "API services[11-13, 47, 48, 74, 75]" msgch_papers_A.md WWW/2025/pirates-of-charity-exploring-donation-based-abuses-in-social-media-platforms "We acknowledge that our conservative filtering approach may have excluded some donation scam accounts. However, as pioneers in the large-scale study of fraudulent donation scams, our goal was to build a solid foundation " msgch_papers_A.md WWW/2025/pirates-of-charity-exploring-donation-based-abuses-in-social-media-platforms "we identified 832 scammers using various techniques to deceive users into making fraudulent donations" msgch_papers_A.md USENIX/2024/the-imitation-game-exploring-brand-impersonation-attacks-on-social-media-platfor "We then combined the brand domain's second-level domain (2LD) name with eight popular keywords, namely rewards, recover, hack, support, help, assist, contact, and team to create a search query for account collection." msgch_papers_A.md USENIX/2024/the-imitation-game-exploring-brand-impersonation-attacks-on-social-media-platfor "We acknowledge that our data filtration is conservative and we might have overlooked accounts that could be involved in brand impersonation or other types of scams." msgch_papers_A.md USENIX/2024/the-imitation-game-exploring-brand-impersonation-attacks-on-social-media-platfor "it is also worth mentioning that prominent brands do not use Telegram as their preferred communication channel. However, Telegram is a popular communication medium among fraudsters, and we expected to find scam accounts " msgch_papers_A.md USENIX/2024/the-imitation-game-exploring-brand-impersonation-attacks-on-social-media-platfor "we found 349,411 squatted accounts targeting 2,625 of 2,847 major international brands" msgch_papers_B.md IMC/2020/demystifying-the-messaging-platforms-ecosystem-through-the-lens-of-twitter "We search for the occurrences of the above URL patterns between April 8 and May 15, 2020 on Twitter, using two dierent approaches: (a) using Twitter's Search API [67] every hour, and (b) using Twitter's Streaming API [6" msgch_papers_B.md IMC/2020/demystifying-the-messaging-platforms-ecosystem-through-the-lens-of-twitter "We empirically find that the limit for WhatsApp is between 250 and 300 groups per user, while on Discord it is up to 100 servers." msgch_papers_B.md IMC/2020/demystifying-the-messaging-platforms-ecosystem-through-the-lens-of-twitter "for Telegram we find the phone numbers of a substantially fewer number of users—509 phone numbers corresponding to 0.68% of the discovered Telegram users." msgch_papers_B.md WWW/2019/mis-information-dissemination-in-whatsapp-gathering-analyzing-and-countermeasure "a national truck drivers' strike (May 21st to June 2nd, 2018); and (ii) the first round of the 2018 Brazilian general elections campaign (August 16th to October 7th, 2018)" msgch_papers_B.md WWW/2021/short-is-the-road-that-leads-from-fear-to-hate-fear-speech-in-indian-whatsapp-gr "Ethics note: We established strict ethics guidelines throughout the project. The Committee on the Use of Humans as Experimental Subjects at MIT approved the data collection as exempt." msgch_papers_B.md WWW/2021/short-is-the-road-that-leads-from-fear-to-hate-fear-speech-in-indian-whatsapp-gr "We explicitly trained our annotators to be aware of the disturbing nature of social media messages and to take regular breaks from the annotation." msgch_papers_B.md USENIX/2025/characterizing-and-detecting-propaganda-spreading-accounts-on-telegram "This dataset comprises 17.3M labeled messages from 13 political and news-oriented channels." msgch_papers_B.md USENIX/2025/characterizing-and-detecting-propaganda-spreading-accounts-on-telegram "channel-level moderation can be performed by human moderators who detect and clean propaganda activities by banning propaganda accounts and deleting propaganda messages." msgch_papers_B.md WWW/2025/exposing-cross-platform-coordinated-inauthentic-activity-in-the-run-up-to-the-20 "we identified 33 highly coordinated channels cosharing URLs to web domains" / "we identified 57 coordinated Telegram channels" msgch_papers_B.md WWW/2024/getting-bored-of-cyberwar-exploring-the-role-of-low-level-cybercrime-actors-in-t "We collect 441 announcements with 57 757 replies and 900k emoji reactions posted in the channel from its inception until 30 June 2022 using Telethon, which interacts with official Telegram APIs to fully capture messages " msgch_papers_B.md WWW/2024/getting-bored-of-cyberwar-exploring-the-role-of-low-level-cybercrime-actors-in-t "using Telethon, which interacts with official Telegram APIs to fully capture messages and metadata" msgch_papers_B.md IEEE-SP/2025/learning-from-censored-experiences-social-media-discussions-around-censorship-ci "we focused on evaluating the content shared on VPN-focused Telegram channels ... We utilized the Telegram API [20] and Telethon client [21] to collect posts from 34 unique Telegram channels mentioned in tweets from our T" msgch_papers_B.md IEEE-SP/2025/learning-from-censored-experiences-social-media-discussions-around-censorship-ci "Telemetrio is an online database of Telegram channels across various languages and categories, and for our purposes, we specifically looked for channels whose primary language was Persian and contained the term "VPN" in " msgch_papers_B.md IEEE-SP/2025/learning-from-censored-experiences-social-media-discussions-around-censorship-ci "the 81 channels shared 1,459 unique VPN installation files" and "the channels distributed over 2,453 files, enabling direct connections to proxy servers" msgch_papers_B.md IEEE-SP/2025/learning-from-censored-experiences-social-media-discussions-around-censorship-ci "from September 18th, 2022, to January 31st, 2023" / "Our temporal study, conducted over 20 weeks" msgch_papers_B.md IEEE-SP/2025/learning-from-censored-experiences-social-media-discussions-around-censorship-ci "we also closely observed the overall responses to the files shared on those channels" / "To identify if a shared file/IP (as part of a proxy) was malicious, we used the VirusTotal API [100]." msgch_papers_B.md IEEE-SP/2025/learning-from-censored-experiences-social-media-discussions-around-censorship-ci "We utilized the Telegram API [20] and Telethon client [21] to collect posts from 34 unique Telegram channels" msgch_papers_B.md IEEE-SP/2025/learning-from-censored-experiences-social-media-discussions-around-censorship-ci "We used two data sources, i.e., Twitter and Telegram; however, we acknowledge that using other sources, such as Facebook, Reddit, etc., could also be beneficial in providing newer insights." msgch_papers_B.md IEEE-SP/2025/learning-from-censored-experiences-social-media-discussions-around-censorship-ci "1,730 IPs (5.9%) were detected as malicious" (out of proxy-address files); "31 out of 690 VPN configurations were malicious (4.4%)"; "HTTP Injections were 221 out of 1,763 (12.5%)" msgch_papers_B.md USENIX/2023/strategies-and-vulnerabilities-of-participants-in-venezuelan-influence-operation "Telegram Groups. During participant recruitment, we have identified and joined six Telegram groups (TuiterosDeChavez, Tuiteros Patriotas, TuiterosActivos, Twiteros Patriotas, Twiteros Activos, Bonos de la Patria) used by" msgch_papers_B.md USENIX/2023/strategies-and-vulnerabilities-of-participants-in-venezuelan-influence-operation "Our recruitment protocol identifies active operatives, by starting with a seed set of communication groups." msgch_papers_B.md NDSS/2020/practical-traffic-analysis-attacks-on-secure-messaging-applications "our adversary-controlled client uses their APIs to record SIM communications of target channels, while for WhatsApp, we manually send messages through its Desktop version (as it does not have an API)." msgch_papers_B.md WWW/2020/the-pod-people-understanding-manipulation-of-social-media-popularity-via-recipro "The earliest-created pod in our dataset became active in late 2016." (data collection performed "On February 26th, 2019" msgch_papers_B.md WWW/2020/the-pod-people-understanding-manipulation-of-social-media-popularity-via-recipro "we used the Telegram API to download the group metadata and all the messages from these public Telegram groups" msgch_papers_C.md USENIX/2024/moderating-illicit-online-image-promotion-for-unsafe-user-generated-content-game "In the case of Discord, we obtained illicit promotional images of UGCGs utilizing a server listing platform [73]. This platform enabled our entry into a variety of Roblox game servers on Discord, leading to the discovery" msgch_papers_C.md USENIX/2024/moderating-illicit-online-image-promotion-for-unsafe-user-generated-content-game "This platform enabled our entry into a variety of Roblox game servers on Discord, leading to the discovery of 210 instances for image-based illicit promotions of UGCGs." msgch_papers_C.md USENIX/2024/moderating-illicit-online-image-promotion-for-unsafe-user-generated-content-game "leading to the discovery of 210 instances for image-based illicit promotions of UGCGs. Among these, 92 images were classified as unsafe, while the remaining 118 were safe." msgch_papers_C.md USENIX/2024/moderating-illicit-online-image-promotion-for-unsafe-user-generated-content-game "leading to the discovery of 210 instances for image-based illicit promotions of UGCGs. Among these, 92 images were classified as unsafe, while the remaining 118 were safe." msgch_papers_C.md WWW/2023/a-prompt-log-analysis-of-text-to-image-generation-systems "The Midjourney dataset [39] is obtained by crawling message records from the Midjourney Discord community over a period of four weeks (June 20 – July 17, 2022). This dataset contains approximately 250K records, with user" msgch_papers_C.md WWW/2023/a-prompt-log-analysis-of-text-to-image-generation-systems "This dataset contains approximately 250K records, with user-input prompts, URLs of generated images, usernames, user IDs, message timestamps, and other Discord message metadata." msgch_papers_C.md USENIX/2025/bots-can-snoop-uncovering-and-mitigating-privacy-risks-of-bots-in-group-chats "Although these public datasets have limitations in representing private group chat dynamics, they enable a glance at potential privacy concerns" msgch_papers_C.md USENIX/2025/bots-can-snoop-uncovering-and-mitigating-privacy-risks-of-bots-in-group-chats "consultation with the Gophers community confirmed it as the open-source project "Discord Gophers Bot"" msgch_papers_C.md IMC/2022/exploring-the-security-and-privacy-risks-of-chatbots-in-messaging-services "Our research fully abides by the ethical principles guidelines outlined in the Menlo and Belmont Report. In particular, our system does not intentionally interact with humans nor collects data containing personal identif" msgch_papers_C.md IMC/2022/exploring-the-security-and-privacy-risks-of-chatbots-in-messaging-services "our findings reveal the inherent risks chatbots pose to users' security and privacy (55% of bots asking for administrator permissions, lack of traceability, improper use of those permissions)" msgch_papers_C.md IMC/2022/exploring-the-security-and-privacy-risks-of-chatbots-in-messaging-services "we find that 55% of chatbots from a leading Discord repository request the "administrator" permission, and only 4.35% of chatbots with permissions actually provide a privacy policy." msgch_papers_C.md USENIX/2025/investigating-the-impact-of-online-community-involvement-on-safety-practices-and "Although PWUD is also active in other online communities such as Telegram groups and self-constructed forums, for ethical reasons, we limited our online data collection to publicly accessible platforms." msgch_papers_C.md WWW/2025/detecting-and-understanding-the-promotion-of-illicit-goods-and-services-on-twitt "Using our PIP contact extractor, we have successfully extracted a total of 212,689 unique contacts across the Twitter platform from all the PIPs and PIP account profiles." msgch_papers_C.md WWW/2025/detecting-and-understanding-the-promotion-of-illicit-goods-and-services-on-twitt "led to the discovery of 37,621 distinct IM accounts, including 9,644 Telegram accounts, 11,561 WeChat accounts, 12,702 QQ accounts, 225 WhatsApp accounts, and 3,489 LINE accounts." msgch_papers_C.md CCS/2019/the-art-and-craft-of-fraudulent-app-promotion-in-google-play "Recruiting WhatsApp/Facebook groups need to aggressively accept new collaborators. We verified that these communication channels are easy to infiltrate." msgch_papers_C.md CCS/2019/the-art-and-craft-of-fraudulent-app-promotion-in-google-play "We have a Facebook group of more than 500 people, from different locations in Bangladesh, collected from various freelance groups in Facebook." msgch_papers_C.md USENIX/2026/cracks-in-the-walled-garden-dissecting-the-gray-market-of-unauthorized-ios-app-d "We selected RedNote (Xiaohongshu) [38] as our primary data source for identifying self-signing service providers, as it is one of the most active text-based social media platforms in China." msgch_papers_C.md USENIX/2026/cracks-in-the-walled-garden-dissecting-the-gray-market-of-unauthorized-ios-app-d "Using keywords such as "iOS certificate" and "customized V (WeChat)", we manually inspected posts and comments to extract publicly advertised signing sites." msgch_papers_C.md NDSS/2024/like-comment-get-scammed-characterizing-comment-scams-on-media-platforms "Because we interacted with users (the scammers), we sought and secured IRB approval from our institution. Our experiment does not involve any risky methods or physical contact with scammers; Instead, we focus solely on o" msgch_papers_C.md NDSS/2024/like-comment-get-scammed-characterizing-comment-scams-on-media-platforms "we obtained a waiver from IRB regarding the debriefing of scammers at the end of our conversations." msgch_papers_D.md NDSS/2021/all-the-numbers-are-us-large-scale-abuse-of-contact-discovery-in-mobile-messengers "Telegram allows each account to add a maximum of 5,000 contacts, irrespective of the rate. Once this limit is exceeded, each account is limited to 100 new numbers per day." msgch_papers_D.md NDSS/2021/all-the-numbers-are-us-large-scale-abuse-of-contact-discovery-in-mobile-messengers "49.6 % have a publicly available profile picture and 89.7 % have a public About text." (Telegram, random subset of 150,000 users; ligature "fi" msgch_papers_D.md NDSS/2021/all-the-numbers-are-us-large-scale-abuse-of-contact-discovery-in-mobile-messengers "There is also additional management information (such as the Telegram ID), which we do not detail here." (adjacent to: "the number of common groups" msgch_papers_D.md NDSS/2021/all-the-numbers-are-us-large-scale-abuse-of-contact-discovery-in-mobile-messengers "all profile pictures of the user (up to 100), and the number of common groups." (ligature "fi" msgch_papers_D.md NDSS/2026/hey-there-you-are-using-whatsapp-enumerating-three-billion-accounts-for-security-and-privacy "we developed libphonegen, a phone number generator leveraging the existing libphonenumber’s national number format data, and generated 646 B mobile phone numbers for the 245 ISO 3166-1 countries" msgch_papers_D.md NDSS/2026/hey-there-you-are-using-whatsapp-enumerating-three-billion-accounts-for-security-and-privacy "To refine our candidate set, we applied a hitlist-based approach to identify popular number ranges, reducing the respective amount to 480 M numbers." msgch_papers_D.md NDSS/2026/hey-there-you-are-using-whatsapp-enumerating-three-billion-accounts-for-security-and-privacy "The measurements were conducted in multiple rounds between the middle of December 2024 and the middle of April 2025." msgch_papers_D.md NDSS/2026/connecting-the-dots-an-investigative-study-on-linking-private-user-data-across-messaging-apps "we have responsibly disclosed our findings to both KakaoTalk and Tinder. KakaoTalk acknowledged our concerns, expressed appreciation for the report, and committed to deploying a fix." msgch_papers_D.md NDSS/2012/guess-who-s-texting-you-evaluating-the-security-of-smartphone-messaging-applicat "we split the entire number range of the San Diego area code 619 into chunks of 5000 phone numbers each and simulated a standard address book upload as performed by WhatsApp during device registration." msgch_papers_D.md NDSS/2012/guess-who-s-texting-you-evaluating-the-security-of-smartphone-messaging-applicat "All tested applications except HeyTell allow the user to upload the entire address book to the system’s server and compare the contained phone numbers to already registered phone numbers stored on the server. The server " msgch_papers_D.md NDSS/2012/guess-who-s-texting-you-evaluating-the-security-of-smartphone-messaging-applicat "It is possible to request V oypi users in the address book of other users. To this end, a simple HTTP request with the phone number of the victim is sent to the server: http://msg.voypi.com/myphone_v1/getusers.php? phone" msgch_papers_D.md NDSS/2012/guess-who-s-texting-you-evaluating-the-security-of-smartphone-messaging-applicat "the WhatsApp server did not prevent us from uploading ten million phone numbers and returned 21095 valid phone numbers that are using the WhatsApp application as well as their status messages. The entire process finished " msgch_papers_D.md NDSS/2012/guess-who-s-texting-you-evaluating-the-security-of-smartphone-messaging-applicat "we set up a SSL proxy that acted as a man-in-themiddle and intercepted requests to HTTPS servers." (the missing hyphen in "man-in-themiddle" msgch_papers_D.md NDSS/2012/guess-who-s-texting-you-evaluating-the-security-of-smartphone-messaging-applicat "We further used SSLsniff [14] by Moxie Marlinspike to read SSL- protected traffic that is not sent over HTTPS (e.g. XMPP)." (the space in "SSL- protected" msgch_papers_D.md NDSS/2012/guess-who-s-texting-you-evaluating-the-security-of-smartphone-messaging-applicat "HeyTell does not support upload of a whole address book for enumeration, but enumeration can be done number by number by requesting to send a voice message for every single number in the address book. This, however, is r" msgch_papers_D.md NDSS/2012/guess-who-s-texting-you-evaluating-the-security-of-smartphone-messaging-applicat "We practically demonstrated an attacker’s capability to enumerate any number of active WhatsApp accounts with a given area code (US area code 619 in our example, which corresponds to Southern San Diego, CA)." msgch_papers_D.md NDSS/2012/guess-who-s-texting-you-evaluating-the-security-of-smartphone-messaging-applicat "active phone numbers of the area code 619 start at 200000. We believe that the mobile number range starts above this value, but have not independently confirmed that." (registered-number distribution within the area code;"
Bibliography entries added
- bib_additions_messaging_channels.bib
@inproceedings{marjanov2026_stayin, author = {Marjanov, Tina and Tsuchiya, Taro and Ioannidis, Konstantinos and Hughes, Jack and Christin, Nicolas and Hutchings, Alice}, title = {Stayin' Alive: How Global Stolen Data Markets Thrive on Telegram}, booktitle = {Proceedings of the USENIX Security Symposium}, year = {2026}, series = {USENIX Security 2026}, url = {https://www.usenix.org/conference/usenixsecurity26/presentation/marjanov}, } @inproceedings{gao2026_doxing, author = {Gao, Yiran and Xia, Pengcheng and Wang, Liu and Liu, Tianming and Wang, Haoyu}, title = {Doxing-as-a-Service: Demystifying the Chinese Online Doxing Ecosystem}, booktitle = {Proceedings of the ACM Web Conference}, year = {2026}, series = {TheWebConf 2026}, doi = {10.1145/3774904.3792296}, } @inproceedings{xu2019_anatomy, author = {Xu, Jiahua and Livshits, Benjamin}, title = {The Anatomy of a Cryptocurrency Pump-and-Dump Scheme}, booktitle = {Proceedings of the USENIX Security Symposium}, year = {2019}, series = {USENIX Security 2019}, url = {https://www.usenix.org/conference/usenixsecurity19/presentation/xu-jiahua}, } @inproceedings{sun2021_having, author = {Sun, Zhibo and Oest, Adam and Zhang, Penghui and Rubio-Medrano, Carlos and Bao, Tiffany and Wang, Ruoyu and Zhao, Ziming and Shoshitaishvili, Yan and Doupé, Adam and Ahn, Gail-Joon}, title = {Having Your Cake and Eating It: An Analysis of Concession-Abuse-as-a-Service}, booktitle = {Proceedings of the USENIX Security Symposium}, year = {2021}, series = {USENIX Security 2021}, url = {https://www.usenix.org/conference/usenixsecurity21/presentation/sun-zhibo}, } @inproceedings{he2025_unmasking, author = {He, Bowen and Hu, Yufeng and Chen, Zhuo and Chen, Yuan and Yu, Ting and Chang, Rui and Wu, Lei and Zhou, Yajin}, title = {Unmasking the Shadow Economy: A Deep Dive into Drainer-as-a-Service Phishing on Ethereum}, booktitle = {Proceedings of the ACM Internet Measurement Conference}, year = {2025}, series = {IMC 2025}, doi = {10.1145/3730567.3764476}, } @inproceedings{weyns2026_mirai, author = {Weyns, Maarten and Ferrero, Dario and Beek, Stefan Op de and Wagner, Daniel and Smaragdakis, Georgios and Griffioen, Harm}, title = {From Mirai to Gorilla: Deep Dive into a Long-Lasting DDoS-for-Hire Botnet}, booktitle = {Proceedings of the USENIX Security Symposium}, year = {2026}, series = {USENIX Security 2026}, url = {https://www.usenix.org/conference/usenixsecurity26/presentation/weyns}, } @inproceedings{acharya2025_pirates, author = {Acharya, Bhupendra and Lazzaro, Dario and Cinà, Antonio Emanuele and Holz, Thorsten}, title = {Pirates of Charity: Exploring Donation-based Abuses in Social Media Platforms}, booktitle = {Proceedings of the ACM Web Conference}, year = {2025}, series = {TheWebConf 2025}, doi = {10.1145/3696410.3714634}, } @inproceedings{hoseini2020_demystifying, author = {Hoseini, Mohamad and Melo, Philipe and Junior, Manoel and Benevenuto, Fabrício and Chandrasekaran, Balakrishnan and Feldmann, Anja and Zannettou, Savvas}, title = {Demystifying the Messaging Platforms' Ecosystem Through the Lens of Twitter}, booktitle = {Proceedings of the ACM Internet Measurement Conference}, year = {2020}, series = {IMC 2020}, doi = {10.1145/3419394.3423651}, } @inproceedings{resende2019_information, author = {Resende, Gustavo and Melo, Philipe F. and Sousa, Hugo and Messias, Johnnatan and Vasconcelos, Marisa and Almeida, Jussara M. and Benevenuto, Fabrício}, title = {(Mis)Information Dissemination in WhatsApp: Gathering, Analyzing and Countermeasures}, booktitle = {Proceedings of the ACM Web Conference}, year = {2019}, series = {TheWebConf 2019}, doi = {10.1145/3308558.3313688}, } @inproceedings{saha2021_short, author = {Saha, Punyajoy and Mathew, Binny and Garimella, Kiran and Mukherjee, Animesh}, title = {"Short is the Road that Leads from Fear to Hate": Fear Speech in Indian WhatsApp Groups}, booktitle = {Proceedings of the ACM Web Conference}, year = {2021}, series = {TheWebConf 2021}, doi = {10.1145/3442381.3450137}, } @inproceedings{kireev2025_characterizing, author = {Kireev, Klim and Mykhno, Yevhen and Troncoso, Carmela and Overdorf, Rebekah}, title = {Characterizing and Detecting Propaganda-Spreading Accounts on Telegram}, booktitle = {Proceedings of the USENIX Security Symposium}, year = {2025}, series = {USENIX Security 2025}, url = {https://www.usenix.org/conference/usenixsecurity25/presentation/kireev}, } @inproceedings{vu2024_easy, author = {Vu, Anh V. and Hutchings, Alice and Anderson, Ross J.}, title = {No Easy Way Out: the Effectiveness of Deplatforming an Extremist Forum to Suppress Hate and Harassment}, booktitle = {Proceedings of the IEEE Symposium on Security and Privacy}, year = {2024}, series = {IEEE S&P 2024}, doi = {10.1109/sp54263.2024.00007}, } @inproceedings{vu2024_getting, author = {Vu, Anh V. and Thomas, Daniel R. and Collier, Ben and Hutchings, Alice and Clayton, Richard and Anderson, Ross J.}, title = {Getting Bored of Cyberwar: Exploring the Role of Low-level Cybercrime Actors in the Russia-Ukraine Conflict}, booktitle = {Proceedings of the ACM Web Conference}, year = {2024}, series = {TheWebConf 2024}, doi = {10.1145/3589334.3645401}, } @inproceedings{vafa2025_learning, author = {Vafa, Elham Pourabbas and Singhal, Mohit and Thota, Poojitha and Roy, Sayak Saha}, title = {Learning from Censored Experiences: Social Media Discussions around Censorship Circumvention Technologies}, booktitle = {Proceedings of the IEEE Symposium on Security and Privacy}, year = {2025}, series = {IEEE S&P 2025}, doi = {10.1109/sp61157.2025.00062}, } @inproceedings{recabarren2023_strategies, author = {Recabarren, Ruben and Carbunar, Bogdan and Hernandez, Nestor and Shafin, Ashfaq Ali}, title = {Strategies and Vulnerabilities of Participants in Venezuelan Influence Operations}, booktitle = {Proceedings of the USENIX Security Symposium}, year = {2023}, series = {USENIX Security 2023}, url = {https://www.usenix.org/conference/usenixsecurity23/presentation/recabarren}, } @inproceedings{aliapoulios2021_characterization, author = {Aliapoulios, Maxwell and Take, Kejsi and Ramakrishna, Prashanth and Borkan, Daniel and Goldberg, Beth and Sorensen, Jeffrey and Turner, Anna and Greenstadt, Rachel and Lauinger, Tobias and McCoy, Damon}, title = {A large-scale characterization of online incitements to harassment across platforms}, booktitle = {Proceedings of the ACM Internet Measurement Conference}, year = {2021}, series = {IMC 2021}, doi = {10.1145/3487552.3487852}, } @inproceedings{bahramali2020_practical, author = {Bahramali, Alireza and Houmansadr, Amir and Soltani, Ramin and Goeckel, Dennis and Towsley, Don}, title = {Practical Traffic Analysis Attacks on Secure Messaging Applications}, booktitle = {Proceedings of the Network and Distributed System Security Symposium}, year = {2020}, series = {NDSS 2020}, url = {https://www.ndss-symposium.org/ndss-paper/practical-traffic-analysis-attacks-on-secure-messaging-applications/}, } @inproceedings{weerasinghe2020_people, author = {Weerasinghe, Janith and Flanigan, Bailey and Stein, Aviel J. and McCoy, Damon and Greenstadt, Rachel}, title = {The Pod People: Understanding Manipulation of Social Media Popularity via Reciprocity Abuse}, booktitle = {Proceedings of the ACM Web Conference}, year = {2020}, series = {TheWebConf 2020}, doi = {10.1145/3366423.3380256}, } @inproceedings{shen2024_anything, author = {Shen, Xinyue and Chen, Zeyuan and Backes, Michael and Shen, Yun and Zhang, Yang}, title = {"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models}, booktitle = {Proceedings of the ACM SIGSAC Conference on Computer and Communications Security}, year = {2024}, series = {CCS 2024}, doi = {10.1145/3658644.3670388}, } @inproceedings{yu2024_listen, author = {Yu, Zhiyuan and Liu, Xiaogeng and Liang, Shunning and Cameron, Zach and Xiao, Chaowei and Zhang, Ning}, title = {Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models}, booktitle = {Proceedings of the USENIX Security Symposium}, year = {2024}, series = {USENIX Security 2024}, url = {https://www.usenix.org/conference/usenixsecurity24/presentation/yu-zhiyuan}, } @inproceedings{guo2024_moderating, author = {Guo, Keyan and Utkarsh, Ayush and Ding, Wenbo and Ondracek, Isabelle and Zhao, Ziming and Freeman, Guo and Vishwamitra, Nishant and Hu, Hongxin}, title = {Moderating Illicit Online Image Promotion for Unsafe User Generated Content Games Using Large Vision-Language Models}, booktitle = {Proceedings of the USENIX Security Symposium}, year = {2024}, series = {USENIX Security 2024}, url = {https://www.usenix.org/conference/usenixsecurity24/presentation/guo-keyan}, } @inproceedings{schrittwieser2012_guess, author = {Schrittwieser, Sebastian and Frühwirt, Peter and Kieseberg, Peter and Leithner, Manuel and Mulazzani, Martin and Huber, Markus and Weippl, Edgar}, title = {Guess Who’s Texting You? Evaluating the Security of Smartphone Messaging Applications}, booktitle = {Proceedings of the Network and Distributed System Security Symposium}, year = {2012}, series = {NDSS 2012}, url = {https://www.ndss-symposium.org/ndss2012/ndss-2012-programme/guess-whos-texting-you-evaluating-security-smartphone-messaging-applications/}, } @inproceedings{li2025_investigating, author = {Li, Jiliang and Lu, Nora Sinong and Hanimann, Isaak and Si, Janice Jianing and Cheng, Dazhao and Zhou, Xiaobo and Wang, Kanye Ye}, title = {Investigating the Impact of Online Community Involvement on Safety Practices and Perceived Risks Among People Who Use Drugs}, booktitle = {Proceedings of the USENIX Security Symposium}, year = {2025}, series = {USENIX Security 2025}, url = {https://www.usenix.org/conference/usenixsecurity25/presentation/li-jiliang}, } @inproceedings{albrecht2021_collective, author = {Albrecht, Martin R. and Blasco, Jorge and Jensen, Rikke Bjerg and Mareková, Lenka}, title = {Collective Information Security in Large-Scale Urban Protests: the Case of Hong Kong}, booktitle = {Proceedings of the USENIX Security Symposium}, year = {2021}, series = {USENIX Security 2021}, url = {https://www.usenix.org/conference/usenixsecurity21/presentation/albrecht}, } @inproceedings{arunasalam2024_security, author = {Arunasalam, Arjun and Farrukh, Habiba and Tekcan, Eliz and Celik, Z. Berkay}, title = {Understanding the Security and Privacy Implications of Online Toxic Content on Refugees}, booktitle = {Proceedings of the USENIX Security Symposium}, year = {2024}, series = {USENIX Security 2024}, url = {https://www.usenix.org/conference/usenixsecurity24/presentation/arunasalam}, } @inproceedings{chou2025_bots, author = {Chou, Kai-Hsiang and Lin, Yi-Min and Wang, Yi-An and Li, Jonathan Weiping and Kim, Tiffany Hyun-Jin and Hsiao, Hsu-Chun}, title = {Bots can Snoop: Uncovering and Mitigating Privacy Risks of Bots in Group Chats}, booktitle = {Proceedings of the USENIX Security Symposium}, year = {2025}, series = {USENIX Security 2025}, url = {https://www.usenix.org/conference/usenixsecurity25/presentation/chou}, } @inproceedings{wang2025_detecting, author = {Wang, Hongyu and Li, Ying and Huang, Ronghong and Mi, Xianghang}, title = {Detecting and Understanding the Promotion of Illicit Goods and Services on Twitter}, booktitle = {Proceedings of the ACM Web Conference}, year = {2025}, series = {TheWebConf 2025}, doi = {10.1145/3696410.3714550}, }
References
- [1]
- Vu, Anh V.; Collier, Ben; Thomas, Daniel R.; Kristoff, John; Clayton, Richard; Hutchings, Alice (2025): "Assessing the Aftermath: the Effects of a Global Takedown against DDoS-for-hire Services", in: Proceedings of the USENIX Security Symposium. (Link)
- [2]
- Gao, Yiran; Xia, Pengcheng; Wang, Liu; Liu, Tianming; Wang, Haoyu (2026): "Doxing-as-a-Service: Demystifying the Chinese Online Doxing Ecosystem", in: Proceedings of the ACM Web Conference. (DOI)
- [3]
- Acharya, Bhupendra; Lazzaro, Dario; Cinà, Antonio Emanuele; Holz, Thorsten (2025): "Pirates of Charity: Exploring Donation-based Abuses in Social Media Platforms", in: Proceedings of the ACM Web Conference. (DOI)
- [4]
- Saha, Punyajoy; Mathew, Binny; Garimella, Kiran; Mukherjee, Animesh (2021): ""Short is the Road that Leads from Fear to Hate": Fear Speech in Indian WhatsApp Groups", in: Proceedings of the ACM Web Conference. (DOI)
- [5]
- Hoseini, Mohamad; Melo, Philipe; Junior, Manoel; Benevenuto, Fabrício; Chandrasekaran, Balakrishnan; Feldmann, Anja; Zannettou, Savvas (2020): "Demystifying the Messaging Platforms' Ecosystem Through the Lens of Twitter", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [6]
- Xu, Jiahua; Livshits, Benjamin (2019): "The Anatomy of a Cryptocurrency Pump-and-Dump Scheme", in: Proceedings of the USENIX Security Symposium. (Link)
- [7]
- Vu, Anh V.; Hutchings, Alice; Anderson, Ross J. (2024): "No Easy Way Out: the Effectiveness of Deplatforming an Extremist Forum to Suppress Hate and Harassment", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
- [8]
- Saha Roy, Sayak; Pourabbas Vafa, Elham; Khanmohamaddi, Kobra; Nilizadeh, Shirin (2025): "DarkGram: A Large-Scale Analysis of Cybercriminal Activity Channels on Telegram", in: Proceedings of the USENIX Security Symposium. (Link)
- [9]
- Kireev, Klim; Mykhno, Yevhen; Troncoso, Carmela; Overdorf, Rebekah (2025): "Characterizing and Detecting Propaganda-Spreading Accounts on Telegram", in: Proceedings of the USENIX Security Symposium. (Link)
- [10]
- Shen, Xinyue; Chen, Zeyuan; Backes, Michael; Shen, Yun; Zhang, Yang (2024): ""Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
