| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| provenance:design:platforms:messaging_channels [2026/09/27 12:06] – Fix namespace-relative links to design and roadmap; correct the override count (two overridden, one kept). Authored by Claude. karel.kubicek.claude | provenance:design:platforms:messaging_channels [2026/09/27 12:36] (current) – Generic review logged (16 findings, 15 accepted, 1 in part); mistakes list extended; regenerated verifier (68 spans, 83 needles) and external checks (80 OK). Authored by Claude. karel.kubicek.claude |
|---|
| * **The draft over-generalised two paper quotes.** It credited "two criminology-trained groups" with a consent-waiver argument that only one paper makes in those terms, and wrote "no paper here describes entering a group by deception" without saying that one of the 26 interacted with the operators it studied. Both rewritten against the text. | * **The draft over-generalised two paper quotes.** It credited "two criminology-trained groups" with a consent-waiver argument that only one paper makes in those terms, and wrote "no paper here describes entering a group by deception" without saying that one of the 26 interacted with the operators it studied. Both rewritten against the text. |
| * **The external-check script silently skipped its last section on the first run.** ''set -u'' aborted on an unset ''$1'' inside the Crossref loop, and the run printed nine FAILED lines above it — all nine were my regexes (tag-stripped text searched for values that live in HTML attributes), not the sources. Fixed; the script now reaches its final ''FAILED checks: N'' line, and exits with N. | * **The external-check script silently skipped its last section on the first run.** ''set -u'' aborted on an unset ''$1'' inside the Crossref loop, and the run printed nine FAILED lines above it — all nine were my regexes (tag-stripped text searched for values that live in HTML attributes), not the sources. Fixed; the script now reaches its final ''FAILED checks: N'' line, and exits with N. |
| | * **Two hand codes behind headline figures were wrong until review.** {[vu2025_assessing]} also monitored Discord channels (Discord 5 → 6), and {[gao2026_doxing]}'s only denominator caveat concerns a secondary sample, so "all nine core papers acknowledge the denominator" was eight. Three more codes (two join modes, one ethics review) were recoded; see [[#Review log]]. |
| | * **The external-source sub-agent said TeraGram had no DOI yet.** It is in the ICWSM 2026 proceedings; the external-currency reviewer caught it. |
| | * **The page missed an application route.** It said no platform offers researchers a way in while its own parent page documents that Meta's Content Library covers WhatsApp Channels. The generic reviewer caught it by reading the two pages together. |
| * **A located note quote is not a published quote.** The readers' notes carry 606 quotes; 501 (82.7%) were located verbatim. The page quotes only strings that the page's own verifier locates, and the notes stay unpublished. | * **A located note quote is not a published quote.** The readers' notes carry 606 quotes; 501 (82.7%) were located verbatim. The page quotes only strings that the page's own verifier locates, and the notes stay unpublished. |
| |
| * **The inclusion rule counts metadata.** //Practical Traffic Analysis Attacks on Secure Messaging Applications// (NDSS 2020) joined over 1,000 public Telegram channels to record message timing and sizes for a traffic model. The reader coded it OUT (traffic analysis); I coded it IN-section, because the rule says "messages, members, media or metadata" and the rule was written first. A reader who prefers "content only" should subtract one: 25 papers, 16 section. | * **The inclusion rule counts metadata.** //Practical Traffic Analysis Attacks on Secure Messaging Applications// (NDSS 2020) joined over 1,000 public Telegram channels to record message timing and sizes for a traffic model. The reader coded it OUT (traffic analysis); I coded it IN-section, because the rule says "messages, members, media or metadata" and the rule was written first. A reader who prefers "content only" should subtract one: 25 papers, 16 section. |
| * **Two other reader verdicts were overridden**, each recorded in its verdict note in ''msgch_fold.mjs'': the NDSS 2024 scam-baiting paper (reader IN-section → CONTEXT dm: one group entered on a scammer's invitation is incidental) and the USENIX 2025 WeChat-groups paper (reader CONTEXT user-study → OUT recruitment: the groups were joined to recruit interviewees). One was questioned and kept: the CCS 2020 impersonation-as-a-service paper (reader OUT; it collected a handful of examples from a Telegram channel reached through the marketplace it studied, which the rule could be read to cover; kept OUT as incidental). | * **Two other reader verdicts were overridden**, each recorded in its verdict note in ''msgch_fold.mjs'': the NDSS 2024 scam-baiting paper (reader IN-section → CONTEXT dm: one group entered on a scammer's invitation is incidental) and the USENIX 2025 WeChat-groups paper (reader CONTEXT user-study → OUT recruitment: the groups were joined to recruit interviewees). One was questioned and kept: the CCS 2020 impersonation-as-a-service paper (reader OUT; it collected a handful of examples from a Telegram channel reached through the marketplace it studied, which the rule could be read to cover; kept OUT as incidental). |
| * **Public-channel reading counts as the route.** Eight papers read public Telegram channels without saying whether they joined. The page keeps them in the population and reports the join mode separately, rather than defining them out — a channel read through Telethon or the web preview is still content from inside the platform's group layer, found by a seed, with the same denominator problem. | * **Public-channel reading counts as the route.** Six papers say they read public Telegram channels without saying whether they joined, and six more do not say how they got in at all. The page keeps them in the population and reports the join mode separately, rather than defining them out — a channel read through Telethon or the web preview is still content from inside the platform's group layer, found by a seed, with the same denominator problem. |
| * **Commercial scraper services count.** Two papers bought Telegram channel data from Apify/Telemetrio. The route is the same; the seed is the vendor's and cannot be described, which the page says. | * **Commercial scraper services count.** Two papers bought Telegram channel data from Apify/Telemetrio. The route is the same; the seed is the vendor's and cannot be described, which the page says. |
| * **"Current" verdicts.** "Historical in these venues" for WhatsApp group monitoring rests on a corpus fact (no WhatsApp group collection after 2021 in the seven venues) and one outside methods paper proposing donation instead; it is **not** a claim that nobody does it anywhere. "Historical as a route" for Twitter link harvesting rests on the X API change documented on [[design:platforms]]. "Forbidden" for Discord user-account automation rests on Discord's own guidelines, support article and developer policy, fetched 2026-09-27. | * **"Current" verdicts.** "Historical in these venues" for WhatsApp group monitoring rests on a corpus fact (no WhatsApp group collection after 2021 in the seven venues) and one outside methods paper proposing donation instead; it is **not** a claim that nobody does it anywhere. "Historical as a route" for Twitter link harvesting rests on the X API change documented on [[design:platforms]]. "Forbidden" for Discord user-account automation rests on Discord's own guidelines, support article and developer policy, fetched 2026-09-27. |
| | N of 26 for every hand code | the 26 IN papers | section E | | | N of 26 for every hand code | the 26 IN papers | section E | |
| | per 1,000 corpus papers | all corpus papers in the year bucket | section D | | | per 1,000 corpus papers | all corpus papers in the year bucket | section D | |
| | 17 of 26 state a review outcome vs 33.8% | the 26 (all empirical) vs 5,118 empirical papers | section G | | | 17 of 26 state a review outcome vs 33.8% | the 26 (all empirical) vs 5,118 empirical papers — the extraction's reading; the hand codes give 16, the difference being {[acharya2025_pirates]} | section G | |
| | 5 / 3 / 9 on request / restricted / public | the 26 vs 5,118 empirical papers | section G | | | 5 / 3 / 9 on request / restricted / public | the 26 vs 5,118 empirical papers | section G | |
| | 24 vs 21 Telegram papers | ''%%platforms_report.mjs --list Telegram%%'' vs the 21 Telegram IN papers | section I | | | 24 vs 21 Telegram papers | ''%%platforms_report.mjs --list Telegram%%'' vs the 21 Telegram IN papers | section I | |
| ===== Quote and figure checks ===== | ===== Quote and figure checks ===== |
| |
| * **Page quotes**: ''verify_messaging_channels_figures.mjs'' pulls **every** ''%%//"…"//%%'' span out of the page source (62 at the last run) and requires each to be located in the paper of a citekey on the same line, or to be an EXTERNAL span whose external check printed OK, or to be on a two-entry NOT_A_QUOTE list. One span needed a ligature fallback (the IMC 2020 PDF has an unmapped ''ff'' glyph in "different" in both renderings), and one a de-hyphenation fallback ("self- constructed" across a line break); both routes are labelled in the output. | * **Page quotes**: ''verify_messaging_channels_figures.mjs'' pulls **every** ''%%//"…"//%%'' span out of the page source (68 at the last run) and requires each to be located in the paper of a citekey on the same line, or to be an EXTERNAL span whose external check printed OK, or to be on a two-entry NOT_A_QUOTE list. One span needed a ligature fallback (the IMC 2020 PDF has an unmapped ''ff'' glyph in "different" in both renderings), and one a de-hyphenation fallback ("self- constructed" across a line break); both routes are labelled in the output. |
| * **Per-paper figures**: 74 needles, located in ''.cols'' or, for ten, only in the ''pypdf'' re-extraction. Three mutated needles (250→350, 0.8%→0.9%, 4,709→4,790) must not be found and are not. | * **Per-paper figures**: 83 needles, located in ''.cols'' or, for twelve, only in the ''pypdf'' re-extraction. Three mutated needles (250→350, 0.8%→0.9%, 4,709→4,790) must not be found and are not. |
| * **Two numbers are verified in the form the PDF prints them**: {[gao2026_doxing]}'s "411, 707" and {[saha2021_short]}'s "8, 000 … 1, 000 … 3, 000" carry a stray space inside the number in every rendering. | * **Two numbers are verified in the form the PDF prints them**: {[gao2026_doxing]}'s "411, 707" and {[saha2021_short]}'s "8, 000 … 1, 000 … 3, 000" carry a stray space inside the number in every rendering. |
| * **Reader notes**: ''msgch_notes_quotecheck.py'' (output below) — 501 of 606 quotes located (82.7%). The misses are dominated by readers stitching two sentences with an ellipsis, against the brief, and by column splices. No page quote depends on an unlocated note quote. | * **Reader notes**: ''msgch_notes_quotecheck.py'' (output below) — 501 of 606 quotes located (82.7%). The misses are dominated by readers stitching two sentences with an ellipsis, against the brief, and by column splices. No page quote depends on an unlocated note quote. |
| ===== Review log ===== | ===== Review log ===== |
| |
| Four reviewers were run on the frozen snapshot of 2026-09-27 (rev 1790510475 of the content page and this page as first published). Their findings are logged here once they return. | Four reviewers, each told that the author's context may not be exhaustive and each handed the page, the provenance page, the report script and its output, the verifier and the external checks. The three focused passes ran in parallel on the frozen snapshot (content rev 1790510757, provenance rev 1790510789); nothing was edited while they ran. Their findings were applied together, then the generic pass ran on the corrected pages. Findings files: ''notes/msgch_review_*.md'' in the workdir. |
| | |
| | ==== Pass 1 — figures against the script (model: sonnet) ==== |
| | |
| | Re-ran the report: byte-identical to the committed output. 5 findings. |
| | |
| | ^ # ^ Finding ^ Disposition ^ |
| | | F1 | high — the page says {[vu2025_assessing]} monitored booter Discord channels, and the paper says so, but the hand code listed Telegram only; the Discord count (5) and the multi-platform count (2) were therefore one short | **accepted**. Code fixed; Discord 6, multi-platform 3; tip box and currency table corrected | |
| | | F2 | medium — the tooling sentence enumerated every category except the PumpOlymp API (1 paper), so its parts did not add up to the report | **accepted**, category added | |
| | | F3 | low — the verifier's footer claimed the TGStat Russia tile and the Telemetr.io Web Archive numbers were re-fetched, but no check covered them | **accepted**: three checks added to the external script (Russia tile, both Telemetr.io numbers from the capture) | |
| | | F4 | low — the page said "at least three mentions of a platform" where the code sums the three platforms (3 borderline candidates, none in the population) | **accepted**, page reworded to match the code rather than the code changed | |
| | | F5 | low — the needle for {[acharya2025_pirates]}'s 85,402 did not contain the number | **accepted**, needle now "(76,111/85,402) from Telegram were filtered" | |
| | |
| | ==== Pass 2 — citations and quotes (model: sonnet) ==== |
| | |
| | All 37 citekeys resolve; no duplicate keys or DOIs; all 15 URL-only entries (13 USENIX, 2 NDSS) plus ''wang2025_detecting'' match the papers' author blocks. 14 findings (4 medium, 10 low). |
| | |
| | ^ # ^ Finding ^ Disposition ^ |
| | | C1 | medium — the 196-channel takedown was presented as the outcome of reporting the 339 monitored channels; it is the outcome for 196 //new// channels found through links on Telegram and Facebook, while the monitored channels' outcome is "only 64 channels (19%) were removed" | **accepted**; both outcomes now stated, with their denominators, in //What to Read First// and the results table | |
| | | C2 | medium — ''join = public-read'' for {[xu2019_anatomy]} rests on the //organiser's// public channel, not on what the researchers did | **accepted**, recoded ''not-stated'' | |
| | | C3 | medium — same for {[vu2024_easy]} ("Both channels permit public access" describes the channels) | **accepted**, recoded ''not-stated''; the join counts became 11 / 6 / 2 / 1 / 6 and the page now explains the distinction | |
| | | C4 | medium — ''irb = not-required'' for {[acharya2025_pirates]}: the paper says only that it "did not directly involve interaction with any human subjects" | **accepted**, recoded ''not-stated''; "say nothing about review" 9 → 10 in three places; the page now explains why the extraction's 17 is one higher than the hand codes' 16 | |
| | | C5 | medium — ''denominator = acknowledged'' for {[gao2026_doxing]}: its only caveat is about a secondary 15-group sample; for the main top-100 dataset it calls the corpus "representative" | **accepted**, recoded ''not-stated''. This broke a headline: "all nine core papers acknowledge" became **eight of nine**, and the page says which one does not | |
| | | C6 | low — "refugee support group" is the page's gloss | **accepted**, now "a refugee participants' WhatsApp group" | |
| | | C7 | low — the NDSS 2024 scam-baiting paper literally satisfies the IN rule for one group | **rejected, already recorded**: kept CONTEXT dm as a documented judgement call; the page says it entered one group on a scammer's invitation | |
| | | C8–C14 | low — seven client-table clauses had no re-fetch in the external script (TDLib tags, Pyrogram README wording, Kurigram fork and currency, GramJS → teleproto on npm, yowsup's Python ceiling, discord.py upload date, DiscordChatExporter 2.48) | **accepted**: eight checks added; all pass. The page's "re-fetches every row" is now true | |
| | |
| | ==== Pass 3 — external currency (model: sonnet) ==== |
| | |
| | Independently re-fetched every tool state, term quote, API limit, designation, directory and dataset. 2 findings. |
| | |
| | ^ # ^ Finding ^ Disposition ^ |
| | | E1 | medium — TeraGram is no longer only an arXiv preprint: Crossref records it in the ICWSM 2026 proceedings (doi 10.1609/icwsm.v20i1.42783, issued 2026-05-25) | **accepted**; footnote cites the DOI, and a Crossref check was added. The external-source sub-agent had reported "arXiv; no DOI yet" — its search was wrong, which is exactly what this pass exists to catch | |
| | | E2 | low — Discord's help page conditions never-expiring invites on Community servers | **accepted**, qualifier added | |
| | |
| | ==== Pass 4 — generic (model: fable) ==== |
| | |
| | Ran on the corrected pages (content rev 1790511893, provenance rev 1790511895). 16 findings (1 high, 6 medium, 9 low); it also re-derived several of the fixes above and found them sound. |
| | |
| | ^ # ^ Finding ^ Disposition ^ |
| | | G1 | high — the page said none of the three platforms has a research route and called WhatsApp Channels "new and untested", while [[design:platforms]] documents that Meta's Content Library covers WhatsApp Channels | **accepted**. Re-fetched Meta's page (headless browser): "public content archive from Facebook, Instagram and WhatsApp Channels". The intro box, the terms section, the currency table and the open questions now say that Channels have an application route, that groups do not, and that no paper here used it; an external check was added | |
| | | G2 | medium — "specific, recent" terms applied to 2019–2024 papers when the page cannot date the clause | **accepted**: "recent" dropped; "read against today's terms", and the page says whether the clauses existed at collection time was not established | |
| | | G3 | medium — the artefact sentence gave base rates only for the two rows where group studies exceed the corpus | **accepted**: public 34.6% vs 47.7% added; the sentence now says they share less openly and shift to gated release | |
| | | G4 | medium — "the only compliant path [on Discord] is a bot an admin invites" ignores a person reading as a member, which the page's own codes record | **accepted**, rewritten in the terms section and the open questions | |
| | | G5 | medium — "still possible at scale in 2025" generalised one platform's pre-mitigation window to the whole route | **accepted**: now "capped on most platforms", with WhatsApp's window, Meta's mitigation statement and the capped platforms named | |
| | | G6 | medium — "what the careful papers converge on" rests on one paper for two bullets, and one is contradicted by {[roy2025_darkgram]}, which sent 11,800 collected posts to the GPT-4 API | **accepted**: reframed as what individual papers did; the third-party-service bullet is marked contested, cites both papers, and states the page's position | |
| | | G7 | medium — "what the 2020–2024 Discord papers describe or imply" attributes user-account automation to six papers when only one is coded that way | **accepted**: names {[hoseini2020_demystifying]} only | |
| | | G8 | low — Twitter as a seed surface is still used (2025); only the free API is historical | **accepted**, currency row split accordingly | |
| | | G9 | low — "known official channels" vs the fourth n/a paper reached through a forum advert | **accepted**: "known groups or channels" | |
| | | G10 | low — the recruitment row named two platforms where the 19 include Telegram and WeChat; two cells had no citation | **accepted**: platforms listed; the two uncited cells point to the verdict list | |
| | | G11 | low — the seed table omitted the "not stated" row (2 papers) | **accepted**, row added | |
| | | G12 | low — "almost all of it is Telegram" (81%) | **accepted**: "four papers in five" | |
| | | G13 | low — the Discord 100-server cap rests on a 2020 measurement while Telegram's is documented | **partly accepted**: both measured caps are now dated "in 2020" and the Discord one is marked "not re-checked against Discord's current documentation"; Discord's support page was not re-fetched for a current figure | |
| | | G14 | low — the page explains Discord bots' limits but not Telegram bots' | **accepted**: a table row on bot accounts, with Telegram's own privacy-mode sentence, and an external check | |
| | | G15 | low — message-level attrition (deletions between back-fill and live polling) not mentioned | **accepted**: {[kireev2025_characterizing]} added — export history "does not contain deleted messages", and moderators removed from below 20% to over 80% of propaganda messages | |
| | | G16 | low — no read-first entry for Discord or for the commoner "one source among several" case | **accepted**: {[shen2024_anything]} and {[vu2025_assessing]} named under //What to Read First// | |
| | |
| | Not re-run after these fixes: the generic pass's own findings were applied as stated, every changed figure re-checked by the number guard, and every new quote by the verifier (68 spans, 83 needles, 0 not located) and the external script. A second generic pass was not run. |
| | |
| | ==== What the review layer was worth ==== |
| | |
| | 37 findings over four passes; 35 accepted in full, 1 in part, 1 rejected as an already-recorded judgement call. The two that changed headline figures came from the focused passes (the missing Discord code, and the one core paper that does not acknowledge its denominator). The one that changed what the page tells a reader to do — an application route to WhatsApp Channels the page had missed — came from the generic pass, reading this page against its parent. |
| |
| ===== The report script ===== | ===== The report script ===== |
| |
| == D. The 26 IN papers: platform, year, venue == | == D. The 26 IN papers: platform, year, venue == |
| platform (a paper can name several): Telegram 21; Discord 5; WhatsApp 3 | platform (a paper can name several): Telegram 21; Discord 6; WhatsApp 3 |
| papers with more than one platform: 2 | papers with more than one platform: 3 |
| bucket corpus papers IN IN core IN per 1,000 corpus papers Telegram IN | bucket corpus papers IN IN core IN per 1,000 corpus papers Telegram IN |
| ---------- ------------- -- ------- -------------------------- ----------- | ---------- ------------- -- ------- -------------------------- ----------- |
| -- join (of 26; multi-valued fields do not sum): | -- join (of 26; multi-valued fields do not sum): |
| 11 42.3% joined | 11 42.3% joined |
| 8 30.8% public-read | 6 23.1% not-stated |
| 4 15.4% not-stated | 6 23.1% public-read |
| 2 7.7% scraper-service | 2 7.7% scraper-service |
| 1 3.8% vendor | 1 3.8% vendor |
| -- irb (of 26; multi-valued fields do not sum): | -- irb (of 26; multi-valued fields do not sum): |
| 10 38.5% approved | 10 38.5% approved |
| 9 34.6% not-stated | 10 38.5% not-stated |
| 3 11.5% not-required | |
| 2 7.7% no-board-named | 2 7.7% no-board-named |
| | 2 7.7% not-required |
| 1 3.8% exempt | 1 3.8% exempt |
| 1 3.8% other-part-only | 1 3.8% other-part-only |
| 1 3.8% users anonymised | 1 3.8% users anonymised |
| -- denominator (of 26; multi-valued fields do not sum): | -- denominator (of 26; multi-valued fields do not sum): |
| 15 57.7% acknowledged | 14 53.8% acknowledged |
| 7 26.9% not-stated | 8 30.8% not-stated |
| 4 15.4% n/a (known group or channel) | 4 15.4% n/a (known group or channel) |
| |
| join=scraper-service: 2 papers, 2 from 2024-2026, 2 from 2022-2026; years 2024, 2025 | join=scraper-service: 2 papers, 2 from 2024-2026, 2 from 2022-2026; years 2024, 2025 |
| platforms=WhatsApp: 3 papers, 0 from 2024-2026, 0 from 2022-2026; years 2019, 2020, 2021 | platforms=WhatsApp: 3 papers, 0 from 2024-2026, 0 from 2022-2026; years 2019, 2020, 2021 |
| platforms=Discord: 5 papers, 3 from 2024-2026, 3 from 2022-2026; years 2020, 2021, 2024, 2024, 2024 | platforms=Discord: 6 papers, 4 from 2024-2026, 4 from 2022-2026; years 2020, 2021, 2024, 2024, 2024, 2025 |
| platforms=Telegram: 21 papers, 13 from 2024-2026, 14 from 2022-2026; years 2019, 2020, 2020, 2020, 2021, 2021, 2021, 2023, 2024, 2024, 2024, 2025, 2025, 2025, 2025, 2025, 2025, 2025, 2026, 2026, 2026 | platforms=Telegram: 21 papers, 13 from 2024-2026, 14 from 2022-2026; years 2019, 2020, 2020, 2020, 2021, 2021, 2021, 2023, 2024, 2024, 2024, 2025, 2025, 2025, 2025, 2025, 2025, 2025, 2026, 2026, 2026 |
| core papers by bucket and object: | core papers by bucket and object: |
| -- CORE papers only (the ones whose main dataset is group/channel data): | -- CORE papers only (the ones whose main dataset is group/channel data): |
| discovery (of 9): directory 4; in-app-search 3; web-search 3; other-platform-links 2; snowball 2; prior-list 1 | discovery (of 9): directory 4; in-app-search 3; web-search 3; other-platform-links 2; snowball 2; prior-list 1 |
| join (of 9): joined 6; public-read 2; not-stated 1 | join (of 9): joined 6; not-stated 2; public-read 1 |
| tooling (of 9): Telegram API (client unnamed) 5; Telethon 2; phones + Garimella-Tyson tool 2; PumpOlymp API 1; WhatsApp Web client 1; Discord API (user account) 1 | tooling (of 9): Telegram API (client unnamed) 5; Telethon 2; phones + Garimella-Tyson tool 2; PumpOlymp API 1; WhatsApp Web client 1; Discord API (user account) 1 |
| irb (of 9): approved 3; not-stated 3; not-required 1; no-board-named 1; exempt 1 | irb (of 9): approved 3; not-stated 3; not-required 1; no-board-named 1; exempt 1 |
| denominator (of 9): acknowledged 9 | denominator (of 9): acknowledged 8; not-stated 1 |
| |
| == F. Per-paper table (IN, 26) == | == F. Per-paper table (IN, 26) == |
| year venue role platforms found collected messages discovery join tooling slug | year venue role platforms found collected messages discovery join tooling slug |
| ---- ------- ------- ------------------------- ----------- ---------------- -------------------- ------------------------------------------- --------------- ---------------------------------------------------------------------------- ------------------------------------------------ | ---- ------- ------- ------------------------- ----------- ---------------- -------------------- ------------------------------------------- --------------- ---------------------------------------------------------------------------- ------------------------------------------------ |
| 2019 USENIX core Telegram 300+ 300+ n/s prior-list public-read Telegram API (client unnamed)+PumpOlymp API the-anatomy-of-a-cryptocurrency-pump-and-dump-sc | 2019 USENIX core Telegram 300+ 300+ n/s prior-list not-stated Telegram API (client unnamed)+PumpOlymp API the-anatomy-of-a-cryptocurrency-pump-and-dump-sc |
| 2019 WWW core WhatsApp 3,444 141 + 364 121,781 + 789,914 web-search joined phones + Garimella-Tyson tool mis-information-dissemination-in-whatsapp-gather | 2019 WWW core WhatsApp 3,444 141 + 364 121,781 + 789,914 web-search joined phones + Garimella-Tyson tool mis-information-dissemination-in-whatsapp-gather |
| 2020 IMC core WhatsApp+Telegram+Discord 351,535 616 8,255,069 other-platform-links joined WhatsApp Web client+Telegram API (client unnamed)+Discord API (user account) demystifying-the-messaging-platforms-ecosystem-t | 2020 IMC core WhatsApp+Telegram+Discord 351,535 616 8,255,069 other-platform-links joined WhatsApp Web client+Telegram API (client unnamed)+Discord API (user account) demystifying-the-messaging-platforms-ecosystem-t |
| 2023 USENIX section Telegram 6 6 n/s in-app-search+snowball joined manual strategies-and-vulnerabilities-of-participants-i | 2023 USENIX section Telegram 6 6 n/s in-app-search+snowball joined manual strategies-and-vulnerabilities-of-participants-i |
| 2024 CCS section Discord 20 6 n/s directory not-stated not-stated do-anything-now-characterizing-and-evaluating-in | 2024 CCS section Discord 20 6 n/s directory not-stated not-stated do-anything-now-characterizing-and-evaluating-in |
| 2024 IEEE-SP section Telegram 2 2 525k known-official-channel public-read Telethon no-easy-way-out-the-effectiveness-of-deplatformi | 2024 IEEE-SP section Telegram 2 2 525k known-official-channel not-stated Telethon no-easy-way-out-the-effectiveness-of-deplatformi |
| 2024 USENIX section Telegram n/s n/s 133,399 in-app-search scraper-service commercial scraper (Apify, Telemetrio) the-imitation-game-exploring-brand-impersonation | 2024 USENIX section Telegram n/s n/s 133,399 in-app-search scraper-service commercial scraper (Apify, Telemetrio) the-imitation-game-exploring-brand-impersonation |
| 2024 USENIX section Discord n/s 2 n/s not-stated not-stated Selenium dont-listen-to-me-understanding-and-exploring-ja | 2024 USENIX section Discord n/s 2 n/s not-stated not-stated Selenium dont-listen-to-me-understanding-and-exploring-ja |
| 2025 IMC section Telegram n/s n/s n/s in-app-search+other-platform-links joined manual unmasking-the-shadow-economy-a-deep-dive-into-dr | 2025 IMC section Telegram n/s n/s n/s in-app-search+other-platform-links joined manual unmasking-the-shadow-economy-a-deep-dive-into-dr |
| 2025 USENIX core Telegram 4,709 339 64,801 directory not-stated Telegram API (client unnamed) darkgram-a-large-scale-analysis-of-cybercriminal | 2025 USENIX core Telegram 4,709 339 64,801 directory not-stated Telegram API (client unnamed) darkgram-a-large-scale-analysis-of-cybercriminal |
| 2025 USENIX section Telegram n/s 52 34,438 prior-list public-read Telethon assessing-the-aftermath-the-effects-of-a-global- | 2025 USENIX section Telegram+Discord n/s 52 34,438 prior-list public-read Telethon assessing-the-aftermath-the-effects-of-a-global- |
| 2025 USENIX core Telegram n/s 13 17.3M directory+in-app-search joined Telethon characterizing-and-detecting-propaganda-spreadin | 2025 USENIX core Telegram n/s 13 17.3M directory+in-app-search joined Telethon characterizing-and-detecting-propaganda-spreadin |
| 2025 WWW section Telegram n/s n/s 85,402 in-app-search scraper-service commercial scraper (Apify, Telemetrio) pirates-of-charity-exploring-donation-based-abus | 2025 WWW section Telegram n/s n/s 85,402 in-app-search scraper-service commercial scraper (Apify, Telemetrio) pirates-of-charity-exploring-donation-based-abus |
| // discovery: directory | in-app-search | other-platform-links | web-search | snowball | prior-list | // discovery: directory | in-app-search | other-platform-links | web-search | snowball | prior-list |
| // (a third party's list or vendor) | known-official-channel | not-stated | // (a third party's list or vendor) | known-official-channel | not-stated |
| // join: joined | public-read (read a public channel; joining not stated) | scraper-service | | // join: joined | public-read (the paper says it read a public channel, not whether it joined) | |
| // vendor | not-stated | // scraper-service | vendor | not-stated (the paper does not say how it got in; "Both channels |
| | // permit public access" describes the channel, not the researchers — coded not-stated) |
| // history: yes (back-filled messages from before collection began) | mixed | not-stated | // history: yes (back-filled messages from before collection began) | mixed | not-stated |
| // irb: approved | exempt | not-required | no-board-named | other-part-only | not-stated | // irb: approved | exempt | not-required | no-board-named | other-part-only | not-stated |
| // passive: stated = the paper says it did not post, interact or contact members | // passive: stated = the paper says it did not post, interact or contact members |
| // denominator: acknowledged = the paper says groups found are not groups that exist, or names its seed bias | // denominator: acknowledged = the paper says groups found are not groups that exist, or names its seed bias, |
| | // for the dataset the page reports (a caveat about a secondary sub-sample does not count) |
| export const HAND_FIELDS = ['platforms', 'object', 'discovery', 'join', 'history', 'members', 'tooling', 'limits', 'irb', 'passive', 'minimise', 'denominator', 'found', 'collected', 'messages']; | export const HAND_FIELDS = ['platforms', 'object', 'discovery', 'join', 'history', 'members', 'tooling', 'limits', 'irb', 'passive', 'minimise', 'denominator', 'found', 'collected', 'messages']; |
| export const HAND = { | export const HAND = { |
| "USENIX/2025/darkgram-a-large-scale-analysis-of-cybercriminal-activity-channels-on-telegram": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["directory"], "join": "not-stated", "history": "not-stated", "members": "admin-only lists, not collected", "tooling": ["Telegram API (client unnamed)"], "limits": ["rate-limit", "channels-removed"], "irb": "not-required", "passive": "stated", "minimise": ["no payload download", "PII anonymised in release"], "denominator": "acknowledged", "found": "4,709", "collected": "339", "messages": "64,801"}, | "USENIX/2025/darkgram-a-large-scale-analysis-of-cybercriminal-activity-channels-on-telegram": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["directory"], "join": "not-stated", "history": "not-stated", "members": "admin-only lists, not collected", "tooling": ["Telegram API (client unnamed)"], "limits": ["rate-limit", "channels-removed"], "irb": "not-required", "passive": "stated", "minimise": ["no payload download", "PII anonymised in release"], "denominator": "acknowledged", "found": "4,709", "collected": "339", "messages": "64,801"}, |
| "USENIX/2026/stayin-alive-how-global-stolen-data-markets-thrive-on-telegram": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["in-app-search", "other-platform-links", "snowball"], "join": "joined", "history": "yes", "members": "posters only", "tooling": ["Telethon"], "limits": ["rate-limit", "channels-removed"], "irb": "approved", "passive": "stated", "minimise": ["file type and size limits", "no de-anonymisation", "controlled-access release"], "denominator": "acknowledged", "found": "21k", "collected": "1,521", "messages": "~14M"}, | "USENIX/2026/stayin-alive-how-global-stolen-data-markets-thrive-on-telegram": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["in-app-search", "other-platform-links", "snowball"], "join": "joined", "history": "yes", "members": "posters only", "tooling": ["Telethon"], "limits": ["rate-limit", "channels-removed"], "irb": "approved", "passive": "stated", "minimise": ["file type and size limits", "no de-anonymisation", "controlled-access release"], "denominator": "acknowledged", "found": "21k", "collected": "1,521", "messages": "~14M"}, |
| "WWW/2026/doxing-as-a-service-demystifying-the-chinese-online-doxing-ecosystem": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["directory", "in-app-search"], "join": "joined", "history": "yes", "members": "no", "tooling": ["Telegram API (client unnamed)"], "limits": ["channels-removed"], "irb": "no-board-named", "passive": "stated", "minimise": ["masked before storage"], "denominator": "acknowledged", "found": "312", "collected": "100", "messages": "25,972"}, | "WWW/2026/doxing-as-a-service-demystifying-the-chinese-online-doxing-ecosystem": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["directory", "in-app-search"], "join": "joined", "history": "yes", "members": "no", "tooling": ["Telegram API (client unnamed)"], "limits": ["channels-removed"], "irb": "no-board-named", "passive": "stated", "minimise": ["masked before storage"], "denominator": "not-stated", "found": "312", "collected": "100", "messages": "25,972"}, |
| "USENIX/2019/the-anatomy-of-a-cryptocurrency-pump-and-dump-scheme": {"platforms": ["Telegram"], "object": "crypto manipulation", "discovery": ["prior-list"], "join": "public-read", "history": "yes", "members": "counts only", "tooling": ["Telegram API (client unnamed)", "PumpOlymp API"], "limits": ["channels-removed"], "irb": "not-stated", "passive": "not-stated", "minimise": [], "denominator": "acknowledged", "found": "300+", "collected": "300+", "messages": "n/s"}, | "USENIX/2019/the-anatomy-of-a-cryptocurrency-pump-and-dump-scheme": {"platforms": ["Telegram"], "object": "crypto manipulation", "discovery": ["prior-list"], "join": "not-stated", "history": "yes", "members": "counts only", "tooling": ["Telegram API (client unnamed)", "PumpOlymp API"], "limits": ["channels-removed"], "irb": "not-stated", "passive": "not-stated", "minimise": [], "denominator": "acknowledged", "found": "300+", "collected": "300+", "messages": "n/s"}, |
| "USENIX/2021/catching-phishers-by-their-bait-investigating-the-dutch-phishing-landscape-throu": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["snowball"], "join": "public-read", "history": "not-stated", "members": "no", "tooling": ["manual"], "limits": [], "irb": "not-stated", "passive": "not-stated", "minimise": [], "denominator": "not-stated", "found": "n/s", "collected": "50", "messages": "n/s"}, | "USENIX/2021/catching-phishers-by-their-bait-investigating-the-dutch-phishing-landscape-throu": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["snowball"], "join": "public-read", "history": "not-stated", "members": "no", "tooling": ["manual"], "limits": [], "irb": "not-stated", "passive": "not-stated", "minimise": [], "denominator": "not-stated", "found": "n/s", "collected": "50", "messages": "n/s"}, |
| "USENIX/2021/having-your-cake-and-eating-it-an-analysis-of-concession-abuse-as-a-service": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["other-platform-links"], "join": "joined", "history": "not-stated", "members": "member profiles", "tooling": ["not-stated"], "limits": [], "irb": "approved", "passive": "not-stated", "minimise": ["anonymised"], "denominator": "n/a (known group or channel)", "found": "1", "collected": "1", "messages": "17,898"}, | "USENIX/2021/having-your-cake-and-eating-it-an-analysis-of-concession-abuse-as-a-service": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["other-platform-links"], "join": "joined", "history": "not-stated", "members": "member profiles", "tooling": ["not-stated"], "limits": [], "irb": "approved", "passive": "not-stated", "minimise": ["anonymised"], "denominator": "n/a (known group or channel)", "found": "1", "collected": "1", "messages": "17,898"}, |
| "IMC/2025/unmasking-the-shadow-economy-a-deep-dive-into-drainer-as-a-service-phishing-on-e": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["in-app-search", "other-platform-links"], "join": "joined", "history": "not-stated", "members": "no", "tooling": ["manual"], "limits": [], "irb": "no-board-named", "passive": "not-stated", "minimise": [], "denominator": "acknowledged", "found": "n/s", "collected": "n/s", "messages": "n/s"}, | "IMC/2025/unmasking-the-shadow-economy-a-deep-dive-into-drainer-as-a-service-phishing-on-e": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["in-app-search", "other-platform-links"], "join": "joined", "history": "not-stated", "members": "no", "tooling": ["manual"], "limits": [], "irb": "no-board-named", "passive": "not-stated", "minimise": [], "denominator": "acknowledged", "found": "n/s", "collected": "n/s", "messages": "n/s"}, |
| "USENIX/2026/from-mirai-to-gorilla-deep-dive-into-a-long-lasting-ddos-for-hire-botnet": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["known-official-channel"], "join": "public-read", "history": "not-stated", "members": "counts only", "tooling": ["not-stated"], "limits": ["channels-removed"], "irb": "not-stated", "passive": "not-stated", "minimise": [], "denominator": "n/a (known group or channel)", "found": "1", "collected": "1", "messages": "n/s"}, | "USENIX/2026/from-mirai-to-gorilla-deep-dive-into-a-long-lasting-ddos-for-hire-botnet": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["known-official-channel"], "join": "public-read", "history": "not-stated", "members": "counts only", "tooling": ["not-stated"], "limits": ["channels-removed"], "irb": "not-stated", "passive": "not-stated", "minimise": [], "denominator": "n/a (known group or channel)", "found": "1", "collected": "1", "messages": "n/s"}, |
| "USENIX/2025/assessing-the-aftermath-the-effects-of-a-global-takedown-against-ddos-for-hire-s": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["prior-list"], "join": "public-read", "history": "not-stated", "members": "no", "tooling": ["Telethon"], "limits": ["channels-removed"], "irb": "approved", "passive": "not-stated", "minimise": ["quotes paraphrased", "aggregate analysis"], "denominator": "acknowledged", "found": "n/s", "collected": "52", "messages": "34,438"}, | "USENIX/2025/assessing-the-aftermath-the-effects-of-a-global-takedown-against-ddos-for-hire-s": {"platforms": ["Telegram", "Discord"], "object": "cybercrime market", "discovery": ["prior-list"], "join": "public-read", "history": "not-stated", "members": "no", "tooling": ["Telethon"], "limits": ["channels-removed"], "irb": "approved", "passive": "not-stated", "minimise": ["quotes paraphrased", "aggregate analysis"], "denominator": "acknowledged", "found": "n/s", "collected": "52", "messages": "34,438"}, |
| "WWW/2025/pirates-of-charity-exploring-donation-based-abuses-in-social-media-platforms": {"platforms": ["Telegram"], "object": "scams and impersonation", "discovery": ["in-app-search"], "join": "scraper-service", "history": "not-stated", "members": "no", "tooling": ["commercial scraper (Apify, Telemetrio)"], "limits": [], "irb": "not-required", "passive": "stated", "minimise": [], "denominator": "not-stated", "found": "n/s", "collected": "n/s", "messages": "85,402"}, | "WWW/2025/pirates-of-charity-exploring-donation-based-abuses-in-social-media-platforms": {"platforms": ["Telegram"], "object": "scams and impersonation", "discovery": ["in-app-search"], "join": "scraper-service", "history": "not-stated", "members": "no", "tooling": ["commercial scraper (Apify, Telemetrio)"], "limits": [], "irb": "not-stated", "passive": "stated", "minimise": [], "denominator": "not-stated", "found": "n/s", "collected": "n/s", "messages": "85,402"}, |
| "USENIX/2024/the-imitation-game-exploring-brand-impersonation-attacks-on-social-media-platfor": {"platforms": ["Telegram"], "object": "scams and impersonation", "discovery": ["in-app-search"], "join": "scraper-service", "history": "not-stated", "members": "no", "tooling": ["commercial scraper (Apify, Telemetrio)"], "limits": [], "irb": "not-stated", "passive": "not-stated", "minimise": [], "denominator": "acknowledged", "found": "n/s", "collected": "n/s", "messages": "133,399"}, | "USENIX/2024/the-imitation-game-exploring-brand-impersonation-attacks-on-social-media-platfor": {"platforms": ["Telegram"], "object": "scams and impersonation", "discovery": ["in-app-search"], "join": "scraper-service", "history": "not-stated", "members": "no", "tooling": ["commercial scraper (Apify, Telemetrio)"], "limits": [], "irb": "not-stated", "passive": "not-stated", "minimise": [], "denominator": "acknowledged", "found": "n/s", "collected": "n/s", "messages": "133,399"}, |
| "IMC/2020/demystifying-the-messaging-platforms-ecosystem-through-the-lens-of-twitter": {"platforms": ["WhatsApp", "Telegram", "Discord"], "object": "the messaging ecosystem itself", "discovery": ["other-platform-links"], "join": "joined", "history": "mixed", "members": "member lists (phone numbers hashed)", "tooling": ["WhatsApp Web client", "Telegram API (client unnamed)", "Discord API (user account)"], "limits": ["join-cap", "invite-expiry"], "irb": "approved", "passive": "not-stated", "minimise": ["phone numbers hashed"], "denominator": "acknowledged", "found": "351,535", "collected": "616", "messages": "8,255,069"}, | "IMC/2020/demystifying-the-messaging-platforms-ecosystem-through-the-lens-of-twitter": {"platforms": ["WhatsApp", "Telegram", "Discord"], "object": "the messaging ecosystem itself", "discovery": ["other-platform-links"], "join": "joined", "history": "mixed", "members": "member lists (phone numbers hashed)", "tooling": ["WhatsApp Web client", "Telegram API (client unnamed)", "Discord API (user account)"], "limits": ["join-cap", "invite-expiry"], "irb": "approved", "passive": "not-stated", "minimise": ["phone numbers hashed"], "denominator": "acknowledged", "found": "351,535", "collected": "616", "messages": "8,255,069"}, |
| "USENIX/2025/characterizing-and-detecting-propaganda-spreading-accounts-on-telegram": {"platforms": ["Telegram"], "object": "politics and misinformation", "discovery": ["directory", "in-app-search"], "join": "joined", "history": "yes", "members": "no", "tooling": ["Telethon"], "limits": [], "irb": "approved", "passive": "stated", "minimise": ["users anonymised", "deletion on request"], "denominator": "acknowledged", "found": "n/s", "collected": "13", "messages": "17.3M"}, | "USENIX/2025/characterizing-and-detecting-propaganda-spreading-accounts-on-telegram": {"platforms": ["Telegram"], "object": "politics and misinformation", "discovery": ["directory", "in-app-search"], "join": "joined", "history": "yes", "members": "no", "tooling": ["Telethon"], "limits": [], "irb": "approved", "passive": "stated", "minimise": ["users anonymised", "deletion on request"], "denominator": "acknowledged", "found": "n/s", "collected": "13", "messages": "17.3M"}, |
| "WWW/2025/exposing-cross-platform-coordinated-inauthentic-activity-in-the-run-up-to-the-20": {"platforms": ["Telegram"], "object": "politics and misinformation", "discovery": ["in-app-search"], "join": "not-stated", "history": "not-stated", "members": "no", "tooling": ["Telegram API (client unnamed)"], "limits": [], "irb": "not-stated", "passive": "not-stated", "minimise": [], "denominator": "not-stated", "found": "n/s", "collected": "15,537", "messages": "4,309,880"}, | "WWW/2025/exposing-cross-platform-coordinated-inauthentic-activity-in-the-run-up-to-the-20": {"platforms": ["Telegram"], "object": "politics and misinformation", "discovery": ["in-app-search"], "join": "not-stated", "history": "not-stated", "members": "no", "tooling": ["Telegram API (client unnamed)"], "limits": [], "irb": "not-stated", "passive": "not-stated", "minimise": [], "denominator": "not-stated", "found": "n/s", "collected": "15,537", "messages": "4,309,880"}, |
| "IEEE-SP/2024/no-easy-way-out-the-effectiveness-of-deplatforming-an-extremist-forum-to-suppres": {"platforms": ["Telegram"], "object": "extremism and harassment", "discovery": ["known-official-channel"], "join": "public-read", "history": "yes", "members": "no", "tooling": ["Telethon"], "limits": [], "irb": "approved", "passive": "not-stated", "minimise": [], "denominator": "n/a (known group or channel)", "found": "2", "collected": "2", "messages": "525k"}, | "IEEE-SP/2024/no-easy-way-out-the-effectiveness-of-deplatforming-an-extremist-forum-to-suppres": {"platforms": ["Telegram"], "object": "extremism and harassment", "discovery": ["known-official-channel"], "join": "not-stated", "history": "yes", "members": "no", "tooling": ["Telethon"], "limits": [], "irb": "approved", "passive": "not-stated", "minimise": [], "denominator": "n/a (known group or channel)", "found": "2", "collected": "2", "messages": "525k"}, |
| "WWW/2024/getting-bored-of-cyberwar-exploring-the-role-of-low-level-cybercrime-actors-in-t": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["known-official-channel"], "join": "public-read", "history": "yes", "members": "counts only", "tooling": ["Telethon"], "limits": [], "irb": "approved", "passive": "not-stated", "minimise": ["aggregate reporting"], "denominator": "n/a (known group or channel)", "found": "1", "collected": "1", "messages": "441 + 57,757 replies"}, | "WWW/2024/getting-bored-of-cyberwar-exploring-the-role-of-low-level-cybercrime-actors-in-t": {"platforms": ["Telegram"], "object": "cybercrime market", "discovery": ["known-official-channel"], "join": "public-read", "history": "yes", "members": "counts only", "tooling": ["Telethon"], "limits": [], "irb": "approved", "passive": "not-stated", "minimise": ["aggregate reporting"], "denominator": "n/a (known group or channel)", "found": "1", "collected": "1", "messages": "441 + 57,757 replies"}, |
| "IEEE-SP/2025/learning-from-censored-experiences-social-media-discussions-around-censorship-ci": {"platforms": ["Telegram"], "object": "censorship circumvention", "discovery": ["directory", "other-platform-links"], "join": "public-read", "history": "not-stated", "members": "no", "tooling": ["Telethon"], "limits": [], "irb": "not-stated", "passive": "not-stated", "minimise": ["usernames masked"], "denominator": "acknowledged", "found": "n/s", "collected": "81", "messages": "54K"}, | "IEEE-SP/2025/learning-from-censored-experiences-social-media-discussions-around-censorship-ci": {"platforms": ["Telegram"], "object": "censorship circumvention", "discovery": ["directory", "other-platform-links"], "join": "public-read", "history": "not-stated", "members": "no", "tooling": ["Telethon"], "limits": [], "irb": "not-stated", "passive": "not-stated", "minimise": ["usernames masked"], "denominator": "acknowledged", "found": "n/s", "collected": "81", "messages": "54K"}, |
| OK paper.cols.txt line 27 roy2025_darkgram | OK paper.cols.txt line 27 roy2025_darkgram |
| "this approach might have introduced potential biases by omitting smaller or newly emerging channels" | "this approach might have introduced potential biases by omitting smaller or newly emerging channels" |
| | OK paper.cols.txt line 27 roy2025_darkgram |
| | "only 64 channels (19%) were removed" |
| OK paper.cols.txt line 27 roy2025_darkgram | OK paper.cols.txt line 27 roy2025_darkgram |
| "led to the removal of all 196 channels, with a median response time of 4 days" | "led to the removal of all 196 channels, with a median response time of 4 days" |
| OK paper.cols.txt line 28 kireev2025_characterizing | OK paper.cols.txt line 28 kireev2025_characterizing |
| "given the small selection, we cannot make any statement about the pervasiveness of this propaganda activity in Telegram" | "given the small selection, we cannot make any statement about the pervasiveness of this propaganda activity in Telegram" |
| OK paper.cols.txt line 61 hoseini2020_demystifying | OK paper.cols.txt line 62 hoseini2020_demystifying |
| "presumably owing to Discord group URLs automatically expiring after a day" | "presumably owing to Discord group URLs automatically expiring after a day" |
| OK EXTERNAL line 61 Discord invites default to 7 days | OK EXTERNAL line 62 Discord invites default to 7 days |
| "a 7 days access link by default" | "a 7 days access link by default" |
| OK EXTERNAL line 62 Discord Server Discovery needs 1,000 members | OK EXTERNAL line 63 Discord Server Discovery needs 1,000 members |
| "at least 1,000 members" | "at least 1,000 members" |
| OK EXTERNAL line 62 TGStat self-reported catalogue size | OK EXTERNAL line 63 TGStat self-reported catalogue size |
| "More than 2 864 885 channels and groups" | "More than 2 864 885 channels and groups" |
| OK EXTERNAL line 63 durov/345: problematic content no longer accessible in Search | OK EXTERNAL line 64 durov/345: problematic content no longer accessible in Search |
| "All the problematic content we identified in Search is no longer accessible" | "All the problematic content we identified in Search is no longer accessible" |
| OK EXTERNAL line 73 Telegram documents the public-channel web preview | OK EXTERNAL line 74 Telegram bots: privacy mode on by default unless added as admin |
| | "Privacy mode is enabled by default for all bots, except bots that were added to a group as admins" |
| | OK EXTERNAL line 75 Telegram documents the public-channel web preview |
| "The contents of public channels can be seen on the Web without a Telegram account" | "The contents of public channels can be seen on the Web without a Telegram account" |
| OK paper.cols.txt line 80 kireev2025_characterizing | OK paper.cols.txt line 82 kireev2025_characterizing |
| "never interacted with the channels" | "never interacted with the channels" |
| OK paper.cols.txt line 81 hoseini2020_demystifying | OK paper.cols.txt line 83 hoseini2020_demystifying |
| "translates to a need for hundreds of phones and SIM cards to join all discovered groups" | "translates to a need for hundreds of phones and SIM cards to join all discovered groups" |
| OK EXTERNAL line 81 Telegram: unofficial-client logins put under observation | OK EXTERNAL line 83 Telegram: unofficial-client logins put under observation |
| "all accounts that log in using unofficial Telegram API clients are automatically put under observation" | "all accounts that log in using unofficial Telegram API clients are automatically put under observation" |
| OK paper.cols.txt line 87 resende2019_information | OK paper.cols.txt line 89 resende2019_information |
| "we are not aware of an approach that would allow us to assess the representativeness of our data as even the total number of groups available in the country is " | "we are not aware of an approach that would allow us to assess the representativeness of our data as even the total number of groups available in the country is " |
| OK paper.cols.txt line 88 saha2021_short | OK paper.cols.txt line 90 saha2021_short |
| "Given that WhatsApp does not provide an API or tools to access the data, there is no way of knowing the representativeness of our dataset" | "Given that WhatsApp does not provide an API or tools to access the data, there is no way of knowing the representativeness of our dataset" |
| OK paper.cols.txt line 88 saha2021_short | OK paper.cols.txt line 90 saha2021_short |
| "a convenience sample" | "a convenience sample" |
| OK paper.cols.txt+lig line 89 hoseini2020_demystifying | OK paper.cols.txt+lig line 91 hoseini2020_demystifying |
| "The use of Twitter as the only data source for discovering public groups of the different messaging platforms potentially introduces some bias in our sample" | "The use of Twitter as the only data source for discovering public groups of the different messaging platforms potentially introduces some bias in our sample" |
| OK paper.cols.txt line 95 marjanov2026_stayin | OK paper.cols.txt line 97 marjanov2026_stayin |
| "inactive/banned" | "inactive/banned" |
| OK paper.cols.txt line 95 marjanov2026_stayin | OK paper.cols.txt line 97 marjanov2026_stayin |
| "the most influential predictor of channel durability" | "the most influential predictor of channel durability" |
| OK paper.cols.txt line 95 marjanov2026_stayin (WARN: located in a citekey other than the nearest) | OK paper.cols.txt line 97 marjanov2026_stayin (WARN: located in a citekey other than the nearest) |
| "might miss short-lived channels due to the retroactive nature of data collection" | "might miss short-lived channels due to the retroactive nature of data collection" |
| OK paper.cols.txt line 97 arunasalam2024_security | OK paper.cols.txt line 97 kireev2025_characterizing |
| | "The data collected through "Export chat history" does not contain deleted messages" |
| | OK paper.cols.txt line 99 arunasalam2024_security |
| "that can only be joined via invitation" | "that can only be joined via invitation" |
| OK paper.cols.txt line 97 arunasalam2024_security | OK paper.cols.txt line 99 arunasalam2024_security |
| "an unintentional leak of the group's "invite link"" | "an unintentional leak of the group's "invite link"" |
| OK NOT-A-QUOTE line 100 template of the sentence NOT to write, in the denominator box | OK NOT-A-QUOTE line 102 template of the sentence NOT to write, in the denominator box |
| "we collected N messages from Telegram" | "we collected N messages from Telegram" |
| OK EXTERNAL line 112 Telethon README: moved to Codeberg | OK EXTERNAL line 114 Telethon README: moved to Codeberg |
| "Moved to https://codeberg.org/Lonami/Telethon" | "Moved to https://codeberg.org/Lonami/Telethon" |
| OK EXTERNAL line 112 Telethon README (reST links rendered): be careful not to break ToS | OK EXTERNAL line 114 Telethon README (reST links rendered): be careful not to break ToS |
| "be careful not to break Telegram's ToS or Telegram can ban the account" | "be careful not to break Telegram's ToS or Telegram can ban the account" |
| OK EXTERNAL line 116 whatsapp-web.js README: WhatsApp does not allow unofficial clients | OK EXTERNAL line 118 whatsapp-web.js README: WhatsApp does not allow unofficial clients |
| "WhatsApp does not allow bots or unofficial clients on their platform, so this shouldn't be considered totally safe" | "WhatsApp does not allow bots or unofficial clients on their platform, so this shouldn't be considered totally safe" |
| OK EXTERNAL line 119 DiscordChatExporter README: automating user accounts is against Discord TOS | OK EXTERNAL line 121 DiscordChatExporter README: automating user accounts is against Discord TOS |
| "automating user accounts is against Discord TOS and may result in you getting banned" | "automating user accounts is against Discord TOS and may result in you getting banned" |
| OK EXTERNAL line 125 Telegram ToS: scraping prohibited for all users and third parties | OK EXTERNAL line 125 Meta Content Library covers WhatsApp Channels |
| | "provide[s] comprehensive access to the public content archive from Facebook, Instagram and WhatsApp Channels" |
| | OK EXTERNAL line 127 Telegram ToS: scraping prohibited for all users and third parties |
| "Telegram additionally prohibits data scraping as part of its Content Licensing and AI Scraping Terms, which apply to all users, businesses, and third-party serv" | "Telegram additionally prohibits data scraping as part of its Content Licensing and AI Scraping Terms, which apply to all users, businesses, and third-party serv" |
| OK EXTERNAL line 125 Telegram Content Licensing: access beyond ordinary use prohibited | OK EXTERNAL line 127 Telegram Content Licensing: access beyond ordinary use prohibited |
| "Access to user-generated content for any purpose other than ordinary, legitimate, and intended use of the Telegram platform as its user is prohibited" | "Access to user-generated content for any purpose other than ordinary, legitimate, and intended use of the Telegram platform as its user is prohibited" |
| OK EXTERNAL line 125 Telegram API ToS 1.5: no ML training on API data | OK EXTERNAL line 127 Telegram API ToS 1.5: no ML training on API data |
| "prohibited from using, accessing or aggregating data obtained from the Telegram platform to train, fine-tune or otherwise engage in the development" | "prohibited from using, accessing or aggregating data obtained from the Telegram platform to train, fine-tune or otherwise engage in the development" |
| OK EXTERNAL line 125 Telegram self-reports below the VLOP threshold | OK EXTERNAL line 127 Telegram self-reports below the VLOP threshold |
| "significantly fewer than 45 million" | "significantly fewer than 45 million" |
| OK EXTERNAL line 126 WhatsApp ToS: no access or collection through automated means | OK EXTERNAL line 128 WhatsApp ToS: no access or collection through automated means |
| "through automated or other means" | "through automated or other means" |
| OK EXTERNAL line 126 WhatsApp ToS: no collecting information about users in impermissible ways | OK EXTERNAL line 128 WhatsApp ToS: no collecting information about users in impermissible ways |
| "collect information of or about our users in any impermissible or unauthorized manner" | "collect information of or about our users in any impermissible or unauthorized manner" |
| OK EXTERNAL line 126 EC: private messaging stays excluded | OK EXTERNAL line 128 EC: private messaging stays excluded |
| "private messaging service" | "private messaging service" |
| OK EXTERNAL line 127 Discord ToS: no scraping without written consent | OK EXTERNAL line 129 Discord ToS: no scraping without written consent |
| "scraping our services without our written consent" | "scraping our services without our written consent" |
| OK EXTERNAL line 127 Discord Community Guidelines: no self-bots | OK EXTERNAL line 129 Discord Community Guidelines: no self-bots |
| "Do not use self-bots or user-bots" | "Do not use self-bots or user-bots" |
| OK EXTERNAL line 127 Discord Developer Policy: do not mine or scrape | OK EXTERNAL line 129 Discord Developer Policy: do not mine or scrape |
| "Do not mine or scrape any data" | "Do not mine or scrape any data" |
| OK EXTERNAL line 132 FLOOD_WAIT_X: a wait of X seconds is required | OK EXTERNAL line 134 FLOOD_WAIT_X: a wait of X seconds is required |
| "A wait of X seconds is required" | "A wait of X seconds is required" |
| OK paper.cols.txt line 132 roy2025_darkgram | OK paper.cols.txt line 134 roy2025_darkgram |
| "at 10-minute intervals to capture new posts while adhering to the API rate limits" | "at 10-minute intervals to capture new posts while adhering to the API rate limits" |
| OK paper.cols.txt line 133 vu2025_assessing | OK paper.cols.txt line 135 vu2025_assessing |
| "often banned rapidly by Discord" | "often banned rapidly by Discord" |
| OK pypdf line 140 schrittwieser2012_guess | OK pypdf line 142 schrittwieser2012_guess |
| "the WhatsApp server did not prevent us from uploading ten million phone numbers and returned 21095 valid phone numbers" | "the WhatsApp server did not prevent us from uploading ten million phone numbers and returned 21095 valid phone numbers" |
| OK paper.cols.txt line 141 hagen2021_numbers | OK paper.cols.txt line 143 hagen2021_numbers |
| "accounts get banned when excessively using the contact discovery service" | "accounts get banned when excessively using the contact discovery service" |
| OK pypdf line 143 gegenhuber2026_there | OK pypdf line 145 gegenhuber2026_there |
| "We encountered no rate limits, our accounts were not banned from the platform" | "We encountered no rate limits, our accounts were not banned from the platform" |
| OK paper.cols.txt line 157 kireev2025_characterizing | OK paper.cols.txt line 159 kireev2025_characterizing |
| "never interacted with the channels" | "never interacted with the channels" |
| OK paper.cols.txt line 163 marjanov2026_stayin | OK paper.cols.txt line 165 marjanov2026_stayin |
| "We do not lie or pretend to be an interested buyer to be admitted into groups. We also do not attempt to join any groups that require payments or vouching by an" | "We do not lie or pretend to be an interested buyer to be admitted into groups. We also do not attempt to join any groups that require payments or vouching by an" |
| OK paper.cols.txt line 163 he2025_unmasking | OK paper.cols.txt line 165 he2025_unmasking |
| "joined several related Telegram groups" | "joined several related Telegram groups" |
| OK paper.cols.txt line 163 he2025_unmasking | OK paper.cols.txt line 165 he2025_unmasking |
| "communicated with operators, acquired wallet drainers" | "communicated with operators, acquired wallet drainers" |
| OK paper.cols.txt line 164 vu2025_assessing | OK paper.cols.txt line 166 vu2025_assessing |
| "sending thousands of messages could be regarded as spamming" | "sending thousands of messages could be regarded as spamming" |
| OK paper.cols.txt line 165 roy2025_darkgram | OK paper.cols.txt line 167 roy2025_darkgram |
| "we did not download or read the payload files" | "we did not download or read the payload files" |
| OK pypdf line 166 marjanov2026_stayin | OK pypdf line 168 marjanov2026_stayin |
| "to avoid sending stolen and potentially sensitive data to third-party servers" | "to avoid sending stolen and potentially sensitive data to third-party servers" |
| OK pypdf+dehyph line 167 li2025_investigating | OK paper.cols.txt line 168 roy2025_darkgram |
| | "we utilized the GPT-4 API" |
| | OK pypdf+dehyph line 169 li2025_investigating |
| "Although PWUD is also active in other online communities such as Telegram groups and self-constructed forums, for ethical reasons, we limited our online data co" | "Although PWUD is also active in other online communities such as Telegram groups and self-constructed forums, for ethical reasons, we limited our online data co" |
| OK paper.cols.txt line 169 albrecht2021_collective | OK paper.cols.txt line 171 albrecht2021_collective |
| "all participants in our study also assumed police monitoring of the public Telegram groups" | "all participants in our study also assumed police monitoring of the public Telegram groups" |
| OK EXTERNAL line 171 Barbosa and Milan 2019: avoid by all means covert bypasses | OK EXTERNAL line 173 Barbosa and Milan 2019: avoid by all means covert bypasses |
| "by all means covert bypasses" | "by all means covert bypasses" |
| OK paper.cols.txt line 260 hoseini2020_demystifying | OK paper.cols.txt line 262 hoseini2020_demystifying |
| "exposes at least one social media account for 30% of the Discord users we monitored" | "exposes at least one social media account for 30% of the Discord users we monitored" |
| OK paper.cols.txt line 263 roy2025_darkgram | OK paper.cols.txt line 265 roy2025_darkgram |
| | "only 64 channels (19%) were removed" |
| | OK paper.cols.txt line 265 roy2025_darkgram |
| "led to the removal of all 196 channels, with a median response time of 4 days" | "led to the removal of all 196 channels, with a median response time of 4 days" |
| OK paper.cols.txt line 264 gao2026_doxing | OK paper.cols.txt line 266 gao2026_doxing |
| "over 300,000 unique individuals" | "over 300,000 unique individuals" |
| OK pypdf line 266 saha2021_short | OK pypdf line 268 saha2021_short |
| "8% of these fear speech users are also admins in the groups where they post fear speech" | "8% of these fear speech users are also admins in the groups where they post fear speech" |
| spans: 62; by route: EXTERNAL 23, paper.cols.txt 32, paper.cols.txt+lig 1, NOT-A-QUOTE 1, pypdf 4, pypdf+dehyph 1 | spans: 68; by route: EXTERNAL 25, paper.cols.txt 36, paper.cols.txt+lig 1, NOT-A-QUOTE 1, pypdf 4, pypdf+dehyph 1 |
| |
| == B. Per-paper figures == | == B. Per-paper figures == |
| OK paper.cols.txt gao2026_doxing | window | OK paper.cols.txt gao2026_doxing | window |
| "from May 30, 2025 to August 25, 2025" | "from May 30, 2025 to August 25, 2025" |
| | OK paper.cols.txt gao2026_doxing | "representative" (the paper's own word for its top 100) |
| | "we selected the top 100 channels by subscriber count as our representative corpus" |
| | OK paper.cols.txt acharya2025_pirates | the sentence the extraction reads as a review decision |
| | "Our research did not directly involve interaction with any human subjects" |
| | OK paper.cols.txt roy2025_darkgram | 64 of 339 removed |
| | "only 64 channels (19%) were removed" |
| | OK paper.cols.txt roy2025_darkgram | 196 new channels found via Telegram and Facebook |
| | "takedown 196 new CACs shared on Telegram and Facebook" |
| | OK paper.cols.txt vu2025_assessing | Discord channels monitored too |
| | "Booters may use Discord; we also monitored these channels" |
| | OK paper.cols.txt kireev2025_characterizing | over 80% removed by moderators |
| | "more than 80% of propaganda messages removed" |
| | OK paper.cols.txt kireev2025_characterizing | below 20% removed |
| | "ranging from below 20%" |
| | OK paper.cols.txt guo2024_moderating | manual collection |
| | "constrained by a manually collected dataset" |
| | OK paper.cols.txt roy2025_darkgram | 11,800 posts to the GPT-4 API |
| | "across all the 11,800 posts identified by the classifier, we utilized the GPT-4 API" |
| OK paper.cols.txt xu2019_anatomy | 300+ channels | OK paper.cols.txt xu2019_anatomy | 300+ channels |
| "we trace the message history of over 300 Telegram channels" | "we trace the message history of over 300 Telegram channels" |
| "Telegram (133,399)" | "Telegram (133,399)" |
| OK paper.cols.txt acharya2025_pirates | 85,402 Telegram posts (denominator of the filter) | OK paper.cols.txt acharya2025_pirates | 85,402 Telegram posts (denominator of the filter) |
| "from Telegram were filtered" | "(76,111/85,402) from Telegram were filtered" |
| OK paper.cols.txt cinus2025_exposing | 15,537 channels | OK paper.cols.txt cinus2025_exposing | 15,537 channels |
| "Telegram 15,537 4,309,880" | "Telegram 15,537 4,309,880" |
| OK paper.cols.txt WEAK yu2024_listen | yu2024: Discord named (weak by design; role checked in the hand codes) | OK paper.cols.txt WEAK yu2024_listen | yu2024: Discord named (weak by design; role checked in the hand codes) |
| "Discord" | "Discord" |
| needles: 74; by route: paper.cols.txt 62, pypdf 12; weak (<20 chars): 4 | needles: 83; by route: paper.cols.txt 71, pypdf 12; weak (<20 chars): 4 |
| |
| === EXTERNAL FIGURES (non-corpus; each re-fetched by external_checks_messaging_channels.sh) === | === EXTERNAL FIGURES (non-corpus; each re-fetched by external_checks_messaging_channels.sh — see its labels) === |
| Telegram: channels_limit_default 500, channels_limit_premium 1000, recommended_channels_limit_default 10 (core.telegram.org/api/config) | Telegram: channels_limit_default 500, channels_limit_premium 1000, recommended_channels_limit_default 10 (core.telegram.org/api/config) |
| Telegram: one api_id per phone number; People Nearby removed 2024-09-06; search clean-up and disclosure change 2024-09-23 | Telegram: one api_id per phone number; People Nearby removed 2024-09-06; search clean-up and disclosure change 2024-09-23 |
| gh pyrogram/pyrogram $TMP/pyrogram.json | gh pyrogram/pyrogram $TMP/pyrogram.json |
| check 'Pyrogram GitHub repo is archived' $TMP/pyrogram.json '"archived": true' | check 'Pyrogram GitHub repo is archived' $TMP/pyrogram.json '"archived": true' |
| | get https://raw.githubusercontent.com/pyrogram/pyrogram/master/README.md $TMP/pyrogram_readme.txt |
| | check 'Pyrogram README: no longer maintained' $TMP/pyrogram_readme.txt 'no longer maintained' |
| get https://pypi.org/pypi/kurigram/json $TMP/kurigram.json | get https://pypi.org/pypi/kurigram/json $TMP/kurigram.json |
| python3 -c "import json;d=json.load(open('$TMP/kurigram.json'));v=d['info']['version'];print(' kurigram PyPI latest',v,d['releases'][v][0]['upload_time'])" | python3 -c "import json;d=json.load(open('$TMP/kurigram.json'));v=d['info']['version'];print(' kurigram PyPI latest',v,d['releases'][v][0]['upload_time'])" |
| check_raw 'Kurigram (Pyrogram fork) is on PyPI' $TMP/kurigram.json '"name": ?"[Kk]urigram"' | check_raw 'Kurigram (Pyrogram fork) is on PyPI' $TMP/kurigram.json '"name": ?"[Kk]urigram"' |
| | check_raw 'Kurigram released in 2026' $TMP/kurigram.json '"upload_time": ?"2026-' |
| | gh kurigram-org/kurigram $TMP/kurigram_gh.json |
| | check 'Kurigram repo not archived' $TMP/kurigram_gh.json '"archived": false' |
| | check_raw 'Kurigram describes itself as a Pyrogram fork' $TMP/kurigram.json '[Ff]ork of [Pp]yrogram|[Pp]yrogram fork' |
| | get https://registry.npmjs.org/telegram $TMP/npm_telegram.json |
| | check_raw 'npm telegram (GramJS) deprecated in favour of teleproto' $TMP/npm_telegram.json '"deprecated":"[^"]*teleproto' |
| gh gram-js/gramjs $TMP/gramjs.json | gh gram-js/gramjs $TMP/gramjs.json |
| check 'GramJS GitHub repo is archived' $TMP/gramjs.json '"archived": true' | check 'GramJS GitHub repo is archived' $TMP/gramjs.json '"archived": true' |
| get https://raw.githubusercontent.com/tdlib/td/master/CMakeLists.txt $TMP/tdlib_cmake.txt | get https://raw.githubusercontent.com/tdlib/td/master/CMakeLists.txt $TMP/tdlib_cmake.txt |
| check 'TDLib master declares a 1.8.x version' $TMP/tdlib_cmake.txt 'project\(TDLib VERSION 1\.8\.\d+' | check 'TDLib master declares a 1.8.x version' $TMP/tdlib_cmake.txt 'project\(TDLib VERSION 1\.8\.\d+' |
| | curl -sL -m 40 "${AUTH[@]}" "https://api.github.com/repos/tdlib/td/tags?per_page=100" -o $TMP/tdlib_tags.json |
| | python3 - $TMP/tdlib_tags.json <<'PY' && echo "OK TDLib: newest tag by version is v1.8.0" || { echo "FAILED TDLib: newest tag by version is v1.8.0"; } |
| | import json, sys, re |
| | tags = [t['name'] for t in json.load(open(sys.argv[1]))] |
| | vs = sorted((tuple(int(x) for x in re.findall(r'\d+', t)), t) for t in tags if re.match(r'v\d', t)) |
| | print(' > tags:', len(tags), 'newest by version:', vs[-1][1]) |
| | sys.exit(0 if vs[-1][1] == 'v1.8.0' else 1) |
| | PY |
| | [ $? -eq 0 ] || FAILS=$((FAILS+1)) |
| gh wwebjs/whatsapp-web.js $TMP/wwebjs.json | gh wwebjs/whatsapp-web.js $TMP/wwebjs.json |
| check 'whatsapp-web.js lives at wwebjs/whatsapp-web.js, not archived' $TMP/wwebjs.json '"full_name": "wwebjs/whatsapp-web.js".*"archived": false' | check 'whatsapp-web.js lives at wwebjs/whatsapp-web.js, not archived' $TMP/wwebjs.json '"full_name": "wwebjs/whatsapp-web.js".*"archived": false' |
| get "https://api.github.com/repos/tgalal/yowsup/commits?per_page=1" $TMP/yowsup_commits.json | get "https://api.github.com/repos/tgalal/yowsup/commits?per_page=1" $TMP/yowsup_commits.json |
| check 'yowsup: last default-branch commit is from 2021' $TMP/yowsup_commits.json '"date": ?"2021-' | check 'yowsup: last default-branch commit is from 2021' $TMP/yowsup_commits.json '"date": ?"2021-' |
| | get https://raw.githubusercontent.com/tgalal/yowsup/master/README.md $TMP/yowsup_readme.txt |
| | check_raw 'yowsup README: requires python <= 3.7' $TMP/yowsup_readme.txt 'python>=2\.7,<=3\.7' |
| get https://pypi.org/pypi/discord.py/json $TMP/dpy.json | get https://pypi.org/pypi/discord.py/json $TMP/dpy.json |
| check 'discord.py PyPI latest 2.x' $TMP/dpy.json '"version": ?"2\.\d+\.\d+"' | check 'discord.py PyPI latest 2.x' $TMP/dpy.json '"version": ?"2\.\d+\.\d+"' |
| | python3 -c "import json;d=json.load(open('$TMP/dpy.json'));v=d['info']['version'];print(' discord.py',v,d['releases'][v][0]['upload_time'])" |
| | check_raw 'discord.py 2.7.1 uploaded March 2026' $TMP/dpy.json '"2\.7\.1": ?\[\{.*?"upload_time": ?"2026-03-' |
| get https://raw.githubusercontent.com/Tyrrrz/DiscordChatExporter/master/Readme.md $TMP/dce.txt | get https://raw.githubusercontent.com/Tyrrrz/DiscordChatExporter/master/Readme.md $TMP/dce.txt |
| [ -s $TMP/dce.txt ] || get https://raw.githubusercontent.com/Tyrrrz/DiscordChatExporter/prime/Readme.md $TMP/dce.txt | [ -s $TMP/dce.txt ] || get https://raw.githubusercontent.com/Tyrrrz/DiscordChatExporter/prime/Readme.md $TMP/dce.txt |
| check 'DiscordChatExporter README: automating user accounts is against Discord TOS' $TMP/dce.txt 'automating user accounts is against Discord TOS and may result in you getting banned' | check 'DiscordChatExporter README: automating user accounts is against Discord TOS' $TMP/dce.txt 'automating user accounts is against Discord TOS and may result in you getting banned' |
| | curl -sL -m 40 "${AUTH[@]}" https://api.github.com/repos/Tyrrrz/DiscordChatExporter/releases/latest -o $TMP/dce_rel.json |
| | check 'DiscordChatExporter latest release 2.48 (August 2026)' $TMP/dce_rel.json '"tag_name": ?"2\.48".*"published_at": ?"2026-08-' |
| |
| echo; echo "== 2. Terms ==" | echo; echo "== 2. Terms ==" |
| pw https://support.discord.com/hc/en-us/articles/360030843331-Enabling-Server-Discovery $TMP/dc_disc.txt | pw https://support.discord.com/hc/en-us/articles/360030843331-Enabling-Server-Discovery $TMP/dc_disc.txt |
| check 'Discord Server Discovery needs 1,000 members' $TMP/dc_disc.txt 'at least 1,000 members' | check 'Discord Server Discovery needs 1,000 members' $TMP/dc_disc.txt 'at least 1,000 members' |
| | |
| | pw https://transparency.meta.com/researchtools/meta-content-library/ $TMP/meta_mcl.html |
| | check_raw 'Meta Content Library covers WhatsApp Channels' $TMP/meta_mcl.html 'provide comprehensive access to the public content archive from Facebook, Instagram and WhatsApp Channels' |
| | get https://core.telegram.org/bots/features $TMP/tg_botfeat.html |
| | check 'Telegram bots: privacy mode on by default unless added as admin' $TMP/tg_botfeat.html 'Privacy mode is enabled by default for all bots, except bots that were added to a group as admins' |
| |
| echo; echo "== 3. Telegram API limits ==" | echo; echo "== 3. Telegram API limits ==" |
| get https://tgstat.com/ $TMP/tgstat.html | get https://tgstat.com/ $TMP/tgstat.html |
| check_raw 'TGStat self-reported catalogue size' $TMP/tgstat.html 'More than [0-9 ]+ channels and groups' | check_raw 'TGStat self-reported catalogue size' $TMP/tgstat.html 'More than [0-9 ]+ channels and groups' |
| | check 'TGStat country tile: Russia ~1.69 million channels' $TMP/tgstat.html 'Russia Channels 1 69\d \d{3}' |
| | get https://web.archive.org/web/20260926223639/https://telemetr.io/en $TMP/telemetr_wb.html |
| | check_raw 'Telemetr.io (Web Archive 2026-09-26): 11M+ channels' $TMP/telemetr_wb.html 'Over 11M\+ channels' |
| | check 'Telemetr.io (Web Archive 2026-09-26): 7M+ channels' $TMP/telemetr_wb.html 'has 7M\+ channels' |
| pw https://disboard.org/ $TMP/disboard.txt | pw https://disboard.org/ $TMP/disboard.txt |
| check 'Disboard is a self-listing directory' $TMP/disboard.txt 'list/find Discord servers' | check 'Disboard is a self-listing directory' $TMP/disboard.txt 'list/find Discord servers' |
| check_raw 'TeraGram arXiv first author Golovin' $TMP/teragram.html 'citation_author" content="Golovin, Anastasia"' | check_raw 'TeraGram arXiv first author Golovin' $TMP/teragram.html 'citation_author" content="Golovin, Anastasia"' |
| check 'TeraGram arXiv: 5.9 billion messages, 712 thousand channels and groups' $TMP/teragram.html 'over 5\.9 billion messages dating from 2015 to 2025, collected from 712 thousand channels and groups' | check 'TeraGram arXiv: 5.9 billion messages, 712 thousand channels and groups' $TMP/teragram.html 'over 5\.9 billion messages dating from 2015 to 2025, collected from 712 thousand channels and groups' |
| | curl -sL -m 30 https://api.crossref.org/works/10.1609/icwsm.v20i1.42783 -o $TMP/tg_cr.json |
| | check_raw 'TeraGram published at ICWSM 2026 (Crossref)' $TMP/tg_cr.json 'TeraGram: A Structured Longitudinal Dataset of the Telegram Messenger.*Web and Social Media' |
| get https://aoir.org/reports/ethics3.pdf $TMP/aoir.pdf | get https://aoir.org/reports/ethics3.pdf $TMP/aoir.pdf |
| python3 - $TMP/aoir.pdf <<'PY' && echo "OK AoIR 3.0: approved 6 October 2019, no messaging-group passage" || { echo "FAILED AoIR 3.0"; exit 1; } | python3 - $TMP/aoir.pdf <<'PY' && echo "OK AoIR 3.0: approved 6 October 2019, no messaging-group passage" || { echo "FAILED AoIR 3.0"; exit 1; } |
| |
| <file text external_checks_messaging_channels-output.txt> | <file text external_checks_messaging_channels-output.txt> |
| run: 2026-09-27T11:57Z | run: 2026-09-27T12:35Z |
| |
| == 1. Client tooling (GitHub / Codeberg / PyPI / npm / Go proxy) == | == 1. Client tooling (GitHub / Codeberg / PyPI / npm / Go proxy) == |
| > scussions": false, "forks_count": 1383, "mirror_url": null, "archived": true, "disabled": false, "open_issues_count": 279, "license": { | > scussions": false, "forks_count": 1383, "mirror_url": null, "archived": true, "disabled": false, "open_issues_count": 279, "license": { |
| OK Pyrogram GitHub repo is archived | OK Pyrogram GitHub repo is archived |
| | GET https://raw.githubusercontent.com/pyrogram/pyrogram/master/README.md -> 2324 bytes |
| | > on • Releases • News ## Pyrogram > [!NOTE] > The project is no longer maintained or supported. Thanks for appreciating it. > Elegant, modern |
| | OK Pyrogram README: no longer maintained |
| GET https://pypi.org/pypi/kurigram/json -> 54193 bytes | GET https://pypi.org/pypi/kurigram/json -> 54193 bytes |
| kurigram PyPI latest 2.2.26 2026-09-12T17:45:06 | kurigram PyPI latest 2.2.26 2026-09-12T17:45:06 |
| > mail":"Danipulok <danipulok@gmail.com>","name":"Kurigram","package_url":"https://pypi.org/project | > mail":"Danipulok <danipulok@gmail.com>","name":"Kurigram","package_url":"https://pypi.org/project |
| OK Kurigram (Pyrogram fork) is on PyPI | OK Kurigram (Pyrogram fork) is on PyPI |
| | > requires_python":">=3.8","size":5457052,"upload_time":"2026-01-28T12:35:37","upload_time_iso_8601":" |
| | OK Kurigram released in 2026 |
| | > iscussions": false, "forks_count": 220, "mirror_url": null, "archived": false, "disabled": false, "open_issues_count": 23, "license": { " |
| | OK Kurigram repo not archived |
| | > s\n\nKurigram is an actively maintained pyrogram fork for Python designed as a drop-in replac |
| | OK Kurigram describes itself as a Pyrogram fork |
| | GET https://registry.npmjs.org/telegram -> 1088391 bytes |
| | > e":"kixxauth","email":"kris@kixx.name"},"deprecated":"This package is archived and no longer maintained. Development continues in teleproto (https://npmjs.com/package/teleproto), a largely compatible, actively maintained fork. See the migration guide at https://docs.teleproto.dev/migrating-from-gramjs","_npmVersion |
| | OK npm telegram (GramJS) deprecated in favour of teleproto |
| > discussions": true, "forks_count": 236, "mirror_url": null, "archived": true, "disabled": false, "open_issues_count": 323, "license": { | > discussions": true, "forks_count": 236, "mirror_url": null, "archived": true, "disabled": false, "open_issues_count": 323, "license": { |
| OK GramJS GitHub repo is archived | OK GramJS GitHub repo is archived |
| > cmake_minimum_required(VERSION 3.10 FATAL_ERROR) project(TDLib VERSION 1.8.67 LANGUAGES CXX C) if (NOT DEFINED CMAKE_MODULE_PATH) set(CMA | > cmake_minimum_required(VERSION 3.10 FATAL_ERROR) project(TDLib VERSION 1.8.67 LANGUAGES CXX C) if (NOT DEFINED CMAKE_MODULE_PATH) set(CMA |
| OK TDLib master declares a 1.8.x version | OK TDLib master declares a 1.8.x version |
| > EwOlJlcG9zaXRvcnkxNzEwNzI5Njc=", "name": "whatsapp-web.js", "full_name": "wwebjs/whatsapp-web.js", "private": false, "owner": { "login": "wwebjs", "id": 87630360, "node_id": "MDEyOk9yZ2FuaXphdGlvbjg3NjMwMzYw", "avatar_url": "https://avatars.githubusercontent.com/u/87630360?v=4", "gravatar_id": "", "url": "https://api.github.com/users/wwebjs", "html_url": "https://github.com/wwebjs", "followers_url": "https://api.github.com/users/wwebjs/followers", "following_url": "https://api.github.com/users/wwebjs/following{/other_user}", "gists_url": "https://api.github.com/users/wwebjs/gists{/gist_id}", "starred_url": "https://api.github.com/users/wwebjs/starred{/owner}{/repo}", "subscriptions_url": "https://api.github.com/users/wwebjs/subscriptions", "organizations_url": "https://api.github.com/users/wwebjs/orgs", "repos_url": "https://api.github.com/users/wwebjs/repos", "events_url": "https://api.github.com/users/wwebjs/events{/privacy}", "received_events_url": "https://api.github.com/users/wwebjs/received_events", "type": "Organization", "user_view_type": "public", "site_admin": false }, "html_url": "https://github.com/wwebjs/whatsapp-web.js", "description": "A WhatsApp client library for NodeJS that connects through the WhatsApp Web browser app", "fork": false, "url": "https://api.github.com/repos/wwebjs/whatsapp-web.js", "forks_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/forks", "keys_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/keys{/key_id}", "collaborators_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/collaborators{/collaborator}", "teams_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/teams", "hooks_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/hooks", "issue_events_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/issues/events{/number}", "events_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/events", "assignees_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/assignees{/user}", "branches_url": "https://api.github.com/repos/wwebjs/whatsapp-web.js/branches{/branch}", "tags_url": "https://api.github.com/ |