| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| design:platforms:messaging_channels [2026/09/27 12:05] – Add the reconciliation with design:platforms' Telegram row (17 shared, 4 here only, 7 there only). Authored by Claude. karel.kubicek.claude | design:platforms:messaging_channels [2026/09/27 12:36] (current) – Apply the generic review: Meta Content Library route for WhatsApp Channels; terms read against today's wording; public-release base rate; Discord compliant paths; enumeration verdict per platform; contested LLM-service practice (DarkGram vs Stayin' Alive) karel.kubicek.claude |
|---|
| **Four things to take away before you design anything.** | **Four things to take away before you design anything.** |
| |
| - **Your population is the groups your seed could see.** A group enters your study because a link to it was posted on another platform, a directory listed it, in-app search returned it, or another group mentioned it. Each of those seeds favours large, public, long-lived, self-advertising groups. All nine papers in these venues whose main dataset is group or channel content say so, and none can estimate how many groups exist. See [[#Finding Groups: the Seed Is the Sample]]. | - **Your population is the groups your seed could see.** A group enters your study because a link to it was posted on another platform, a directory listed it, in-app search returned it, or another group mentioned it. Each of those seeds favours large, public, long-lived, self-advertising groups. Eight of the nine papers in these venues whose main dataset is group or channel content say so, and none can estimate how many groups exist. See [[#Finding Groups: the Seed Is the Sample]]. |
| - **Joining is the data-collection act, and each platform shows you something different.** WhatsApp gives you messages only from your joining date; Telegram and Discord gave {[hoseini2020_demystifying]} the history back to the group's creation. Telegram admins can hide member lists (visible in 24 of 100 joined groups); WhatsApp shows members' phone numbers. Accounts are capped: Telegram documents a default of **500** channels and supergroups per account, and {[hoseini2020_demystifying]} measured WhatsApp's cap at 250–300 groups and Discord's at 100 servers. See [[#Joining: What You Can Then See]]. | - **Joining is the data-collection act, and each platform shows you something different.** WhatsApp gives you messages only from your joining date; Telegram and Discord gave {[hoseini2020_demystifying]} the history back to the group's creation. Telegram admins can hide member lists (visible in 24 of 100 joined groups); WhatsApp shows members' phone numbers. Accounts are capped: Telegram documents a default of **500** channels and supergroups per account, and {[hoseini2020_demystifying]} measured WhatsApp's cap at 250–300 groups and Discord's at 100 servers. See [[#Joining: What You Can Then See]]. |
| - **The terms of all three forbid most of what this literature did, and the tooling has moved.** Telegram's content-licensing terms prohibit access to user-generated content "//for any purpose other than ordinary, legitimate, and intended use of the Telegram platform as its user//"; WhatsApp forbids unofficial clients; Discord bans self-bots. Telethon — the most-named client here — left GitHub for Codeberg, Pyrogram and GramJS are archived, and yowsup has not had a commit since 2021. See [[#Client Tooling, Terms and Bans]]. | - **Read against today's terms, all three forbid most of what this literature did, and the tooling has moved.** Telegram's content-licensing terms prohibit access to user-generated content "//for any purpose other than ordinary, legitimate, and intended use of the Telegram platform as its user//"; WhatsApp forbids unofficial clients; Discord bans self-bots. Telethon — the most-named client here — left GitHub for Codeberg, Pyrogram and GramJS are archived, and yowsup has not had a commit since 2021. The one application route is Meta's Content Library, which covers **WhatsApp Channels** (not groups). See [[#Client Tooling, Terms and Bans]]. |
| - **The members did not join a study.** Of the **26** papers in these venues that collected from inside groups, **10** report an ethics board approval, **7** state that they never posted or interacted, and **9** say nothing about review at all. [[Practices:Ethics]] has no section on this route yet. See [[#Ethics: Recording Groups Whose Members Did Not Join a Study]]. | - **The members did not join a study.** Of the **26** papers in these venues that collected from inside groups, **10** report an ethics board approval, **7** state that they never posted or interacted, and **10** say nothing about review at all. [[Practices:Ethics]] has no section on this route yet. See [[#Ethics: Recording Groups Whose Members Did Not Join a Study]]. |
| </WRAP> | </WRAP> |
| |
| <WRAP tip> | <WRAP tip> |
| **The literature is smaller than the name counts suggest, and almost all of it is Telegram.** The gap pass that proposed this page counted papers naming Telegram, WhatsApp or Discord at least ten times in full text: 40, 41 and 13. Reading them, only **23** of those **76** papers collect from inside a group or channel — precision **30.3%** — and the WhatsApp count is worst: **4 of 41** (9.8%), because a paper that names WhatsApp ten times is usually about its encryption, its users, or its traffic. A wider candidate set and a hand audit of all 147 candidates find **26** papers that use this route in the seven venues — **21** on Telegram, **5** on Discord, **3** on WhatsApp — and **none before 2019**. The audit, with every rejected candidate, is in [[#Use in Publications]]. | **The literature is smaller than the name counts suggest, and four papers in five are Telegram.** The gap pass that proposed this page counted papers naming Telegram, WhatsApp or Discord at least ten times in full text: 40, 41 and 13. Reading them, only **23** of those **76** papers collect from inside a group or channel — precision **30.3%** — and the WhatsApp count is worst: **4 of 41** (9.8%), because a paper that names WhatsApp ten times is usually about its encryption, its users, or its traffic. A wider candidate set and a hand audit of all 147 candidates find **26** papers that use this route in the seven venues — **21** on Telegram, **6** on Discord, **3** on WhatsApp — and **none before 2019**. The audit, with every rejected candidate, is in [[#Use in Publications]]. |
| </WRAP> | </WRAP> |
| |
| ===== What to Read First ===== | ===== What to Read First ===== |
| |
| Five papers, chosen because each shows a different part of the route rather than a different topic. | Five papers, chosen because each shows a different part of the route rather than a different topic. All five are Telegram or WhatsApp; for Discord, start with {[shen2024_anything]} (servers found through Disboard), and for the commoner case where group data is one source among several, {[vu2025_assessing]} (booter channels as one of eight datasets, each cross-checked against the others). |
| |
| * {[hoseini2020_demystifying]} (IMC 2020) — **all three platforms side by side**, and still the only paper here that compares them. Harvested **351,535** group URLs from **2,234,128** tweets over 38 days, then joined **616** WhatsApp, Telegram and Discord groups. Read it for the per-platform differences in what a joined member can see, for the join caps it measured, and for how much personal data a group exposes: "//For WhatsApp, even without an account, we could collect an impressive number of over 34K phone numbers. Moreover, after joining groups, we obtain another 20K phone numbers.//" | * {[hoseini2020_demystifying]} (IMC 2020) — **all three platforms side by side**, and still the only paper here that compares them. Harvested **351,535** group URLs from **2,234,128** tweets over 38 days, then joined **616** WhatsApp, Telegram and Discord groups. Read it for the per-platform differences in what a joined member can see, for the join caps it measured, and for how much personal data a group exposes: "//For WhatsApp, even without an account, we could collect an impressive number of over 34K phone numbers. Moreover, after joining groups, we obtain another 20K phone numbers.//" |
| * {[resende2019_information]} (TheWebConf 2019) — **WhatsApp monitoring with physical phones**, the design the 2018–2021 WhatsApp literature shared. Searched Google, Twitter and Facebook for ''%%chat.whatsapp.com%%'' links, found **3,444**, of which only **1,828** still worked, and joined 141 and 364 groups with cell phones for two Brazilian events. Read it for the attrition between "link found" and "group monitored", and for its honest limitation: the number of groups "//was constrained by the available devices and their resources (memory)//". | * {[resende2019_information]} (TheWebConf 2019) — **WhatsApp monitoring with physical phones**, the design the 2018–2021 WhatsApp literature shared. Searched Google, Twitter and Facebook for ''%%chat.whatsapp.com%%'' links, found **3,444**, of which only **1,828** still worked, and joined 141 and 364 groups with cell phones for two Brazilian events. Read it for the attrition between "link found" and "group monitored", and for its honest limitation: the number of groups "//was constrained by the available devices and their resources (memory)//". |
| * {[marjanov2026_stayin]} (USENIX Security 2026) — **snowball discovery done carefully**, at the largest scale here. **21k** candidate channels from keyword search, cybercriminal forums and recurrent snowballing; roughly half examined; **1,521** joined, 448 of them private; about **14 million** messages over a year. Read it for its joining rules — "//We do not lie or pretend to be an interested buyer to be admitted into groups//" — and for the check every channel study should copy: only **0.8%** of its stolen-data channels appear in the large public Telegram datasets. | * {[marjanov2026_stayin]} (USENIX Security 2026) — **snowball discovery done carefully**, at the largest scale here. **21k** candidate channels from keyword search, cybercriminal forums and recurrent snowballing; roughly half examined; **1,521** joined, 448 of them private; about **14 million** messages over a year. Read it for its joining rules — "//We do not lie or pretend to be an interested buyer to be admitted into groups//" — and for the check every channel study should copy: only **0.8%** of its stolen-data channels appear in the large public Telegram datasets. |
| * {[roy2025_darkgram]} (USENIX Security 2025) — **a directory as the seed**, and the bias it states. Took **4,709** English channels with 10,000 or more followers from Telemetr.io, kept **339**, and polled them every ten minutes through the official API. Read it for the bias sentence — "//this approach might have introduced potential biases by omitting smaller or newly emerging channels//" — and for disclosure as a measurable outcome: reporting newly found channels to Telegram "//led to the removal of all 196 channels, with a median response time of 4 days//". | * {[roy2025_darkgram]} (USENIX Security 2025) — **a directory as the seed**, and the bias it states. Took **4,709** English channels with 10,000 or more followers from Telemetr.io, kept **339**, and polled them every ten minutes through the official API. Read it for the bias sentence — "//this approach might have introduced potential biases by omitting smaller or newly emerging channels//" — and for disclosure as a measurable outcome, which came out two ways: of the 339 channels it had monitored and reported, "//only 64 channels (19%) were removed//"; of 196 //new// channels its classifier later found through links shared on Telegram and Facebook, reporting "//led to the removal of all 196 channels, with a median response time of 4 days//". |
| * {[kireev2025_characterizing]} (USENIX Security 2025) — **a registered client, back-filled history, and an ethics section that answers the questions**. A Telethon client "//officially registered as such on the Telegram website//", Telegram's export-history call for the past (it "//returned all messages from either the past 36 months or up to a limit//"), real-time collection for the present, **17.3M** messages from 13 hand-picked channels — and the plain admission that "//given the small selection, we cannot make any statement about the pervasiveness of this propaganda activity in Telegram//". | * {[kireev2025_characterizing]} (USENIX Security 2025) — **a registered client, back-filled history, and an ethics section that answers the questions**. A Telethon client "//officially registered as such on the Telegram website//", Telegram's export-history call for the past (it "//returned all messages from either the past 36 months or up to a limit//"), real-time collection for the present, **17.3M** messages from 13 hand-picked channels — and the plain admission that "//given the small selection, we cannot make any statement about the pervasiveness of this propaganda activity in Telegram//". |
| |
| ^ What you do with the messenger ^ Example in these venues ^ Where it belongs ^ | ^ What you do with the messenger ^ Example in these venues ^ Where it belongs ^ |
| | **Enter groups, channels or servers and record what is inside** | {[marjanov2026_stayin]}, {[saha2021_short]}, {[shen2024_anything]} | this page | | | **Enter groups, channels or servers and record what is inside** | {[marjanov2026_stayin]}, {[saha2021_short]}, {[shen2024_anything]} | this page | |
| | Harvest invite links or group handles elsewhere, and never enter | Telegram contacts parsed out of tweets {[wang2025_detecting]}; Telegram links in token contracts | this page's discovery section — a link census is not a group census | | | Harvest invite links or group handles elsewhere, and never enter | Telegram contacts parsed out of tweets {[wang2025_detecting]}; Telegram links in token contracts (verdict list on the provenance page) | this page's discovery section — a link census is not a group census | |
| | Enumerate **accounts** through contact discovery | 3,546,479,731 WhatsApp accounts {[gegenhuber2026_there]} | [[#Account Enumeration Is a Different Route]] | | | Enumerate **accounts** through contact discovery | 3,546,479,731 WhatsApp accounts {[gegenhuber2026_there]} | [[#Account Enumeration Is a Different Route]] | |
| | Reuse a dataset someone else collected from groups | Pushshift Telegram and DISCO reused by {[chou2025_bots]} | [[Design:Existing datasets]], plus the coverage warning in [[#The Denominator: Groups Found, Not Groups That Exist]] | | | Reuse a dataset someone else collected from groups | Pushshift Telegram and DISCO reused by {[chou2025_bots]} | [[Design:Existing datasets]], plus the coverage warning in [[#The Denominator: Groups Found, Not Groups That Exist]] | |
| | Recruit participants in groups or servers | 19 user studies that posted calls in Discord servers or WhatsApp groups | [[Design:User studies]] — ask the moderators first, as the careful ones did | | | Recruit participants in groups or servers | 19 user studies that posted calls in messaging groups or servers (Discord, WhatsApp, Telegram, WeChat) — listed with the other verdicts on the provenance page | [[Design:User studies]] — ask the moderators first, as the careful ones did | |
| | Interview people about their groups | Hong Kong protesters on large public versus small private groups {[albrecht2021_collective]} | [[Design:User studies]]; its findings belong in your ethics section | | | Interview people about their groups | Hong Kong protesters on large public versus small private groups {[albrecht2021_collective]} | [[Design:User studies]]; its findings belong in your ethics section | |
| | Study the protocol, the app or its traffic | MTProto cryptanalysis, traffic analysis, push-notification leaks | not a platform measurement | | | Study the protocol, the app or its traffic | MTProto cryptanalysis, traffic analysis, push-notification leaks | not a platform measurement | |
| | **Web search** for invite-link patterns | 3 | {[resende2019_information]}, {[saha2021_short]} | groups whose links were indexed by a search engine | | | **Web search** for invite-link patterns | 3 | {[resende2019_information]}, {[saha2021_short]} | groups whose links were indexed by a search engine | |
| | **The known official channel** of the thing you study | 3 | {[vu2024_easy]}, {[vu2024_getting]} | no sampling at all — the population is one or two channels, and that is fine if you say so | | | **The known official channel** of the thing you study | 3 | {[vu2024_easy]}, {[vu2024_getting]} | no sampling at all — the population is one or two channels, and that is fine if you say so | |
| | | **not stated** | 2 | {[bahramali2020_practical]}, {[yu2024_listen]} | — | |
| |
| Four things about seeds that the papers learned the hard way. | Four things about seeds that the papers learned the hard way. |
| |
| * **A link is not a group.** {[resende2019_information]} found 3,444 WhatsApp links and 1,828 worked; {[weerasinghe2020_people]} collected 38,000 group URLs over three snowball iterations; a classifier kept 4,425 as pod-related, **873** of those were currently active public Telegram groups, and **432** of those were the engagement "pods" it studied. {[marjanov2026_stayin]} examined roughly half of its 21k candidates. Report every stage of that funnel, not the last number. | * **A link is not a group.** {[resende2019_information]} found 3,444 WhatsApp links and 1,828 worked; {[weerasinghe2020_people]} collected 38,000 group URLs over three snowball iterations; a classifier kept 4,425 as pod-related, **873** of those were currently active public Telegram groups, and **432** of those were the engagement "pods" it studied. {[marjanov2026_stayin]} examined roughly half of its 21k candidates. Report every stage of that funnel, not the last number. |
| * **Expiry shapes a link harvest.** {[hoseini2020_demystifying]} found more Discord URLs than WhatsApp or Telegram ones "//presumably owing to Discord group URLs automatically expiring after a day//", so users re-posted fresh links. That was 2020. Discord's own help page now says an invite shows "//a 7 days access link by default//", and never-expiring invites exist((https://support.discord.com/hc/en-us/articles/208866998-Invites-101 — fetched 2026-09-27 with a headless browser; ''%%curl%%'' gets HTTP 403. The date the default changed was not established.)). A 2020 link-volume comparison across platforms does not carry over to 2026. | * **Expiry shapes a link harvest.** {[hoseini2020_demystifying]} found more Discord URLs than WhatsApp or Telegram ones "//presumably owing to Discord group URLs automatically expiring after a day//", so users re-posted fresh links. That was 2020. Discord's own help page now says an invite shows "//a 7 days access link by default//", and in Community servers an invite can be set never to expire((https://support.discord.com/hc/en-us/articles/208866998-Invites-101 — fetched 2026-09-27 with a headless browser; ''%%curl%%'' gets HTTP 403. The date the default changed was not established.)). A 2020 link-volume comparison across platforms does not carry over to 2026. |
| * **Directories have entry rules.** Discord's own Server Discovery lists only servers with "//at least 1,000 members//" that are at least eight weeks old((https://support.discord.com/hc/en-us/articles/360030843331-Enabling-Server-Discovery — fetched 2026-09-27 with a headless browser.)); Disboard orders listings by how recently the owner bumped them; TGStat claims "//More than 2 864 885 channels and groups//" and its own country tiles put about 1.69 million of the channels in Russia((https://tgstat.com/ — fetched 2026-09-27. Telemetr.io, which {[roy2025_darkgram]} and {[vafa2025_learning]} used, was Cloudflare-walled to every tool we tried; its own page, read from a Web Archive capture of 2026-09-26, gives two inconsistent catalogue sizes, "11M+" and "7M+".)). A directory is a sampling frame with a language skew and a popularity floor; name it and give its threshold. | * **Directories have entry rules.** Discord's own Server Discovery lists only servers with "//at least 1,000 members//" that are at least eight weeks old((https://support.discord.com/hc/en-us/articles/360030843331-Enabling-Server-Discovery — fetched 2026-09-27 with a headless browser.)); Disboard orders listings by how recently the owner bumped them; TGStat claims "//More than 2 864 885 channels and groups//" and its own country tiles put about 1.69 million of the channels in Russia((https://tgstat.com/ — fetched 2026-09-27. Telemetr.io, which {[roy2025_darkgram]} and {[vafa2025_learning]} used, was Cloudflare-walled to every tool we tried; its own page, read from a Web Archive capture of 2026-09-26, gives two inconsistent catalogue sizes, "11M+" and "7M+".)). A directory is a sampling frame with a language skew and a popularity floor; name it and give its threshold. |
| * **In-app search was edited under you.** On 23 September 2024 Telegram announced that moderators had cleaned up search: "//All the problematic content we identified in Search is no longer accessible//"((https://t.me/durov/345 — Pavel Durov's channel post of 2024-09-23, fetched 2026-09-27.)). A keyword-search seed for abusive content run before and after that date is sampling two different search engines. Telegram's recommendation call is also capped: by default it returns at most **10** similar channels to a non-Premium account((https://core.telegram.org/api/config — ''%%recommended_channels_limit_default%%'', fetched 2026-09-27.)), which bounds any snowball built on it. | * **In-app search was edited under you.** On 23 September 2024 Telegram announced that moderators had cleaned up search: "//All the problematic content we identified in Search is no longer accessible//"((https://t.me/durov/345 — Pavel Durov's channel post of 2024-09-23, fetched 2026-09-27.)). A keyword-search seed for abusive content run before and after that date is sampling two different search engines. Telegram's recommendation call is also capped: by default it returns at most **10** similar channels to a non-Premium account((https://core.telegram.org/api/config — ''%%recommended_channels_limit_default%%'', fetched 2026-09-27.)), which bounds any snowball built on it. |
| | Message history | **from your joining date only** {[hoseini2020_demystifying]} | back to creation for groups joined by {[hoseini2020_demystifying]}; channels via an export-history call that returned up to 36 months {[kireev2025_characterizing]}; public channels readable without joining (below) | back to channel creation {[hoseini2020_demystifying]} | | | Message history | **from your joining date only** {[hoseini2020_demystifying]} | back to creation for groups joined by {[hoseini2020_demystifying]}; channels via an export-history call that returned up to 36 months {[kireev2025_characterizing]}; public channels readable without joining (below) | back to channel creation {[hoseini2020_demystifying]} | |
| | Member list | yes, with **phone numbers** (store a hash, as {[hoseini2020_demystifying]} did) | admin can hide it: available in **24 of 100** joined groups {[hoseini2020_demystifying]}; a channel's subscriber list is admin-only {[roy2025_darkgram]} | partly: members could be listed for 49% of the users in the joined servers, and at least one linked social-media account was exposed for 30% of monitored users {[hoseini2020_demystifying]} | | | Member list | yes, with **phone numbers** (store a hash, as {[hoseini2020_demystifying]} did) | admin can hide it: available in **24 of 100** joined groups {[hoseini2020_demystifying]}; a channel's subscriber list is admin-only {[roy2025_darkgram]} | partly: members could be listed for 49% of the users in the joined servers, and at least one linked social-media account was exposed for 30% of monitored users {[hoseini2020_demystifying]} | |
| | Account cap | 250–300 groups per account, measured by {[hoseini2020_demystifying]} | **500** channels and supergroups by default, 1,000 with Premium, documented by Telegram((https://core.telegram.org/api/config — ''%%channels_limit_default%%'' and ''%%channels_limit_premium%%'', fetched 2026-09-27. These are client defaults; the live value is served by ''%%help.getAppConfig%%''. Exceeding it returns ''%%CHANNELS_TOO_MUCH%%'' (https://core.telegram.org/method/channels.joinChannel).)) | 100 servers, measured by {[hoseini2020_demystifying]} | | | Account cap | 250–300 groups per account, measured by {[hoseini2020_demystifying]} in 2020 | **500** channels and supergroups by default, 1,000 with Premium, documented by Telegram((https://core.telegram.org/api/config — ''%%channels_limit_default%%'' and ''%%channels_limit_premium%%'', fetched 2026-09-27. These are client defaults; the live value is served by ''%%help.getAppConfig%%''. Exceeding it returns ''%%CHANNELS_TOO_MUCH%%'' (https://core.telegram.org/method/channels.joinChannel).)) | 100 servers, measured by {[hoseini2020_demystifying]} in 2020; not re-checked against Discord's current documentation | |
| | | A bot account instead of a user account | not available | a bot cannot add itself: an admin adds it, and in groups it sees only commands and replies unless it is an admin or its privacy mode is off — "//Privacy mode is enabled by default for all bots, except bots that were added to a group as admins//"((https://core.telegram.org/bots/features — fetched 2026-09-27.)) | a bot must be added by a server admin; it cannot join from an invite link | |
| | Reading without joining | no | **yes for public channels**: "//The contents of public channels can be seen on the Web without a Telegram account//", at ''%%t.me/s/<channel>%%''((https://telegram.org/tour/channels — fetched 2026-09-27. Telegram documents the preview on its tour page only; there is no technical documentation, and the page served 20 posts with a ''%%?before=%%'' pagination link when fetched.)) | no; a bot account needs an admin to add it | | | Reading without joining | no | **yes for public channels**: "//The contents of public channels can be seen on the Web without a Telegram account//", at ''%%t.me/s/<channel>%%''((https://telegram.org/tour/channels — fetched 2026-09-27. Telegram documents the preview on its tour page only; there is no technical documentation, and the page served 20 posts with a ''%%?before=%%'' pagination link when fetched.)) | no; a bot account needs an admin to add it | |
| |
| Across the 26 papers, **11** say they joined as members, **8** read public channels without saying whether they joined, **2** bought the data from a commercial scraper service, **1** received it from a vendor, and **4** do not say how they got in. Only **8** say whether they back-filled history — **7** did, and {[hoseini2020_demystifying]} did on two platforms of three. That silence matters, because a study that joined in March and one that back-filled to the group's creation have different time windows for the same group. | Across the 26 papers, **11** say they joined as members, **6** say they read public channels without saying whether they joined, **2** bought the data from a commercial scraper service, **1** received it from a vendor, and **6** do not say how they got in — two of those describe their channels as open to the public, which is a property of the channel, not a statement of what the researchers did. Only **8** say whether they back-filled history — **7** did, and {[hoseini2020_demystifying]} did on two platforms of three. That silence matters, because a study that joined in March and one that back-filled to the group's creation have different time windows for the same group. |
| |
| Three consequences for a design. | Three consequences for a design. |
| * "//The use of Twitter as the only data source for discovering public groups of the different messaging platforms potentially introduces some bias in our sample//" {[hoseini2020_demystifying]} | * "//The use of Twitter as the only data source for discovering public groups of the different messaging platforms potentially introduces some bias in our sample//" {[hoseini2020_demystifying]} |
| |
| Hand-coded: **all 9** papers whose main dataset is group or channel content acknowledge the seed bias or the unknown population; across all 26, **15** do, **4** study one or two known official channels where the question does not arise, and **7** — all of them using the route as one source among several — say nothing. | Hand-coded: **8 of the 9** papers whose main dataset is group or channel content acknowledge the seed bias or the unknown population. The ninth, {[gao2026_doxing]}, calls its top 100 channels by subscribers "representative" and caveats only a secondary sample of groups. Across all 26, **14** acknowledge it, **4** study one or two known groups or channels where the question does not arise, and **8** say nothing — {[gao2026_doxing]} and seven papers that use the route as one source among several. |
| |
| Three further things a reviewer will ask about. | Three further things a reviewer will ask about. |
| |
| * **Survivorship.** Groups and channels disappear during the study, and the disappearance is not random. {[marjanov2026_stayin]} ends its year with **79%** of its non-gateway stolen-data channels "//inactive/banned//", finds Telegram's September 2024 policy change to be "//the most influential predictor of channel durability//", and warns that it "//might miss short-lived channels due to the retroactive nature of data collection//". {[xu2019_anatomy]} found 43 of its pump-and-dump channels already deleted. Record when each group was last reachable, and run your collection often enough that a banned channel's messages are already on disk. | * **Survivorship.** Groups and channels disappear during the study, and the disappearance is not random. {[marjanov2026_stayin]} ends its year with **79%** of its non-gateway stolen-data channels "//inactive/banned//", finds Telegram's September 2024 policy change to be "//the most influential predictor of channel durability//", and warns that it "//might miss short-lived channels due to the retroactive nature of data collection//". {[xu2019_anatomy]} found 43 of its pump-and-dump channels already deleted. Record when each group was last reachable, and run your collection often enough that a banned channel's messages are already on disk. Messages disappear too: a back-filled history is the set of messages that survived. {[kireev2025_characterizing]} ran real-time collection alongside the export because "//The data collected through "Export chat history" does not contain deleted messages//", and the deletions were its subject — moderators removed from below 20% to over 80% of propaganda messages, depending on the channel. |
| * **The big public datasets do not cover niche communities.** The Pushshift Telegram dataset (27.8K channels, 317M messages, ICWSM 2020)((Baumgartner, Zannettou, Squire and Blackburn, "The Pushshift Telegram Dataset", ICWSM 2020, https://doi.org/10.1609/icwsm.v14i1.7348 — Crossref record checked 2026-09-27. It contains only public channels, as {[chou2025_bots]} notes.)), TGDataset (120,979 channels, KDD 2025)((La Morgia, Mei and Mongardini, "TGDataset: Collecting and Exploring the Largest Telegram Channels Dataset", KDD 2025, https://doi.org/10.1145/3690624.3709397 — Crossref record checked 2026-09-27.)) and TeraGram (5.9 billion messages from 712 thousand channels and groups, 2015–2025)((Golovin et al., "TeraGram: A Structured Longitudinal Dataset of the Telegram Messenger", arXiv 2605.15956, 15 May 2026, marked as accepted to ICWSM 2026 — abstract fetched 2026-09-27.)) are the obvious shortcut. {[marjanov2026_stayin]} checked: only six of its stolen-data channels — **0.8%** — appear in them. A reused dataset is a population the dataset's builder sampled, from their seed. | * **The big public datasets do not cover niche communities.** The Pushshift Telegram dataset (27.8K channels, 317M messages, ICWSM 2020)((Baumgartner, Zannettou, Squire and Blackburn, "The Pushshift Telegram Dataset", ICWSM 2020, https://doi.org/10.1609/icwsm.v14i1.7348 — Crossref record checked 2026-09-27. It contains only public channels, as {[chou2025_bots]} notes.)), TGDataset (120,979 channels, KDD 2025)((La Morgia, Mei and Mongardini, "TGDataset: Collecting and Exploring the Largest Telegram Channels Dataset", KDD 2025, https://doi.org/10.1145/3690624.3709397 — Crossref record checked 2026-09-27.)) and TeraGram (5.9 billion messages from 712 thousand channels and groups, 2015–2025)((Golovin et al., "TeraGram: A Structured Longitudinal Dataset of the Telegram Messenger", ICWSM 2026, https://doi.org/10.1609/icwsm.v20i1.42783 — Crossref record checked 2026-09-27; the figures are from the arXiv abstract, arXiv 2605.15956.)) are the obvious shortcut. {[marjanov2026_stayin]} checked: only six of its stolen-data channels — **0.8%** — appear in them. A reused dataset is a population the dataset's builder sampled, from their seed. |
| * **Found through a link is not the same as public.** A group whose invite link circulates is "public" to anyone who has the link; its members may not think so. {[arunasalam2024_security]} records a refugee support group, "//that can only be joined via invitation//", infiltrated after "//an unintentional leak of the group's "invite link"//". If your seed is a leaked link, your sample includes groups whose members believe they are closed. | * **Found through a link is not the same as public.** A group whose invite link circulates is "public" to anyone who has the link; its members may not think so. {[arunasalam2024_security]} records a refugee participants' WhatsApp group, "//that can only be joined via invitation//", infiltrated after "//an unintentional leak of the group's "invite link"//". If your seed is a leaked link, your sample includes groups whose members believe they are closed. |
| |
| <WRAP todo> | <WRAP todo> |
| ===== Client Tooling, Terms and Bans ===== | ===== Client Tooling, Terms and Bans ===== |
| |
| The papers name their client far less often than a crawl paper names its browser. Hand-coded across the 26: **Telethon** 6, the official Telegram API with no client named 7, **WhatsApp** through physical phones and the Garimella–Tyson collection tool 2, the WhatsApp Web client 1, the Discord API with a user account 1, a **commercial scraper service** (Apify, Telemetrio) 2, a vendor's crawler 1, Selenium 1, **by hand** 4, and **not stated** 3. The extraction schema agrees on all seven papers that name a client library. | The papers name their client far less often than a crawl paper names its browser. Hand-coded across the 26: **Telethon** 6, the official Telegram API with no client named 7, **WhatsApp** through physical phones and the Garimella–Tyson collection tool 2, the WhatsApp Web client 1, the Discord API with a user account 1, a **commercial scraper service** (Apify, Telemetrio) 2, a pump-and-dump aggregator's API (PumpOlymp) 1, a vendor's crawler 1, Selenium 1, **by hand** 4, and **not stated** 3. The extraction schema agrees on all seven papers that name a client library. |
| |
| ==== The clients, as of 2026-09-27 ==== | ==== The clients, as of 2026-09-27 ==== |
| ==== What the terms say ==== | ==== What the terms say ==== |
| |
| None of the three has a research exception, and all three forbid most of what this literature did. That is not new — it is the same position as scraping in [[Design:Platforms]] — but the wording is specific, recent and worth quoting in your ethics section rather than paraphrasing. | None of the three has a research exception in its terms, and read against today's wording all three forbid most of what this literature did — whether the clauses quoted below existed when the 2019–2024 papers collected was not established. That is the same position as scraping in [[Design:Platforms]], but the wording is specific and worth quoting in your ethics section rather than paraphrasing. The one application route is for **WhatsApp Channels**: Meta's Content Library "//provide[s] comprehensive access to the public content archive from Facebook, Instagram and WhatsApp Channels//" to approved researchers((https://transparency.meta.com/researchtools/meta-content-library/ — fetched 2026-09-27 with a headless browser; eligibility and the independent review are on [[Design:Platforms]].)) — WhatsApp groups are not in it, and no paper here has used it. |
| |
| * **Telegram.** The general terms say "//Telegram additionally prohibits data scraping as part of its Content Licensing and AI Scraping Terms, which apply to all users, businesses, and third-party services accessing the platform//", and those terms say "//Access to user-generated content for any purpose other than ordinary, legitimate, and intended use of the Telegram platform as its user is prohibited//"((https://telegram.org/tos and https://telegram.org/tos/content-licensing — fetched 2026-09-27. The content-licensing page carries no date, and when it was introduced was not established.)). The API terms add that you are "//prohibited from using, accessing or aggregating data obtained from the Telegram platform to train, fine-tune or otherwise engage in the development//" of machine-learning models((https://core.telegram.org/api/terms, §1.5 — fetched 2026-09-27.)) — which, read literally, covers training a classifier on messages you collected. Telegram is not designated under the EU Digital Services Act: its own page reports "//significantly fewer than 45 million//" EU recipients((https://telegram.org/tos/eu-dsa — fetched 2026-09-27; Telegram is absent from the Commission's designation list on the same date.)), so the vetted-researcher route of [[Design:Platforms]] does not reach it. | * **Telegram.** The general terms say "//Telegram additionally prohibits data scraping as part of its Content Licensing and AI Scraping Terms, which apply to all users, businesses, and third-party services accessing the platform//", and those terms say "//Access to user-generated content for any purpose other than ordinary, legitimate, and intended use of the Telegram platform as its user is prohibited//"((https://telegram.org/tos and https://telegram.org/tos/content-licensing — fetched 2026-09-27. The content-licensing page carries no date, and when it was introduced was not established.)). The API terms add that you are "//prohibited from using, accessing or aggregating data obtained from the Telegram platform to train, fine-tune or otherwise engage in the development//" of machine-learning models((https://core.telegram.org/api/terms, §1.5 — fetched 2026-09-27.)) — which, read literally, covers training a classifier on messages you collected. Telegram is not designated under the EU Digital Services Act: its own page reports "//significantly fewer than 45 million//" EU recipients((https://telegram.org/tos/eu-dsa — fetched 2026-09-27; Telegram is absent from the Commission's designation list on the same date.)), so the vetted-researcher route of [[Design:Platforms]] does not reach it. |
| * **WhatsApp.** The terms forbid using the service "//through automated or other means//" in unauthorised ways, including to "//collect information of or about our users in any impermissible or unauthorized manner//"((https://www.whatsapp.com/legal/terms-of-service — fetched 2026-09-27 with a headless browser; ''%%curl%%'' gets HTTP 400. The EEA version is worded the same on these clauses.)). One thing has changed in the researcher's favour: the Commission designated WhatsApp a very large online platform on 26 January 2026 **because of Channels**, with "//private messaging service//" explicitly out of scope((https://digital-strategy.ec.europa.eu/en/news/commission-designates-whatsapp-very-large-online-platform-under-digital-services-act — fetched 2026-09-27.)). Public WhatsApp Channels are therefore inside the DSA's data-access regime; WhatsApp groups are not. Whether WhatsApp has handled an Article 40 request was not checked. | * **WhatsApp.** The terms forbid using the service "//through automated or other means//" in unauthorised ways, including to "//collect information of or about our users in any impermissible or unauthorized manner//"((https://www.whatsapp.com/legal/terms-of-service — fetched 2026-09-27 with a headless browser; ''%%curl%%'' gets HTTP 400. The EEA version is worded the same on these clauses.)). One thing has changed in the researcher's favour: the Commission designated WhatsApp a very large online platform on 26 January 2026 **because of Channels**, with "//private messaging service//" explicitly out of scope((https://digital-strategy.ec.europa.eu/en/news/commission-designates-whatsapp-very-large-online-platform-under-digital-services-act — fetched 2026-09-27.)). Public WhatsApp Channels are therefore inside the DSA's data-access regime; WhatsApp groups are not. Whether WhatsApp has handled an Article 40 request was not checked. |
| * **Discord.** The terms forbid "//scraping our services without our written consent//", the community guidelines say "//Do not use self-bots or user-bots//", and the developer policy says "//Do not mine or scrape any data//" and forbids training models on message content((https://discord.com/terms, https://discord.com/guidelines and https://support-dev.discord.com/hc/en-us/articles/8563934450327-Discord-Developer-Policy — fetched 2026-09-27. The developer policy does not mention research at all; the only research language anywhere is the terms' "written consent", and no public procedure for obtaining it was found.)). An automated user account that joins servers from invite links — what the 2020–2024 Discord papers here describe or imply — is the thing all three documents prohibit. | * **Discord.** The terms forbid "//scraping our services without our written consent//", the community guidelines say "//Do not use self-bots or user-bots//", and the developer policy says "//Do not mine or scrape any data//" and forbids training models on message content((https://discord.com/terms, https://discord.com/guidelines and https://support-dev.discord.com/hc/en-us/articles/8563934450327-Discord-Developer-Policy — fetched 2026-09-27. The developer policy does not mention research at all; the only research language anywhere is the terms' "written consent", and no public procedure for obtaining it was found.)). An automated user account that joins servers from invite links — what {[hoseini2020_demystifying]} describes — is the thing all three documents prohibit. A person reading a server as an ordinary member is not what they describe; it does not scale, and it is how {[guo2024_moderating]} appears to have worked. |
| |
| ==== Bans and limits, as the papers report them ==== | ==== Bans and limits, as the papers report them ==== |
| ^ Practice ^ Papers (of 26) ^ Example ^ | ^ Practice ^ Papers (of 26) ^ Example ^ |
| | ethics board **approval** reported | 10 | {[marjanov2026_stayin]} (department ethics committee), {[kireev2025_characterizing]}, {[hoseini2020_demystifying]} | | | ethics board **approval** reported | 10 | {[marjanov2026_stayin]} (department ethics committee), {[kireev2025_characterizing]}, {[hoseini2020_demystifying]} | |
| | **exempt**, or **not required** as not human-subjects research | 1 + 3 | {[saha2021_short]} (exempt), {[roy2025_darkgram]} ("deemed not to require" review) | | | **exempt**, or **not required** as not human-subjects research | 1 + 2 | {[saha2021_short]} (exempt), {[roy2025_darkgram]} ("deemed not to require" review) | |
| | institutional standard invoked, **no board named** | 2 | {[gao2026_doxing]}, {[he2025_unmasking]} | | | institutional standard invoked, **no board named** | 2 | {[gao2026_doxing]}, {[he2025_unmasking]} | |
| | approval reported for a **different part** of the study | 1 | {[yu2024_listen]} — the approval covers its user study, not the Discord collection | | | approval reported for a **different part** of the study | 1 | {[yu2024_listen]} — the approval covers its user study, not the Discord collection | |
| | **nothing stated** | 9 | including both 2019 core papers | | | **nothing stated** | 10 | including both 2019 core papers | |
| | says it **never posted, interacted or contacted** members | 7 | {[kireev2025_characterizing]}: collection accounts "//never interacted with the channels//" | | | says it **never posted, interacted or contacted** members | 7 | {[kireev2025_characterizing]}: collection accounts "//never interacted with the channels//" | |
| |
| The extraction schema, reading the same 26 papers independently, finds a stated review outcome in **17 of 26 (65.4%)**, against **33.8%** of all 5,118 empirical papers in the corpus. Group studies report ethics review about twice as often as the average paper here, and still leave a third silent. | The extraction schema, reading the same 26 papers independently, finds a stated review outcome in **17 of 26 (65.4%)** — one more than the hand codes, because it reads {[acharya2025_pirates]}'s sentence that the research "did not directly involve interaction with any human subjects" as a review decision — against **33.8%** of all 5,118 empirical papers in the corpus. Group studies report ethics review about twice as often as the average paper here, and still leave more than a third silent. |
| |
| What the careful papers converge on, stated so you can adopt or argue with it: | What individual careful papers did, stated so you can adopt or argue with it — most bullets rest on one or two papers, and one is contested: |
| |
| * **Enter only the way any member would.** "//We do not lie or pretend to be an interested buyer to be admitted into groups. We also do not attempt to join any groups that require payments or vouching by an existing member.//" {[marjanov2026_stayin]} None of the 26 says it entered a group by deception. One went further than observing: {[he2025_unmasking]} "//joined several related Telegram groups//" of wallet-drainer operators, "//communicated with operators, acquired wallet drainers//", and did so from anonymised accounts without paying anything. If your design needs interaction, it is a different ethics case from lurking, and your review should see it as one. | * **Enter only the way any member would.** "//We do not lie or pretend to be an interested buyer to be admitted into groups. We also do not attempt to join any groups that require payments or vouching by an existing member.//" {[marjanov2026_stayin]} None of the 26 says it entered a group by deception. One went further than observing: {[he2025_unmasking]} "//joined several related Telegram groups//" of wallet-drainer operators, "//communicated with operators, acquired wallet drainers//", and did so from anonymised accounts without paying anything. If your design needs interaction, it is a different ethics case from lurking, and your review should see it as one. |
| * **Consent is waived, not assumed, and the waiver is argued.** {[marjanov2026_stayin]} cites the British Society of Criminology's guidance that consent may be waived for publicly accessible online communities studied as collective patterns; {[vu2025_assessing]} did not seek consent because "//sending thousands of messages could be regarded as spamming//", analysed collectively, and paraphrased every quote. Neither argument covers a small group whose members expect privacy. | * **Consent is waived, not assumed, and the waiver is argued.** {[marjanov2026_stayin]} cites the British Society of Criminology's guidance that consent may be waived for publicly accessible online communities studied as collective patterns; {[vu2025_assessing]} did not seek consent because "//sending thousands of messages could be regarded as spamming//", analysed collectively, and paraphrased every quote. Neither argument covers a small group whose members expect privacy. |
| * **Minimise at collection, not at publication.** Hash phone numbers ({[hoseini2020_demystifying]}); map names to identifiers and discard them ({[resende2019_information]}); do not download payloads that contain other people's data ({[roy2025_darkgram]}: "//we did not download or read the payload files//"); cap file types and sizes ({[marjanov2026_stayin]}); paraphrase quotes so a member cannot be found by searching ({[vu2025_assessing]}). | * **Minimise at collection, not at publication.** Hash phone numbers ({[hoseini2020_demystifying]}); map names to identifiers and discard them ({[resende2019_information]}); do not download payloads that contain other people's data ({[roy2025_darkgram]}: "//we did not download or read the payload files//"); cap file types and sizes ({[marjanov2026_stayin]}); paraphrase quotes so a member cannot be found by searching ({[vu2025_assessing]}). |
| * **Keep collected messages away from third-party services.** {[marjanov2026_stayin]} classified messages with a local model "//to avoid sending stolen and potentially sensitive data to third-party servers//". Sending group messages to a hosted LLM moves them to a fourth party the members never heard of, and — on Telegram — sits badly with the terms quoted above. | * **Keep collected messages away from third-party services — contested.** {[marjanov2026_stayin]} classified messages with a local model "//to avoid sending stolen and potentially sensitive data to third-party servers//". {[roy2025_darkgram]}, one of the five papers to read first, did the opposite: "//we utilized the GPT-4 API//" to classify 11,800 collected posts. Sending group messages to a hosted model moves them to a party the members never heard of, and on Telegram sits badly with the terms quoted above; this page sides with the local model, but the field has not settled it. |
| * **Declining is a legitimate design.** {[li2025_investigating]} studied people who use drugs and wrote: "//Although PWUD is also active in other online communities such as Telegram groups and self-constructed forums, for ethical reasons, we limited our online data collection to publicly accessible platforms.//" | * **Declining is a legitimate design.** {[li2025_investigating]} — not one of the 26, because it declined — studied people who use drugs and wrote: "//Although PWUD is also active in other online communities such as Telegram groups and self-constructed forums, for ethical reasons, we limited our online data collection to publicly accessible platforms.//" |
| |
| **What members expect depends on the group.** Hong Kong protesters interviewed by {[albrecht2021_collective]} distinguished large public Telegram groups — "//all participants in our study also assumed police monitoring of the public Telegram groups//" — from small groups of people who knew each other. The refugee group of {[arunasalam2024_security]} was invitation-only and experienced an outsider's arrival as an attack. Size, how the link circulates and what the group is for are the variables; "public" is not one bit. | **What members expect depends on the group.** Hong Kong protesters interviewed by {[albrecht2021_collective]} distinguished large public Telegram groups — "//all participants in our study also assumed police monitoring of the public Telegram groups//" — from small groups of people who knew each other. The refugee group of {[arunasalam2024_security]} was invitation-only and experienced an outsider's arrival as an attack. Size, how the link circulates and what the group is for are the variables; "public" is not one bit. |
| A paper counts if it **collects data from inside messaging-platform groups, channels or servers** — Telegram, WhatsApp, Discord, or another messenger's group feature — as a measurement source: it finds groups or channels, joins, subscribes, reads their public preview or has a scraper do so, and records messages, members, media or metadata. It is **core** if that data is its main dataset and **section** if it is one source among several. Harvesting invite links without entering, enumerating accounts, reusing someone else's group dataset, recruiting participants in groups, and studying the protocol or app are recorded but do not count. | A paper counts if it **collects data from inside messaging-platform groups, channels or servers** — Telegram, WhatsApp, Discord, or another messenger's group feature — as a measurement source: it finds groups or channels, joins, subscribes, reads their public preview or has a scraper do so, and records messages, members, media or metadata. It is **core** if that data is its main dataset and **section** if it is one source among several. Harvesting invite links without entering, enumerating accounts, reusing someone else's group dataset, recruiting participants in groups, and studying the protocol or app are recorded but do not count. |
| |
| The candidate set is the union of three probes over full text and the extraction schema — a platform name ten or more times, route vocabulary (invite links, ''%%t.me%%'', ''%%chat.whatsapp.com%%'', ''%%discord.gg%%'', Telethon, TDLib, TGStat, Disboard and similar) with at least three mentions of a platform, or a platform named in a paper's title, population sources, measured phenomena or tools — giving **144** papers, plus **3** added by hand from a recall probe. The 44 that looked like group studies or close neighbours were read in full; the other 103 were decided from the sentence around every platform-name and route-vocabulary hit. Every one of the 147 has a verdict: | The candidate set is the union of three probes over full text and the extraction schema — a platform name ten or more times, route vocabulary (invite links, ''%%t.me%%'', ''%%chat.whatsapp.com%%'', ''%%discord.gg%%'', Telethon, TDLib, TGStat, Disboard and similar) with at least three mentions of the three platforms combined, or a platform named in a paper's title, population sources, measured phenomena or tools — giving **144** papers, plus **3** added by hand from a recall probe. The 44 that looked like group studies or close neighbours were read in full; the other 103 were decided from the sentence around every platform-name and route-vocabulary hit. Every one of the 147 has a verdict: |
| |
| ^ Verdict ^ Papers ^ | ^ Verdict ^ Papers ^ |
| | **79%** of stolen-data channels inactive or banned by the end of the year; only **0.8%** of them present in the public Telegram datasets {[marjanov2026_stayin]} | 1,282 non-gateway stolen-data channels of 1,521 joined, August 2024 – August 2025 | | | **79%** of stolen-data channels inactive or banned by the end of the year; only **0.8%** of them present in the public Telegram datasets {[marjanov2026_stayin]} | 1,282 non-gateway stolen-data channels of 1,521 joined, August 2024 – August 2025 | |
| | **78.37K** propaganda messages (1.8% of the dataset) sent by **6,250** accounts (2.2% of accounts) {[kireev2025_characterizing]} | 17.3M messages from 13 political and news channels | | | **78.37K** propaganda messages (1.8% of the dataset) sent by **6,250** accounts (2.2% of accounts) {[kireev2025_characterizing]} | 17.3M messages from 13 political and news channels | |
| | Reporting newly detected channels "//led to the removal of all 196 channels, with a median response time of 4 days//" {[roy2025_darkgram]} | channels found by its classifier during a three-month live run | | | Of the channels it monitored and reported, "//only 64 channels (19%) were removed//"; of 196 new channels its classifier found later, reporting "//led to the removal of all 196 channels, with a median response time of 4 days//" {[roy2025_darkgram]} | 339 monitored channels; 196 channels found through links shared on Telegram and Facebook during a three-month live run | |
| | Personal identity information of "//over 300,000 unique individuals//" exposed in three months {[gao2026_doxing]} | 411,707 messages from five doxing query groups, May–August 2025 | | | Personal identity information of "//over 300,000 unique individuals//" exposed in three months {[gao2026_doxing]} | 411,707 messages from five doxing query groups, May–August 2025 | |
| | **11,728** URLs shared during the strike and **92,654** during the election campaign {[resende2019_information]} | 141 and 364 joined political groups, Brazil 2018 | | | **11,728** URLs shared during the strike and **92,654** during the election campaign {[resende2019_information]} | 141 and 364 joined political groups, Brazil 2018 | |
| |
| ^ Method ^ Verdict ^ Evidence ^ | ^ Method ^ Verdict ^ Evidence ^ |
| | Harvesting invite links from **Twitter/X search and streaming** | **historical as a route** | {[hoseini2020_demystifying]} (2020) used the free Twitter APIs, which are now metered; see [[Design:Platforms]]. Harvesting links from other surfaces — forums, GitHub, token contracts — is current | | | Harvesting invite links through the **free Twitter/X APIs** | **historical** | {[hoseini2020_demystifying]} (2020) used the free Search and Streaming APIs, which are now metered; see [[Design:Platforms]]. Searching Twitter, forums or GitHub for links as a seed is current — {[he2025_unmasking]} and {[vafa2025_learning]} (2025) | |
| | **Directories** (TGStat, Telemetr.io, Disboard) as the seed | **current** | 7 papers, 6 of them 2024–2026; TGStat and Disboard answered on 2026-09-27, Telemetr.io only behind a Cloudflare challenge | | | **Directories** (TGStat, Telemetr.io, Disboard) as the seed | **current** | 7 papers, 6 of them 2024–2026; TGStat and Disboard answered on 2026-09-27, Telemetr.io only behind a Cloudflare challenge | |
| | **In-app search and snowballing** | **current, with a changed search engine** | in-app search in 8 papers, 7 of them 2024–2026; {[marjanov2026_stayin]}, {[gao2026_doxing]}; Telegram moderated its search in September 2024 and caps recommendations at 10 per query for non-Premium accounts | | | **In-app search and snowballing** | **current, with a changed search engine** | in-app search in 8 papers, 7 of them 2024–2026; {[marjanov2026_stayin]}, {[gao2026_doxing]}; Telegram moderated its search in September 2024 and caps recommendations at 10 per query for non-Premium accounts | |
| | **WhatsApp groups joined with physical phones** and the Garimella–Tyson tool | **historical in these venues** | the design of {[resende2019_information]} and {[saha2021_short]}; no paper here collects WhatsApp group content after 2021. Outside these venues, the proposed replacement is data donation (Garimella and Chauchard, 2025)((Garimella and Chauchard, "WhatsApp Explorer: A data donation tool to facilitate research on WhatsApp", //Mobile Media & Communication// 13(3), 2025, https://doi.org/10.1177/20501579251326809 — Crossref record checked 2026-09-27.)) | | | **WhatsApp groups joined with physical phones** and the Garimella–Tyson tool | **historical in these venues** | the design of {[resende2019_information]} and {[saha2021_short]}; no paper here collects WhatsApp group content after 2021. Outside these venues, the proposed replacement is data donation (Garimella and Chauchard, 2025)((Garimella and Chauchard, "WhatsApp Explorer: A data donation tool to facilitate research on WhatsApp", //Mobile Media & Communication// 13(3), 2025, https://doi.org/10.1177/20501579251326809 — Crossref record checked 2026-09-27.)) | |
| | **WhatsApp Channels** through the DSA | **new and untested** | WhatsApp designated a VLOP because of Channels on 26 January 2026; zero papers here | | | **WhatsApp Channels** through Meta's Content Library or the DSA | **current routes, untested here** | Channels are in the Content Library; WhatsApp designated a VLOP because of Channels on 26 January 2026; zero papers here use either | |
| | **Telethon** over the official API | **current; moved to Codeberg** | 6 papers, all 2024–2026; 1.45.0 released 10 September 2026 | | | **Telethon** over the official API | **current; moved to Codeberg** | 6 papers, all 2024–2026; 1.45.0 released 10 September 2026 | |
| | **Pyrogram**, **GramJS**, **yowsup** | **superseded or dead** | archived, archived, no commit since 2021 | | | **Pyrogram**, **GramJS**, **yowsup** | **superseded or dead** | archived, archived, no commit since 2021 | |
| | **Unofficial WhatsApp clients** (whatsapp-web.js, Baileys, whatsmeow) | **current, and forbidden by WhatsApp's terms** | maintained; used at scale for enumeration {[gegenhuber2026_there]} | | | **Unofficial WhatsApp clients** (whatsapp-web.js, Baileys, whatsmeow) | **current, and forbidden by WhatsApp's terms** | maintained; used at scale for enumeration {[gegenhuber2026_there]} | |
| | **Discord user-account automation** | **forbidden** | self-bots banned by Discord's guidelines; none of the five Discord papers here discusses it | | | **Discord user-account automation** | **forbidden** | self-bots banned by Discord's guidelines; none of the six Discord papers here discusses it | |
| | **Commercial scraper services** for Telegram (Apify, Telemetrio) | **current, opaque** | {[acharya2024_imitation]}, {[acharya2025_pirates]}; you inherit the service's seed and cannot describe it | | | **Commercial scraper services** for Telegram (Apify, Telemetrio) | **current, opaque** | {[acharya2024_imitation]}, {[acharya2025_pirates]}; you inherit the service's seed and cannot describe it | |
| | **Reusing public Telegram datasets** (Pushshift 2020, TGDataset 2025, TeraGram 2026) | **current, low coverage of niche communities** | 0.8% overlap with a stolen-data population {[marjanov2026_stayin]} | | | **Reusing public Telegram datasets** (Pushshift 2020, TGDataset 2025, TeraGram 2026) | **current, low coverage of niche communities** | 0.8% overlap with a stolen-data population {[marjanov2026_stayin]} | |
| | **LLM classification** of collected messages | **new, 2025–2026** | GPT-4 in {[roy2025_darkgram]}; a local Gemma 3 model in {[marjanov2026_stayin]}, chosen so that messages stay on the researchers' machines | | | **LLM classification** of collected messages | **new, 2025–2026** | GPT-4 in {[roy2025_darkgram]}; a local Gemma 3 model in {[marjanov2026_stayin]}, chosen so that messages stay on the researchers' machines | |
| | **Account enumeration** via contact discovery | **a different route; still possible at scale in 2025** | {[gegenhuber2026_there]} found no rate limits; see above | | | **Account enumeration** via contact discovery | **a different route; capped on most platforms** | WhatsApp had no effective rate limit in {[gegenhuber2026_there]}'s 2024–25 window and Meta says it is testing mitigations; Telegram, Signal and KakaoTalk capped or banned in {[hagen2021_numbers]} and {[kang2026_connecting]} | |
| |
| ===== What to Report ===== | ===== What to Report ===== |
| |
| - **The seed**, by name — which directory, which search terms, which platform's links, which snowball rule — and **every stage of the funnel**: candidates found, candidates examined, reachable, joined, collected. | - **The seed**, by name — which directory, which search terms, which platform's links, which snowball rule — and **every stage of the funnel**: candidates found, candidates examined, reachable, joined, collected. |
| - **How you entered**: joined as a member, read a public preview, a bot added by an admin, or a scraper service (and which). **4 of 26** papers here do not say. | - **How you entered**: joined as a member, read a public preview, a bot added by an admin, or a scraper service (and which). **6 of 26** papers here do not say. |
| - **The time window per group**: from joining, or back-filled to what date. **18 of 26** do not say whether they back-filled. | - **The time window per group**: from joining, or back-filled to what date. **18 of 26** do not say whether they back-filled. |
| - **What you recorded**: messages, member lists, media, reactions — and what you deliberately did not. | - **What you recorded**: messages, member lists, media, reactions — and what you deliberately did not. |
| - **Your accounts**: how many, how the numbers were obtained, which client and version, and any bans or flood-waits. | - **Your accounts**: how many, how the numbers were obtained, which client and version, and any bans or flood-waits. |
| - **Attrition**: groups that disappeared, links that expired, admins who removed you — with dates. | - **Attrition**: groups that disappeared, links that expired, admins who removed you — with dates. |
| - **Ethics review and the terms position** you took, and whether you ever posted or interacted. **9 of 26** say nothing about review. | - **Ethics review and the terms position** you took, and whether you ever posted or interacted. **10 of 26** say nothing about review. |
| - **What you can share**: channel identifiers, message identifiers and code even where the messages themselves cannot be released. By the extraction's reading, **5 of 26** release data **on request** and **3** under restriction — against 1.7% and 1.4% of all empirical papers — and 9 publicly; see [[:Artifacts]]. | - **What you can share**: channel identifiers, message identifiers and code even where the messages themselves cannot be released. By the extraction's reading, **9 of 26** (34.6%) release data publicly against **47.7%** of all empirical papers, while **5** release it only on request and **3** under restriction, against 1.7% and 1.4% — group studies share less openly and shift to gated release; see [[:Artifacts]]. |
| |
| ===== Open Questions ===== | ===== Open Questions ===== |
| <WRAP todo> | <WRAP todo> |
| * **How many groups exist, and what fraction a seed reaches?** No paper here estimates it. A capture-recapture design across two independent seeds (a directory and a link harvest, say) would give the first defensible denominator for any platform. | * **How many groups exist, and what fraction a seed reaches?** No paper here estimates it. A capture-recapture design across two independent seeds (a directory and a link harvest, say) would give the first defensible denominator for any platform. |
| * **Is Discord research still possible without breaking Discord's rules?** The only compliant path is a bot an admin invites, which selects servers whose admins agree. Nobody in these venues has published a study built that way and characterised the selection. | * **What does Discord research look like within Discord's rules?** The compliant paths are a person reading as an ordinary member, which does not scale, and a bot an admin invites, which selects servers whose admins agree. Nobody in these venues has published a study built the second way and characterised that selection. |
| * **What does an Article 40 request for WhatsApp Channels look like?** WhatsApp has been a designated platform since January 2026 because of Channels; zero papers use the route. | * **What does a WhatsApp Channels study through the Content Library or an Article 40 request look like?** Both routes exist for Channels; zero papers here use either, and neither reaches WhatsApp groups. |
| * **Does the September 2024 Telegram policy change split every longitudinal Telegram series?** {[marjanov2026_stayin]} found it the strongest predictor of channel survival; any series that crosses it is measuring two regimes. | * **Does the September 2024 Telegram policy change split every longitudinal Telegram series?** {[marjanov2026_stayin]} found it the strongest predictor of channel survival; any series that crosses it is measuring two regimes. |
| * **What should [[Practices:Ethics]] say about joining groups?** Enter only as any member could; argue the consent waiver; minimise at collection; keep messages off third-party services; treat invite-only groups as private. Those five are what the careful papers here converge on, and none is written down on that page yet. | * **What should [[Practices:Ethics]] say about joining groups?** Enter only as any member could; argue the consent waiver; minimise at collection; keep messages off third-party services; treat invite-only groups as private. Those five are what individual careful papers here did — one of them contested by another — and none is written down on that page yet. |
| </WRAP> | </WRAP> |
| |