This is an old revision of the document!
Table of Contents
Research of Specific Large Platforms
Most pages in Design assume your population is a sample of the web: a ranking stands in for “websites”, and you draw from it. This page is about the other case — when the population is one platform. You are not sampling the web any more; you are sampling whatever that platform lets you see, through a channel it controls and can close. The sampling frame, the instrument and the counterparty are the same organisation.
That changes four things, and this page is organised around them: how you get in (and which routes still exist in 2026), what the denominator of what you got actually is, what it costs and what the limits are, and what the rules of engagement are — terms of service, account bans, the login wall, and since December 2025 a European regulator that fines a platform for forbidding research scraping.
This is not a page about any one company's API surface; those change faster than a wiki does. It is a page about the design decisions that a platform study forces on you and a sample-of-the-web study does not. Where a fact about a specific service is load-bearing it is dated and linked to the primary source, and the provenance page records what was verified and what could not be.
The access route is the design decision, and the routes changed under the field. In the publication corpus (seven venues, 2010–2026, 5,859 extracted papers), 897 papers name a large platform as the subject of measurement — 15.7% of the 5,712 that drew any study population. Of those 897, 383 (42.7%) name an existing dataset as a source of their data, and 208 (23.2%) name one and no primary collection of their own at all.
Meanwhile the official researcher routes that replaced the open APIs are almost entirely absent from this literature. Across the 5,855 papers with full text: the Meta Content Library appears once, and that once is a parenthetical noting that CrowdTangle “now replaced by the Meta Content Library” [1Saha Roy, Sayak; Pourabbas Vafa, Elham; Khanmohamaddi, Kobra; Nilizadeh, Shirin (2025): "DarkGram: A Large-Scale Analysis of Cybercriminal Activity Channels on Telegram", in: Proceedings of the USENIX Security Symposium. (Link)]. The TikTok Research API appears in zero papers. The YouTube Researcher Program appears in two. The X/Twitter Academic Research track appears in eight. By contrast Pushshift — a third-party Reddit archive — appears in 35, and Ad Library in 69.
Part of that absence is mechanical, and the page says so rather than letting you infer otherwise: these programmes are 2023-and-later launches read through a corpus whose 2025–2026 slice is under-covered by construction, and the venues that publish most platform work of this kind — ICWSM, CHI, FAccT, the communications journals — are not in the corpus at all. The right reading is “there is no worked example in these seven venues”, not “the programmes do not work”.
Either way the consequence for you is the same. If you plan a platform study in 2026 through an official research programme, assume you are an early adopter: there is no methods section in these venues to copy, and the application lead time is weeks, not days.
What to Read First
Four papers, chosen because each demonstrates a different access route rather than a different topic:
- [2Bouchaud, Paul; Liénard, Jean F. (2024): "Beyond the Guidelines: Assessing Meta's Political Ad Moderation in the EU", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] (IMC 2024) — an ad-transparency archive as the instrument. Uses “a comprehensive dataset provided by Meta under the European Digital Services Act, encompassing all ads targeting EU countries” and finds that only 7.7% of undeclared political ads were moderated as political, while 60.4% of the ads Meta did moderate did not match Meta's own criteria. Read it for how a regulator-mandated archive becomes a measurement instrument, and for the double-sided error analysis that an archive study needs.
- [3Vombatkere, Karan; Mousavi, Sepehr; Zannettou, Savvas; Roesner, Franziska; Gummadi, Krishna P. (2024): "TikTok and the Art of Personalization: Investigating Exploration and Exploitation on Social Media Feeds", in: Proceedings of the ACM Web Conference. (DOI)] (TheWebConf 2024) — sock-puppet accounts on a feed you cannot query. Five bot accounts plus a donated dataset of 347 real users; concludes that TikTok “exploits real users' interests in between 30% and 50% of all recommended videos in the first thousand videos”. Read it for how you measure a recommender with no API at all.
- [4Karnam, Sai Keerthana; Dash, Abhisek; Das, Antariksh; Mousavi, Sepehr; Bechtold, Stefan; Gummadi, Krishna P.; Mukherjee, Animesh; Weber, Ingmar; Zannettou, Savvas (2026): "Setting the Course, but Forgetting to Steer: Analyzing Compliance with GDPR's Right of Access to Data by Instagram, TikTok, and Youtube", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] (IEEE S&P 2026) — the data subject's right of access as the instrument. Sock-puppet accounts on Instagram, TikTok and YouTube, then GDPR Art. 15 requests, then a comparison of the returned data package against the ground truth the authors themselves generated. All three platforms failed to report purpose, recipients and retention period.
- [5Gegenhuber, Gabriel K.; Frenzel, Philipp E.; Günther, Maximilian; Ullrich, Johanna; Judmayer, Aljosha (2026): "Hey there! You are using WhatsApp: Enumerating Three Billion Accounts for Security and Privacy", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] (NDSS 2026) — enumeration at platform scale, and the ethics that go with it. Probes 63 billion candidate phone numbers and discovers 3,546,479,731 WhatsApp accounts across 245 countries, 57% of them with a public profile picture. Read it for what a platform-scale study has to say about disclosure and harm before a programme committee will accept it.
If you only read one thing about the rules, read the Commission's December 2025 decision fining X €120 million: it is the first DSA non-compliance decision, and one of the three grounds is research data access.1)
Is Your Study a Platform Study?
The platform's name appearing in your methods section does not make it one. A name can play any of these roles, and only the first is what this page is about:
| Role the platform plays | Example | Where it belongs |
|---|---|---|
| Subject — you measured the platform, its store, its API, its ad archive, or content you got off it | 3.5 billion WhatsApp accounts enumerated [5Gegenhuber, Gabriel K.; Frenzel, Philipp E.; Günther, Maximilian; Ullrich, Johanna; Judmayer, Aljosha (2026): "Hey there! You are using WhatsApp: Enumerating Three Billion Accounts for Security and Privacy", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] | this page |
| Recruitment channel — it supplied participants or annotators | Amazon Mechanical Turk; “professional networks, Reddit, Twitter, Slack, and Upwork” | User studies |
| Infrastructure — it supplied compute, storage or a vantage point | AWS, EC2, Lambda | Crawling location |
| Hosted model — it supplied a classifier or a speech API | Amazon Rekognition, Transcribe, Connect Voice ID | Website classification, IP classification |
| Ranking list — “Alexa” the retired top-sites list | Alexa top 1M | Website selection |
| Benchmark corpus named after a platform — you trained on it | Amazon Reviews, Facebook Hateful Memes, Twitter15/16 | not a platform measurement at all |
| Operator — you are the platform, measuring your own systems | [6Cohn-Gordon, Katriel; Damaskinos, Georgios; Neto, Divino; Cordova, Joshi; Reitz, Benoît; Strahs, Benjamin; Obenshain, Daniel; Pearce, Paul; Papagiannis, Ioannis (2020): "DELF: Safeguarding deletion correctness in Online Social Networks", in: Proceedings of the USENIX Security Symposium. (Link)], [7Schlinker, Brandon; Cunha, Ítalo S.; Chiu, Yi-Ching; Sundaresan, Srikanth; Katz-Bassett, Ethan (2019): "Internet Performance from Facebook's Edge", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] | this page, with the caveat below |
This distinction is not pedantry; it is the single largest source of error in counting this literature. In our corpus the string “Amazon” appears as a study-population source in 412 papers where it means the Alexa ranking list, 179 where it means Mechanical Turk, and 23 where it means AWS — against 39 where it means an Amazon platform. See Use in Publications.
The operator row deserves its own warning. A paper written with or inside the platform gets data nobody else can get and loses the ability to publish the population: [6Cohn-Gordon, Katriel; Damaskinos, Georgios; Neto, Divino; Cordova, Joshi; Reitz, Benoît; Strahs, Benjamin; Obenshain, Daniel; Pearce, Paul; Papagiannis, Ioannis (2020): "DELF: Safeguarding deletion correctness in Online Social Networks", in: Proceedings of the USENIX Security Symposium. (Link)] is a Facebook-authored study of Facebook's own deletion pipeline. Such papers are valuable and are not replicable by you. Do not cite one as evidence that a measurement is feasible from outside.
The Ways In, and Which Are Current
Every route below is real and in use somewhere. What differs is whether it is current, what it costs, and what population it yields. Dates are as of 2026-08-27 and verified against primary sources; the provenance page lists each URL, quote and fetch date, and names the claims that could not be verified.
The Status column is one word; the dated evidence for each verdict is in Which Methods Are Current at the foot of the page.
| Route | Status | What it gets you | Papers in our corpus that name it |
|---|---|---|---|
| Open public API, free | historical | the population the platform chose to expose | 91 papers name a Twitter/X streaming or search API, peaking 2022 |
| Paid metered API | current | whatever you can afford | X moved to pay-per-usage; see the arithmetic below |
| Platform research programme | current | a curated, filtered archive inside a controlled environment | Meta Content Library 1, TikTok Research API 0, YouTube Researcher Program 2 |
| Regulator route (DSA Art. 40 vetted researcher) | current | in principle non-public data from a designated VLOP | effectively zero papers; see the audit below |
| Transparency / ad archive | current | only what the archive covers — usually ads that actually ran | 69 papers name an Ad Library or ad archive |
| Third-party archive of platform data | historical | someone else's crawl, with their gaps — the dumps circulate, the live service does not | 35 papers name Pushshift |
| Reuse of a published platform dataset | current | a frozen snapshot you did not design | 383 of 897 name an existing dataset; 208 name one and no primary collection |
| Logged-out scraping | current, contested | the public surface, as a bot sees it | 157 papers name a scraping or automation tool |
| Accounts you created (sock puppets) | current | the logged-in surface, and anything personalised | 14 papers use the term “sock puppet”; 22 are labelled `account-registration` |
| Data donation | current | real users' own data, with consent | 25 papers mention data donation, 11 of them in 2025 |
| Right of access / DSAR | current | what the platform says it holds about a subject | 52 papers, rising from 2022 |
| On-device or client-side extraction | niche | the model or logic actually shipped to users | [8West, Jack; Thiemt, Lea; Ahmed, Shimaa; Bartig, Maggie; Fawaz, Kassem; Banerjee, Suman (2024): "A Picture is Worth 500 Labels: A Case Study of Demographic Disparities in Local Machine Learning Models for Instagram and TikTok", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] extracted the on-device ML models from the Instagram and TikTok apps |
| Operator collaboration | not replicable | everything, and no reproducibility | [6Cohn-Gordon, Katriel; Damaskinos, Georgios; Neto, Divino; Cordova, Joshi; Reitz, Benoît; Strahs, Benjamin; Obenshain, Daniel; Pearce, Paul; Papagiannis, Ioannis (2020): "DELF: Safeguarding deletion correctness in Online Social Networks", in: Proceedings of the USENIX Security Symposium. (Link)], [7Schlinker, Brandon; Cunha, Ítalo S.; Chiu, Yi-Ching; Sundaresan, Srikanth; Katz-Bassett, Ethan (2019): "Internet Performance from Facebook's Edge", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] |
The metered API changes the study design, not just the budget
X's API is now pay-per-usage with no subscription: $0.005 per Post read, $0.010 per User read, and “Pay-per-usage plans are capped at 3 million Post reads per monthly billing cycle” with Enterprise above that.2) One million posts is therefore $5,000 and the monthly ceiling on the self-serve route is 3 million posts. A study that used to be “collect the 1% stream for six months” is now a budget line and a cap.
This is not hypothetical, and the field has already adapted. [9Nguyen, Hoang Dai; Dhungana, Sumit; Itha, Madhulika; Vadrevu, Phani (2025): ""Please don't send that bot anything": A Mixed-methods Study of Personal Impersonation Attacks Targeting Digital Payments on Social Media", in: Proceedings of the USENIX Security Symposium. (Link)] (USENIX Security 2025) built its collection pipeline around the read cap: it polls a counts endpoint, which “does not add to our monthly tweet read limit”, and only fetches actual posts when the count is non-zero — because “the total number of tweets that can be downloaded is very small for the Basic tier”. That is still a sound pattern, but the tier it names is no longer what a new project gets: X's own changelog records that pay-per-usage “officially launched” on 6 February 2026 as the self-serve model, with the older Basic and Pro plans “remain[ing] available” only to existing subscribers, and then a further price change on 20 April 2026 introducing “Owned Reads” at $0.001 per resource for a developer's own account data.3) The pricing model of the single most-measured platform in this corpus changed twice inside 2026. Date it, and re-check it before you submit.
The other side of the same story: [10Galeazzi, Alessandro; Paudel, Pujan; Conti, Mauro; Cristofaro, Emiliano De; Stringhini, Gianluca (2026): "Revealing The Secret Power: How Algorithms Can Influence Content Visibility on Twitter/X", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] (NDSS 2026) states that academic access to the X API has “been restricted since June 2023” and therefore works from two previously published tweet datasets rather than collecting anything. Both papers are correct methodology for their moment. Neither is a template you can lift.
The research programmes exist; nobody in these venues has published through them
- Meta Content Library and Content Library API — covers public content from Facebook, Instagram, WhatsApp Channels and (in the web tool) Threads. Applications are reviewed not by Meta but independently by the Secure Data Access Center (CASD) in France. Eligibility is an accredited, degree-granting, not-for-profit academic institution, or a not-for-profit research institution.4)
- TikTok Research Tools — accounts, content and Shops data; open to “Academic institutions in the US, EEA, UK or Switzerland”, to a “Not-for-profit and/or independent research institution, organization, association, or body in the EU” — immediately followed on the same page by “We are currently beta testing this service with select researchers in the US, UK, Switzerland, Norway, Iceland and Liechtenstein”, so read the non-academic pathway as a beta rather than as open — and (for youth-safety research only) to Brazil. Stated turnaround “within 4 weeks”. The separate Commercial Content Library covers advertising and, as of this fetch, “we are ONLY including data from EU countries”.5)
- YouTube Researcher Program — “scaled, expanded access to global video metadata across the entire public YouTube corpus via our Data API”, for students, research staff and faculty at accredited institutions. Separate from, and not satisfied by, the ordinary Data API quota-increase form.6)
Every one of these is a controlled-access environment with an eligibility test, an application, and a review. Budget the lead time in your project plan, and note that eligibility is institutional: a researcher at a commercial lab is excluded from all three.
The regulator route: new, and untested in this literature
For services the European Commission has designated as very large online platforms (VLOPs), Article 40 of the Digital Services Act creates a vetted-researcher data-access right. The implementing instrument is Commission Delegated Regulation (EU) 2025/2050 of 1 July 2025, which lays down “the technical conditions and procedures under which providers of very large online platforms and of very large online search engines are to share data with vetted researchers” and enters into force on the twentieth day after its publication in the Official Journal — which was 9 October 2025, so it has been in force since 29 October 2025.7) Vetting is done by a national Digital Services Coordinator, and applications go through the Commission's DSA Data Access Portal, which states “You can send applications as of 29 October 2025” and carries a public Research projects register of “ongoing research projects conducted by vetted researchers who have access to data under Article 40”.8) The current designation list is maintained by the Commission and moves: it was last updated 24 July 2026 and now includes WhatsApp Ireland alongside Facebook, Instagram, TikTok, YouTube, X and the Amazon Store.9)
Do not read “the DSA gives researchers access” as “papers now use it.” A probe for /Article 40|vetted researcher/ over the 5,855 full texts returns 19 papers. We hand-checked all 19: exactly one is on point — [11McCrosky, Jesse; Malla, Ranadheer; Tanskanen, Aapo; Camargo, Chico Q. (2026): "Does This Button Work? Investigating YouTube's Ineffective User Controls", in: Proceedings of the ACM Web Conference. (DOI)], discussing proposed vetted-researcher mandates. The rest are ACM reference numbers (“Article 40, 10 pages”), GDPR Article 40 codes of conduct, or artefacts released “to vetted researchers” on request. The narrower probe /Digital Services Act/ finds 25 papers (2023:1, 2024:7, 2025:6, 2026:11) — and in nearly all of them the DSA is legal background, not the door they came through. The exception is [2Bouchaud, Paul; Liénard, Jean F. (2024): "Beyond the Guidelines: Assessing Meta's Political Ad Moderation in the EU", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], which used a dataset Meta supplied under the DSA.
Data donation and right-of-access: the two routes that are actually growing
These are the two routes where the corpus shows a real upward trend, and both work by going through the user rather than the platform.
| Route | Papers (of 5,855 with full text) | Per year |
|---|---|---|
| mentions data donation | 25 | 2019:1, 2023:2, 2024:5, 2025:11, 2026:6 |
| mentions right of access / subject access request / DSAR | 52 | 2016:2, 2020:3, 2021:3, 2022:10, 2023:6, 2024:11, 2025:8, 2026:9 |
Two 2026 papers show the shape. [11McCrosky, Jesse; Malla, Ranadheer; Tanskanen, Aapo; Camargo, Chico Q. (2026): "Does This Button Work? Investigating YouTube's Ineffective User Controls", in: Proceedings of the ACM Web Conference. (DOI)] evaluates YouTube's “Don't recommend channel” and “Not interested” controls using 22,722 donating participants from Mozilla's YouTube Regrets project, and finds a control-group Bad Recommendation rate of “about 2.3% — meaning that roughly one in every 43 recommendations was a Bad Recommendation”, with no control eliminating them after four weeks. [12Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)] analyses 48,511 ad explanations across Facebook, Instagram, YouTube and X, collected through the Who Targets Me donation extension, and finds that “98.9% of all [YouTube] explanation texts cite only the main targeting form” without the attributes actually used.
Donation buys you the logged-in, personalised surface without creating a single fake account. It costs you control of the population: your denominator is now whoever installed the extension, which is a convenience sample with an activist skew. Say so, and see Biases.
The Denominator Problem
This is the part a sample-of-the-web study does not have, and the part reviewers now ask about. Every platform access route hands you a filtered population, and the filter is usually documented. Read it, and put it in your paper as the denominator.
- A research archive is not the platform. The Meta Content Library includes Facebook posts to Pages, groups and events “as well as posts that appear on public profiles that are either verified or that have 100 or more followers”; Instagram business and creator accounts plus personal accounts “verified or that have 100 or more followers”; Threads public profiles with 100 or more followers.10) A prevalence computed there is a prevalence among accounts above a follower threshold, not among users.
- An ad archive contains ads that ran. It cannot tell you what was rejected, and it can be wrong in both directions about what is political — which is exactly [2Bouchaud, Paul; Liénard, Jean F. (2024): "Beyond the Guidelines: Assessing Meta's Political Ad Moderation in the EU", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]'s finding, on both sides at once.
- A sampled stream is a sample. [13Kaleli, Beliz; Kondracki, Brian; Egele, Manuel; Nikiforakis, Nick; Stringhini, Gianluca (2021): "To Err.Is Human: Characterizing the Threat of Unintended URLs in Social Media", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] is explicit that its dataset came from “the 1% streaming API that Twitter provides to vetted researchers” and that consequently “all the numbers that we presented in this paper are lower bounds”. Say which stream and which sampling rate, and say which direction the bias runs.
- We could not establish what Reddit currently offers a researcher. Every Reddit-owned domain refused our fetches, so this page makes no claim about Reddit's API pricing, rate limits or researcher terms — only about what the corpus shows (63 papers measuring Reddit, 14 naming a Reddit API or PRAW, 35 naming Pushshift). Check it yourself before planning around it, and see platforms for exactly what failed.
- A third-party archive has the archiver's gaps, not yours. Pushshift is still the most-named Reddit source in this corpus (35 papers, e.g. [14Zannettou, Savvas; Caulfield, Tristan; Blackburn, Jeremy; Cristofaro, Emiliano De; Sirivianos, Michael; Stringhini, Gianluca; Suarez-Tangil, Guillermo (2018): "On the Origins of Memes by Means of Fringe Web Communities", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]), and its coverage window and completeness are properties of Pushshift, not of Reddit. It is also no longer a route you can open: Reddit's own moderator help page says Pushshift API access “will be reinstated for verified Reddit moderators”, that each one “need[s] explicit approval from Reddit”, and that “the use of Pushshift will be limited to moderation use cases only”.11)
- A reused dataset freezes a platform that has since changed. 383 of the 897 platform-subject papers name an existing dataset (
temporal.modeis multi-valued, so 175 of those also collected something themselves; 208 collected nothing). If that is you, date the snapshot and say what changed on the platform since — the four false positives in our Twitter/X audit were all papers using Twitter-derived benchmark corpora with no relationship to the platform at collection time. - A cross-platform study has one denominator per platform, and they are not comparable. [15Acharya, Bhupendra; Lazzaro, Dario; López-Morales, Efrén; Oest, Adam; Saad, Muhammad; Cinà, Antonio Emanuele; Schönherr, Lea; Holz, Thorsten (2024): "The Imitation Game: Exploring Brand Impersonation Attacks on Social Media Platforms", in: Proceedings of the USENIX Security Symposium. (Link)] looked for brand-impersonation accounts targeting the top 10K Tranco brands on X, Instagram, Telegram and YouTube at once; [16Beluri, Mario; Acharya, Bhupendra; Khodayari, Soheil; Stivala, Giada; Pellegrino, Giancarlo; Holz, Thorsten (2025): "Exploration of the Dynamics of Buy and Sale of Social Media Accounts", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] tracked accounts advertised for sale across five platforms and found blocking efficacy ranging from 5.02% (YouTube) up to 48% — its own summary is that “TikTok and Instagram demonstrated the highest detection efficacy at 48%, whereas YouTube and Facebook showed the lowest efficacy at just 5%”. A cross-platform rate is only meaningful if you say what you could see on each platform, because the routes in differ and so does the visible surface. A gap between two platforms may be a gap between two access routes.
- The logged-out surface is not the platform either. Of the 260 platform-subject papers that report a crawl configuration, 145 (55.8%) are labelled as using no authentication at all, 22 as registering an account, 4 as logging in manually, and none as using automated login or SSO. Whatever those studies measured, it is mostly what an anonymous visitor sees. These are schema labels, not audited ones, and Registration is explicit about why that matters for this exact field:
crawlConfigcarries one evidence quote for the whole object, so theauthenticationlabel cannot be checked against its quote, and a hand-audit of another field of that same object,consentAction, found 19.4% false positives (consent). We did not audit these 145. Treat the shape as real and the precise share as unverified.
Rate Limits, Quotas, Bans and the Login Wall
Platform papers talk about limits and rules markedly more than the corpus average, and increasingly so:
| Full-text probe | Platform-subject (of 897) | All papers (of 5,855) | Platform-subject in 2025 |
|---|---|---|---|
| terms of service / acceptable use | 134 (14.9%) | 455 (7.8%) | 24 of 108 (22.2%) |
| rate limit | 107 (11.9%) | 563 (9.6%) | 15 of 108 (13.9%) |
| anti-bot / bot detection / a named bot-management vendor | 94 (10.5%) | 666 (11.4%) | 19 of 108 (17.6%) |
| CAPTCHA | 62 (6.9%) | 313 (5.3%) | 7 of 108 (6.5%) |
| residential proxy / proxy pool / a named proxy vendor | 38 (4.2%) | 275 (4.7%) | 5 of 108 (4.6%) |
| login wall / logged-in session / authenticated crawl | 30 (3.3%) | 123 (2.1%) | 2 of 108 (1.9%) |
| sock puppet | 14 (1.6%) | 26 (0.4%) | 3 of 108 (2.8%) |
Read the ToS row as the headline: at 14.9%, a platform paper is roughly twice as likely as the average paper in these seven venues (7.8%) to discuss terms of service, and in 2025 more than one in five did. Against a narrower and fairer baseline — papers that ran a crawl, where Registration reports 11.5% — the ratio is about 1.3×. Either way the direction is the same, and it is the direction the reviewing is moving.
Four practical consequences:
- Design for the limit, not around it. [9Nguyen, Hoang Dai; Dhungana, Sumit; Itha, Madhulika; Vadrevu, Phani (2025): ""Please don't send that bot anything": A Mixed-methods Study of Personal Impersonation Attacks Targeting Digital Payments on Social Media", in: Proceedings of the USENIX Security Symposium. (Link)]'s counts-endpoint trigger is the pattern: find the cheap query that tells you whether the expensive query is worth making. Report the limit you worked under, because it bounds your recall.
- The limit can be the finding. [17Xue, Diwen; Ramesh, Reethika; S, Valdik S.; Evdokimov, Leonid; Viktorov, Andrey; Jain, Arham; Wustrow, Eric; Basso, Simone; Ensafi, Roya (2021): "Throttling Twitter: an emerging censorship technique in Russia", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] measured Russian ISPs throttling Twitter and showed the mechanism was SNI-based: “throttling is triggered upon observing Twitter-related domains (*.twimg.com, twitter.com, t.co) in the SNI”. If you are being throttled and can characterise by what, that is a result.
- Expect the platform to fight your crawler, and say whether it did. Only 14 of 260 platform-subject papers with a crawl configuration (5.4%) say anything about robots.txt — barely above the 4.7% baseline for all papers with a crawl configuration. If the platform served you a bot challenge, that is a measurement result about the platform, not an embarrassment.
- Sock puppets are the highest-risk route. They get you the personalised surface — [3Vombatkere, Karan; Mousavi, Sepehr; Zannettou, Savvas; Roesner, Franziska; Gummadi, Krishna P. (2024): "TikTok and the Art of Personalization: Investigating Exploration and Exploitation on Social Media Feeds", in: Proceedings of the ACM Web Conference. (DOI)] and [4Karnam, Sai Keerthana; Dash, Abhisek; Das, Antariksh; Mousavi, Sepehr; Bechtold, Stefan; Gummadi, Krishna P.; Mukherjee, Animesh; Weber, Ingmar; Zannettou, Savvas (2026): "Setting the Course, but Forgetting to Steer: Analyzing Compliance with GDPR's Right of Access to Data by Instagram, TikTok, and Youtube", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] both need them — and they are the fact pattern most likely to breach terms of service and, in the US, to have been litigated. Get ethics review first (see Ethics), keep the accounts minimal, and write down what you did: only 32.5% of the 841 empirical platform-subject papers state an ethics-review outcome — essentially at the 33.8% corpus baseline (no test was run, and the platform papers are inside that baseline). That is not good enough for a study whose method is creating fake accounts on someone else's service.
Terms of service are no longer only the platform's weapon
The traditional advice was “the ToS forbids automated access, so scraping is a risk you take.” As of December 2025 that is only half the picture in the EU. The Commission's first DSA non-compliance decision fined X €120 million, and one of the three grounds was researcher data access: “X's terms of service prohibit eligible researchers from independently accessing its public data, including through scraping”, and “X's processes for researchers' access to public data impose unnecessary barriers”.12) A designated VLOP's blanket anti-research ToS clause is now itself an enforcement target.
This does not mean scraping is lawful because a regulator dislikes the clause. It means the legal position is moving and is different in the EU and the US, and it is not settled by anything on this page. The provenance page records the US case law we looked at, which of it we could verify against a primary source and which we could not, and why we did not put unverified case outcomes here. Talk to your institution.
Use in Publications
All figures below come from the corpus of seven venues (CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P), 2010–2026, 5,859 extracted papers. The script is scripts/platforms_report.mjs; its unedited output, and every query, is on platforms.
The inclusion rule, and its measured precision
A paper counts as measuring platform P if P is matched, in a subject role, in at least one of: the title, a population source list, a measured phenomenon, or the name of a tool it used or produced. Roles that do not count: recruitment channel, infrastructure, hosted model, ranking list, benchmark corpus, operator-internal, token collision. For a paper with human participants, a source-list match alone does not count — that is a recruitment channel.
This is a candidate set with regex recall, not a curated set, so we hand-audited samples of it against the papers' own full text and publish the precision:
| Family | Candidates | Audited | Genuine | Precision | What the false positives were |
|---|---|---|---|---|---|
| TikTok | 7 | 7 (all) | 4 — [8West, Jack; Thiemt, Lea; Ahmed, Shimaa; Bartig, Maggie; Fawaz, Kassem; Banerjee, Suman (2024): "A Picture is Worth 500 Labels: A Case Study of Demographic Disparities in Local Machine Learning Models for Instagram and TikTok", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)], [3Vombatkere, Karan; Mousavi, Sepehr; Zannettou, Savvas; Roesner, Franziska; Gummadi, Krishna P. (2024): "TikTok and the Art of Personalization: Investigating Exploration and Exploitation on Social Media Feeds", in: Proceedings of the ACM Web Conference. (DOI)], [18Simko, Lucy; Hutchinson, Adryana; Isaac, Alvin; Fries, Evan; Sherr, Micah; Aviv, Adam J. (2024): ""Modern problems require modern solutions": Community-Developed Techniques for Online Exam Proctoring Evasion", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], [4Karnam, Sai Keerthana; Dash, Abhisek; Das, Antariksh; Mousavi, Sepehr; Bechtold, Stefan; Gummadi, Krishna P.; Mukherjee, Animesh; Weber, Ingmar; Zannettou, Savvas (2026): "Setting the Course, but Forgetting to Steer: Analyzing Compliance with GDPR's Right of Access to Data by Instagram, TikTok, and Youtube", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] | 57% | two interview studies where TikTok is the topic, one tool-list collision |
| Twitter/X | 142 | 15 (every 10th) | 11 | 73% | all four were reused Twitter-derived benchmark corpora (Twitter15/16, CrisisMMD), and all four were 2022 or later |
| Meta | 151 | 16 (every 10th) | 11 | 69% | Facebook Hateful Memes benchmark, Facebook-as-an-identity-provider example, related-work mention |
| Amazon | 72 | 11 (every 7th) | 4 | 36% | Amazon Reviews benchmarks, devices bought on Amazon, a hosted speech service |
Treat the ranking below as a ranking. Do not quote a family's count as a precise number of papers without the precision above attached to it — and note that these samples are small: at n = 15 the 95% interval around 73% is roughly ±20 points. They are coarse corrections, not measurements.
Which platforms the field measures
Of the 5,712 papers that drew a study population, 897 (15.7%) name a large platform as the subject of measurement. A paper can name several, so the rows below do not sum to 897 — [15Acharya, Bhupendra; Lazzaro, Dario; López-Morales, Efrén; Oest, Adam; Saad, Muhammad; Cinà, Antonio Emanuele; Schönherr, Lea; Holz, Thorsten (2024): "The Imitation Game: Exploring Brand Impersonation Attacks on Social Media Platforms", in: Proceedings of the USENIX Security Symposium. (Link)] alone counts for four:
| Rank | Platform family | Papers | Share of 5,712 |
|---|---|---|---|
| 1 | Google (Play / Search / Ads) | 323 | 5.7% |
| 2 | Meta (Facebook / Instagram / WhatsApp) | 151 | 2.6% |
| 3 | Twitter / X | 142 | 2.5% |
| 4 | Amazon | 72 | 1.3% |
| 5 | YouTube | 63 | 1.1% |
| 6 | 63 | 1.1% | |
| 7 | Apple App Store | 56 | 1.0% |
| 8 | WeChat / Weibo / QQ | 50 | 0.9% |
| 9 | Telegram | 24 | 0.4% |
| 10 | Yelp | 17 | 0.3% |
| 11 | 16 | 0.3% | |
| 12 | Flickr | 11 | 0.2% |
| 13= | Steam | 10 | 0.2% |
| 13= | Mastodon | 10 | 0.2% |
| 15= | Discord | 8 | 0.1% |
| 15= | Netflix | 8 | 0.1% |
| 17= | TikTok | 7 | 0.1% |
| 17= | Twitch | 7 | 0.1% |
| 19 | eBay | 6 | 0.1% |
Rows 1 and 7 — Google and the Apple App Store — are largely app-store work, which is closer to Mobile and app measurement than to this page. Note also what this page does not cover: search-engine result auditing and the Google ads ecosystem are a real body of work inside these venues and have no page on this wiki yet; they are out of scope here and named in Open Questions. Note the last rows: TikTok is 7 papers, of which 4 survived a full audit. There is no body of TikTok measurement in these seven venues to systematise.
Amazon means five different things, and mostly not the shop
This is the fold that decides the numbers, so it is published in full. Counting the roles that a platform name plays in a paper's stated population source:
| What “Amazon”/“Alexa” means in a population source list | Papers |
|---|---|
| the Alexa top-sites ranking list (retired 2022; see Website selection) | 412 |
| Amazon Mechanical Turk as a recruitment platform (see User studies) | 179 |
| AWS / EC2 / S3 / Lambda as infrastructure (see Crawling location) | 23 |
| a hosted ML service (Rekognition, Transcribe, Polly, Connect Voice ID) | 3 |
| an actual Amazon platform — the storefront, the Alexa skill store, Fire TV | 39 |
A substring match on “Amazon” therefore over-counts Amazon platform measurement by roughly sixteen to one. An earlier iteration of our own script did exactly this and reported 176 Amazon papers.
Two other pages count these same names and get different numbers. Both are right; they answer different questions. Website selection reports 463 papers for Alexa, folded over the 1,143 that drew a web-unit population. The 412 above is the subset our role rule assigns to the ranking list, over all 5,492 papers with a stated population source. Reconciling them: 486 papers name “Alexa” in a stated population source at all, 464 of those also have a web-unit population tuple — which is the figure Website selection is reporting — and 412 is what survives after Alexa-as-skill-store strings are diverted to the subject role. Likewise our 179 Mechanical Turk papers is a count of population.sourceList strings; User studies publishes three different bounds on the same thing — 139 by the schema enum, 92 by quote, 279 by full text — and warns against reading any of them as usage. Neither fold has been reconciled with the other, and this page does not claim its number is the better one.
The Twitter/X curve, and what it does and does not show
Three-year buckets, as a share of that bucket's whole corpus, so the growth of the corpus itself does not create a trend:
| Bucket | Corpus | Twitter/X as subject | Meta as subject | TikTok as subject |
|---|---|---|---|---|
| 2010–2012 | 386 | 17 (4.4%) | 14 (3.6%) | 0 |
| 2013–2015 | 481 | 15 (3.1%) | 20 (4.2%) | 0 |
| 2016–2018 | 667 | 13 (1.9%) | 17 (2.5%) | 0 |
| 2019–2021 | 1,185 | 34 (2.9%) | 34 (2.9%) | 0 |
| 2022–2024 | 1,955 | 42 (2.1%) | 39 (2.0%) | 5 (0.3%) |
| 2025–2026 (provisional) | 1,185 | 21 (1.8%) | 27 (2.3%) | 2 (0.2%) |
Read this table carefully; it is easy to over-claim from.
- 2025–2026 is provisional: CCS 2026 and IMC 2026 have not been held, and IEEE S&P 2026 and TheWebConf 2026 abstracts are not in OpenAlex, so selection under-covers those venue-years by construction. See corpus.
- The decline from 4.4% to 1.8% is not cleanly attributable to the API closing: the share was already down to 1.9% in 2016–2018, years before the 2023 restriction.
- The recent counts are inflated relative to the earlier ones by the failure mode our audit found: all four Twitter/X false positives were 2022 or later, and all were benchmark-corpus reuse. The honest statement is that primary collection from X gave way to reuse of frozen corpora, which is a claim about the composition of the count, not its size.
What the field went through to get the data
Access routes named in tool lists, over the 897 platform-subject papers. Free-text tool names matched by regex — a ranking, never a precise share:
| Route named | Papers | Share of 897 |
|---|---|---|
| any API-shaped tool | 236 | 26.3% |
| a Twitter/X streaming or search API | 91 | 10.1% |
| a custom scraper or crawler | 86 | 9.6% |
| generic browser automation (Selenium, Puppeteer, Playwright, …) | 80 | 8.9% |
| YouTube Data API | 18 | 2.0% |
| Reddit API / PRAW | 14 | 1.6% |
| Pushshift | 13 | 1.4% |
| Facebook/Meta Graph or Marketing API | 7 | 0.8% |
| Instagram API or scraper | 6 | 0.7% |
| Meta / Facebook Ad Library | 5 | 0.6% |
| CrowdTangle | 5 | 0.6% |
TikTok-Api — the unofficial scraper library, not the Research API | 1 | 0.1% |
| named no tool for it at all | 552 | 61.5% |
That last row is the reporting gap on this page: for 61.5% of platform-subject papers, no named instrument survives into this corpus's tool lists. Read that as an upper bound on the gap rather than as “three in five do not state it”: extraction recall on free-text tool names is imperfect, and we did not hand-audit the 552. Even as an upper bound it is the largest single reporting hole the page found, because the route is the design decision.
The reproducibility cost
Platform-subject papers release artefacts slightly less often than the corpus average, and withhold slightly more often:
| `artifacts.availability` | Platform-subject empirical (841) | All empirical (5,118) |
|---|---|---|
| public | 381 (45.3%) | 2,439 (47.7%) |
| none-mentioned | 332 (39.5%) | 1,964 (38.4%) |
| restricted | 17 (2.0%) | 73 (1.4%) |
| explicitly withheld | 10 (1.2%) | 52 (1.0%) |
The differences are small, and given the ~90% run-to-run stability of this field they are indicative rather than measured. What is not small is the reason: two 2025–2026 papers state outright that their data source no longer exists. [19Biswas, Md. Rafiul; Bessghaier, Mabrouka; Ibrahim, Shimaa; Mikros, George K.; Zaghouani, Wajdi (2026): "Longitudinal Trends in Global Climate Change Discourse on Facebook", in: Proceedings of the ACM Web Conference. (DOI)] collected 299,329 Facebook posts through CrowdTangle and records that “the use of the CrowdTangle API poses reproducibility challenges, as the tool was discontinued in 2024 and limits future data collection or temporal expansion”; [20Cinus, Federico; Minici, Marco; Luceri, Luca; Ferrara, Emilio (2025): "Exposing Cross-Platform Coordinated Inauthentic Activity in the Run-Up to the 2024 U.S. Election", in: Proceedings of the ACM Web Conference. (DOI)] refers to “(the now defunct) Crowdtangle”. CrowdTangle was withdrawn on 14 August 2024, and Meta's own page now reads “As of August 14, 2024, CrowdTangle is no longer available”.13) See Artifacts for what to deposit when the collection channel may not outlive the paper.
What the field measures on platforms
`detection.phenomenon` is free text and only ~20% stable run-to-run, so this is a fold into families, published with its residue:
| Family | Papers (of 897) |
|---|---|
| tracking / third-party data flows | 109 (12.2%) |
| network or infrastructure performance | 98 (10.9%) |
| spam / abuse / fraud / fake accounts | 79 (8.8%) |
| app, store or SDK analysis | 79 (8.8%) |
| advertising, targeting, ad delivery | 70 (7.8%) |
| account security, hijacking, authentication | 63 (7.0%) |
| misinformation and content moderation | 42 (4.7%) |
| privacy settings and user disclosure | 35 (3.9%) |
| deletion, data-subject rights, compliance | 30 (3.3%) |
| recommendation, personalisation, feeds | 24 (2.7%) |
| unmapped residue | 415 (46.3%) |
The residue is nearly half. Do not read this table as a partition of the field; read it as ten families that are definitely present, with a large remainder that our fold could not classify. The residue is printed in full on platforms.
Measured results you can cite
The corpus carries the papers' own measured prevalences with their own denominators. A sample, each verified against the paper's full text:
| Finding | Denominator the paper used |
|---|---|
| Only 7.7% of undeclared political ads were moderated as political by Meta; 60.4% of the ads Meta did moderate did not match its own criteria [2Bouchaud, Paul; Liénard, Jean F. (2024): "Beyond the Guidelines: Assessing Meta's Political Ad Moderation in the EU", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] | 29.5 M ads from the Meta Ad Library API, 16 EU countries |
| 3,546,479,731 WhatsApp accounts discovered; 57% with a public profile picture, 66% of a 500,000-image sample containing a detectable face [5Gegenhuber, Gabriel K.; Frenzel, Philipp E.; Günther, Maximilian; Ullrich, Johanna; Judmayer, Aljosha (2026): "Hey there! You are using WhatsApp: Enumerating Three Billion Accounts for Security and Privacy", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] | 63.2 bn candidate phone numbers enumerated |
| TikTok “exploits real users' interests in between 30% and 50% of all recommended videos in the first thousand videos” [3Vombatkere, Karan; Mousavi, Sepehr; Zannettou, Savvas; Roesner, Franziska; Gummadi, Krishna P. (2024): "TikTok and the Art of Personalization: Investigating Exploration and Exploitation on Social Media Feeds", in: Proceedings of the ACM Web Conference. (DOI)] | 347 donating users, 4.9 M videos, plus 5 bot accounts |
| 17,842 Amazon products restricted from shipping to at least one world region; 1.1% of 796,081 sampled books restricted to at least one of four Middle Eastern countries [21Knockel, Jeffrey; Dałek, Jakub; Aljizawi, Noura; Ahmed, Mohamed; Meletti, Levi; Lau, Justin (2026): "Banned Books: Analysis of Censorship on Amazon.com", Proceedings on Privacy Enhancing Technologies 2026(3):200-214. (DOI)] | Common Crawl-derived Amazon product set |
| Platform blocking of accounts advertised for sale worked on 19.71% of them overall, and very unevenly: YouTube 5.02%, Facebook 5.70%, X 18.67%, Instagram 46.41%, TikTok 816 of 1,700 — the paper's own summary is “TikTok and Instagram demonstrated the highest detection efficacy at 48%, whereas YouTube and Facebook showed the lowest efficacy at just 5%” [16Beluri, Mario; Acharya, Bhupendra; Khodayari, Soheil; Stivala, Giada; Pellegrino, Giancarlo; Holz, Thorsten (2025): "Exploration of the Dynamics of Buy and Sale of Social Media Accounts", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] | 11,457 visible accounts from 11 marketplaces |
| 136,009 Twitter users' Mastodon accounts identified across 2,879 instances; 96% of migrants joined the largest quartile of instances [22He, Jiahui; Zia, Haris Bin; Castro, Ignacio; Raman, Aravindh; Sastry, Nishanth; Tyson, Gareth (2023): "Flocking to Mastodon: Tracking the Great Twitter Migration", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] | 15,886 Mastodon instances, 1.02 M crawled accounts |
| Over 90% of targetable Facebook identities in the US had at least one data-broker-provided attribute (Australia 81.3%, UK 74.4%) [23Venkatadri, Giridhari; Sapiezynski, Piotr; Redmiles, Elissa M.; Mislove, Alan; Goga, Oana; Mazurek, Michelle L.; Gummadi, Krishna P. (2019): "Auditing Offline Data Brokers via Facebook's Advertising Platform", in: Proceedings of the ACM Web Conference. (DOI)] | the Facebook advertising interface, 7 countries |
| On X, posts containing external links had a median visibility score an order of magnitude below those without (“0.0069 vs. 0.084” for two named accounts) [10Galeazzi, Alessandro; Paudel, Pujan; Conti, Mauro; Cristofaro, Emiliano De; Stringhini, Gianluca (2026): "Revealing The Secret Power: How Algorithms Can Influence Content Visibility on Twitter/X", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] | 17 M + 35 M tweets from two published datasets |
Methodology and limitations of these figures
Every query, the report script and its unedited output, the folds with their residue, the four hand-audits, the quote checks and the external sources with their fetch dates are on platforms. Corpus-level caveats — the seven-venue scope, the provisional 2025–2026 slice, field stability — are on corpus.
Three limits worth restating here. The corpus is seven security, privacy and measurement venues; the platform-measurement literature also lives in ICWSM, CHI, FAccT and communications journals, and none of those are in it, so “the field” on this page means those seven venues. The platform families are regex candidate sets with the precision measured above, not curated sets. And the full-text probes are keyword probes: a paper that hit a rate limit and did not use the words “rate limit” is invisible to them.
Which Methods Are Current
| Method | Verdict | Evidence |
|---|---|---|
| Free open platform APIs (Twitter 1%/streaming, Facebook Graph for research, CrowdTangle) | historical | X academic access restricted since June 2023 [10Galeazzi, Alessandro; Paudel, Pujan; Conti, Mauro; Cristofaro, Emiliano De; Stringhini, Gianluca (2026): "Revealing The Secret Power: How Algorithms Can Influence Content Visibility on Twitter/X", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]; CrowdTangle withdrawn 14 August 2024; two 2025–2026 papers call it defunct [20Cinus, Federico; Minici, Marco; Luceri, Luca; Ferrara, Emilio (2025): "Exposing Cross-Platform Coordinated Inauthentic Activity in the Run-Up to the 2024 U.S. Election", in: Proceedings of the ACM Web Conference. (DOI)] [19Biswas, Md. Rafiul; Bessghaier, Mabrouka; Ibrahim, Shimaa; Mikros, George K.; Zaghouani, Wajdi (2026): "Longitudinal Trends in Global Climate Change Discourse on Facebook", in: Proceedings of the ACM Web Conference. (DOI)] |
| Metered pay-per-use API | current | X pay-per-usage, $0.005/post read, 3 M/month cap, verified 2026-08-27 |
| Platform research programmes (Meta Content Library via CASD, TikTok Research Tools, YouTube Researcher Program) | current, but you will be first | 1, 0 and 2 papers respectively in 5,855 |
| DSA Art. 40 vetted-researcher route | current, newly operational, untested here | Delegated Reg. (EU) 2025/2050; 1 on-point mention in 5,855 papers |
| Ad-transparency archives | current and productive | 69 papers; [2Bouchaud, Paul; Liénard, Jean F. (2024): "Beyond the Guidelines: Assessing Meta's Political Ad Moderation in the EU", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], [12Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)], [24Capozzi, Arthur; Morales, Gianmarco De Francisci; Mejova, Yelena; Monti, Corrado; Panisson, André (2023): "The Thin Ideology of Populist Advertising on Facebook during the 2019 EU Elections", in: Proceedings of the ACM Web Conference. (DOI)] |
| Data donation | current and rising | 25 papers, 11 of them 2025; [11McCrosky, Jesse; Malla, Ranadheer; Tanskanen, Aapo; Camargo, Chico Q. (2026): "Does This Button Work? Investigating YouTube's Ineffective User Controls", in: Proceedings of the ACM Web Conference. (DOI)], [12Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)] |
| Right of access / DSAR as an instrument | current and rising | 52 papers; [4Karnam, Sai Keerthana; Dash, Abhisek; Das, Antariksh; Mousavi, Sepehr; Bechtold, Stefan; Gummadi, Krishna P.; Mukherjee, Animesh; Weber, Ingmar; Zannettou, Savvas (2026): "Setting the Course, but Forgetting to Steer: Analyzing Compliance with GDPR's Right of Access to Data by Instagram, TikTok, and Youtube", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] |
| Sock-puppet accounts for personalised surfaces | current, high-risk, unavoidable for feed studies | 14 papers use the term; [3Vombatkere, Karan; Mousavi, Sepehr; Zannettou, Savvas; Roesner, Franziska; Gummadi, Krishna P. (2024): "TikTok and the Art of Personalization: Investigating Exploration and Exploitation on Social Media Feeds", in: Proceedings of the ACM Web Conference. (DOI)], [4Karnam, Sai Keerthana; Dash, Abhisek; Das, Antariksh; Mousavi, Sepehr; Bechtold, Stefan; Gummadi, Krishna P.; Mukherjee, Animesh; Weber, Ingmar; Zannettou, Savvas (2026): "Setting the Course, but Forgetting to Steer: Analyzing Compliance with GDPR's Right of Access to Data by Instagram, TikTok, and Youtube", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] |
| Logged-out scraping | current, contested; a designated VLOP was fined for forbidding it | 157 papers name a scraper; EC decision 5 December 2025 |
| Third-party archives (Pushshift) | the live service is closed to researchers; the historical dumps are what the field is actually using | 35 papers, still 9 in 2025 |
| Reuse of a published platform dataset | current and very common — often a symptom, not a choice | 383 of 897 name one; 208 have no primary collection at all |
| On-device / client-side model extraction | current, niche, powerful | [8West, Jack; Thiemt, Lea; Ahmed, Shimaa; Bartig, Maggie; Fawaz, Kassem; Banerjee, Suman (2024): "A Picture is Worth 500 Labels: A Case Study of Demographic Disparities in Local Machine Learning Models for Instagram and TikTok", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] |
| Amazon Product Advertising API 5.0 | retired | deprecated in favour of the Creators API; calls now return HTTP 403 AccessDeniedException14) |
What to Report
A platform study should let a reader reconstruct the population you actually saw. Report:
- The route in, by name and version — which API, which programme, which archive, which extension, which dataset. 61.5% of platform-subject papers in this corpus do not.
- The date window of collection, and what changed on the platform inside it. Platforms are not stationary; a six-month window can span a product change.
- The filter the route applies — follower thresholds, sampling rate, ad-archive scope, region restriction — as an explicit denominator, not as a footnote.
- The limit you worked under and how you designed around it, with your recall consequence.
- Authentication state: logged out, an account you registered, a donated session. 34.2% of platform papers with a crawl configuration do not say.
- Whether the platform pushed back: bot challenges, CAPTCHAs, throttling, bans. This is a result.
- The account policy, if you made accounts: how many, how they were named, what they did, whether they interacted with real users, and what happened to them afterwards.
- Ethics review and the terms-of-service position you took, and on whose advice. Only 32.5% of empirical platform-subject papers state a review outcome.
- What you can deposit and what you cannot, and why — see Artifacts.
Should There Be Facebook, Twitter, TikTok and Amazon Pages?
start has promised four sub-pages — Facebook, Twitter, TikTok, Amazon — since before this page existed. We are not writing them, and the red links have been removed rather than left dangling. The reasoning, so a future editor can overturn it on evidence:
| Promised page | Corpus support | Decision |
|---|---|---|
| TikTok | 7 candidate papers, 4 genuine after a full audit, none before 2023 | No page. Not because TikTok is unimportant — because the TikTok measurement literature is mostly outside these seven venues, so a page built from this corpus would be four papers and would misrepresent the field. The four are named above. |
| Amazon | 72 candidates, 36% precision; the genuine ones split between the Alexa voice/skill ecosystem ([25Cheng, Long; Wilson, Christin; Liao, Song; Young, Jeffrey; Dong, Daniel; Hu, Hongxin (2020): "Dangerous Skills Got Certified: Measuring the Trustworthiness of Skill Certification in Voice Personal Assistant Platforms", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], [26Liao, Song; Aldeen, Mohammed; Yan, Jingwen; Cheng, Long; Luo, Xiapu; Cai, Haipeng; Hu, Hongxin (2024): "Understanding GDPR Non-Compliance in Privacy Policies of Alexa Skills in European Marketplaces", in: Proceedings of the ACM Web Conference. (DOI)]) and the retail storefront ([21Knockel, Jeffrey; Dałek, Jakub; Aljizawi, Noura; Ahmed, Mohamed; Meletti, Levi; Lau, Justin (2026): "Banned Books: Analysis of Censorship on Amazon.com", Proceedings on Privacy Enhancing Technologies 2026(3):200-214. (DOI)]) | No page. “Amazon” is not one measurement object, and the Alexa skill-store work is closer to app-store measurement than to this page, though Mobile and app measurement does not cover skill stores yet. |
| Twitter / X | 142 candidates, 73% precision — the largest coherent body | No page. The access route that produced nearly all of it no longer exists, so a page would be a history of a closed API. What survives is the access-route material above. |
| Facebook / Meta | 151 candidates, 69% precision | No page. The method-bearing parts are already elsewhere: the pixel and server-side flows on Server side tracking and Requests, the SDKs on Mobile and app measurement, consent on consent. The ad archive is the one genuinely distinct instrument, and it is not Facebook-specific. |
A per-company page is the wrong axis anyway. What a platform study reuses across companies is the route in and the denominator that route implies; what does not transfer is the company's current API surface, which is precisely the part that goes stale fastest. This page is organised by route for that reason.
The one page that the evidence does support is not on the promised list: ad-transparency archives as a measurement instrument — 69 papers, a live regulatory driver, real double-sided error characteristics [2Bouchaud, Paul; Liénard, Jean F. (2024): "Beyond the Guidelines: Assessing Meta's Political Ad Moderation in the EU", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], and a 2026 cross-platform comparison [12Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)]. It is recorded as a candidate rather than linked, so this page does not create the very problem it just cleaned up.
Open Questions
- Has anyone published a DSA Art. 40 vetted-researcher study yet? One on-point mention in 5,855 papers. If the first such paper lands in 2027, its methods section becomes the template this page currently cannot give.
- What does a Meta Content Library study look like? Zero examples in these venues. In particular: how do you report a denominator that is defined by a follower threshold, and can you publish anything from inside a controlled computing environment that satisfies an artefact badge?
- What is the real cost of a metered-API study? Nobody has published the arithmetic. A short note with a worked budget — reads, cap, and what you had to drop — would be worth more than another dataset.
- Does the field's move to reused corpora change its findings? 383 of 897 platform papers name an existing dataset and 208 collected nothing of their own. Whether the conclusions drawn from a 2019 Twitter corpus still describe X in 2026 is an answerable question nobody in this corpus asks.
- Search-engine and ads-ecosystem auditing has no page on this wiki, and it is the largest measured family here — Google is rank 1 with 323 papers. SERP audits and ad-delivery audits share this page's problems (no API, sock puppets, personalisation) but have their own instruments. Out of scope here; worth its own page.
- Ad-transparency archives deserve their own page (see above). It needs the archive-by-archive coverage comparison that [12Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)] starts.
Related Pages
- Automated measurements — crawl, scan, or app analysis. A platform study is usually a crawl with an unusual counterparty.
- Website selection — because “Alexa” on this page is 412 papers using a retired ranking list, not Amazon.
- User studies — because “Amazon” on this page is 179 papers using Mechanical Turk. Crowdworkers labelling your platform data are annotation, not a user study.
- Crawling location — the vantage point, and why a datacenter IP gets a different platform than a residential one.
- Mobile and app measurement — app stores, SDKs, static versus dynamic analysis, and certificate pinning. It does not currently cover voice-assistant skill stores or the on-device model-extraction route [8West, Jack; Thiemt, Lea; Ahmed, Shimaa; Bartig, Maggie; Fawaz, Kassem; Banerjee, Suman (2024): "A Picture is Worth 500 Labels: A Case Study of Demographic Disparities in Local Machine Learning Models for Instagram and TikTok", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] used; those are described here instead.
- Longitudinal — repeating a platform measurement when the platform, not just the web, moved between waves.
- Registration — the mechanics of the accounts a sock-puppet study needs.
- Interaction — driving a feed once you are in.
- Requests, Server side tracking — the platform as a third party on someone else's site, which is a different study.
- Ethics — fake accounts, enumeration, and disclosure.
- Artifacts — depositing data from a channel that may not outlive the paper.
- Biases — donation panels, follower thresholds, and convenience samples.
- platforms — every query, fold, audit and external check behind this page.
References
- [1]
- Saha Roy, Sayak; Pourabbas Vafa, Elham; Khanmohamaddi, Kobra; Nilizadeh, Shirin (2025): "DarkGram: A Large-Scale Analysis of Cybercriminal Activity Channels on Telegram", in: Proceedings of the USENIX Security Symposium. (Link)
- [2]
- Bouchaud, Paul; Liénard, Jean F. (2024): "Beyond the Guidelines: Assessing Meta's Political Ad Moderation in the EU", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [3]
- Vombatkere, Karan; Mousavi, Sepehr; Zannettou, Savvas; Roesner, Franziska; Gummadi, Krishna P. (2024): "TikTok and the Art of Personalization: Investigating Exploration and Exploitation on Social Media Feeds", in: Proceedings of the ACM Web Conference. (DOI)
- [4]
- Karnam, Sai Keerthana; Dash, Abhisek; Das, Antariksh; Mousavi, Sepehr; Bechtold, Stefan; Gummadi, Krishna P.; Mukherjee, Animesh; Weber, Ingmar; Zannettou, Savvas (2026): "Setting the Course, but Forgetting to Steer: Analyzing Compliance with GDPR's Right of Access to Data by Instagram, TikTok, and Youtube", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
- [5]
- Gegenhuber, Gabriel K.; Frenzel, Philipp E.; Günther, Maximilian; Ullrich, Johanna; Judmayer, Aljosha (2026): "Hey there! You are using WhatsApp: Enumerating Three Billion Accounts for Security and Privacy", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
- [6]
- Cohn-Gordon, Katriel; Damaskinos, Georgios; Neto, Divino; Cordova, Joshi; Reitz, Benoît; Strahs, Benjamin; Obenshain, Daniel; Pearce, Paul; Papagiannis, Ioannis (2020): "DELF: Safeguarding deletion correctness in Online Social Networks", in: Proceedings of the USENIX Security Symposium. (Link)
- [7]
- Schlinker, Brandon; Cunha, Ítalo S.; Chiu, Yi-Ching; Sundaresan, Srikanth; Katz-Bassett, Ethan (2019): "Internet Performance from Facebook's Edge", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [8]
- West, Jack; Thiemt, Lea; Ahmed, Shimaa; Bartig, Maggie; Fawaz, Kassem; Banerjee, Suman (2024): "A Picture is Worth 500 Labels: A Case Study of Demographic Disparities in Local Machine Learning Models for Instagram and TikTok", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
- [9]
- Nguyen, Hoang Dai; Dhungana, Sumit; Itha, Madhulika; Vadrevu, Phani (2025): ""Please don't send that bot anything": A Mixed-methods Study of Personal Impersonation Attacks Targeting Digital Payments on Social Media", in: Proceedings of the USENIX Security Symposium. (Link)
- [10]
- Galeazzi, Alessandro; Paudel, Pujan; Conti, Mauro; Cristofaro, Emiliano De; Stringhini, Gianluca (2026): "Revealing The Secret Power: How Algorithms Can Influence Content Visibility on Twitter/X", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
- [11]
- McCrosky, Jesse; Malla, Ranadheer; Tanskanen, Aapo; Camargo, Chico Q. (2026): "Does This Button Work? Investigating YouTube's Ineffective User Controls", in: Proceedings of the ACM Web Conference. (DOI)
- [12]
- Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)
- [13]
- Kaleli, Beliz; Kondracki, Brian; Egele, Manuel; Nikiforakis, Nick; Stringhini, Gianluca (2021): "To Err.Is Human: Characterizing the Threat of Unintended URLs in Social Media", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
- [14]
- Zannettou, Savvas; Caulfield, Tristan; Blackburn, Jeremy; Cristofaro, Emiliano De; Sirivianos, Michael; Stringhini, Gianluca; Suarez-Tangil, Guillermo (2018): "On the Origins of Memes by Means of Fringe Web Communities", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [15]
- Acharya, Bhupendra; Lazzaro, Dario; López-Morales, Efrén; Oest, Adam; Saad, Muhammad; Cinà, Antonio Emanuele; Schönherr, Lea; Holz, Thorsten (2024): "The Imitation Game: Exploring Brand Impersonation Attacks on Social Media Platforms", in: Proceedings of the USENIX Security Symposium. (Link)
- [16]
- Beluri, Mario; Acharya, Bhupendra; Khodayari, Soheil; Stivala, Giada; Pellegrino, Giancarlo; Holz, Thorsten (2025): "Exploration of the Dynamics of Buy and Sale of Social Media Accounts", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [17]
- Xue, Diwen; Ramesh, Reethika; S, Valdik S.; Evdokimov, Leonid; Viktorov, Andrey; Jain, Arham; Wustrow, Eric; Basso, Simone; Ensafi, Roya (2021): "Throttling Twitter: an emerging censorship technique in Russia", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [18]
- Simko, Lucy; Hutchinson, Adryana; Isaac, Alvin; Fries, Evan; Sherr, Micah; Aviv, Adam J. (2024): ""Modern problems require modern solutions": Community-Developed Techniques for Online Exam Proctoring Evasion", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
- [19]
- Biswas, Md. Rafiul; Bessghaier, Mabrouka; Ibrahim, Shimaa; Mikros, George K.; Zaghouani, Wajdi (2026): "Longitudinal Trends in Global Climate Change Discourse on Facebook", in: Proceedings of the ACM Web Conference. (DOI)
- [20]
- Cinus, Federico; Minici, Marco; Luceri, Luca; Ferrara, Emilio (2025): "Exposing Cross-Platform Coordinated Inauthentic Activity in the Run-Up to the 2024 U.S. Election", in: Proceedings of the ACM Web Conference. (DOI)
- [21]
- Knockel, Jeffrey; Dałek, Jakub; Aljizawi, Noura; Ahmed, Mohamed; Meletti, Levi; Lau, Justin (2026): "Banned Books: Analysis of Censorship on Amazon.com", Proceedings on Privacy Enhancing Technologies 2026(3):200-214. (DOI)
- [22]
- He, Jiahui; Zia, Haris Bin; Castro, Ignacio; Raman, Aravindh; Sastry, Nishanth; Tyson, Gareth (2023): "Flocking to Mastodon: Tracking the Great Twitter Migration", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [23]
- Venkatadri, Giridhari; Sapiezynski, Piotr; Redmiles, Elissa M.; Mislove, Alan; Goga, Oana; Mazurek, Michelle L.; Gummadi, Krishna P. (2019): "Auditing Offline Data Brokers via Facebook's Advertising Platform", in: Proceedings of the ACM Web Conference. (DOI)
- [24]
- Capozzi, Arthur; Morales, Gianmarco De Francisci; Mejova, Yelena; Monti, Corrado; Panisson, André (2023): "The Thin Ideology of Populist Advertising on Facebook during the 2019 EU Elections", in: Proceedings of the ACM Web Conference. (DOI)
- [25]
- Cheng, Long; Wilson, Christin; Liao, Song; Young, Jeffrey; Dong, Daniel; Hu, Hongxin (2020): "Dangerous Skills Got Certified: Measuring the Trustworthiness of Skill Certification in Voice Personal Assistant Platforms", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
- [26]
- Liao, Song; Aldeen, Mohammed; Yan, Jingwen; Cheng, Long; Luo, Xiapu; Cai, Haipeng; Hu, Hongxin (2024): "Understanding GDPR Non-Compliance in Privacy Policies of Alexa Skills in European Marketplaces", in: Proceedings of the ACM Web Conference. (DOI)
curl); the page is stamped “Updated 1 year ago”. A paper citing Pushshift today is citing a historical dump, and should say which one and when it was obtained.help.crowdtangle.com no longer resolves in DNS.