User Tools

Site Tools


design:user_studies

This is an old revision of the document!


User Studies

You are about to measure the web, and a crawl is not going to answer the question. You need humans: someone to take a survey, sit in a lab, donate telemetry, click a consent banner, or tell you whether a piece of security advice is even comprehensible. That is the other top-level Design branch on this site — the sibling of Automated measurements — and it is not an HCI textbook. It is what these seven security, privacy and web venues actually do when they recruit people, where they recruit them, what they forget to report, and which of those choices is already dead.

Amazon Mechanical Turk permanently closes on 30 September 2026. Amazon's own acceptable-use-policy page is the primary source: “Mechanical Turk will close September 30, 2026” and “Amazon Mechanical Turk will permanently close on September 30, 2026.”1) The requester site still carries the older banner that it is “no longer accepting new customers” (in force since 30 July 2026).2) A 2026 paper that “uses MTurk” is using a platform that will not exist when the camera-ready is due. Do not copy 2010–2012 pay scales into a 2026 protocol.

CHI, SOUPS and CSCW are not in this corpus. Every figure below is a claim about CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf and IEEE S&P, 2010–2026, 5,859 extracted papers. Usable-security work also publishes at SOUPS and CHI. Treating these seven venues as “the field” undercounts the literature a reviewer from those communities will expect you to have read. See Conferences and Corpus.

This page is for a PhD student who already knows how to crawl and now has to put people in the loop. What belongs: which recruitment channels these seven venues actually use; which of those channels is current in 2026; the annotation trap that turns Mechanical Turk labelers into fake participants; what the papers forget to say (compensation, region, recruitment); what to report so a reviewer can reconstruct the sample. What does not: how to write a survey, how to run a t-test, how the Belmont Report is structured, or how to moderate a focus group. Those are textbooks. Ethics review, preregistration, hypothesis tests and inter-rater agreement already have pages — this one links them rather than restating them.

Queries, folds, the unedited report and the quote checks are on user_studies. Corpus-wide caveats: Corpus.

What to Read First

Eight papers will orient you faster than the tables, and each is here for a methodological reason.

Paper Why it is first
Wei et al., USENIX Security 2024, SoK (or SoLK?) [1Wei, Miranda; Mink, Jaron; Eiger, Yael; Kohno, Tadayoshi; Redmiles, Elissa M.; Roesner, Franziska (2024): "SoK (or SoLK?): On the Quantitative Study of Sociodemographic Factors and Computer Security Behaviors", in: Proceedings of the USENIX Security Symposium. (Link)] The sociodemographic SoK, and a warning about WEIRD samples. Reviewed existing scholarship on sociodemographics and secure behavior (151 papers) and then a focused subset of 47, plus a Facebook in-feed survey of 16,829 users across 16 countries. They position the work as a systemization of a lack of knowledge. Read it before you write “our MTurk sample generalises”.
Klemmer et al., IEEE S&P 2025, Transparency in Usable Privacy and Security Research [2Klemmer, Jan H.; Schmüser, Juliane; Lowens, Byron M.; Fischer, Fabian; Schmüser, Lea; Schaub, Florian; Fahl, Sascha (2025): "Transparency in Usable Privacy and Security Research: Scholars' Perspectives, Practices, and Recommendations", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] What 24 UPS researchers actually do. Interviewed 24 usable-privacy-and-security scholars. Compensation was optional ($25 or local equivalent); 16 accepted. Read it for the reporting culture, not for a recruitment recipe.
Zeber et al., TheWebConf 2020, The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing [3Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] A crawl is not a user. Compared crawls to Firefox Pioneer telemetry: “Approximately 52,000 users participated, all of whom were using Firefox versions 67 or 68 in the en-US localization”, over a dataset of 30 million site visits. For a majority of the top domains they inspected, “the number of third parties reported by the crawler is well into the right tail” of the human distribution. If your claim is about people, a crawl is a surrogate with a known direction of bias.
Berke et al., PETS 2025, How Unique is Whose Web Browser? [4Berke, Alex; Calacci, Dan; Mahari, Robert; Yabe, Takahiro; Larson, Kent; Pentland, Sandy (2025): "How Unique is Whose Web Browser? The Role of Demographics in Browser Fingerprinting Among US Users", Proceedings on Privacy Enhancing Technologies 2025(1):720-758. (DOI)] Current Prolific-scale uniqueness. “Participants were recruited via the online platform Prolific with a male/female gender balance and were offered $0.60 for an estimated 2 minute survey.” 12,461 total participants; browser-attribute data with informed consent from 8,400 US participants. This is what a 2025 recruited-crowd study at n≥10,000 looks like. It is not Mechanical Turk.
Redmiles et al., USENIX Security 2020, A Comprehensive Quality Evaluation of Security and Privacy Advice on the Web [5Redmiles, Elissa M.; Warford, Noel; Jayanti, Amritha; Koneru, Aravind; Kross, Sean; Morales, Miraida; Stevens, Rock; Mazurek, Michelle L. (2020): "A Comprehensive Quality Evaluation of Security and Privacy Advice on the Web", in: Proceedings of the USENIX Security Symposium. (Link)] Advice quality, three samples. Distinguished Paper. 50 MTurk workers wrote search queries; 1,586 US users recruited via Cint evaluated actionability and comprehensibility; 41 qualified experts scored the advice. The paper is a measurement of the web's security-advice ecosystem (1,264 documents, 374 unique recommended behaviours) that needed humans to say whether any of it is usable.
Kelley et al., IEEE S&P 2012, Guess Again (and Again and Again) [6Kelley, Patrick Gage; Komanduri, Saranga; Mazurek, Michelle L.; Shay, Richard; Vidas, Timothy; Bauer, Lujo; Christin, Nicolas; Cranor, Lorrie Faith; L´opez, Julio C. (2012): "Guess Again (and Again and Again): Measuring Password Strength by Simulating Password-Cracking Algorithms", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] Historical MTurk, not a pay template. “From August 2010 to January 2011, we advertised a two-part study on Mechanical Turk, paying between 25 and 55 cents for the first part and between 50 and 70 cents for the second part.” Plus 280 Carnegie Mellon email users. Do not copy those rates. They are a 2010–2011 fact, and they sit below every current platform minimum this page could verify.
Pei et al., TheWebConf 2020, Attention Please [7Pei, Weiping; Mayer, Arthur; Tu, Kaylynn; Yue, Chuan (2020): "Attention Please: Your Attention Check Questions in Survey Studies Can Be Automatically Answered", in: Proceedings of the ACM Web Conference. (DOI)] Attention checks can be auto-answered. This paper is not in the human-subjects population: it proposes an attack, it does not recruit. Quote: attackers “can leverage deep learning techniques to pass attention check questions automatically.” A checkbox that “the Turkers passed the attention check” is not evidence they were paying attention.
Herbert et al., USENIX Security 2025, Digital Security Perceptions and Practices Around the World [8Herbert, Franziska; Munyendo, Collins W.; Hielscher, Jonas; Becker, Steffen; Zou, Yixin (2025): "Digital Security Perceptions and Practices Around the World: A WEIRD versus Non-WEIRD Comparison", in: Proceedings of the USENIX Security Symposium. (Link)] A 12-country panel, WEIRD vs not. Commissioned Kantar: n=12,351, twelve countries. This is what “we recruited a panel” means at the scale the corpus's giants actually reach — and it is not Prolific and not MTurk.

Which Population Is Yours

Three schema signals look like “user study” and they disagree. This page's population is one of them, named, and the disagreement is the finding rather than something to union away.

Slice Definition Papers Share of 5,859
humanSubjects participants[] is non-empty — the page population 1,357 23.2%
studyTypes includes user-study the extractor tagged the study type 1,149 19.6%
studyTypes includes interview-or-survey 755 12.9%
population.unit is human-participants 979 16.7%
humanAnnotation[] non-empty someone labelled data, not a participant 3,318 56.6%
annotatorType crowdworkers crowdworkers as labelers 73 1.2%

Every figure below is on humanSubjects (N=1,357) unless it says otherwise.

The type tag and the participants array disagree in both directions: 1,125 papers have both, 232 have participants[] but no user-study type, 24 have the type and an empty participants[]. The type-only papers are crawls, telemetry, or crowd annotation the extractor filed as a study type. Unioning the two would invent a population that is neither. The item brief's “965 papers” was computed on the old 4,322-paper corpus and is stale; the current count is 1,357.

149 of the 1,357 (11.0%) also crawled; 149 of 1,120 crawled papers (13.3%) also recruited. Automated measurement and a user study are two instruments that co-occur, not a partition. See Automated measurements.

The annotation trap

Crowdworkers who label data are annotators, not participants. Of 3,318 papers with a humanAnnotation[] record, 73 set annotatorType to crowdworkers. 40 of those 73 also have participants[]; 33 do not. Those 33 are hired labelers. Counting them as human subjects double-counts Mechanical Turk and invents a consent story the paper did not tell. Inter-rater agreement, open coding and the rest of the annotation apparatus live on Interrater agreement.

Where the work is published

As a share of each venue's own output (not as a share of 1,357):

Venue humanSubjects all extracted papers Share of that venue
PoPETs 215 510 42.2%
IEEE S&P 208 767 27.1%
USENIX Security 367 1,410 26.0%
NDSS 144 701 20.5%
TheWebConf 170 843 20.2%
CCS 186 990 18.8%
IMC 67 638 10.5%

PoPETs is the user-study venue in this corpus as a share of its own output. IMC is the least — the same split Conferences uses to tell you not to send a survey to IMC because IMC is in the name. By volume, USENIX Security publishes the most of it (367 papers).

The share of each year-window that is human-subjects work has been rising slowly: 16.4% (2010–2013), 20.2% (2014–2017), 22.9% (2018–2021), 24.9% (2022–2024), 25.6% (2025–2026*). The 2025–2026 window is provisional — CCS 2026 and IMC 2026 have not been held, and IEEE S&P / WWW 2026 abstracts are under-selected by construction. Do not read the last window as a complete year.

35 of 1,357 records are posters or very short (slug poster-, title starting Poster:, or pages⇐4). Dropping them moves the population to 1,322 and does not reorder recruitment.

How the Field Recruits

The schema has a recruitment enum. 903 of 1,357 (66.5%) state a channel; 454 (33.5%) produce only a sentinel. Sentinel-only is the finding: one human-subjects paper in three does not name how people were recruited. Sentinels are never counted as a channel.

recruitment (schema enum) Papers Share of 1,357
other-crowd-platform 193 14.2%
professional-network 189 13.9%
university-pool 174 12.8%
mechanical-turk 139 10.2%
social-media 119 8.8%
snowball 91 6.7%
in-person 68 5.0%
prolific 61 4.5%
panel-company 42 3.1%

Shares do not sum to 100%: a paper can use two channels, and 454 name none. Do not read “mechanical-turk 139” as “how often the field used MTurk.” The enum undercounts. Many quotes that name Mechanical Turk or Prolific in so many words were filed as other-crowd-platform. Recovering platform names from the evidence quotes and from paper.cols.txt (an upper bound: related-work citations fire) gives:

Platform Schema enum Quote recovery Quote+full text Share of 1,357 (full text)
Mechanical Turk 139 92 279 20.6%
Prolific 61 102 194 14.3%
CloudResearch 1 10 0.7%
CrowdFlower / Figure Eight / Appen 3 18 1.3%
Upwork / Freelancer 11 31 2.3%
Clickworker 3 6 0.4%
Wenjuanxing / Credamo 5 7 0.5%
Respondi 3 4 0.3%
Microworkers 1 2 0.1%
Toloka 2 2 0.1%

Schema is a lower bound. Full-text recovery is an upper bound. Schema MTurk is 49.8% of the recovered-MTurk set; schema Prolific is 31.4% of recovered-Prolific. Union of recovered MTurk∪Prolific: 409 of 1,357 (30.1%); 64 papers hit both. Of the 193 other-crowd-platform papers, 168 name one of the families above and 25 name nothing recoverable — listed in full on the provenance page.

Qualtrics full-text hits 124 of 1,357 human-subjects papers. It is the survey tool, not a recruitment family, and it is not in the table.

Which platform is current

Date the method from the recovered series, not from the schema. Schema Prolific is 0 until the 2018–2021 window; recovered Prolific is already 27 papers in that window because other-crowd-platform ate it.

Window HS schema MTurk recovered MTurk schema Prolific recovered Prolific
2010–2013 84 14 23 0 0
2014–2017 155 20 42 0 3
2018–2021 329 60 104 2 27
2022–2024 486 38 86 25 94
2025–2026* 303 7 24 34 70

Per calendar year, recovered Prolific overtakes recovered MTurk in 2024 (39 vs 28), and the gap is wide in the provisional 2025–2026 slice (46 vs 19 in 2025; 24 vs 5 in 2026). Schema-only, the crossover is later: 2025–2026* schema Prolific 34 vs schema MTurk 7. Report both bounds. After 30 September 2026 the MTurk column is historical.

University-pool, professional-network, social-media, snowball, in-person and panel-company are all still live channels. Professional-network is the second schema enum (189) and grew in 2022–2024 (76 of 486). Panel-company is small in the schema (42) and is how the n≥10,000 survey giants actually get people — Kantar [8Herbert, Franziska; Munyendo, Collins W.; Hielscher, Jonas; Becker, Steffen; Zou, Yixin (2025): "Digital Security Perceptions and Practices Around the World: A WEIRD versus Non-WEIRD Comparison", in: Proceedings of the USENIX Security Symposium. (Link)], Cint [5Redmiles, Elissa M.; Warford, Noel; Jayanti, Amritha; Koneru, Aravind; Kross, Sean; Morales, Miraida; Stevens, Rock; Mazurek, Michelle L. (2020): "A Comprehensive Quality Evaluation of Security and Privacy Advice on the Web", in: Proceedings of the USENIX Security Symposium. (Link)].

Platforms in 2026

A ranking of what the literature did is not advice about what to do now.

Channel Status on 2026-08-27 Use it?
Mechanical Turk Permanently closes 30 September 2026. Closed to new customers since 30 July 2026. Operated since 2005 (21 years). No, for any study that will still be in the field after that date. Existing requesters: read Amazon's own FAQ from the AUP banner, then leave.
Prolific Live. Absolute minimum £6 / $8 per hour; recommended £9 / $12 per hour. Academic / non-profit platform fee 33.3% of participant rewards; corporate 42.8%, charged on top of rewards.3) Yes, if a crowd sample is the right population. Pay the recommended rate unless you have a reason and say so.
CloudResearch Live. Homepage: “Survey Platform & Participant Recruitment”. The recovered-full-text count in this corpus is 10 of 1,357 (TurkPrime-era papers included). A named MTurk-adjacent alternative. Confirm on their site that the product you want still exists; this page does not certify a migration path.
University pool / SONA / course credit Live. Schema 174 of 1,357. Right when your claim is about students; wrong when you then write “users”.
Professional network Live. Schema 189. How Klemmer et al. recruited 24 UPS researchers [2Klemmer, Jan H.; Schmüser, Juliane; Lowens, Byron M.; Fischer, Fabian; Schmüser, Lea; Schaub, Florian; Fahl, Sascha (2025): "Transparency in Usable Privacy and Security Research: Scholars' Perspectives, Practices, and Recommendations", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]. Right for expert samples.
Panel company (Kantar, Cint, …) Live. Schema 42; this is how the cross-country n≥10k surveys happen. Right when you need a specified national sample and can pay for it.
In-product / donation / telemetry Live where the product is. Mozilla Rally is dead (GitHub org archived July 2024) — see Ethics. Firefox Pioneer, the panel behind Zeber et al., is not something you can join in 2026. Right when you already have users. Not a recruitment method you can start from zero.
Qualtrics Live, as a survey tool. 124 of 1,357 human-subjects papers name it in full text. Do not list it under recruitment.

The script at the foot of this page re-fetches the MTurk, Prolific and CloudResearch URLs and exits 1 if a dated claim has disappeared. Re-run it before you copy a rate into a protocol.

How Large Are These Samples?

Paper-level maximum n (the largest participants[].n on the paper): stated by 1,284 of 1,357 (94.6%). Unstated: 73. Among those 1,284: min 1, 25th percentile 20, median 43, 75th percentile 210, 90th percentile 820, 99th percentile 13,000, max 253238367.

Bucket (paper-level max n) Papers Share of 1,284 with n stated
1–19 319 24.8%
20–49 340 26.5%
50–99 141 11.0%
100–499 295 23.0%
500–999 84 6.5%
1,000–9,999 87 6.8%
≥10,000 18 1.4%

659 of 1,284 (51.3%) are under 50. Median excluding the 18 giants: 41 (N=1,266). The giants do not move the median, they sit in the tail.

n≥10,000 is not “a big survey”. Hand-classified from the extraction's own evidence quotes (the report exits 1 if this map and the n≥10,000 set diverge):

Role Papers What it is
in-product 7 Existing users of Facebook, Chrome, or a live site. Not recruited. The largest is 253238367 Facebook users who visited in Aug–Oct 2010.
donation 4 Opt-in telemetry, extension, or data-donation. Zeber et al. [3Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] is 52,000 Firefox Pioneer users.
panel 3 Commissioned survey panel. Herbert et al. [8Herbert, Franziska; Munyendo, Collins W.; Hielscher, Jonas; Becker, Steffen; Zou, Yixin (2025): "Digital Security Perceptions and Practices Around the World: A WEIRD versus Non-WEIRD Comparison", in: Proceedings of the USENIX Security Symposium. (Link)] is Kantar, 12,351, 12 countries.
university-field 1 A captive institutional population (phishing simulation).
recruited-crowd 3 Actually recruited on MTurk or Prolific at this scale.

The three recruited-crowd giants: Kelley et al. [6Kelley, Patrick Gage; Komanduri, Saranga; Mazurek, Michelle L.; Shay, Richard; Vidas, Timothy; Bauer, Lujo; Christin, Nicolas; Cranor, Lorrie Faith; L´opez, Julio C. (2012): "Guess Again (and Again and Again): Measuring Password Strength by Simulating Password-Cracking Algorithms", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] (MTurk, 12,000, 2010–2011); Park et al. [9Park, Sungkyu; Park, Jamie Yejean; Chin, Hyojin; Kang, Jeong-han; Cha, Meeyoung (2021): "An Experimental Study to Understand User Experience and Perception Bias Occurred by Fact-checking Messages", in: Proceedings of the ACM Web Conference. (DOI)] (MTurk, 11,145, TheWebConf 2021); Berke et al. [4Berke, Alex; Calacci, Dan; Mahari, Robert; Yabe, Takahiro; Larson, Kent; Pentland, Sandy (2025): "How Unique is Whose Web Browser? The Role of Demographics in Browser Fingerprinting Among US Users", Proceedings on Privacy Enhancing Technologies 2025(1):720-758. (DOI)] (Prolific, 12,461, PETS 2025). Facebook's 253238367 is not a recruited sample. Neither is a Chrome in-product survey, nor Utz et al. [10Utz, Christine; Degeling, Martin; Fahl, Sascha; Schaub, Florian; Holz, Thorsten (2019): "(Un)informed Consent: Studying GDPR Consent Notices in the Field", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] (CCS 2019): “a between-subjects study with 82,890 real website visitors of a German e-commerce website”, plus 110 voluntary survey responses. If you write n=82890 for a consent-banner field study, say that they were visitors, not recruits.

Compensation and Regions

A compensation string is present on any tuple for 746 of 1,357 (55.0%). A non-sentinel region is present for 665 of 1,357 (49.0%). Both match Ethics's sibling accounting of the same fields. One human-subjects paper in two does not say what people were paid; one in two does not say where they were.

Compensation folded from free text (single-label, ordered: unpaid before course-credit before raffle before hourly before gift-card before a stated amount; residue printed on the provenance page):

Kind Papers Share of 1,357
stated-amount 371 27.3%
gift-card 149 11.0%
unpaid 84 6.2%
hourly 57 4.2%
raffle 35 2.6%
course-credit 6 0.4%

139 compensation tuples did not map; 611 papers have no compensation string at all. An hourly rate you can check against Prolific's £6 / $8 floor is the exception, not the rule.

Regions folded through the same geo.mjs used on Crawling location. A paper naming two countries counts in both rows. Shares of 1,357:

Country Papers Share of 1,357
United States 416 30.7%
Germany 104 7.7%
United Kingdom 89 6.6%
China 55 4.1%
India 46 3.4%
Canada 26 1.9%
Australia 25 1.8%
Netherlands 22 1.6%
Switzerland 22 1.6%
Pakistan 21 1.5%

416 of the 665 papers that state a region (62.6%) name the United States. The US share of all 1,357 is 30.7% only because half the papers state no region at all. Wei et al. [1Wei, Miranda; Mink, Jaron; Eiger, Yael; Kohno, Tadayoshi; Redmiles, Elissa M.; Roesner, Franziska (2024): "SoK (or SoLK?): On the Quantitative Study of Sociodemographic Factors and Computer Security Behaviors", in: Proceedings of the USENIX Security Symposium. (Link)] exist because this is a problem. Herbert et al. [8Herbert, Franziska; Munyendo, Collins W.; Hielscher, Jonas; Becker, Steffen; Zou, Yixin (2025): "Digital Security Perceptions and Practices Around the World: A WEIRD versus Non-WEIRD Comparison", in: Proceedings of the USENIX Security Symposium. (Link)] exist because a 12-country Kantar panel is what it takes to not be a US/DE/UK convenience sample.

Ethics Review, Study Kind, Platform

Ethics numbers on this population must match Ethics. They do: of 1,357 human-subjects papers, 1,022 (75.3%) state a review outcome; 896 (66.0%) are approved or exempt; 335 (24.7%) are silent (none-mentioned 277 + no record 58). The reporting rate has moved: 28.6% of 2010–2013 human-subjects papers stated a review outcome, 85.2% of 2022–2024, 86.8% of 2025–2026*. Being in the human-subjects column is what makes the apparatus fire — crawling papers without participants sit at 21.3%. That contrast, the venue ethics bodies, and what “not human subjects” does not get you, are that page's job. Ramulu et al. [11Ramulu, Harshini Sri; Schmitt, Helen; Rerich, Bogdan; Rodriguez, Rachel Gonzalez; Kohno, Tadayoshi; Acar, Yasemin (2025): "Ethics in Computer Security Research: A Data-Driven Assessment of the Past, the Present, and the Possible Future", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] is the independent 2024 four-venue hand-read. This page does not re-derive it.

Study kind is multi-label (a paper can run a survey and an interview). Of 1,357: survey 511 (37.7%), lab-study 373 (27.5%), interview 349 (25.7%), field-study 229 (16.9%), usability-test 99 (7.3%), other 79 (5.8%), focus-group 26 (1.9%), diary-study 1. Stated 1,357, sentinel-only 0.

The hybrid 149 (crawled and recruited) are slightly more field-study and usability-test than the rest (field-study 34 of 149, 22.8%; usability-test 19 of 149, 12.8%). Utz et al. [10Utz, Christine; Degeling, Martin; Fahl, Sascha; Schaub, Florian; Holz, Thorsten (2019): "(Un)informed Consent: Studying GDPR Consent Notices in the Field", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] is the type specimen: a live site, real visitors, a consent notice as the instrument, and a small voluntary survey on top. Zeber et al. [3Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] is the other type specimen: a crawl compared to donated human telemetry.

platforms[] on human-subjects papers, multi-label: offline 496 (36.6%), other-online-service 458 (33.8%), web 393 (29.0%), mobile 333 (24.5%), iot 117 (8.6%). Offline (lab / interview) is the modal platform. This page is not only web-user studies; it is user studies in these seven venues. A lab password study and a Prolific fingerprinting survey are both in N=1,357.

What to Report

A reviewer reconstructing your sample should not have to guess. Of 1,357 human-subjects papers in these venues, 33.5% do not name the recruitment channel, 45.0% do not state compensation, 51.0% do not state a region. Do not be those papers.

  1. Who was recruited, through what channel, on which platform. “Online” is a sentinel. Name Prolific, the university pool, Kantar, the in-product surface, or the professional list. If they were visitors of a live site rather than recruits, say that — it changes what n means [10Utz, Christine; Degeling, Martin; Fahl, Sascha; Schaub, Florian; Holz, Thorsten (2019): "(Un)informed Consent: Studying GDPR Consent Notices in the Field", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)].
  2. n, with the unit. 12,461 Prolific respondents is not 8,400 consented browser-attribute files [4Berke, Alex; Calacci, Dan; Mahari, Robert; Yabe, Takahiro; Larson, Kent; Pentland, Sandy (2025): "How Unique is Whose Web Browser? The Role of Demographics in Browser Fingerprinting Among US Users", Proceedings on Privacy Enhancing Technologies 2025(1):720-758. (DOI)]. 82,890 visitors is not 110 survey respondents [10Utz, Christine; Degeling, Martin; Fahl, Sascha; Schaub, Florian; Holz, Thorsten (2019): "(Un)informed Consent: Studying GDPR Consent Notices in the Field", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)]. 52,000 telemetry donors is not 50 MTurk workers.
  3. What they were paid, in an amount a reader can convert to an hourly rate. Prolific's absolute floor on 2026-08-27 is £6 / $8 per hour; recommended is £9 / $12. Kelley et al.'s 25–70 cents [6Kelley, Patrick Gage; Komanduri, Saranga; Mazurek, Michelle L.; Shay, Richard; Vidas, Timothy; Bauer, Lujo; Christin, Nicolas; Cranor, Lorrie Faith; L´opez, Julio C. (2012): "Guess Again (and Again and Again): Measuring Password Strength by Simulating Password-Cracking Algorithms", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] is a 2010–2011 fact, not a template. If compensation was optional, say how many took it [2Klemmer, Jan H.; Schmüser, Juliane; Lowens, Byron M.; Fischer, Fabian; Schmüser, Lea; Schaub, Florian; Fahl, Sascha (2025): "Transparency in Usable Privacy and Security Research: Scholars' Perspectives, Practices, and Recommendations", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)].
  4. Where they were, at least at country level. 62.6% of papers that state a region name the United States. If your sample is WEIRD, Wei et al. [1Wei, Miranda; Mink, Jaron; Eiger, Yael; Kohno, Tadayoshi; Redmiles, Elissa M.; Roesner, Franziska (2024): "SoK (or SoLK?): On the Quantitative Study of Sociodemographic Factors and Computer Security Behaviors", in: Proceedings of the USENIX Security Symposium. (Link)] is the paper a reviewer will cite back at you. If it is not, Herbert et al. [8Herbert, Franziska; Munyendo, Collins W.; Hielscher, Jonas; Becker, Steffen; Zou, Yixin (2025): "Digital Security Perceptions and Practices Around the World: A WEIRD versus Non-WEIRD Comparison", in: Proceedings of the USENIX Security Symposium. (Link)] is the current existence proof that a 12-country panel is possible.
  5. The ethics determination, not a vibe. The venues that carry this literature will reject on ethics with the technical review untouched — see Ethics.
  6. Attention checks, and what they actually show. Pei et al. [7Pei, Weiping; Mayer, Arthur; Tu, Kaylynn; Yue, Chuan (2020): "Attention Please: Your Attention Check Questions in Survey Studies Can Be Automatically Answered", in: Proceedings of the ACM Web Conference. (DOI)] demonstrated they can be passed automatically. Report the check, the fail rate, and that a passed check is not proof of attention.
  7. If you also crawled: which claims come from the crawl and which from the people. Zeber et al. [3Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] is the reason the two are not interchangeable.
  8. Preregistration if you have hypotheses about people — already normal on this side of the house, almost absent on the crawl side. See Study preregistration. Tests, corrections, and clustered/paired designs: Hypothesis testing, Pvalue corrections. Open coding: Interrater agreement.

Live check: the platforms this page named

Re-run before you copy a rate. Exits 1 if a fetch fails or a needle the page depends on has disappeared. Needles are substrings, not regex — a $ in a regex is end-of-string, which is how a first version of this script reported Prolific as missing $8 while the HTML sat in the file. A 33.3 needle hits a CSS gradient; the check therefore demands 33.3% of participant rewards.

platform_status.py
#!/usr/bin/env python3
"""Live status of the recruitment platforms a 2026 web-measurement paper uses.
 
Mechanical Turk's own acceptable-use-policy page is the primary source for the
close date. requester.mturk.com still carries the older "no longer accepting
new customers" banner. Prolific's researcher-help pages carry the minimum
hourly rates. CloudResearch is the named MTurk-adjacent alternative.
 
    uv run python pages/platform_status.py
 
Exits 1 if any fetch fails, or if a page that is supposed to carry a dated
claim no longer does — that is the page going stale, not a network blip.
Prints FAILED on the failing check. No API keys.
"""
from __future__ import annotations
 
import re
import sys
import urllib.error
import urllib.request
 
UA = (
    "measuretheweb-platform-status/1.0 "
    "(https://measuretheweb.org/design/user_studies)"
)
 
CHECKS = [
    (
        "mturk-aup",
        "https://www.mturk.com/acceptable-use-policy",
        [
            r"September 30, 2026",
            r"permanently close",
        ],
    ),
    (
        "mturk-requester",
        "https://requester.mturk.com/",
        [
            r"no longer accepting new customers",
        ],
    ),
    (
        "prolific-pricing",
        "https://researcher-help.prolific.com/en/articles/445239-what-is-your-pricing",
        [
            "£6.00 / $8.00",
            "33.3% of participant rewards",
            "42.8% of participant rewards",
        ],
    ),
    (
        "prolific-pay",
        "https://researcher-help.prolific.com/en/articles/445266-how-much-should-i-pay-participants",
        [
            r"£6",
            r"$8",
            r"£9",
            r"$12",
        ],
    ),
    (
        "cloudresearch",
        "https://www.cloudresearch.com/",
        [
            r"CloudResearch",
        ],
    ),
]
 
 
def fetch(url: str) -> str:
    req = urllib.request.Request(url, headers={"User-Agent": UA})
    try:
        with urllib.request.urlopen(req, timeout=40) as resp:
            status = resp.status
            body = resp.read().decode("utf-8", errors="replace")
    except urllib.error.HTTPError as e:
        raise RuntimeError(f"{url}: HTTP {e.code}") from e
    except urllib.error.URLError as e:
        raise RuntimeError(f"{url}: {e.reason}") from e
    if status != 200:
        raise RuntimeError(f"{url}: HTTP {status}")
    return body
 
 
def strip_html(html: str) -> str:
    html = re.sub(r"(?is)<script[^>]*>.*?</script>", " ", html)
    html = re.sub(r"(?is)<style[^>]*>.*?</style>", " ", html)
    html = re.sub(r"(?s)<[^>]+>", " ", html)
    html = html.replace("&nbsp;", " ").replace("&amp;", "&")
    html = re.sub(r"\s+", " ", html)
    return html
 
 
def snippet(text: str, pat: str, width: int = 140) -> str:
    m = re.search(pat, text, flags=re.I)
    if m is None:
        return ""
    i = m.start()
    return text[max(0, i - 40) : i + width]
 
 
def main() -> int:
    failed = 0
    for name, url, needles in CHECKS:
        try:
            raw = fetch(url)
        except RuntimeError as e:
            print(f"FAILED  {name}  fetch  {e}")
            failed += 1
            continue
        # Search the raw HTML as well as the stripped text. Prolific's
        # Intercom help pages keep the rates inside JSON/HTML entities that
        # a naive tag-strip can scramble, and a regex needle containing '$'
        # is end-of-string, not a dollar sign — search as a literal.
        text = strip_html(raw)
        blob = raw + "\n" + text
        missing = [n for n in needles if n not in blob]
        if missing:
            print(f"FAILED  {name}  missing {missing}  {url}")
            print(f"         first 200 stripped chars: {text[:200]!r}")
            failed += 1
            continue
        print(f"ok      {name}  {url}")
        for n in needles:
            src = blob if n in blob else text
            print(f"         {n!r} -> {snippet(src, re.escape(n))!r}")
    if failed:
        print(f"FAILED  {failed} of {len(CHECKS)} checks")
        return 1
    print(f"ok      {len(CHECKS)}/{len(CHECKS)} checks")
    return 0
 
 
if __name__ == "__main__":
    sys.exit(main())

Output of uv run python pages/platform_status.py, run 2026-08-27 (exits 0, 5/5):

ok      mturk-aup  https://www.mturk.com/acceptable-use-policy
         'September 30, 2026' -> 'er-headline">Mechanical Turk will close September 30, 2026.</div>\n      <div class="maintenance-banner-body">We regularly evaluate our programs, tools, and services and make adjust'
         'permanently close' -> 'r 30, 2026. Amazon Mechanical Turk will permanently close on September 30, 2026. For Workers and Requesters currently using the service, please visit our <a href="/help">FAQ</a> pa'
ok      mturk-requester  https://requester.mturk.com/
         'no longer accepting new customers' -> '9888;&#xFE0F; Amazon Mechanical Turk is no longer accepting new customers. We recommend existing customers migrate to a third-party solution. More information is available <a href='
ok      prolific-pricing  https://researcher-help.prolific.com/en/articles/445239-what-is-your-pricing
         '£6.00 / $8.00' -> 'ft"><p><b>Absolute minimum allowed:</b> £6.00 / $8.00 per hour</p></div></li></ul></div><div class="intercom-interblocks-paragraph no-margin intercom-interblocks-align-left"><p>You'
         '33.3% of participant rewards' -> 'rticipant rewards (corporate customers) 33.3% of participant rewards (academia and non-profits) Example: If you pay $70.00 to participants, your total cost will be $100 (corporate)'
         '42.8% of participant rewards' -> 'ercentage on top of participant rewards 42.8% of participant rewards (corporate customers) 33.3% of participant rewards (academia and non-profits) Example: If you pay $70.00 to par'
ok      prolific-pay  https://researcher-help.prolific.com/en/articles/445266-how-much-should-i-pay-participants
         '£6' -> '><p>We enforce a minimum hourly rate of £6 / $8, and recommend paying at least £9 / $12 per hour. However, the right reward for your study depends on several factors, including the'
         '$8' -> 'e enforce a minimum hourly rate of £6 / $8, and recommend paying at least £9 / $12 per hour. However, the right reward for your study depends on several factors, including the effo'
         '£9' -> ' £6 / $8, and recommend paying at least £9 / $12 per hour. However, the right reward for your study depends on several factors, including the effort required, the participants you’'
         '$12' -> ' $8, and recommend paying at least £9 / $12 per hour. However, the right reward for your study depends on several factors, including the effort required, the participants you’re ta'
ok      cloudresearch  https://www.cloudresearch.com/
         'CloudResearch' -> 'application/xml" title="Sitemap"><title>CloudResearch: Survey Platform & Participant Recruitment</title><link rel="canonical" href="https://www.cloudresearch.com/"><meta name="desc'
ok      5/5 checks

Methodology and limitations of these figures

Population: participants[] non-empty on data/extract/run1/extractions.jsonl, 5,859 papers, seven venues, 2010–2026. That is OVERVIEW.md's humanSubjects. Figures are paper-counts, never tuples. Sentinels are never answers. Recruitment and study-kind are schema enums; crowd-platform names are recovered with an ordered regex family list in scripts/us_fold.mjs (multi-label) from evidence quotes and from paper.cols.txt. Full-text recovery is an upper bound because a related-work citation of Mechanical Turk fires the same regex. Compensation is a single-label fold over free text; 139 tuples did not map. Regions go through scripts/geo.mjs. The 18 papers with n≥10,000 are a hand map keyed by paper id; the report fails if the map and the set diverge. Poster sensitivity: 35 records, ranking unchanged. 4 corpus-wide papers lack paper.cols.txt, 1 of them in humanSubjects. Every published figure was re-derived on this sitting; none was carried from the item brief (which still says 965 / 53.0% / 47.8% from the 4,322-paper corpus).

The working log — every query, the unedited report, the 25-paper unnamed other-crowd residue, the quote checks, the sources rejected, and the review log — is user_studies. Corpus-level caveats: Corpus.

  • Automated measurements — the other top-level Design branch. 149 papers sit in both.
  • Ethics — review outcomes, venue ethics bodies, dead donation panels (Rally, Citizen Browser). The 75.3% / 66.0% / 24.7% split on this page is that page's human-subjects column.
  • Study preregistration — already normal for participant studies; almost absent for crawls.
  • Hypothesis testing, Pvalue corrections, Regression — the tests this population actually runs.
  • Interrater agreement — crowdworkers as labelers, not as participants.
  • Fingerprinting — uniqueness is why you cannot skip humans; Berke et al. is on both pages.
  • Conferences — SOUPS and CHI are where a usable-security paper goes; they are absent here. PoPETs 42.2% vs IMC 10.5% is that page's pick-your-venue table.
  • Corpus — seven venues, the funnel, 2025–2026 provisional.
[1]
Wei, Miranda; Mink, Jaron; Eiger, Yael; Kohno, Tadayoshi; Redmiles, Elissa M.; Roesner, Franziska (2024): "SoK (or SoLK?): On the Quantitative Study of Sociodemographic Factors and Computer Security Behaviors", in: Proceedings of the USENIX Security Symposium. (Link)
[2]
Klemmer, Jan H.; Schmüser, Juliane; Lowens, Byron M.; Fischer, Fabian; Schmüser, Lea; Schaub, Florian; Fahl, Sascha (2025): "Transparency in Usable Privacy and Security Research: Scholars' Perspectives, Practices, and Recommendations", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[3]
Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
[4]
Berke, Alex; Calacci, Dan; Mahari, Robert; Yabe, Takahiro; Larson, Kent; Pentland, Sandy (2025): "How Unique is Whose Web Browser? The Role of Demographics in Browser Fingerprinting Among US Users", Proceedings on Privacy Enhancing Technologies 2025(1):720-758. (DOI)
[5]
Redmiles, Elissa M.; Warford, Noel; Jayanti, Amritha; Koneru, Aravind; Kross, Sean; Morales, Miraida; Stevens, Rock; Mazurek, Michelle L. (2020): "A Comprehensive Quality Evaluation of Security and Privacy Advice on the Web", in: Proceedings of the USENIX Security Symposium. (Link)
[6]
Kelley, Patrick Gage; Komanduri, Saranga; Mazurek, Michelle L.; Shay, Richard; Vidas, Timothy; Bauer, Lujo; Christin, Nicolas; Cranor, Lorrie Faith; L´opez, Julio C. (2012): "Guess Again (and Again and Again): Measuring Password Strength by Simulating Password-Cracking Algorithms", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[7]
Pei, Weiping; Mayer, Arthur; Tu, Kaylynn; Yue, Chuan (2020): "Attention Please: Your Attention Check Questions in Survey Studies Can Be Automatically Answered", in: Proceedings of the ACM Web Conference. (DOI)
[8]
Herbert, Franziska; Munyendo, Collins W.; Hielscher, Jonas; Becker, Steffen; Zou, Yixin (2025): "Digital Security Perceptions and Practices Around the World: A WEIRD versus Non-WEIRD Comparison", in: Proceedings of the USENIX Security Symposium. (Link)
[9]
Park, Sungkyu; Park, Jamie Yejean; Chin, Hyojin; Kang, Jeong-han; Cha, Meeyoung (2021): "An Experimental Study to Understand User Experience and Perception Bias Occurred by Fact-checking Messages", in: Proceedings of the ACM Web Conference. (DOI)
[10]
Utz, Christine; Degeling, Martin; Fahl, Sascha; Schaub, Florian; Holz, Thorsten (2019): "(Un)informed Consent: Studying GDPR Consent Notices in the Field", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[11]
Ramulu, Harshini Sri; Schmitt, Helen; Rerich, Bogdan; Rodriguez, Rachel Gonzalez; Kohno, Tadayoshi; Acar, Yasemin (2025): "Ethics in Computer Security Research: A Data-Driven Assessment of the Past, the Present, and the Possible Future", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
1)
Amazon Mechanical Turk, Acceptable Use Policy. Checked 2026-08-27.
2)
Amazon Mechanical Turk, requester.mturk.com. Checked 2026-08-27.
You could leave a comment if you were logged in.
design/user_studies.1787840639.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki