User Tools

Site Tools


provenance:programming:registration

Provenance: Programming:Registration

Back to Automating Login and Registration. Corpus-wide selection and extraction notes are on corpus. This is the page-specific query log.

Run record

  • Run date: 2026-08-27 (UTC).
  • Authoring agent: Cursor Grok 4.6, executing the drain item programming:registration (improve; 2.5 KB stub) directly rather than via a headless claude -p session. Review: three focused GPT 5.6 Sol medium passes on the frozen first draft (findings applied below), then one generic pass with no checklist after those fixes (see Review log).
  • Corpus at run time: 5,859 extracted papers, 2010–2026, CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P. Read-only inputs under /workspace/publications_dataset/data/.
  • Read first: data/extract/OVERVIEW.md, the task spec, live programming:registration stub (rev 1739892438, 2,544 bytes), live programming:interaction (rev 1787833760 — underway, not edited), Ethics, Shepherd MADWeb PDF, Kubicek WWW 2024 author PDF.
  • This is an extension of the stub. Title kept. No write to the publication mount. Wiki saves through scripts/dw.mjs with –if-rev on existing pages.

Why this page, not an overlap

Neighbour What it already answers What this page takes
interaction depth, typing vs fill(), schema 81/857 published as “got past a login” Shepherd, SSO-as-instrument, automated registration, newsletter, CAPTCHA-as-blocker, ToS of accounts. Corrects 81 → 17 without editing the neighbour
ethics 92/697 test accounts; 17/697 robots.txt/ToS/AUP; 129/1,120 mention ToS cites those figures, does not re-derive
stateful_stateless the profile login writes not session vs stateless

The stub named Shepherd, Dimova 2023, Ghasemisharif 2018, Zhou 2014, Cookie Hunter, Kubicek 2024, Englehardt 2018, Mathur, Senol 2022, Chatzimpyrros 2019. Zhou 2014 is SSOScan. Mathur is the 2023 Big Data & Society version of the 2020 political-email project. Chatzimpyrros is LNCS 2020. Shepherd / Mathur / Chatzimpyrros are out of corpus and labelled so.

Population and queries

All counts are papers unless labelled otherwise. Sentinels are not answers.

Query Denominator Result
POPULATIONS.webCrawled 5,859 857
POPULATIONS.crawled 5,859 1,120
authentication stated (non-sentinel) 857 web 634 (74.0%)
same 1,120 crawled 779 (69.6%) — OVERVIEW.md row
schema logged (four non-none values) 857 81 (9.5%)
schema sso 857 web 0
schema sso 1,120 crawled 1 (DarkFleece, mobile)
ROLE over the 81 81 81/81; missing/stale key → FAILURE and exit 1
INSTRUMENT = scale-register ∪ scale-login 857 17 (2.0%)
SCHEMA_MISS (closed list) 4 papers, not unioned into 81
Kubicek WWW 2024 in extractions.jsonl NO. In .meta with DOI 10.1145/3589334.3645709. No paper.cols.txt

Folding and residue

ROLE is a single-label hand map, not a regex fold. Residue of the 81 is 0 by construction: the report refuses to finish if the map and the schema set diverge. Tally: scale-register 8, scale-login 9, manual-session 9, lab-scanner 14, sso-study 2, newsletter 1, marketplace 6, social 10, other 11, mislabel 11.

Rejected probes (not published as counts): a loose register sweep (domain registration, “we did not register”, related work); lowercase shepherd (homograph). The capital-S probe hits 17 papers and was read, not counted as a tool. SCHEMA_MISS is not unioned into the 81.

Evidence quotes checked

scripts/verify_registration_figures.mjs: 63 needles, 0 missing, 6 shorter than 20 characters (flagged). Against paper.cols.txt, not crawlConfig.evidence.quote (one quote for the whole object).

Load-bearing pairings that were easy to get wrong:

  • PETS 2022 59% / 5.5% from the body paragraph (of websites that sent mail / unsolicited marketing with no confirmation), not the abstract.
  • Login-policies 39% from ground-truth manual accounts, not of 45.0K.
  • Kubicek 25.7% (169,765) is of 660,202 unique domains, not of 504,509 loaded (that ratio is 33.7%). The PDF writes the 25.7% “among the loaded websites”.

External and industry verification

Source Load-bearing fact Verification 2026-08-27 Decision
Cloudflare Turnstile index.md CAPTCHA-free; Managed / Non-interactive / Invisible; WCAG 2.2 AA; updated 2026-08-14 fetched Used
Turnstile widget index.md challenges.cloudflare.com; updated 2026-04-16 fetched Used
Google reCAPTCHA docs landing deprecated; replacement is Cloud Fraud Defense landing HTTP 200 + “deprecated”; cloud.google.com/recaptcha/docs/overview HTTP 200 + “Fraud Defense” Used; do not say only “still serves 200”
AZcaptcha docs 2023 papers: OCR not humans azcaptcha.com/docs HTTP 200; live text mentions a workers pool Qualify as paper-era
hCaptcha docs still live HTTP 200 Used; HTTP 200, not a share
Shepherd PDFs + Crossref 10.14722/madweb.2020.23008 7,113; 2,759; ≥97% HTTP 200; title match Used; out of corpus
Kubicek author PDF + access page 660,202; 5.9%; 75/20/2/3; farm HTTP 200 Used; extraction-gap
Tranco 82Q3V June 2022 list HTTP 200 Used
Crossref 10.1007/978-3-030-42051-2_7 Chatzimpyrros title match Used; out of corpus
Crossref 10.1177/20539517221145371 Mathur 300K emails title match Used; out of corpus; journal 2023
BugMeNot Shepherd's credential source HTTP 403 from this IP Recorded, not treated as gone

Rejected: Similarweb / wmtips / SEO CAPTCHA-share pages.

What the corpus does not establish

  • A 2026 CAPTCHA-provider mix including Turnstile.
  • How often a Shepherd-style verifier disagrees with “the login form disappeared”.
  • A hand split of the 129 ToS-mentioning crawling papers.
  • Kubicek et al. inside the extraction.

Judgement calls

  • Publish 17, not 81, as the instrument figure; keep 81 visible because the neighbour already printed it. The 17 is a hand split of the 81, not a census of the 776 papers the schema did not mark as logged in.
  • Do not edit programming:interaction in this sitting.
  • Cite Kubicek from the PDF rather than omit the largest registration crawl.
  • Date human CAPTCHA farms; do not present them as current best practice.
  • login_diff.py is Shepherd's last step only (differential markers; token equality, not substring). It is not a login crawler.

Bibliography

15 new keys appended after a fresh live export. Reused: drakonakis2020_cookie, englehardt2018_email, senol2022_leaky, rautenstrauch2024_auth, kaizer2016_characterizing. Not used: ghasemisharif2023_bytecode, dimova2021cname, mathur2019_scale (wrong papers).

Report output

Unedited output of node scripts/report_registration.mjs (2026-08-27). The ROLE check printed OK and the process exited 0. Section I (full-text probes and the capital-S Shepherd hit list) is in scripts/report_registration-output.txt; it is an upper bound and was not published as a count.

==============================================================================
A. POPULATION
==============================================================================
corpus                                              5859
crawled  (crawlConfig OR automated-web-crawl)       1120
webCrawled  (crawled AND platforms includes web)    857   <- this page's denominator
  crawled but NOT web                               263

==============================================================================
B. crawlConfig.authentication — schema, of web crawls
==============================================================================
none 553 (64.5%); not-stated 204 (23.8%); account-registration 42 (4.9%);
manual-login 24 (2.8%); (no crawlConfig object) 19 (2.2%); automated-login 15 (1.8%).
stated (non-sentinel): 634 / 857 = 74.0%
schema "logged in of any kind": 81 / 857 = 9.5%

==============================================================================
C. same field, of all crawled papers (OVERVIEW.md row)
==============================================================================
stated: 779 / 1120 = 69.6%
sso among crawled: 1  (0 among webCrawled)
  USENIX/2024/darkfleece-probing-the-dark-side-of-android-subscription-apps  platforms=["mobile"]

==============================================================================
D. ROLE hand map over the 81
==============================================================================
scale-register 8 (9.9%); scale-login 9 (11.1%); manual-session 9 (11.1%);
lab-scanner 14 (17.3%); sso-study 2 (2.5%); newsletter 1 (1.2%);
marketplace 6 (7.4%); social 10 (12.3%); other 11 (13.6%); mislabel 11 (13.6%).
INSTRUMENT (scale-register ∪ scale-login): 17 / 857 = 2.0%
  of the 81 schema hits: 17 / 81 = 21.0%

==============================================================================
G. SCHEMA_MISS — closed list, not a sweep
==============================================================================
PETS/2018/i-never-signed-up-for-this-privacy-implications-of-email-tracking  auth=none  role=newsletter
USENIX/2018/o-single-sign-off-where-art-thou-an-empirical-analysis-of-single-sign-on-account  auth=none  role=sso-study
USENIX/2022/leaky-forms-a-study-of-email-and-password-exfiltration-before-form-submission  auth=none  role=form-nosubmit
PETS/2023/everybodys-looking-for-ssomething-a-large-scale-evaluation-on-the-privacy-of-oau  auth=none  role=sso-study
SCHEMA_MISS size: 4

==============================================================================
H. Kubicek WWW 2024 — in the bibliographic index, not in the extraction
==============================================================================
in extractions.jsonl: NO
paper.cols.txt: ABSENT. Cited from people.inf.ethz.ch/basin/pubs/www24.pdf
figures (author PDF, 2026-08-27): 660,202 unique; loaded 504,509;
  form on 25.7% (169,765) of 660,202 unique domains (PDF writes "among the loaded websites";
  169,765/504,509 loaded = 33.7%).

OK — ROLE covers the schema set both ways; SCHEMA_MISS keys exist and are not schema-logged.

Review log

Three focused passes on GPT 5.6 Sol medium, as directed for this sitting (the task spec's sonnet/fable split was overridden). They ran in parallel on the frozen first draft (out/freeze_registration/). The content page was not edited while those reviewers were running. Findings were applied, then a generic pass with no checklist on out/freeze_registration_generic/ (agent 5a44afde-2a4d-4fa4-85cb-33c2da9e4358). Every accept/reject is below. No GENERIC_REVIEW placeholder.

GPT 5.6 Sol medium — figures versus script

Finding Disposition
Cookie Hunter table and report joined 25,242 with 13.7% of 168,594. Arithmetic is 15.0%. The paper's 13.7% is registered-and-logged-in of signup domains; 25,242 is a separate accounts-created count they call “almost 12%”. Accepted. Split the table into three rows. Report Z no longer writes “13.7% → 25,242”.
Kubicek “form on 25.7% (169,765) of loaded” is the unique-domain rate: 169,765/660,202 = 25.7%; 169,765/504,509 = 33.7%. The PDF writes the 25.7% “among the loaded websites”. Accepted. Page and report now name the 660,202 denominator and the 33.7% loaded ratio.
login_diff.py substring false positives (accounting, gravatar, logout-policy). Accepted. Path-segment equality plus -link/-btn suffixes. New self-tests; documented run re-pasted.
857, 81, 17, 2.0%, ROLE shares, year buckets, PETS 2022 59%/5.5%, login-policies 39% of ground-truth. Accepted, no change.

GPT 5.6 Sol medium — citations and quotes

Finding Disposition
Cookie Hunter rate conflates two results (same as figures pass). Accepted (same split).
The Cookie Hunter farm-refusal quote had no local [1Drakonakis, Kostas; Ioannidis, Sotiris; Polakis, Jason (2020): "The Cookie Hunter: Automated Black-box Auditing for Web Authentication and Authorization Flaws", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)]. Accepted. Citekey on the quote paragraph.
“Tripwire yes” under human CAPTCHA farms. Paper used DeCaptcher; it does not establish that the solvers were human. Accepted. Tripwire is a third-party solver, not a named human farm.
All 20 citation keys unique; quotes otherwise verbatim. Accepted, no change.

GPT 5.6 Sol medium — external currency

Finding Disposition
developers.google.com/recaptcha HTTP 200 but the page is deprecated in favour of Google Cloud Fraud Defense. Accepted. Page names the deprecation and the replacement URL. Probe checks both.
AZcaptcha “OCR not humans” is a 2023 paper claim; live docs mention a workers pool. Accepted. Dated to those papers; live docs recorded.
Turnstile, hCaptcha, Shepherd/Kubicek PDFs, Tranco 82Q3V, Crossref records. Accepted, no change.

GPT 5.6 Sol medium — generic

Ran after the focused findings were applied. No checklist. The content page was not edited while this reviewer ran.

Finding Disposition
17/857 is presented as corpus prevalence; only the 81 schema hits were hand-read. Four SCHEMA_MISS papers prove false negatives exist. Accepted as a qualification. Intro, WRAP, and methodology now say the 17 is a hand split of the 81, not a census of the 776 unlabelled papers. The published 17 is unchanged.
“Logged-out is a design” overstates: the schema records labels, not motives, and the quote cannot be checked. Accepted. Opening now: default practice; 553 are labelled none.
“There is a consensus that almost nobody reads the ToS of 45.0K sites.” Five papers cannot establish that. Accepted. Now: those papers wrote they did not check; that is five papers, not a vote of 45.0K.
aria-label=“Log out” is not a marker: segments() does not collapse spaces to the keyword log-out. Accepted. Whole-attribute canonicalisation; eighth self-test.
–quotes without a regex is skipped; Shepherd DOI HTTP status is printed and not asserted. Accepted. Bare –quotes now exits 1. DOI must be HTTP 200.
Open-questions box uses a WRAP todo tag the reviewer called lowercase. Rejected. The tag is the required uppercase form; check_wrap.mjs passed. The defect that guard catches is the lowercase wrap tag.
Review log promised a generic heading that was not yet written. Accepted. This heading is that log.

References

[1]
Drakonakis, Kostas; Ioannidis, Sotiris; Polakis, Jason (2020): "The Cookie Hunter: Automated Black-box Auditing for Web Authentication and Authorization Flaws", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
provenance/programming/registration.txt · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki