User Tools

Site Tools


provenance:design:mobile_and_app_measurement:mini_programs

Provenance: Measuring Super-App Mini-Programs

Working notes behind mini_programs: every probe with its population, the hand audit that produced the 16-paper population, the hand codes behind every “N of 16” on the page, the quote and figure checks, the external sources with fetch dates and the ones rejected, and the decisions taken along the way. Corpus-level caveats — the seven-venue scope, the provisional 2025–2026 slice, extraction stability — are on corpus and are not repeated here.

This page is a log, not prose. It is for somebody checking a number.

The run

Item Value
Date 2026-09-27
Corpus data/extract/run1/extractions.jsonl, 5,859 extracted papers, 5,855 with paper.cols.txt (the 2026-08-11 extension, commit 8a6b843). The ~20% free-text stability and 0.9% unlocatable-quote figures on corpus were measured on the older 4,322-paper run; this page uses neither.
Venues CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P, 2010–2026
Scope check sitemap.mjs on 2026-09-27: 195 pages, 0 promised-but-missing, no page on mini-programs or super apps. mobile_and_app_measurement (rev 1789266443, 58,332 B) contains none of “mini”, “WeChat”, “super app”, “Alipay”, “app-in-app”, so there was no overlap to replace with a link, and the parent had no child list. Created a new child page and added a child pointer to the parent.
Report script scripts/report_mini_programs.mjs + scripts/mp_probes.mjs (probes) + scripts/mp_fold.mjs (rule, verdicts, hand codes, papers outside the extraction) — all three and the output are below, unedited
Figure and quote verifier scripts/verify_mini_programs_figures.mjs — output below
External checks scripts/external_checks_mini_programs.sh — script and output below; exit status is the number of FAILED checks
Number guard check_page_numbers.mjs whole-page against the concatenation of the report, verifier and external-check outputs: every figure traces. That guard is a presence check, not a binding check; figures were also read against the output by hand and by the figures reviewer.
Bibliography 27 entries appended to bibliography in two saves (25 at rev 1790526095, 2 Telegram Mini Apps preprints at rev 1790527209 after the external-currency review; 1,162 → 1,189 entries): 19 from bibgen.mjs over the corpus index, 8 hand-written from Crossref and arXiv records for papers outside the index; USENIX and PoPETs author lists from the landing pages' citation_author metadata and author line; bib_dedup_scan.py over the merged file: 0 same-DOI or same-title pairs, 30 same-surname-same-year candidates involving a new key, all 30 different papers by title; no literal @ inside a field
Agents One Opus session (queries, triage of 41 candidates, verdicts, hand codes, drafting, verification). Four Sonnet readers read the 17 papers that looked like mini-program studies (16 candidates + 1 recall hit) into unpublished notes (brief below). One Opus sub-agent did the external-source pass (about 30 verified claims, notes unpublished); every load-bearing claim on the page was then re-fetched by the external-check script. Six review sub-agents in three rounds — three focused Sonnet passes, a Sonnet re-check of their fixes, a generic Fable pass, and a Sonnet re-check of its fixes — see Review log.

Mistakes made in this run

  • Reader quotes were close but not verbatim, and the verifier caught all of them. Four page spans and twelve needles copied from the reading notes failed the first verification: “we identified fine-grained browsing data in 89.7%” (the paper says “we also identified browsing data in 89.7%”), “bias results towards false negatives” (the paper: “These limitation biases results”), a sentence ending in a full stop where the paper continues with a comma, “super-app” where the paper prints “superapp”, “TikTok” where it prints “Tiktok”, “miniCrawler” for “mini-Crawler”, and a lower-case “we” at a sentence start. Every one was rewritten against paper.cols.txt or the pypdf rendering; the notes' own machine check (below) located 273 of 428 note quotes (63.8%) verbatim, which is why the notes stay unpublished.
  • The first bibgen.mjs run reported all 16 corpus papers “NOT IN INDEX” because the key list was passed as one quoted shell argument. Re-run with one argument per key.
  • A duplicate entry was one step from the bibliography. bibgen.mjs generated zhou2025_your for the IEEE S&P 2025 secrets paper; the DOI scan found it already present as zhou2025_secrets. Dropped; the page cites the existing key. chen2026_when collided with an existing key for a different paper and was renamed chen2026_minigames.
  • Seven claims in the first draft overstated the evidence and were corrected before any review: “three crawling routes, two of which the host has since closed or changed” (only MiniCrawler's is reported closed; the search endpoint's current state is unknown); the malware corpus “took two and a half years” (it was collected March 2020 – June 2022, revisited to December 2022); “delisted within about six months” (delisting was measured at the end of 2022 over miniapps collected since 2020); “7 of the 9 deployed-population papers state no host version” (9 of 9 once two n/a codes were recoded not-stated); a language label on two unpackers that no source gave; a sentence about author clustering that no query measured; and “before a single flow was readable”.
  • One external check asserted nothing on its first version. “Tencent publishes no mini-program count” was checked against the Q2 2026 release, which contains zero mentions of “Mini Program” — so “no count next to a mention” was vacuously true. The script now also reads the 2025 annual and Q1 2026 PDFs and requires a positive control (the Weixin/WeChat MAU line) to be present in each before accepting “no count”; the annual release does mention “content-related Mini Programs” twice, with no number.
  • Four hand codes were wrong until review (see Review log): [1Shi, Yizhe; Yang, Zhemin; Liu, Dingyi; Zhong, Kangwei; Dai, Jiarun; Yang, Min (2026): "Better Safe than Sorry: Uncovering the Insecure Resource Management in App-in-App Cloud Services", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] (figures review) and [2Yang, Yuqing; Zhang, Yue; Lin, Zhiqiang (2025): "Understanding Miniapp Malware: Identification, Dissection, and Characterization", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] (generic review) were coded as acknowledging that their sample is not the population, which neither text says — both had caveats about the detector, not the crawl; the WRAP box's “six of the nine” is four; [3Wang, Mona; Lin, Pellaeon; Knockel, Jeffrey; Greenberg, Will; Mayer, Jonathan; Mittal, Prateek (2025): "What WeChat Knows: Pervasive First-Party Tracking in a Billion-User Super-App Ecosystem", in: Proceedings on Privacy Enhancing Technologies. (DOI)] was coded as using Chinese seed keywords, which it attributes to a cited prior method; [4Cai, Yifeng; Zhang, Ziqi; Yao, Mengyu; Liu, Junlin; Zhao, Xiaoke; Fu, Xinyi; Li, Ruoyu; Liu, Zhe; Chen, Xiangqun; Guo, Yao; Li, Ding (2025): "I Can Tell Your Secrets: Inferring Privacy Attributes from Mini-app Interaction History in Super-apps", in: Proceedings of the USENIX Security Symposium. (Link)] was coded “none-stated” for artefacts although it discusses and declines release (new code declined). Only the first moved a printed figure.
  • A quote hid its own citation. The page quoted [5Liu, Kaizheng; Yang, Ming; Ling, Zhen; Zhang, Yue; Lei, Chongqing; Luo, Junzhou; Fu, Xinwen (2024): "RIoTFuzzer: Companion App Assisted Remote Fuzzing for Detecting Vulnerabilities in IoT Devices", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)]'s “most mini-apps employ obfuscation techniques” without the “as reported in [45]” that precedes it; [45] is [6Zhang, Yue; Turkistani, Bayan; Yang, Allen Yuqing; Zuo, Chaoshun; Lin, Zhiqiang (2021): "A Measurement Study of WeChat Mini-Apps", Proceedings of the ACM on Measurement and Analysis of Computing Systems 5(2). (DOI)]. Caught by the citations reviewer; the sentence now credits the measurement to its source.
  • Two readers missed released code. KeySentinel ([7Zhou, Jiawei; Zhang, Zidong; Ying, Lingyun; Chai, Huajun; Cao, Jiuxin; Duan, Haixin (2025): "Hey, Your Secrets Leaked! Detecting and Characterizing Secret Leakage in the Wild", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]) and RIoTFuzzer ([5Liu, Kaizheng; Yang, Ming; Ling, Zhen; Zhang, Yue; Lei, Chongqing; Luo, Junzhou; Fu, Xinwen (2024): "RIoTFuzzer: Companion App Assisted Remote Fuzzing for Detecting Vulnerabilities in IoT Devices", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)]) print their repository URL only in the reference list; the readers coded “not stated”. The report's hand-versus-schema cross-check (section G) flagged both; verified in the text and recoded.

No credential or token was printed, logged or exposed. GH_TOKEN is referenced in the external script only as ${GH_TOKEN:+x} and inside an Authorization header.

No ~~DISCUSSION~~ block on this page, following the default set on ad_archives: comments belong on the content page.

Judgement calls

  • A child page, not a section of the parent. The parent is 58 KB on native-app acquisition and analysis and never mentions mini-programs; the population here (16 papers) has its own acquisition problem, its own tooling and its own privacy question (the host as first party), none of which the parent covers. Everything that transfers (Frida, pinning, rooted handsets) is linked, not repeated.
  • Core versus section counts the host framework as the object. Two readers coded [8Lu, Haoran; Xing, Luyi; Xiao, Yue; Zhang, Yifan; Liao, Xiaojing; Wang, XiaoFeng; Wang, Xueqiang (2020): "Demystifying Resource Management Risks in Emerging Mobile App-in-App Ecosystems", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] and [9Zhang, Lei; Zhang, Zhibo; Liu, Ancong; Cao, Yinzhi; Zhang, Xiaohan; Chen, Yanjun; Zhang, Yuan; Yang, Guangliang; Yang, Min (2022): "Identity Confusion in WebView-based Mobile App-in-app Ecosystems", in: Proceedings of the USENIX Security Symposium. (Link)] IN-section because their unit is a host API or a host binary rather than deployed mini-programs; the rule written first says “mini-programs, or the host's mini-program framework and its APIs”, so both are core. The distinction the readers were reaching for is kept, as the unit hand code (deployed / host-framework / host-traffic / usage-logs), and the page's first table is built on it. A reader who prefers “deployed mini-programs only” has 9 papers.
  • [5Liu, Kaizheng; Yang, Ming; Ling, Zhen; Zhang, Yue; Lei, Chongqing; Luo, Junzhou; Fu, Xinwen (2024): "RIoTFuzzer: Companion App Assisted Remote Fuzzing for Detecting Vulnerabilities in IoT Devices", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] is section, not core (reader: core). Its object is IoT devices; the mini-app layer is what the fuzzer works around. It is also the one IN paper the gap rule misses.
  • Medusa (USENIX 2023) is not added. The recall probe fired on “host app” and app names; the reader found no mini-program, super-app or host-name hit in the text.
  • Four CONTEXT papers, not OUT. WeChat's in-app browser (two papers) and face verification service (one) are host features a mini-program measurement meets; GitHub mini-program source code (one) is a possible, unused population source. None is counted.
  • Papers outside the extraction are cited, never counted. [10Wei, Zhiao; Wang, Chao; Faheem, Haseeb-Ur-Rehman; Xing, Luyi; Aafer, Yousra; Lin, Zhiqiang (2026): "Raising the Flag: Detecting Missing Permission Controls in Mini-Program APIs", in: Proceedings of the USENIX Security Symposium. (Link)] was read in full from the publisher PDF because it is the most recent host-API paper and the only 2026 use of Xposed; [11Shi, Yizhe; Yang, Zhemin; Yang, Yifan; Yang, Yunteng; Yang, Min (2026): "Convenience at a Cost: the Security Risks of Template-Based Development in the App-in-App Ecosystem", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] and [12Yang, Yuqing; Lin, Zhiqiang (2026): "Real or Rogue? Detecting Malicious Miniapps with Deceptive Reporting Interface", in: Proceedings of the ACM Web Conference. (DOI)] are cited from abstracts only. Every “N of 16” excludes all three.
  • Xposed is “still used here”, against the parent page's “superseded”. The parent dates Xposed for general app hooking; in this sub-field it is what MiniCrawler requires and what a 2026 paper used to hook four hosts' bridge classes. The two statements are about different uses and both stand.
  • MiniCrawler is “historical as published” although a 2026 paper names it. The code has not changed since 2021-06-29, pins a 2020-era WeChat, and MiniCAT (2024) reports its batch appid route blocked; [13Chen, Pei; Hong, Geng; Qin, Yicheng; Wang, Huazhe; Wu, Mengying; Yang, Min; Zhao, Ziru; Zhu, Yuanpeng; Su, Tao (2026): "When Fun Turns Toxic: A First Look at Aggressive Advertising in Mini-games", in: Proceedings of the USENIX Security Symposium. (Link)] names it for a crawl it does not date. The page says exactly that rather than calling the method dead or current.
  • No significance test on “3 of 16 state an ethics approval” against 33.8%. With 16 papers the comparison is printed as counts and the page says so.
  • The unit “deployed” includes [14He, Yi; Guan, Yunchao; Lun, Ruoyu; Song, Shangru; Guo, Zhihao; Zhuge, Jianwei; Chen, Jianjun; Wei, Qiang; Wu, Zehui; Yu, Miao; Shi, Hetian; Li, Qi (2024): "Demystifying the Security Implications in IoT Device Rental Services", in: Proceedings of the USENIX Security Symposium. (Link)] (75 rental mini-programs) and [7Zhou, Jiawei; Zhang, Zidong; Ying, Lingyun; Chai, Huajun; Cao, Jiuxin; Duan, Haixin (2025): "Hey, Your Secrets Leaked! Detecting and Characterizing Secret Leakage in the Wild", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] (41,719), both section papers; the deployed-population figures (“9 of 9 state no host version”, “2 of 9 describe Chinese keyword handling”) are over all nine.
  • Hand codes are single-coder. Verdicts and codes were decided by one session from the reader notes and the text; there is no second coder and no agreement figure. The page says so.

Queries, with populations and denominators

Denominators used

Figure on the page Denominator Where
37 papers, 21 from 2024–2026, under the gap rule 5,855 papers with full text report section A
57 candidates, 16 in the population (13 core, 3 section) the candidate set sections A–B
precision / recall of each probe the probe's own set; recall over the 16 section C
9 / 5 / 1 / 1 by unit; hosts; sides the 16 IN papers section D
acquisition, host version, keywords, validation, denominator for deployed studies the 9 deployed-population papers section E
instrumentation, LLM use, review, disclosure, artefacts the 16 sections E, Y
13 of 16 filed as mobile-apps; 13 of 16 in appAnalysis; 15 of 16 with no stated location the 16, read through the extraction schema section G
33.8% and 28.2% review-outcome baselines 5,118 empirical papers; 529 appAnalysis papers section G
“our arithmetic” on per-paper figures the two numbers the paper prints section Y

Probe 1: reproducing the gap pass

The task brief's count — “37 papers, 21 from 2024–2026” with “≥10 hits of /mini-?programs?|WeChat/” — reproduces exactly with that regex case-insensitive over whitespace-collapsed paper.cols.txt. Case-sensitive gives 33 / 18. The regex does not match “mini program” with a space, “mini-app”, “miniapp”, “super app” or “app-in-app”; of the 22 papers it returns that are not in the population, 19 contain none of the ecosystem's vocabulary at all (WeChat as a messenger, payment method, voice-login target or dataset owner).

Probe 2: the candidate set

Defined in mp_probes.mjs (below). A paper is a candidate if any of: (a) the gap rule; (b) the ecosystem vocabulary — mini[- ]?programs?, mini[- ]?apps?, mini[- ]?games?, super[- ]?apps?, app[- ]in[- ]app — at least three times combined; © any hit of the WeChat package and markup names (wxapkg, WXML, WXSS, wx.request / wx.login / wx.getUserInfo, JSSDK); (d) any hit of Telegram Mini Apps / Snap Minis / LINE MINI / Instant Games; (e) the vocabulary or WeChat/Alipay in the title, a population source or unit, a detection phenomenon, a used or produced tool, or a classifier resource or target. 57 papers; channel overlap in section A. The app[- ]in[- ]app family matches “In-App” purchases in Apple privacy-label papers (two homographs, both OUT).

Probe 3: recall outside the candidate set

scripts/_mp_recall.mjs: over the non-candidates, at least three hits of applet / light app / quick app / instant app / H5 app or page / lite app / host app / sub-app / JSBridge and at least three hits of a host name (WeChat, Alipay, Baidu, Douyin, TikTok, Taobao, QQ, Snapchat, Telegram, LINE, Grab, Gojek, Paytm, Kakao). 6 hits: five use “host app” for an app embedding an SDK, and Medusa (in-app QR scanning), which was read in full and is not about mini-programs. The probe is narrow; its recall is not measured.

Probe 4: the venue index

The report's section J reads all 16,864 records in data/corpus2/.meta and matches titles against the vocabulary: 18 titles, 12 in the extraction (all IN), 6 not — three workshop front-matter records (SaTS 2023–2025, screened out correctly), one paper selected but never retrieved (USENIX 2026, empty full-text directory), and two never screened because OpenAlex has no abstracts for IEEE S&P 2026 and TheWebConf 2026. The report throws if a new title appears here without a hand entry in OUTSIDE_EXTRACTION.

The hand audit

  • Inclusion rule, written before any verdict was counted: printed at the top of section B and in mp_fold.mjs.
  • 16 candidates read in full, plus one recall hit (READ_IN_FULL and RECALL_READ in the fold; every IN paper is among them, and the report fails if one is not). Four Sonnet readers, one brief (below), one markdown file each, every factual field with a verbatim quote. Three reader verdicts were overridden (see Judgement calls).
  • 41 candidates decided from sentence contexts: _mp_ctx.mjs printed, for every candidate, up to six contexts around the ecosystem vocabulary, falling back to WeChat/Alipay contexts when there were fewer than three. The dominant OUT codes — WeChat as messenger 10, related-work mention 9, the super app as an ordinary app 5 — are visible in the contexts without reading further. The risk in this step is a mini-program study whose contexts look like a mention; the vocabulary probe (which alone reaches all 16) and the title probe are the defences, and neither is a second reading.
  • Hand codes for the 16 (HAND in the fold): unit, hosts, side, acquisition route, collected and analysed counts, analysis kind, tools, instrumentation, host version stated, account type, Chinese keyword handling, validation, LLM use, ethics review, disclosure, artefacts, and whether the paper acknowledges its sample is not the population. “not-stated” means the paper does not say.
  • Hand versus schema (section G): tools[] names MiniCrawler for exactly the five papers the hand codes do; ethics.notifiedAffectedParties is yes for all 16, matching the hand disclosure codes; ethics.reviewOutcome = approved matches the 3 hand-coded approvals; the extraction reads the IEEE S&P 2025 secrets paper as “explicitly-discussed-no-review”, which its text (“Following ethical guidelines [65], all analyses were conducted locally”) does not say — the hand code stays “none-mentioned”. Artefacts disagreed on two papers until the readers' misses were fixed (see Mistakes made in this run).
  • Hand codes checked against the paper text, not the notes: all 9 denominator codes of the deployed-population papers (after two were found wrong), all 16 irb codes and all 16 artifacts codes (section K's probes plus the schema cross-check; every hit read), and every code a reviewer spot-checked — 41 of the 288 codes systematically, the rest by spot-check. Of the codes checked, 6 were wrong (2 denominator, 3 artifacts, 1 keywords); the other fields' codes (unit, hosts, acquisition, tools, validation) were taken from the reader notes and each is visible, with its paper, in the fold below. That error rate is why the page prints counts, not rates, for hand-coded fields.
  • The invariant: the report throws if a candidate lacks a verdict, a verdict names a non-candidate, a verdict note still says PENDING, an IN paper lacks hand codes, hand codes exist for a non-IN paper, a hand-code record lacks a field, or a venue-index title outside the extraction has no hand entry.

Quote and figure checks

  • Page quotes: verify_mini_programs_figures.mjs pulls every //"…"// span out of the page source and requires each to be located in the paper of a citekey on the same line (the publisher PDF text for [10Wei, Zhiao; Wang, Chao; Faheem, Haseeb-Ur-Rehman; Xing, Luyi; Aafer, Yousra; Lin, Zhiqiang (2026): "Raising the Flag: Detecting Missing Permission Controls in Mini-Program APIs", in: Proceedings of the USENIX Security Symposium. (Link)]), or to be an EXTERNAL span whose external check printed OK. Six spans are found only in the pypdf rendering, where .cols splices two columns across them. Four print a WARN because the nearest citekey on the line is a different paper; each was read and the quote is attributed in the sentence to the right one.
  • Per-paper figures: 101 needles, located in .cols or, for thirteen, only in the pypdf re-extraction. Three mutated needles (40,880→40,881; 41,726→41,727; 170→171) must not be found and are not. The needles under 20 characters are tool names whose presence is the claim.
  • One number is verified in the form the paper prints it: [2Yang, Yuqing; Zhang, Yue; Lin, Zhiqiang (2025): "Understanding Miniapp Malware: Identification, Dissection, and Characterization", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]'s “19, 905” carries a space inside the number in both renderings; the page writes 19,905.
  • NUL bytes: the MiniCAT .cols file contains 141 NUL bytes, which make grep/ugrep return nothing silently; the verifier and the notes check fold them out before matching.
  • Reader notes: mp_notes_quotecheck.py (output below) — 273 of 428 quoted spans located verbatim (63.8%); of the 155 misses, 109 have at least 90% of their word positions covered by 4-word windows present in the text (column splices and reader insertions), 35 between 60% and 90%, and 11 below 60% — those are the readers' own labels in quotation marks (“denominator acknowledged”, “unit = documented API”), not quotes. No page quote depends on an unlocated note quote.
  • Attribution guard: check_attributions.mjs found no “X et al., VENUE YEAR” on the page to check; the page names no authors in its text.

External sources

Every load-bearing external claim is re-fetched by external_checks_mini_programs.sh (script and output below); the output prints the matched context for each. The Alipay and Baidu developer documentation are single-page apps that needed Playwright to render; they are named on the page but no claim rests on their content.

Claim on the page Primary source
wxappUnpacker archived, last commit 2020-04-18; wux1an/wxapkg v2.0.0 (2026-04-16); wedecode newest tag v0.10.6, last commit 2026-08-27 GitHub API (/repos, /commits, /releases/latest, /tags sorted by version, since the API does not sort)
MiniCrawler last commit 2021-06-29; README pins WeChat 7.0.19 and 7.0.20 with Xposed GitHub API, raw README
TaintMini, MiniCAT and APIDiff decline to ship unpackers or crawlers “due to potential legal implications” raw READMEs
MiniSec gated datasets: institutional authentication, written consent, AppSecret set “requires additional consent and agreement” minimalware.github.io
WeChat: whitelisted domains, ICP filing, fixed Referer with appid and version, cloud hosting over the private protocol developers.weixin.qq.com network capability page
code2Session server-side only; error 40226 blocks high-risk users developers.weixin.qq.com server API page
privacy declaration enforcement 2023-09-15 postponed to 2023-10-17; undeclared APIs disabled developers.weixin.qq.com privacy-agreement guide
DevTools vs device difference for wx.getUserProfile (base library 2.10.4–2.16.1) developers.weixin.qq.com API page
MIIT filing notice (number, dates, distribution platforms incl. mini-programs); unfiled mini-programs inaccessible or delisted; filing system searchable by category miit.gov.cn notice; WeChat filing FAQ
WeChat licence clauses 8.2.1.2 and 8.2.1.4; international ToS last modified 2025-11-18 weixin.qq.com/agreement; wechat.com/en/service_terms.html
Telegram Mini Apps: web pages in Telegram; user data; “should not be trusted”; Bot API 8.0 on 2024-11-17 core.telegram.org/bots/webapps
TikTok mini games: code package; launched markets developers.tiktok.com mini-games overview and technical overview
Tencent publishes no mini-program count in 2025–2026 results Tencent 2025 annual and Q1 2026 result PDFs (static.www.tencent.com), Q2 2026 release (PR Newswire copy of Tencent's own text) — each with a positive control
SaTS 2026 at CCS 2026 with an “Agentic/LLM-based techniques” topic superappsec.github.io
SoK: arXiv 2306.07495, no venue on the record; CCS 2026 OAuth preprint accepted arxiv.org abstract pages
SIGMETRICS/POMACS 2021, RAID 2023, ICSE 2023, ASE 2023 records Crossref (title, container, date printed)
Raising the Flag figures (2,067 APIs, 183 (8.85%), four hosts, Xposed, about 55 USD) the USENIX PDF, extracted with pypdf
Real or Rogue? and Convenience at a Cost exist, with their abstracts Semantic Scholar API (the ACM DL answered 403 to curl and a Cloudflare challenge to Playwright; IEEE Xplore is closed access)

Sources rejected

  • Re-uploads of the unveilr unpacker (the original repository returns 404): unknown provenance, not recommended.
  • businessofapps.com, sqmagazine.co.uk, statista — mini-program and WeChat statistics aggregators; no primary source behind their figures. Tencent's own releases were used instead, and they give no count.
  • The ecosystem sizes the corpus papers quote (“about 4 million”, “more than 4.3 million”) cite a “Decoded” blog post and a WalktheChat post respectively; they are reported on the page as what the papers say, not as a measured size.
  • The SaTS 2026 call's “millions of miniapps” — the workshop's own framing, not a measurement.
  • WebFetch summaries — none used as a quote; every quoted string was matched in fetched bytes.

What could not be established

  • Whether the MIIT filing system can be queried at scale, and whether a record carries the appid. Not tested; it would be the only population frame that is not “what our search returned”.
  • Whether MiniCrawler or the reverse-engineered search endpoint still work against a current WeChat. Not tested; MiniCAT's 2024 report of the batch-appid block is the only evidence.
  • Whether WeChat's per-mini-program privacy declaration is fetchable without the developer's admin console.
  • The effective date of WeChat's 2022 avatar/nickname rule change: the announcement is behind a WeChat-scan login wall (curl and Playwright both redirected). The page gives no date.
  • What WeChat's developer terms say about decompiling third-party packages: the external pass found no reverse-engineering clause in the mini-program developer terms; the user licence clauses quoted on the page are about the client.
  • Whether TikTok's international app offers non-game mini-programs, and anything about Snapchat's.
  • Any official 2025–2026 count of mini-programs or their users.
  • A second coder's agreement on the verdicts and hand codes. Not done.

Follow-ups filed and changes to other pages

  • bibliography (rev 1790510451 → 1790526095): 25 entries appended before </bibtex>, 1,162 → 1,187 entries; ?purge=true issued before the page was saved.
  • mobile_and_app_measurement (rev 1789266443 → 1790526198): a one-paragraph child-page pointer after the scope paragraph, a Related Pages entry, and the Xposed row of Which Methods Are Current qualified (“superseded for general app hooking, not for hooking super-app bridges”). The parent had no mention of mini-programs, so nothing was moved out of it.
  • design (rev 1790510524 → 1790526200): table row, “19 pages” → 20, the child-page sentence now names three child pages, “Re-derived 2026-09-27”. report_namespace_overviews.mjs re-run: design 20 20 equal.
  • roadmap (rev 1790511917 → 1790526202): an Assessed row. roadmap (rev 1790511918 → 1790526204): a dated decision entry under 3g.
  • Nothing filed. The one piece of adjacent work the page points at — whether the regulator's filing system can serve as a population frame — is an open research question, not a wiki item.
  • This page and its content page, after review: content rev 1790526119 → 1790527251 (round 1) → 1790528159 (round 2) → final save below; provenance rev 1790526233 → 1790527253 → 1790528161 → final save below. The final revisions are recorded in notes/mp_log.md and on the drain item.

Review log

Four reviewers, each told that the author's context may not be exhaustive and handed the page source, the provenance page, the report script and its output, the verifier and external-check outputs. The three focused passes ran in parallel on frozen snapshots of rev 1790526119 (content) and rev 1790526233 (provenance) in out/mp/review/; fixes were applied together after all three returned. Their findings files are notes/mp_review_figures.md, notes/mp_review_citations.md, notes/mp_review_external.md. My own self-review, written while they ran, is notes/mp_selfreview.md.

Round 1: focused reviewers (Sonnet)

# Reviewer Finding Decision
F1 figures “the ecosystem vocabulary finds all 16 at twice the gap rule's precision” — the ratio is 53.3 / 40.5 = 1.32 Accepted. Now states both precisions; the report's Y block prints the ratio.
F2 figures [1Shi, Yizhe; Yang, Zhemin; Liu, Dingyi; Zhong, Kangwei; Dai, Jiarun; Yang, Min (2026): "Better Safe than Sorry: Uncovering the Insecure Resource Management in App-in-App Cloud Services", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] coded denominator: yes with no supporting sentence; “six of the nine” should be five Accepted. Recoded no; WRAP box now “Five of the nine”. I re-read the paper's discussion: its generality caveat is about hosts, not about the crawl covering the ecosystem.
F3 figures [3Wang, Mona; Lin, Pellaeon; Knockel, Jeffrey; Greenberg, Will; Mayer, Jonathan; Mittal, Prateek (2025): "What WeChat Knows: Pervasive First-Party Tracking in a Billion-User Super-App Ecosystem", in: Proceedings on Privacy Enhancing Technologies. (DOI)] keywords: Chinese rests on a sentence describing a cited method Accepted. Recoded not-stated; the field's definition now says “the paper's own seed terms”. No printed figure moved (the field is tallied over deployed-population papers only).
F4 figures [4Cai, Yifeng; Zhang, Ziqi; Yao, Mengyu; Liu, Junlin; Zhao, Xiaoke; Fu, Xinyi; Li, Ruoyu; Liu, Zhe; Chen, Xiangqun; Guo, Yao; Li, Ding (2025): "I Can Tell Your Secrets: Inferring Privacy Attributes from Mini-app Interaction History in Super-apps", in: Proceedings of the USENIX Security Symposium. (Link)] artifacts: none-stated although it discusses and declines release Accepted. New value declined; the report's artefact tallies and schema cross-check updated so declined still counts as no release (11 of 16 unchanged).
F5 figures three cells in the papers table drop a figure or host list the report has Accepted. Host lists and “1,031 documented APIs” filled in; 1,031 added as a verifier needle.
C1 citations the RIoTFuzzer obfuscation quote omits “as reported in [45]”, and [45] is [6Zhang, Yue; Turkistani, Bayan; Yang, Allen Yuqing; Zuo, Chaoshun; Lin, Zhiqiang (2021): "A Measurement Study of WeChat Mini-Apps", Proceedings of the ACM on Measurement and Analysis of Computing Systems 5(2). (DOI)] Accepted. Sentence credits [6Zhang, Yue; Turkistani, Bayan; Yang, Allen Yuqing; Zuo, Chaoshun; Lin, Zhiqiang (2021): "A Measurement Study of WeChat Mini-Apps", Proceedings of the ACM on Measurement and Analysis of Computing Systems 5(2). (DOI)]'s obfuscation-rate measurement (its abstract, re-fetched by the external script) and quotes RIoTFuzzer citing it. The span is column-spliced in .cols (“as” sits before a figure caption), so it is checked as a needle.
X1 external Ackites/KillWxapkg, the most-starred unpacker, is missing Accepted. Added to the tool table and the currency row as stalling (v2.4.1, 2024-09-20); both facts re-fetched by the script.
X2 external WeMinT's repository bundles wxappUnpacker, contradicting “research groups do not ship unpackers” Accepted. “Most research groups”, WeMinT named as the exception; checked by the script through the GitHub contents API.
X3 external two August 2026 arXiv preprints measure Telegram Mini Apps Accepted. Cited in Beyond WeChat, the currency table and Open Questions, with “not yet in any venue”; both abstract pages re-fetched by the script. The “no paper in these venues” claim stands and is now dated against them.
X4 external the SoK “no journal-ref” check cannot fail when the paper is published without the author updating arXiv Accepted. Replaced with an OpenAlex check that every indexed location is a repository.
X5 external get() never checked the HTTP status Accepted. get() now prints the status and FAILs on non-2xx.
X6 external the parent rates Xposed dead and LSPosed stalled, so “superseded for general hooking” implies a healthier contrast than the parent gives Accepted. Reworded to the parent's own verdicts.
S1–S5 self “every dynamic study hooks the client”; “75 of 81 reachable only as a mini-program”; “serves its corpora”; “every large corpus hooked the client”; “the same failure class as AppSecret leaks” Applied — each over-generalised a paper or a source; see notes/mp_selfreview.md.

Round 1 accepted all 12 reviewer findings and rejected none. The citations reviewer checked all 26 keys, the USENIX/PoPETs author lists and about 20 quotes independently and found one defect, which is what a clean pass looks like when the verifier has already run; the figures reviewer's F2 is the one that changed a headline number and could only be found by reading a paper, not by any guard.

Round 2: re-check of the fixes (Sonnet) and the generic review (Fable)

The re-check diffed the two snapshots, re-fetched every new external claim, mutation-tested the hardened get() and the OpenAlex check, and found 0 defects in the round-1 fixes (notes/mp_review_recheck.md). The generic review (notes/mp_review_generic.md) ran on the same snapshot (rev 1790527251 / 1790527253) and returned 20 findings, 6 marked major.

# Finding Decision
G1 (major) “recall almost nowhere … bounded at best by 100 unflagged” is contradicted by recall 85.56% (500 per host) and 83.55% (a labelled set) Accepted. Vulnerability paragraph, currency row and Open Question rewritten with the counts; recall figures added as needles.
G2 (major) [2Yang, Yuqing; Zhang, Yue; Lin, Zhiqiang (2025): "Understanding Miniapp Malware: Identification, Dissection, and Characterization", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] denominator: yes rests on detector caveats; “five of the nine” is four Accepted. Re-read: its “representative and generalizable” sentence is about countermeasures across hosts. Recoded; WRAP box “Four of the nine”; all nine denominator codes then re-checked against the text (see The hand audit).
G3 (major) “none of the 16 discusses the host's terms” — the 2026 mini-games paper asserts “compliance with platform policies” Accepted. Now “none quotes or analyses the licence”, with the three compliance assertions named; backed by section K's terms probe (every hit read).
G4 (major) “WeChat resists emulators / Unreliable” contradicts the page's own Android 14 emulator paper and the BlueStacks run Accepted. “Stock images do not run it; others did”; currency row “Mixed”.
G5 (major) MiniCAT: the 14,920 timeouts are quoted, not “silent”; the informative denominator is the 26,806 completed (49.8%) Accepted with the reviewer's caveat that the paper does not say whether a timed-out query can flag. Y block prints it.
G6 (major) MiniCrawler “closed” in the intro rests on one 2024 sentence, against later papers that name it Accepted. Intro and currency row now state both sides; “What to Read First” says “its 2024 report”. The reviewer counted three later papers; on re-reading, two name MiniCrawler or its extension for undated crawls (2025, 2026) and one reuses the extended crawler's method — the page says that.
G7 gated datasets “reused by later papers” / “Current, cheapest start” with no corpus use Accepted. “None in the corpus”; currency “Available, untried in these venues”, with the crawl's age.
G8 “quarterly results for 2025–2026” wider than the three releases read Accepted. Names the three documents.
G9 the host “oracle” is a token grab with a third party's leaked secret, recommended without an ethics note Accepted. Said in the vulnerability paragraph, the ethics bullets and the currency row.
G10 975/570 is a 2021 count stated in the present tense Accepted. Dated in the intro and Language section; the 2021 date is a needle.
G11 Telegram photo_url is optional and privacy-gated Accepted. Reworded; the “privacy settings allow” clause is now an external check.
G12 DevTools cannot attach to a third-party mini-program; the bug note is 2021-era Accepted. Added with the remote-debugging and wx.setEnableDebug documentation, both re-fetched by the script.
G13 “measured once” vs two host-side papers Accepted.
G14 the filing system is a lookup by number or name, not a category browse Accepted. “If it can be enumerated”.
G15 the provenance does not say how many hand codes were checked against the paper Accepted. 41 of 288 systematically, 6 wrong among those checked; recorded in The hand audit.
G16 the provenance recorded a fourth review before it ran Accepted. This log and the Agents row were rewritten after the last review returned.
G17 the recall probe's six hits reported as one Accepted.
G18 one table cell carried two tools' states unattributed Accepted.
G19 “nobody knows the size of any ecosystem” — the hosts do Accepted. “No host publishes a count you can use as a denominator.”
G20 storage budgets and the licence's governing-law clause are missing; the 8.2.1.4 reading is the page's own Accepted. 6.29 TB and 126.38 GB (needles) with per-package arithmetic; clause 12.3 (external check); “on our reading”.

All 20 accepted, none rejected. The generic pass found the defects no guard could: four sentences about the literature that the papers contradict (G1, G3, G4, G6), a second instance of the hand-code error the figures reviewer had found one paper over (G2), and advice with an unstated ethics cost (G9). Its fixes were re-verified by the verifier (46 spans, 100 needles, 0 not located), the external script (68 OK, 0 FAILED) and the number guard; they were not sent to a third review round. The round-2 fixes are therefore reviewed only by those guards.

Round 3: re-check of the round-2 fixes (Sonnet)

# Finding Decision
R1 the G1 fix still undercounts: [7Zhou, Jiawei; Zhang, Zidong; Ying, Lingyun; Chai, Huajun; Cao, Jiuxin; Duan, Haixin (2025): "Hey, Your Secrets Leaked! Detecting and Characterizing Secret Leakage in the Wild", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] reports recall on a labelled benchmark (83.38% on WeChat, Table 14); only its 300-detection field check is precision-only Accepted. Vulnerability paragraph, Open Question and currency row now say two papers report recall against a labelled set; 83.38% added as a needle.

The re-check verified every other round-2 fix against the papers and the fetched sources, and checked the provenance's hand-code audit bullet (“41 of 288”, “6 wrong”) against the review record. The R1 fix is a three-phrase change backed by a verifier needle and was not sent to a further round.

The report script

report_mini_programs.mjs
// report_mini_programs.mjs — every corpus figure on design:mobile_and_app_measurement:mini_programs
// ("Measuring Super-App Mini-Programs"), with its denominator. The candidate set comes from
// mp_probes.mjs, the population from the hand verdicts in mp_fold.mjs; this script throws if the two
// diverge in either direction, or if an IN paper has no hand codes.
//   node scripts/report_mini_programs.mjs          > scripts/report_mini_programs-output.txt
//   node scripts/report_mini_programs.mjs --list   (every verdict, one per line)
import fs from 'node:fs';
import path from 'node:path';
import { loadExtractions, POPULATIONS, pct, table, dataRoot } from './lib.mjs';
import { runProbes, channels, isCandidate, GAP_THRESH, VOCAB, key } from './mp_probes.mjs';
import { VERDICTS, RECALL_READ, HAND, HAND_FIELDS, INCLUSION_RULE, CODES, READ_IN_FULL, OUTSIDE_EXTRACTION } from './mp_fold.mjs';
 
const die = (m) => { console.error('FAILURE: ' + m); process.exit(1); };
const rows = loadExtractions();
const byKey = new Map(rows.map((p) => [key(p), p]));
if (rows.length !== 5859) die(`corpus contract: expected 5,859 records, got ${rows.length}`);
const venues = new Set(rows.map((p) => p.venue));
if (venues.size !== 7) die(`corpus contract: expected 7 venues, got ${[...venues]}`);
 
console.log('design:mobile_and_app_measurement:mini_programs — report');
console.log(`corpus: ${rows.length} extraction records, venues ${[...venues].sort().join(', ')}`);
 
// ---------------------------------------------------------------- A. probes
const { hits, missing } = runProbes(rows);
console.log(`\n== A. Probes (full text = paper.cols.txt, whitespace-collapsed; ${hits.length} papers with text, ${missing} records without a text file) ==`);
const gapSet = hits.filter((r) => r.gap >= GAP_THRESH);
const gapCs = hits.filter((r) => (r.p && (fs.readFileSync(path.join(dataRoot(), 'fulltext', String(r.p.year), r.p.venue, r.p.slug, 'paper.cols.txt'), 'latin1').replace(/\s+/g, ' ').match(/mini-?programs?|WeChat/g) || []).length >= GAP_THRESH));
console.log(`gap rule (/mini-?programs?|WeChat/gi >= ${GAP_THRESH}): ${gapSet.length} papers, ${gapSet.filter((r) => r.p.year >= 2024).length} from 2024-2026 (case-sensitive variant: ${gapCs.length}, ${gapCs.filter((r) => r.p.year >= 2024).length})`);
const vRows = Object.keys(VOCAB).map((n) => [n, hits.filter((r) => r.vocab[n] >= 1).length, hits.filter((r) => r.vocab[n] >= 3).length, hits.filter((r) => r.vocab[n] >= 10).length]);
vRows.push(['WeChat or Weixin (name alone)', hits.filter((r) => r.hosts.WeChat >= 1).length, hits.filter((r) => r.hosts.WeChat >= 3).length, hits.filter((r) => r.hosts.WeChat >= 10).length]);
vRows.push(['wxapkg / WXML / WXSS / wx.* API', hits.filter((r) => r.artefact >= 1).length, hits.filter((r) => r.artefact >= 3).length, hits.filter((r) => r.artefact >= 10).length]);
console.log(table(['vocabulary family', 'papers >=1', '>=3', '>=10'], vRows));
const cand = hits.filter(isCandidate);
const candKeys = new Set(cand.map((r) => r.key));
const ch = {};
for (const r of cand) for (const c of channels(r)) ch[c] = (ch[c] || 0) + 1;
console.log(`candidate set: ${cand.length} papers; enters by channel (channels overlap): ${Object.entries(ch).map(([c, n]) => `${c} ${n}`).join('; ')}`);
const only = {};
for (const r of cand) { const c = channels(r); if (c.length === 1) only[c[0]] = (only[c[0]] || 0) + 1; }
console.log(`entered by one channel only: ${Object.entries(only).map(([c, n]) => `${c} ${n}`).join('; ')}`);
 
// ---------------------------------------------------------------- B. verdict invariant
const V = new Map();
for (const v of VERDICTS) { if (V.has(v[0])) die(`duplicate verdict ${v[0]}`); V.set(v[0], v); }
for (const k of candKeys) if (!V.has(k)) die(`candidate without a verdict: ${k}`);
for (const k of V.keys()) if (!candKeys.has(k)) die(`verdict for a paper that is not a candidate: ${k}`);
for (const [k, verdict, code, note] of VERDICTS) {
  if (!['IN', 'CONTEXT', 'OUT'].includes(verdict)) die(`bad verdict ${verdict} for ${k}`);
  if (!(code in CODES[verdict])) die(`unknown ${verdict} code '${code}' for ${k}`);
  if (/PENDING/.test(note)) die(`verdict still pending: ${k}`);
}
for (const [k] of RECALL_READ) { if (candKeys.has(k)) die(`RECALL_READ key is a candidate: ${k}`); if (!byKey.has(k)) die(`RECALL_READ key not in extraction: ${k}`); }
const IN = VERDICTS.filter((v) => v[1] === 'IN');
for (const [k] of IN) if (!HAND[k]) die(`IN paper without hand codes: ${k}`);
for (const k of Object.keys(HAND)) if (!IN.find((v) => v[0] === k)) die(`hand codes for a paper that is not IN: ${k}`);
for (const [k, h] of Object.entries(HAND)) for (const f of HAND_FIELDS) if (!(f in h)) die(`hand codes for ${k} lack field ${f}`);
for (const k of READ_IN_FULL) if (!V.has(k)) die(`READ_IN_FULL key without a verdict: ${k}`);
const inKeys = new Set(IN.map((v) => v[0]));
for (const k of inKeys) if (!READ_IN_FULL.includes(k)) die(`IN paper not read in full: ${k}`);
 
console.log(`\n== B. Verdicts over the ${candKeys.size} candidates ==`);
console.log(INCLUSION_RULE);
const tally = {};
for (const [, verdict, code] of VERDICTS) tally[`${verdict} ${code}`] = (tally[`${verdict} ${code}`] || 0) + 1;
console.log(table(['verdict code', 'papers', 'meaning'], Object.entries(tally).sort((a, b) => a[0].localeCompare(b[0])).map(([vc, n]) => { const [v, c] = vc.split(' '); return [vc, n, CODES[v][c]]; })));
const nIn = IN.length, nCore = IN.filter((v) => v[2] === 'core').length;
const nCtx = VERDICTS.filter((v) => v[1] === 'CONTEXT').length, nOut = VERDICTS.filter((v) => v[1] === 'OUT').length;
console.log(`IN ${nIn} (core ${nCore}, section ${nIn - nCore}); CONTEXT ${nCtx}; OUT ${nOut}; total ${VERDICTS.length}`);
console.log(`read in full: ${READ_IN_FULL.length} candidates + ${RECALL_READ.length} recall-probe paper(s); decided from sentence contexts: ${VERDICTS.length - READ_IN_FULL.length}; every IN paper was read in full`);
for (const [k, v, c, note] of RECALL_READ) console.log(`  recall-probe paper read, not added: ${k} — ${v} ${c}: ${note}`);
 
// ---------------------------------------------------------------- C. precision and recall of the probes
console.log('\n== C. Precision and recall of each probe against the hand verdicts ==');
const pr = (name, set) => {
  const s = [...set];
  const tp = s.filter((k) => inKeys.has(k)).length;
  return [name, s.length, tp, pct(tp, s.length), `${tp}/${nIn}`, pct(tp, nIn)];
};
const vocabSet = hits.filter((r) => r.vocabSum >= 3).map((r) => r.key);
const titleSet = rows.filter((p) => /mini[- ]?programs?|mini[- ]?apps?|miniapps?|mini[- ]?games?|super[- ]?apps?|app[- ]in[- ]app/i.test(p.title)).map(key);
console.log(table(['probe', 'papers', 'IN', 'precision', 'recall', 'recall %'], [
  pr(`gap rule: /mini-?programs?|WeChat/gi >= ${GAP_THRESH}`, gapSet.map((r) => r.key)),
  pr('WeChat or Weixin >= 10 (case-sensitive)', hits.filter((r) => r.hosts.WeChat >= 10).map((r) => r.key)),
  pr('ecosystem vocabulary >= 3', vocabSet),
  pr('title names the ecosystem', titleSet),
  pr('this page: candidate set (union)', candKeys),
]));
const gapKeys = new Set(gapSet.map((r) => r.key));
console.log(`IN papers the gap rule misses: ${IN.filter((v) => !gapKeys.has(v[0])).map((v) => v[0]).join(', ') || 'none'}`);
const gapNotIn = [...gapKeys].filter((k) => !inKeys.has(k));
const gnTally = {};
for (const k of gapNotIn) { const v = V.get(k); const c = `${v[1]} ${v[2]}`; gnTally[c] = (gnTally[c] || 0) + 1; }
console.log(`gap-rule papers that are not IN (${gapNotIn.length}), by verdict: ${Object.entries(gnTally).sort((a, b) => b[1] - a[1]).map(([c, n]) => `${c} ${n}`).join('; ')}`);
const weOnly = gapNotIn.filter((k) => hits.find((h) => h.key === k).vocabSum === 0);
console.log(`of those, papers whose gap-rule hits are all "WeChat" (zero ecosystem vocabulary): ${weOnly.length}`);
 
// ---------------------------------------------------------------- D. shape of the IN set
console.log(`\n== D. The ${nIn} IN papers: year, venue, unit of study, host, side ==`);
const inP = IN.map((v) => byKey.get(v[0]));
const yr = {};
for (const p of inP) yr[p.year] = (yr[p.year] || 0) + 1;
console.log('per year: ' + Object.keys(yr).sort().map((y) => `${y}:${yr[y]}`).join(', ') + '   (2025 and 2026 provisional)');
console.log(`first IN paper: ${Math.min(...inP.map((p) => p.year))}; from 2024-2026: ${inP.filter((p) => p.year >= 2024).length} of ${nIn}; from 2025-2026: ${inP.filter((p) => p.year >= 2025).length}`);
const vn = {};
for (const p of inP) vn[p.venue] = (vn[p.venue] || 0) + 1;
console.log('per venue: ' + Object.entries(vn).sort((a, b) => b[1] - a[1]).map(([v, n]) => `${v} ${n}`).join(', ') + '   (IMC, WWW: 0 unless listed)');
const multiTally = (field, keys = [...inKeys]) => {
  const t = {};
  for (const k of keys) { const vals = HAND[k][field]; for (const v of (Array.isArray(vals) ? vals : [vals])) t[v] = (t[v] || 0) + 1; }
  return Object.entries(t).sort((a, b) => b[1] - a[1] || a[0].localeCompare(b[0]));
};
const show = (field, keys = [...inKeys], label = '') => {
  console.log(`-- ${field}${label} (of ${keys.length}; multi-valued fields do not sum):`);
  for (const [v, n] of multiTally(field, keys)) console.log(`   ${String(n).padStart(3)}  ${pct(n, keys.length).padStart(6)}  ${v}`);
};
for (const f of ['unit', 'hosts', 'side']) show(f);
 
// ---------------------------------------------------------------- E. hand codes: how the population was obtained
const deployed = [...inKeys].filter((k) => HAND[k].unit === 'deployed');
const framework = [...inKeys].filter((k) => HAND[k].unit === 'host-framework');
console.log(`\n== E. How the population was obtained — the ${deployed.length} papers whose unit is deployed mini-programs ==`);
for (const f of ['acquisition', 'analysis', 'hostVersion', 'accounts', 'keywords', 'validation', 'denominator']) show(f, deployed);
console.log(`\n-- the ${framework.length} host-framework papers:`);
for (const f of ['acquisition', 'analysis', 'instrumentation', 'hostVersion']) show(f, framework);
console.log('\n-- all IN papers:');
for (const f of ['instrumentation', 'llm', 'irb', 'disclosure', 'artifacts']) show(f);
 
// ---------------------------------------------------------------- F. per-paper table
console.log(`\n== F. Per-paper table (IN, ${nIn}) ==`);
console.log(table(['year', 'venue', 'role', 'unit', 'hosts', 'collected', 'analysed', 'acquisition', 'slug'],
  IN.map(([k, , code]) => { const p = byKey.get(k), h = HAND[k]; return [p.year, p.venue, code, h.unit, h.hosts.join('+'), h.collected, h.analysed, [].concat(h.acquisition).join('+'), p.slug.slice(0, 44)]; })
    .sort((a, b) => a[0] - b[0] || a[1].localeCompare(b[1]))));
 
// ---------------------------------------------------------------- G. what the extraction says about the same papers
console.log(`\n== G. The extraction schema over the ${nIn} IN papers ==`);
const units = {};
for (const p of inP) for (const u of new Set(p.population.map((x) => x.unit))) units[u] = (units[u] || 0) + 1;
console.log('population[].unit values used (a paper counts once per value): ' + Object.entries(units).sort((a, b) => b[1] - a[1]).map(([u, n]) => `${u} ${n}`).join('; '));
const appA = inP.filter(POPULATIONS.appAnalysis ? POPULATIONS.appAnalysis : (p) => p.studyTypes.includes('mobile-app-analysis'));
console.log(`studyTypes includes mobile-app-analysis (the parent page's appAnalysis slice, 529): ${inP.filter((p) => p.studyTypes.includes('mobile-app-analysis')).length} of ${nIn}`);
console.log(`platforms includes mobile: ${inP.filter((p) => p.platforms.includes('mobile')).length} of ${nIn}; includes web: ${inP.filter((p) => p.platforms.includes('web')).length}`);
const loc = {};
for (const p of inP) { const ls = new Set(p.vantage.flatMap((v) => v.locations)); const s = [...ls].filter((l) => l !== 'not-stated'); const k = p.vantage.length === 0 ? '(no vantage tuple)' : s.length ? 'states a location' : 'not-stated only'; loc[k] = (loc[k] || 0) + 1; }
console.log('vantage[].locations: ' + Object.entries(loc).map(([k, n]) => `${k} ${n}`).join('; '));
const eth = {};
for (const p of inP) { const v = p.ethics === null ? '(no ethics object)' : p.ethics.reviewOutcome; eth[v] = (eth[v] || 0) + 1; }
console.log('ethics.reviewOutcome: ' + Object.entries(eth).map(([k, n]) => `${k} ${n}`).join('; '));
const notif = {};
for (const p of inP) { const v = p.ethics === null ? '(no ethics object)' : p.ethics.notifiedAffectedParties; notif[v] = (notif[v] || 0) + 1; }
console.log('ethics.notifiedAffectedParties: ' + Object.entries(notif).map(([k, n]) => `${k} ${n}`).join('; '));
const emp = rows.filter(POPULATIONS.empirical);
const empStated = emp.filter((p) => p.ethics !== null && p.ethics.reviewOutcome !== 'none-mentioned').length;
console.log(`baseline: empirical papers stating a review outcome: ${empStated}/${emp.length} (${pct(empStated, emp.length)})`);
const mobileApp = rows.filter((p) => p.studyTypes.includes('mobile-app-analysis'));
const maStated = mobileApp.filter((p) => p.ethics !== null && p.ethics.reviewOutcome !== 'none-mentioned').length;
console.log(`baseline: appAnalysis papers stating a review outcome: ${maStated}/${mobileApp.length} (${pct(maStated, mobileApp.length)})`);
const art = {};
for (const p of inP) { const v = p.artifacts === null ? '(no artifacts object)' : p.artifacts.availability; art[v] = (art[v] || 0) + 1; }
console.log('artifacts.availability: ' + Object.entries(art).map(([k, n]) => `${k} ${n}`).join('; '));
for (const [k] of IN) {
  const p = byKey.get(k), h = HAND[k];
  const sIrb = p.ethics === null ? '(none)' : p.ethics.reviewOutcome, sArt = p.artifacts === null ? '(none)' : p.artifacts.availability;
  if ((h.irb === 'approval') !== (sIrb === 'approved')) console.log(`   review hand/schema disagree: ${k} (hand ${h.irb}, schema ${sIrb})`);
  if (['code', 'code+data', 'gated-data'].includes(h.artifacts) !== (sArt === 'public' || sArt === 'restricted')) console.log(`   artifacts hand/schema disagree: ${k} (hand ${h.artifacts}, schema ${sArt})`);
}
const llmSchema = inP.filter((p) => p.classification.some((c) => c.method === 'llm')).length;
console.log(`classification[].method == llm: ${llmSchema} of ${nIn}`);
const toolRe = /minicrawler|mini-crawler|wxappunpacker|unveilr|frida|xposed|jadx|soot|codeql|doublex|jaw\b|esprima|wala|uiautomator|ui automator|burp|pywinauto|jieba|jeb|ida pro|ghidra/i;
const toolT = {};
for (const p of inP) for (const n of new Set(p.tools.filter((t) => (t.usedOrMentioned === 'used' || t.usedOrMentioned === 'produced') && toolRe.test(t.name)).map((t) => t.name.toLowerCase().match(toolRe)[0].replace('mini-crawler', 'minicrawler').replace('ui automator', 'uiautomator')))) toolT[n] = (toolT[n] || 0) + 1;
console.log('tools[] used/produced, folded by name (papers): ' + Object.entries(toolT).sort((a, b) => b[1] - a[1]).map(([t, n]) => `${t} ${n}`).join('; '));
for (const [k] of IN) {
  const h = HAND[k], p = byKey.get(k);
  const handMC = [].concat(h.acquisition).includes('MiniCrawler');
  const schemaMC = p.tools.some((t) => /mini-?crawler/i.test(t.name) && (t.usedOrMentioned === 'used' || t.usedOrMentioned === 'produced'));
  if (handMC !== schemaMC) console.log(`   MiniCrawler hand/schema disagree: ${k} (hand ${handMC}, tools[] used/produced ${schemaMC})`);
}
 
// ---------------------------------------------------------------- H. per-paper measured results (detection[].prevalence)
console.log(`\n== H. detection[].prevalence for the IN papers (model summaries; the page quotes the paper, checked by verify_mini_programs_figures.mjs) ==`);
for (const [k] of IN) { const p = byKey.get(k); console.log(`${p.year} ${p.venue} ${p.slug.slice(0, 60)}`); for (const d of p.detection.slice(0, 6)) console.log(`   - ${d.phenomenon} | ${d.prevalence}`); }
 
// ---------------------------------------------------------------- I. CONTEXT papers, and the papers outside the extraction
console.log('\n== I. CONTEXT papers, listed with code and note ==');
for (const [k, , code, note] of VERDICTS.filter((v) => v[1] === 'CONTEXT').sort((a, b) => a[2].localeCompare(b[2]))) console.log(`  ${code.padEnd(12)} ${k}\n               ${note}`);
console.log('\n== J. Mini-program papers in the venue index that never reached the extraction ==');
const metaDir = path.join(dataRoot(), 'corpus2', '.meta');
const TITLE_RE = /mini[- ]?(program|app|game)|miniapp|super[- ]?app|app[- ]in[- ]app/i;
const idx = [];
for (const f of fs.readdirSync(metaDir).filter((x) => x.endsWith('.json'))) {
  const d = JSON.parse(fs.readFileSync(path.join(metaDir, f), 'utf8'));
  const recs = Array.isArray(d) ? d : (d.papers || d.records);
  for (const r of recs) if (TITLE_RE.test(r.title)) idx.push({ venue: r.venue, year: r.year, slug: r.slug, title: r.title });
}
const idxOut = idx.filter((r) => !rows.some((p) => p.slug === r.slug && p.year === r.year));
console.log(`venue-index titles matching the ecosystem vocabulary: ${idx.length}; in the extraction: ${idx.length - idxOut.length}; not in the extraction: ${idxOut.length}`);
const oeKeys = new Set(OUTSIDE_EXTRACTION.map((o) => o.slug));
for (const r of idxOut) {
  const o = OUTSIDE_EXTRACTION.find((x) => x.slug === r.slug);
  if (!o) die(`venue-index paper outside the extraction with no hand entry: ${r.venue} ${r.year} ${r.slug}`);
  console.log(`  ${r.year} ${r.venue} ${r.title}\n      why absent: ${o.why}\n      used on the page as: ${o.use}`);
}
for (const s of oeKeys) if (!idxOut.some((r) => r.slug === s)) die(`OUTSIDE_EXTRACTION entry no longer outside the extraction: ${s}`);
 
// ---------------------------------------------------------------- Y. arithmetic on figures the papers print
// Each line names the paper and the two printed numbers; the page labels these as our arithmetic.
console.log('\n== Y. Arithmetic on per-paper figures (the page labels these as derived) ==');
const ar = (label, num, den) => console.log(`  ${label}: ${num.toLocaleString('en-US')} / ${den.toLocaleString('en-US')} = ${pct(num, den)}`);
ar('yang2025_miniapp: delisted by end of 2022 / collected by June 2022', 360467, 4595680);
ar('wang2023_uncovering: APIs missing from the English docs / Chinese docs (975 - 570)', 975 - 570, 975);
ar('yang2022_cross: WeChat miniapps lacking appId checks / all crawled WeChat miniapps', 50281, 2571490);
ar('zhang2024_minicat: timeouts / analysed', 14920, 41726);
ar('zhang2024_minicat: potentially vulnerable / all collected', 13349, 44273);
ar('shi2025_skeleton: mini-apps with a leak / analysed', 54728, 402527);
ar('shi2025_skeleton: analysed / collected', 402527, 413775);
ar('shi2026_better: vulnerable / all crawled', 2815, 1248815);
ar('shi2026_better: cloud-using / all crawled', 22695, 1248815);
ar('zhang2023_leak: MK leaks / crawled', 40880, 3450586);
ar('chen2026_minigames: Cocos games / crawled', 2076, 6769);
ar('wang2025_wechat: analysed / attempted', 104, 170);
console.log(`  IN papers releasing code (code, code+data): ${IN.filter(([k]) => ['code', 'code+data'].includes(HAND[k].artifacts)).length} of ${nIn}; any artefact incl. gated data: ${IN.filter(([k]) => ['code', 'code+data', 'gated-data'].includes(HAND[k].artifacts)).length}`);
ar('zhang2024_minicat: potentially vulnerable / analysed minus timed out (41,726 - 14,920)', 13349, 41726 - 14920);
console.log(`  zhang2024_minicat: packages on which the detector completed: ${(41726 - 14920).toLocaleString('en-US')}`);
console.log(`  yang2022_cross: storage per WeChat package: 6.29 TB / 2,571,490 = ${(6.29e12 / 2571490 / 1e6).toFixed(2)} MB (decimal units)`);
console.log(`  zhang2024_minicat: storage per unpacked package: 126.38 GB / 44,273 = ${(126.38e9 / 44273 / 1e6).toFixed(2)} MB (decimal units)`);
console.log(`  vocabulary >= 3 precision / gap-rule precision: ${(16 / 30 / (15 / 37)).toFixed(2)}x (53.3% vs 40.5%)`);
console.log(`  IN papers stating an IRB/ethics-board approval: ${IN.filter(([k]) => HAND[k].irb === 'approval').length} of ${nIn}`);
console.log(`  IN papers using an LLM anywhere in the pipeline: ${IN.filter(([k]) => HAND[k].llm === 'yes').length} of ${nIn} (years ${IN.filter(([k]) => HAND[k].llm === 'yes').map(([k]) => byKey.get(k).year).join(', ')})`);
console.log(`  deployed-population papers naming MiniCrawler: ${deployed.filter((k) => HAND[k].acquisition.includes('MiniCrawler')).map((k) => byKey.get(k).year).sort().join(', ')}; reusing a prior crawl: ${deployed.filter((k) => HAND[k].acquisition.includes('prior-crawl')).map((k) => byKey.get(k).year).sort().join(', ')}`);
console.log(`  papers naming Xposed: ${IN.filter(([k]) => HAND[k].instrumentation.includes('Xposed')).map(([k]) => byKey.get(k).year).join(', ')}; Frida: ${IN.filter(([k]) => HAND[k].instrumentation.includes('Frida')).map(([k]) => byKey.get(k).year).join(', ')}`);
console.log(`  hosts: papers including a host outside WeChat/Baidu/TikTok-Douyin/Alipay/QQ: ${IN.filter(([k]) => HAND[k].hosts.some((h) => h === 'other' || h === 'IoT host')).length}`);
console.log(`  hosts: papers naming a host other than WeChat: ${IN.filter(([k]) => HAND[k].hosts.some((h) => h !== 'WeChat')).length}; WeChat-only: ${IN.filter(([k]) => HAND[k].hosts.length === 1 && HAND[k].hosts[0] === 'WeChat').length}`);
 
// ---------------------------------------------------------------- K. probes behind the page's negative claims about the 16
// Each hand code the page prints as "none" or "not stated" is backed by a probe over the paper's own text; every hit was read.
console.log('\n== K. Probes behind negative claims (distinct matched strings per paper; every hit read by hand) ==');
const KP = { irb: /\bIRB\b|ethics? (review )?(board|committee)|institutional review|review board/gi,
  release: /github\.com\/[\w.-]+|zenodo|gitlab\.com|figshare|osf\.io/gi,
  terms: /terms of (service|use)|licen[cs]e agreement|platform polic(y|ies)|in compliance|ensur\w* compliance|comply with/gi };
for (const [k] of IN) {
  const t = fs.readFileSync(path.join(dataRoot(), 'fulltext', String(byKey.get(k).year), byKey.get(k).venue, byKey.get(k).slug, 'paper.cols.txt'), 'latin1').replace(/\u0000/g, '').replace(/\s+/g, ' ');
  const f = Object.fromEntries(Object.entries(KP).map(([n, re]) => [n, [...new Set((t.match(re) || []).map((x) => x.toLowerCase()))].slice(0, 5)]));
  console.log(`  ${byKey.get(k).year} ${byKey.get(k).venue} ${byKey.get(k).slug.slice(0, 34)} | irb=${HAND[k].irb} ${JSON.stringify(f.irb)} | artifacts=${HAND[k].artifacts} ${JSON.stringify(f.release)} | terms ${JSON.stringify(f.terms)}`);
}
console.log('  reading of the terms hits: the 2026 mini-games paper asserts its crawl "ensur[es] compliance with platform policies"; the NDSS 2026 paper "compliance with all relevant laws and regulations"; the 2024 rental paper complies with "the vendor\'s bug bounty plan"; the other hits are about permission policies, privacy regulation or advertising policies. None quotes or analyses the host\'s user licence.');
 
if (process.argv.includes('--list')) {
  console.log('\n== --list: every verdict ==');
  for (const [k, v, c, note] of VERDICTS) console.log(`${v}\t${c}\t${k}\t${note}`);
}
console.log('\nOK: candidates and verdicts agree in both directions; every IN paper has hand codes.');

The report script's output

report_mini_programs-output.txt
design:mobile_and_app_measurement:mini_programs — report
corpus: 5859 extraction records, venues CCS, IEEE-SP, IMC, NDSS, PETS, USENIX, WWW
 
== A. Probes (full text = paper.cols.txt, whitespace-collapsed; 5855 papers with text, 4 records without a text file) ==
gap rule (/mini-?programs?|WeChat/gi >= 10): 37 papers, 21 from 2024-2026 (case-sensitive variant: 33, 18)
vocabulary family                papers >=1  >=3  >=10
-------------------------------  ----------  ---  ----
mini-program                     32          17   6
mini-app                         26          15   10
mini-game                        5           1    1
super-app                        28          15   12
app-in-app                       33          8    4
WeChat or Weixin (name alone)    171         62   31
wxapkg / WXML / WXSS / wx.* API  9           6    1
candidate set: 57 papers; enters by channel (channels overlap): schema 27; gap 37; vocab 30; artefact 9; other-host 1
entered by one channel only: schema 7; gap 16; vocab 13
 
== B. Verdicts over the 57 candidates ==
INCLUSION RULE (written before any verdict was counted):
IN      — mini-programs (third-party programs that run inside a host "super app" and reach the device
          and the user only through APIs the host provides), or the host's mini-program framework and
          its APIs, are measured or analysed. Sub-coded CORE (the main object of the paper) or SECTION
          (one analysed population among several, e.g. one of three platforms).
CONTEXT — adjacent: mini-program source code on GitHub rather than deployed mini-programs; a host
          app's own feature (its in-app browser, its face verification) rather than its mini-programs.
OUT     — WeChat / Alipay only as a messenger, a payment method, an attack target, a dataset owner or
          a recruitment channel; related-work mentions; homographs ("MiniApps" the token, "In-App" purchases).
verdict code           papers  meaning
---------------------  ------  ------------------------------------------------------------------------------------
CONTEXT host-feature   3       a host app's own feature, not its mini-programs
CONTEXT source-repos   1       mini-program source code on GitHub, not deployed mini-programs
IN core                13      mini-programs or the host framework are the main object
IN section             3       mini-programs are one analysed population among several
OUT homograph          4       the probe matched another sense of the word
OUT host-app-analysis  5       the super app analysed as an ordinary app (ports, ROM privileges, modified binaries)
OUT host-sdk           2       the host's SDK inside ordinary apps
OUT mention            9       related work or passing mention
OUT messenger          10      WeChat/Alipay as a messenger, contact point, payment method or login target
OUT operator-data      4       a dataset supplied by the host operator, no mini-programs
OUT recruitment        3       WeChat groups or accounts used to recruit participants
IN 16 (core 13, section 3); CONTEXT 4; OUT 37; total 57
read in full: 16 candidates + 1 recall-probe paper(s); decided from sentence contexts: 41; every IN paper was read in full
  recall-probe paper read, not added: USENIX/2023/medusa-attack-exploring-security-hazards-of-in-app-qr-code-scanning — OUT unrelated: reader D: native apps' built-in QR-code handlers; no mini-program, super-app or host-name hit in the text (the recall probe fired on "host app" and other app names)
 
== C. Precision and recall of each probe against the hand verdicts ==
probe                                       papers  IN  precision  recall  recall %
------------------------------------------  ------  --  ---------  ------  --------
gap rule: /mini-?programs?|WeChat/gi >= 10  37      15  40.5%      15/16   93.8%
WeChat or Weixin >= 10 (case-sensitive)     31      12  38.7%      12/16   75.0%
ecosystem vocabulary >= 3                   30      16  53.3%      16/16   100.0%
title names the ecosystem                   12      12  100.0%     12/16   75.0%
this page: candidate set (union)            57      16  28.1%      16/16   100.0%
IN papers the gap rule misses: CCS/2024/riotfuzzer-companion-app-assisted-remote-fuzzing-for-detecting-vulnerabilities-i
gap-rule papers that are not IN (22), by verdict: OUT messenger 9; OUT host-app-analysis 4; CONTEXT host-feature 3; OUT host-sdk 2; OUT operator-data 2; OUT mention 2
of those, papers whose gap-rule hits are all "WeChat" (zero ecosystem vocabulary): 19
 
== D. The 16 IN papers: year, venue, unit of study, host, side ==
per year: 2020:1, 2022:2, 2023:3, 2024:3, 2025:5, 2026:2   (2025 and 2026 provisional)
first IN paper: 2020; from 2024-2026: 10 of 16; from 2025-2026: 7
per venue: CCS 6, USENIX 5, NDSS 3, PETS 1, IEEE-SP 1   (IMC, WWW: 0 unless listed)
-- unit (of 16; multi-valued fields do not sum):
     9   56.3%  deployed
     5   31.3%  host-framework
     1    6.3%  host-traffic
     1    6.3%  usage-logs
-- hosts (of 16; multi-valued fields do not sum):
    14   87.5%  WeChat
     7   43.8%  Baidu
     6   37.5%  other
     5   31.3%  TikTok/Douyin
     4   25.0%  Alipay
     2   12.5%  QQ
     1    6.3%  IoT host
-- side (of 16; multi-valued fields do not sum):
    12   75.0%  vulnerability
     4   25.0%  privacy
     2   12.5%  host-API
     1    6.3%  advertising
     1    6.3%  malware
 
== E. How the population was obtained — the 9 papers whose unit is deployed mini-programs ==
-- acquisition (of 9; multi-valued fields do not sum):
     4   44.4%  MiniCrawler
     2   22.2%  prior-crawl
     1   11.1%  audit-set
     1   11.1%  client-cache
     1   11.1%  keyword-search
     1   11.1%  qr-codes
     1   11.1%  search-api
     1   11.1%  store-revisit
-- analysis (of 9; multi-valued fields do not sum):
     8   88.9%  static
     3   33.3%  dynamic
     2   22.2%  manual
-- hostVersion (of 9; multi-valued fields do not sum):
     9  100.0%  not-stated
-- accounts (of 9; multi-valued fields do not sum):
     5   55.6%  own-test-accounts
     3   33.3%  not-stated
     1   11.1%  real-name
-- keywords (of 9; multi-valued fields do not sum):
     7   77.8%  not-stated
     2   22.2%  Chinese
-- validation (of 9; multi-valued fields do not sum):
     5   55.6%  manual-sample
     2   22.2%  manual-all
     1   11.1%  api-oracle
     1   11.1%  ground-truth-set
-- denominator (of 9; multi-valued fields do not sum):
     5   55.6%  no
     4   44.4%  yes
 
-- the 5 host-framework papers:
-- acquisition (of 5; multi-valued fields do not sum):
     3   60.0%  official-docs
     3   60.0%  own-test-miniapps
     1   20.0%  app-store-crawl
     1   20.0%  MiniCrawler
     1   20.0%  purchased-devices
-- analysis (of 5; multi-valued fields do not sum):
     5  100.0%  dynamic
     3   60.0%  static
-- instrumentation (of 5; multi-valued fields do not sum):
     3   60.0%  Frida
     2   40.0%  Xposed
-- hostVersion (of 5; multi-valued fields do not sum):
     2   40.0%  os-only
     2   40.0%  pinned
     1   20.0%  not-stated
 
-- all IN papers:
-- instrumentation (of 16; multi-valued fields do not sum):
     4   25.0%  Frida
     2   12.5%  Xposed
-- llm (of 16; multi-valued fields do not sum):
    14   87.5%  no
     2   12.5%  yes
-- irb (of 16; multi-valued fields do not sum):
    13   81.3%  none-mentioned
     3   18.8%  approval
-- disclosure (of 16; multi-valued fields do not sum):
    13   81.3%  host-vendor
     5   31.3%  bug-bounty
     5   31.3%  developers
     3   18.8%  CVE
     1    6.3%  CERT
-- artifacts (of 16; multi-valued fields do not sum):
     8   50.0%  code
     4   25.0%  none-stated
     2   12.5%  code+data
     1    6.3%  declined
     1    6.3%  gated-data
 
== F. Per-paper table (IN, 16) ==
year  venue    role     unit            hosts                                    collected                             analysed                     acquisition                        slug
----  -------  -------  --------------  ---------------------------------------  ------------------------------------  ---------------------------  ---------------------------------  --------------------------------------------
2020  CCS      core     host-framework  WeChat+QQ+Baidu+TikTok/Douyin+other      11 hosts                              927 sub-app APIs             own-test-miniapps+official-docs    demystifying-resource-management-risks-in-em
2022  CCS      core     deployed        WeChat+Baidu                             2,571,490 WeChat + 148,512 Baidu      all (static)                 MiniCrawler                        cross-miniapp-request-forgery-root-causes-at
2022  USENIX   core     host-framework  WeChat+Alipay+Baidu+TikTok/Douyin+other  6,000 Android apps                    47 super apps                app-store-crawl+own-test-miniapps  identity-confusion-in-webview-based-mobile-a
2023  CCS      core     deployed        WeChat+Baidu                             3,450,586 WeChat + 171,989 Baidu      all (static)                 search-api                         dont-leak-your-keys-understanding-measuring-
2023  CCS      core     host-framework  WeChat+QQ+Baidu+TikTok/Douyin+other      5 host APKs; 267,359 WeChat miniapps  1,829 API candidates         official-docs+MiniCrawler          uncovering-and-exploiting-hidden-apis-in-mob
2023  USENIX   core     host-framework  WeChat                                   1,031 documented APIs                 all, on Windows/Android/iOS  official-docs+own-test-miniapps    one-size-does-not-fit-all-uncovering-and-exp
2024  CCS      core     deployed        WeChat                                   44,273                                41,726                       client-cache                       minicat-understanding-and-detecting-cross-pa
2024  CCS      section  host-framework  IoT host                                 27 devices                            27                           purchased-devices                  riotfuzzer-companion-app-assisted-remote-fuz
2024  USENIX   section  deployed        WeChat                                   75 mini-programs                      75                           keyword-search+qr-codes            demystifying-the-security-implications-in-io
2025  IEEE-SP  section  deployed        WeChat                                   41,719                                41,719                       prior-crawl                        hey-your-secrets-leaked-detecting-and-charac
2025  NDSS     core     deployed        WeChat                                   4,595,680                             all (static)                 MiniCrawler+store-revisit          understanding-miniapp-malware-identification
2025  NDSS     core     deployed        WeChat+Baidu+Alipay+TikTok/Douyin+other  413,775                               402,527                      MiniCrawler                        the-skeleton-keys-a-large-scale-analysis-of-
2025  PETS     core     host-traffic    WeChat                                   170                                   104                          rankings+keyword-search            what-wechat-knows-pervasive-first-party-trac
2025  USENIX   core     usage-logs      Alipay                                   288,895 users                         219,826 users                operator-logs                      i-can-tell-your-secrets-inferring-privacy-at
2026  NDSS     core     deployed        WeChat+TikTok/Douyin+Alipay+Baidu+other  1,248,815                             22,695 using cloud services  prior-crawl                        better-safe-than-sorry-uncovering-the-insecu
2026  USENIX   core     deployed        WeChat+other                             6,769                                 2,076 Cocos                  MiniCrawler+audit-set              when-fun-turns-toxic-a-first-look-at-aggress
 
== G. The extraction schema over the 16 IN papers ==
population[].unit values used (a paper counts once per value): mobile-apps 13; other 5; documents 2; iot-devices 2; human-participants 1; code-repositories 1
studyTypes includes mobile-app-analysis (the parent page's appAnalysis slice, 529): 13 of 16
platforms includes mobile: 15 of 16; includes web: 1
vantage[].locations: not-stated only 15; (no vantage tuple) 1
ethics.reviewOutcome: none-mentioned 12; approved 3; explicitly-discussed-no-review 1
ethics.notifiedAffectedParties: yes 16
baseline: empirical papers stating a review outcome: 1728/5118 (33.8%)
baseline: appAnalysis papers stating a review outcome: 149/529 (28.2%)
artifacts.availability: public 10; none-mentioned 4; explicitly-withheld 1; restricted 1
classification[].method == llm: 0 of 16
tools[] used/produced, folded by name (papers): minicrawler 5; frida 4; burp 3; ida pro 3; xposed 2; soot 2; doublex 1; jeb 1; pywinauto 1; jieba 1; wxappunpacker 1; codeql 1; jadx 1; ghidra 1; esprima 1; wala 1; jaw 1; uiautomator 1
 
== H. detection[].prevalence for the IN papers (model summaries; the page quotes the paper, checked by verify_mini_programs_figures.mjs) ==
2020 CCS demystifying-resource-management-risks-in-emerging-mobile-ap
   - System Resource Exposure flaws | 39 escaped sub-app APIs across 11 host apps
   - Sub-window Deception flaws | 10 flaws in the overall landscape
   - Sub-app Lifecycle Hijacking | 3 Android hosts: Wechat, HostAppA, and DingTalk
   - Android API-permission mapping accuracy | 97.9% at API level 27; 97.1% at API 28; 95% at API 29
   - Documentation permission extraction | 100% precision and 100% recall on 52 mappings
   - Sub-window detection | 5 sub-app APIs associated with UI deception flaws
2022 CCS cross-miniapp-request-forgery-root-causes-attacks-and-vulner
   - cross-miniapp communication | 52,394 (2.04%) WeChat miniapps and 494 (0.33%) Baidu miniapps
   - CMRF vulnerability | 50,281 (95.97%) WeChat and 493 (99.80%) Baidu miniapps
   - security impact of CMRF | 55.05% of WeChat and 7.09% of Baidu miniapps
   - cross-miniapp redirection relationships | 2,907 appIds connected by 4,912 edges
   - CMRF attack case studies | Shopping-for-free attack succeeded
2022 USENIX identity-confusion-in-webview-based-mobile-app-in-app-ecosys
   - identity confusion vulnerabilities | 47 of 47 super-apps vulnerable to at least one type
   - domain name confusion | 15 timing-based and 15 frame-based cases
   - AppID confusion | 38 super-apps
   - capability confusion | 2 super-apps
   - privilege escalation | 38 of 47 super-apps
   - phishing | 31 of 47 super-apps
2023 CCS dont-leak-your-keys-understanding-measuring-and-exploiting-t
   - WeChat master-key leakage | 40,880 of 3,450,586 mini-programs
   - Baidu master-key leakage | 7,476 of 171,989 mini-programs (4.35%)
   - Sensitive-resource access | 24,701 accessed sensitive data and at least one cloud service
   - Leaked-key attacks | All 10 tested Baidu mini-programs were vulnerable
2023 CCS uncovering-and-exploiting-hidden-apis-in-mobile-super-apps
   - undocumented APIs | All five tested super apps contained hidden APIs
   - unchecked hidden APIs | WeChat 7.77%, WeCom 6.75%, Baidu 7.08%, TikTok 26.67%, QQ 12.88%
   - sensitive-resource access | 39 WeChat, 40 WeCom, 8 Baidu, 32 TikTok, and 38 QQ APIs
   - third-party hidden-API usage | 78,974 of 267,359 WeChat miniapps (29.54%)
   - hidden-API exploitability | Attacks covered web access, malware installation, screenshots, phone numbers, and contacts
2023 USENIX one-size-does-not-fit-all-uncovering-and-exploiting-cross-pl
   - API existence discrepancies | 109 APIs
   - API permission discrepancies | 17 APIs
   - API output discrepancies | 22 APIs
   - API output uniqueness | 22 APIs
   - Fingerprintable APIs | 13 APIs
2024 CCS minicat-understanding-and-detecting-cross-page-request-forge
   - MiniCPRF vulnerability | 32.0% (13,349/41,726) potentially vulnerable
   - Potential MiniCPRF attack paths | 119,471 risky pages
   - False positives | 0/100 false positives
   - False negatives | 3/100 false negatives
   - Verified front-end vulnerabilities | 316/400 (79.0%) severe security issues
   - Mini-program popularity | 3,208 mini-programs and 9,007 domains measured
2026 NDSS better-safe-than-sorry-uncovering-the-insecure-resource-mana
   - Insecure cloud-resource management | 2,815 mini-apps (12.40%) and 8,062 insecure cloud operations
   - Hidden cloud capabilities | 4,202 hidden cloud database names in 1,486 mini-apps; 539 mini-apps (36.27%) vulnerable
   - User identity-check flaws | UIC-1: 919 mini-apps; UIC-2: 1,896 mini-apps
   - Sensitive-resource allocation flaws | SRA-1: 106 mini-apps; SRA-2: 993 mini-apps
   - Guessable cloud-storage paths | 709 vulnerable mini-apps and 1,023 sensitive operations
   - Detection accuracy | 2,815 vulnerable mini-apps after filtering
2025 PETS what-wechat-knows-pervasive-first-party-tracking-in-a-billio
   - profile-data exfiltration | 84.3% of 51 profile flows
   - search-query exfiltration | 72.9% of 85 search flows
   - fine-grained browsing-data tracking | 76.0% of 96 browsing flows
   - health Mini Program browsing tracking | 89.7% of 39 health Mini Programs
2025 USENIX i-can-tell-your-secrets-inferring-privacy-attributes-from-mi
   - Privacy-attribute inference | 95.5% accuracy for 16.1% of samples at a 0.9 confidence threshold
   - Mini-app interaction-history awareness | 1 of 31 super-apps mentioned Mini-H; none mentioned Op-H
2026 USENIX when-fun-turns-toxic-a-first-look-at-aggressive-advertising-
   - aggressive advertising | 49.95% (457/915)
   - aggressive advertising bypass | 94.09% relied on dynamic, network-dependent triggers
   - short advertising intervals | 67.39% shorter than Facebook's 30-second threshold
   - aggressive advertising categories | 44.48% interruptive, 7.54% hijacking, 12.13% unstoppable, and 1.64% deceptive
   - MAAD detection accuracy | 94.57% precision and 83.55% recall
   - large-scale MAAD deployment | 877/1,613 games (54.37%)
2025 NDSS understanding-miniapp-malware-identification-dissection-and-
   - evasive miniapp malware | 19,905 malware samples from 4,595,680 collected miniapps
   - code vetting evasion | 18,428 using evaluate; 18,112 using eval
   - content vetting evasion | 34 of 500 sampled miniapps contained content-vetting evasion
   - false-positive detection | 487 correctly identified and 13 false positives among 500
   - malicious payload categories | 5 miniapps manually sampled per category
   - sensitive data access | Approximately a quarter collected location data
2025 NDSS the-skeleton-keys-a-large-scale-analysis-of-credential-leaka
   - credential leakage | 84,491 credential leaks spanning 54,728 mini-apps
   - credential leakage | KeyMagnet analyzed 402,527 of 413,775 collected mini-apps
   - cross-mini-app credential leakage | 47 of 100 sampled WeChat mini-apps leaked other mini-apps' root credentials
   - sensitive information in logs | 1,930 of 11,955 mini-apps recorded sensitive information
2024 USENIX demystifying-the-security-implications-in-iot-device-rental-
   - IoT device and app vulnerabilities | 57 vulnerabilities in 28 products
   - Device vulnerabilities | 23 vulnerabilities in 14 devices
   - App vulnerabilities | 34 vulnerabilities in 23 apps
   - Device serial-number enumeration | 56/92 products (60.9%)
   - Large-scale exploitation | 84% (48/57) of vulnerabilities
   - Device impersonation | 14 of 15 tested devices
2025 IEEE-SP hey-your-secrets-leaked-detecting-and-characterizing-secret-
   - secret leakage across platforms | 24.78% of 4,280 GitHub repositories; 7.47% of 668,847 PyPI packages; 30.08% of 41,719 WeChat MPs
   - secret leakage in files | 11,826 true secrets across 1,806,530 benchmark files
   - secret persistence across PyPI versions | 85.91% retained the same number; 4.09% removed secrets; 8.84% added more
   - internal codebase secret leakage | 858 detected secrets; 38 valid among 104 examined
   - obfuscated-secret detection | 48 of 280 corresponding obfuscated secrets detected, approximately 17.14%
2024 CCS riotfuzzer-companion-app-assisted-remote-fuzzing-for-detecti
   - IoT device vulnerabilities | 11 vulnerabilities across 10 IoT devices
   - Cloud-server verification bypass | average improvement of 76.62%, maximum 362.62%
   - Device crash or denial of service | 11 confirmed after retransmitting exploiting packets three times
 
== I. CONTEXT papers, listed with code and note ==
  host-feature WWW/2024/phishinwebview-analysis-of-anti-phishing-entities-in-mobile-apps-with-webview-ta
               WeChat's in-app browser does not use Google Safe Browsing in China; the host app's web container, not its mini-programs
  host-feature CCS/2025/deep-dive-into-in-app-browsers-uncovering-hidden-pitfalls-in-certificate-validat
               in-app browsers incl. WeChat 8.0.51; WeChat would not launch on Android 11/15 emulators and was run on BlueStacks — a host-app instrumentation pitfall
  host-feature IEEE-SP/2023/understanding-the-in-security-of-cross-side-face-verification-systems-in-mobile
               the host app's face-verification service, which sub-apps can invoke; mini-programs themselves not analysed
  source-repos CCS/2023/understanding-and-detecting-abused-image-hosting-modules-as-malicious-services
               GitHub repositories of mini-programs are one of three categories of abusers of image-hosting upload APIs; deployed mini-programs not analysed
 
== J. Mini-program papers in the venue index that never reached the extraction ==
venue-index titles matching the ecosystem vocabulary: 18; in the extraction: 12; not in the extraction: 6
  2023 CCS SaTS'23: The 1st ACM Workshop on Secure and Trustworthy Superapps.
      why absent: workshop front matter, screened out (correctly)
      used on the page as: evidence that a dedicated CCS workshop exists
  2024 CCS SaTS '24: The 2nd ACM Workshop on Secure and Trustworthy Superapps.
      why absent: workshop front matter, screened out (correctly)
      used on the page as: evidence that a dedicated CCS workshop exists
  2025 CCS SaTS '25: The 3rd ACM Workshop on Security and Privacy of AI-Empowered Mobile Super Apps.
      why absent: workshop front matter, screened out (correctly)
      used on the page as: evidence that a dedicated CCS workshop exists
  2026 IEEE-SP Convenience at a Cost: the Security Risks of Template-Based Development in the App-in-App Ecosystem.
      why absent: IEEE S&P 2026: no abstract in OpenAlex, so never screened (under-selected by construction)
      used on the page as: abstract only (Semantic Scholar); cited, not counted
  2026 USENIX Raising the Flag: Detecting Missing Permission Controls in Mini-Program APIs
      why absent: selected by screening; PDF not retrieved (empty fulltext directory)
      used on the page as: read in full from the USENIX PDF (out/mp/ext/wei_zhiao.pdf); cited, not counted
  2026 WWW Real or Rogue? Detecting Malicious Miniapps with Deceptive Reporting Interface.
      why absent: TheWebConf 2026: no abstract in OpenAlex, so never screened (under-selected by construction)
      used on the page as: abstract only (Semantic Scholar; ACM DL is Cloudflare-walled); cited, not counted
 
== Y. Arithmetic on per-paper figures (the page labels these as derived) ==
  yang2025_miniapp: delisted by end of 2022 / collected by June 2022: 360,467 / 4,595,680 = 7.8%
  wang2023_uncovering: APIs missing from the English docs / Chinese docs (975 - 570): 405 / 975 = 41.5%
  yang2022_cross: WeChat miniapps lacking appId checks / all crawled WeChat miniapps: 50,281 / 2,571,490 = 2.0%
  zhang2024_minicat: timeouts / analysed: 14,920 / 41,726 = 35.8%
  zhang2024_minicat: potentially vulnerable / all collected: 13,349 / 44,273 = 30.2%
  shi2025_skeleton: mini-apps with a leak / analysed: 54,728 / 402,527 = 13.6%
  shi2025_skeleton: analysed / collected: 402,527 / 413,775 = 97.3%
  shi2026_better: vulnerable / all crawled: 2,815 / 1,248,815 = 0.2%
  shi2026_better: cloud-using / all crawled: 22,695 / 1,248,815 = 1.8%
  zhang2023_leak: MK leaks / crawled: 40,880 / 3,450,586 = 1.2%
  chen2026_minigames: Cocos games / crawled: 2,076 / 6,769 = 30.7%
  wang2025_wechat: analysed / attempted: 104 / 170 = 61.2%
  IN papers releasing code (code, code+data): 10 of 16; any artefact incl. gated data: 11
  zhang2024_minicat: potentially vulnerable / analysed minus timed out (41,726 - 14,920): 13,349 / 26,806 = 49.8%
  zhang2024_minicat: packages on which the detector completed: 26,806
  yang2022_cross: storage per WeChat package: 6.29 TB / 2,571,490 = 2.45 MB (decimal units)
  zhang2024_minicat: storage per unpacked package: 126.38 GB / 44,273 = 2.85 MB (decimal units)
  vocabulary >= 3 precision / gap-rule precision: 1.32x (53.3% vs 40.5%)
  IN papers stating an IRB/ethics-board approval: 3 of 16
  IN papers using an LLM anywhere in the pipeline: 2 of 16 (years 2026, 2024)
  deployed-population papers naming MiniCrawler: 2022, 2025, 2025, 2026; reusing a prior crawl: 2025, 2026
  papers naming Xposed: 2020, 2022; Frida: 2023, 2023, 2025, 2024
  hosts: papers including a host outside WeChat/Baidu/TikTok-Douyin/Alipay/QQ: 7
  hosts: papers naming a host other than WeChat: 10; WeChat-only: 6
 
== K. Probes behind negative claims (distinct matched strings per paper; every hit read by hand) ==
  2020 CCS demystifying-resource-management-r | irb=none-mentioned [] | artifacts=code ["github.com/mozillasecurity","github.com/aslody","github.com/itseez"] | terms ["comply with"]
  2022 CCS cross-miniapp-request-forgery-root | irb=none-mentioned [] | artifacts=code ["github.com/osuseclab"] | terms []
  2022 USENIX identity-confusion-in-webview-base | irb=none-mentioned [] | artifacts=none-stated ["github.com/soot-oss"] | terms []
  2023 CCS dont-leak-your-keys-understanding- | irb=none-mentioned [] | artifacts=none-stated [] | terms []
  2023 CCS uncovering-and-exploiting-hidden-a | irb=none-mentioned [] | artifacts=none-stated [] | terms []
  2023 USENIX one-size-does-not-fit-all-uncoveri | irb=none-mentioned [] | artifacts=code ["github.com/1n3"] | terms []
  2024 CCS minicat-understanding-and-detectin | irb=none-mentioned [] | artifacts=code ["github.com/en","github.com/fxsjy","github.com/pywinauto","github.com/system-cpu","github.com/imingyu"] | terms []
  2026 NDSS better-safe-than-sorry-uncovering- | irb=approval ["irb"] | artifacts=code+data ["zenodo","github.com/wemobiledev"] | terms ["ensure compliance"]
  2025 PETS what-wechat-knows-pervasive-first- | irb=none-mentioned [] | artifacts=code+data ["github.com/wemobiledev"] | terms []
  2025 USENIX i-can-tell-your-secrets-inferring- | irb=approval ["irb"] | artifacts=declined [] | terms ["comply with"]
  2026 USENIX when-fun-turns-toxic-a-first-look- | irb=none-mentioned [] | artifacts=code ["github.com/w","zenodo","github.com/cwi-swat","github.com/wala","github.com/swc-project"] | terms ["platform policies","platform policy","terms of use","ensuring compliance"]
  2025 NDSS understanding-miniapp-malware-iden | irb=none-mentioned [] | artifacts=gated-data ["github.com/jkeylu","github.com/tarruda"] | terms ["terms of use"]
  2025 NDSS the-skeleton-keys-a-large-scale-an | irb=approval ["irb"] | artifacts=code ["github.com/keymagnetproject2025","github.com/ant-move","github.com/wala"] | terms []
  2024 USENIX demystifying-the-security-implicat | irb=none-mentioned [] | artifacts=none-stated [] | terms ["comply with"]
  2025 IEEE-SP hey-your-secrets-leaked-detecting- | irb=none-mentioned [] | artifacts=code ["github.com/abbrcode","github.com/david47k","github.com/dwyl","github.com/aoa0","github.com/gitleaks"] | terms []
  2024 CCS riotfuzzer-companion-app-assisted- | irb=none-mentioned [] | artifacts=code ["github.com/androguard","github.com/wireghoul","github.com/iputils","github.com/kzliu2017"] | terms []
  reading of the terms hits: the 2026 mini-games paper asserts its crawl "ensur[es] compliance with platform policies"; the NDSS 2026 paper "compliance with all relevant laws and regulations"; the 2024 rental paper complies with "the vendor's bug bounty plan"; the other hits are about permission policies, privacy regulation or advertising policies. None quotes or analyses the host's user licence.
 
OK: candidates and verdicts agree in both directions; every IN paper has hand codes.

The probes

mp_probes.mjs
// mp_probes.mjs — candidate probes for design:mobile_and_app_measurement:mini_programs
// ("Measuring Super-App Mini-Programs"). Imported by mp_fold.mjs and report_mini_programs.mjs.
// Nothing here is a population: every probe produces a CANDIDATE set, and the population is the
// hand verdict in mp_fold.mjs.
//
// Texts are read one at a time from paper.cols.txt with whitespace collapsed (a PDF line break
// inside "mini program" otherwise undercounts), and never held all at once.
import fs from 'node:fs';
import path from 'node:path';
import { dataRoot } from './lib.mjs';
 
// The 2026-09-22 gap pass's rule as the task brief states it: ">= 10 hits of /mini-?programs?|WeChat/".
// Case-insensitive, which is how its 37 / 21 reproduce exactly. It does NOT match "mini program"
// (space), "mini-app" or "super app" — the vocabulary probes below exist because of that.
export const GAP_RE = /mini-?programs?|WeChat/gi;
export const GAP_THRESH = 10;
 
// The ecosystem's own vocabulary. Each family is counted separately so the report can print which
// names a paper uses; the candidate rule sums them.
export const VOCAB = {
  'mini-program': /\bmini[- ]?programs?\b/gi,
  'mini-app': /\bmini[- ]?apps?\b/gi,
  'mini-game': /\bmini[- ]?games?\b/gi,
  'super-app': /\bsuper[- ]?apps?\b/gi,
  'app-in-app': /\bapp[- ]in[- ]app\b/gi,
};
export const VOCAB_THRESH = 3;
 
// Artefacts of the WeChat framework itself: package format, markup languages, API namespace. Any one
// hit enters, because a paper that unpacked 40,000 packages may still call them "apps".
export const ARTEFACT_RE = /wxapkg|\bWXML\b|\bWXSS\b|\bwx\.(request|login|getUserInfo)|JSSDK|JS-SDK/g;
 
// Non-Chinese host-app programmes, case-sensitive product names.
export const OTHER_HOST_RE = /Telegram Mini Apps?|Snap Minis|LINE MINI|Instant Games/g;
 
// Schema channel: the extraction's own free-text fields.
export const SCHEMA_RE = /mini[- ]?programs?|mini[- ]?apps?|super[- ]?apps?|app[- ]in[- ]app|wechat|alipay|mini[- ]?games?/i;
 
// The host names, for the per-paper host tally (case-sensitive: "line", "grab", "qq" are words).
export const HOSTS = {
  WeChat: /\bWeChat\b|\bWeixin\b/g,
  Alipay: /\bAlipay\b|\bAliPay\b/g,
  Baidu: /\bBaidu\b/g,
  'TikTok/Douyin': /\bTikTok\b|\bDouyin\b/g,
  QQ: /\bQQ\b/g,
};
 
export function textPath(p) {
  return path.join(dataRoot(), 'fulltext', String(p.year), p.venue, p.slug, 'paper.cols.txt');
}
export function readCollapsed(p) {
  const f = textPath(p);
  if (!fs.existsSync(f)) return null;
  return fs.readFileSync(f, 'latin1').replace(/\s+/g, ' ');
}
export const count = (t, re) => (t.match(new RegExp(re.source, re.flags.includes('g') ? re.flags : re.flags + 'g')) || []).length;
 
export function schemaHit(p) {
  const fields = [];
  if (SCHEMA_RE.test(p.title)) fields.push('title');
  for (const x of p.population) if (SCHEMA_RE.test(String(x.sourceList)) || SCHEMA_RE.test(String(x.unit))) fields.push('population');
  for (const x of p.detection) if (SCHEMA_RE.test(String(x.phenomenon))) fields.push('detection');
  for (const x of p.tools) if ((x.usedOrMentioned === 'used' || x.usedOrMentioned === 'produced') && SCHEMA_RE.test(String(x.name))) fields.push('tools');
  for (const x of p.classification) if (SCHEMA_RE.test(String(x.resourceName)) || SCHEMA_RE.test(String(x.target))) fields.push('classification');
  return [...new Set(fields)];
}
 
export const key = (p) => `${p.venue}/${p.year}/${p.slug}`;
 
// One pass over the corpus. Returns per-paper probe counts for every paper with text, plus the
// number of records with no text file, so no paper drops silently out of a denominator.
export function runProbes(rows) {
  const out = [];
  let missing = 0;
  for (const p of rows) {
    const t = readCollapsed(p);
    if (t === null) { missing += 1; continue; }
    const r = { key: key(p), p, gap: count(t, GAP_RE), artefact: count(t, ARTEFACT_RE), otherHost: count(t, OTHER_HOST_RE), schema: schemaHit(p), vocab: {}, hosts: {} };
    for (const [n, re] of Object.entries(VOCAB)) r.vocab[n] = count(t, re);
    for (const [n, re] of Object.entries(HOSTS)) r.hosts[n] = count(t, re);
    r.vocabSum = Object.values(r.vocab).reduce((a, b) => a + b, 0);
    out.push(r);
  }
  return { hits: out, missing };
}
 
export const channels = (r) => {
  const c = [];
  if (r.gap >= GAP_THRESH) c.push('gap');
  if (r.vocabSum >= VOCAB_THRESH) c.push('vocab');
  if (r.artefact >= 1) c.push('artefact');
  if (r.otherHost >= 1) c.push('other-host');
  if (r.schema.length) c.push('schema');
  return c;
};
export const isCandidate = (r) => channels(r).length > 0;

The inclusion rule, the 57 verdicts and the 16 papers' hand codes

mp_fold.mjs
// mp_fold.mjs — inclusion rule, hand verdicts and hand codes behind
// design:mobile_and_app_measurement:mini_programs ("Measuring Super-App Mini-Programs").
// Imported by report_mini_programs.mjs; nothing here is computed, it is the audit surface.
//
// The probes (mp_probes.mjs) produce a CANDIDATE set of 57 papers. Every one has a verdict below,
// and the report throws if a candidate lacks a verdict or a verdict names a non-candidate. One paper
// outside the candidate set was read because a recall probe surfaced it (RECALL_READ).
//
// Verdicts for the papers read in full were made from sub-agent reading notes
// (notes/mp_papers_[A-D].md, quotes machine-checked by mp_notes_quotecheck.py) and decided by the
// main session; reader verdicts that were overridden carry "judgement call" in their note. The
// others were decided from sentence contexts around every vocabulary hit (_mp_ctx.mjs).
 
export const INCLUSION_RULE = `INCLUSION RULE (written before any verdict was counted):
IN      — mini-programs (third-party programs that run inside a host "super app" and reach the device
          and the user only through APIs the host provides), or the host's mini-program framework and
          its APIs, are measured or analysed. Sub-coded CORE (the main object of the paper) or SECTION
          (one analysed population among several, e.g. one of three platforms).
CONTEXT — adjacent: mini-program source code on GitHub rather than deployed mini-programs; a host
          app's own feature (its in-app browser, its face verification) rather than its mini-programs.
OUT     — WeChat / Alipay only as a messenger, a payment method, an attack target, a dataset owner or
          a recruitment channel; related-work mentions; homographs ("MiniApps" the token, "In-App" purchases).`;
 
export const CODES = {
  IN: {
    core: 'mini-programs or the host framework are the main object',
    section: 'mini-programs are one analysed population among several',
  },
  CONTEXT: {
    'source-repos': 'mini-program source code on GitHub, not deployed mini-programs',
    'host-feature': "a host app's own feature, not its mini-programs",
  },
  OUT: {
    messenger: 'WeChat/Alipay as a messenger, contact point, payment method or login target',
    'operator-data': 'a dataset supplied by the host operator, no mini-programs',
    'host-sdk': "the host's SDK inside ordinary apps",
    'host-app-analysis': 'the super app analysed as an ordinary app (ports, ROM privileges, modified binaries)',
    recruitment: 'WeChat groups or accounts used to recruit participants',
    mention: 'related work or passing mention',
    homograph: 'the probe matched another sense of the word',
  },
};
 
// Read because the recall probe (_mp_recall.mjs) surfaced it; not a candidate, so not in VERDICTS.
export const RECALL_READ = [
  ['USENIX/2023/medusa-attack-exploring-security-hazards-of-in-app-qr-code-scanning', 'OUT', 'unrelated', 'reader D: native apps\' built-in QR-code handlers; no mini-program, super-app or host-name hit in the text (the recall probe fired on "host app" and other app names)'],
];
 
// [key, verdict, code, note]
export const VERDICTS = [
  // ---- read in full (notes/mp_papers_[A-D].md)
  ['CCS/2020/demystifying-resource-management-risks-in-emerging-mobile-app-in-app-ecosystems', 'IN', 'core', 'host framework: 927 sub-app APIs across 11 app-in-app hosts, probed with self-built test sub-apps and Xposed (judgement call: reader IN-SECTION because four hosts are browsers; the host framework is the paper\'s main object, which the rule counts as core)'],
  ['CCS/2022/cross-miniapp-request-forgery-root-causes-attacks-and-vulnerability-detection', 'IN', 'core', 'deployed: 2,571,490 WeChat and 148,512 Baidu miniapps via MiniCrawler; CmrfScanner (DoubleX-based AST analysis)'],
  ['USENIX/2022/identity-confusion-in-webview-based-mobile-app-in-app-ecosystems', 'IN', 'core', 'host framework: 47 super apps found by filtering 6,000 Android apps; static + Xposed (judgement call: reader IN-SECTION because the unit is the host binary; the rule counts the host framework as core)'],
  ['CCS/2023/dont-leak-your-keys-understanding-measuring-and-exploiting-the-appsecret-leaks-i', 'IN', 'core', 'deployed: 3,450,586 WeChat mini-programs via a reverse-engineered search API and 14,020 keywords; 40,880 leak their AppSecret'],
  ['CCS/2023/uncovering-and-exploiting-hidden-apis-in-mobile-super-apps', 'IN', 'core', 'host framework: hidden APIs in five V8-based super apps (APIScope: Soot + Frida); 267,359 WeChat miniapps scanned for usage'],
  ['USENIX/2023/one-size-does-not-fit-all-uncovering-and-exploiting-cross-platform-discrepant-ap', 'IN', 'core', 'host framework: WeChat\'s documented miniapp APIs diffed across Windows, Android and iOS builds (APIDiff)'],
  ['CCS/2024/minicat-understanding-and-detecting-cross-page-request-forgery-vulnerabilities-i', 'IN', 'core', 'deployed: 44,273 WeChat mini-programs pulled from the Windows client cache via pywinauto + jieba keywords; wxappUnpacker + CodeQL'],
  ['NDSS/2026/better-safe-than-sorry-uncovering-the-insecure-resource-management-in-app-in-app-cloud-services', 'IN', 'core', 'deployed: 1,248,815 mini-apps from four super apps "adhered to the methods established in previous studies" (the AppSecret and Skeleton Keys crawls); 22,695 use cloud services; ICREMiner with Gemini'],
  ['PETS/2025/what-wechat-knows-pervasive-first-party-tracking-in-a-billion-user-super-app-eco', 'IN', 'core', 'host traffic: MMTLS reverse-engineered (Frida, Jadx, Ghidra); 104 of 170 popular Mini Programs traced by hand; US/Canadian phone numbers'],
  ['USENIX/2025/i-can-tell-your-secrets-inferring-privacy-attributes-from-mini-app-interaction-h', 'IN', 'core', 'operator logs: AliPay\'s own mini-app usage history for 288,895 consenting users; AliPay co-authors and IRB'],
  ['USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games', 'IN', 'core', 'deployed: 6,769 mini-games crawled with miniCrawler from WeChat, Facebook Instant Games, QuickGame; 2,076 Cocos games statically analysed (MAAD)'],
  ['NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization', 'IN', 'core', 'deployed: 4,595,680 WeChat miniapps (2020-2022, MiniCrawler), revisited for delisting; 19,905 malware'],
  ['NDSS/2025/the-skeleton-keys-a-large-scale-analysis-of-credential-leakage-in-mini-apps', 'IN', 'core', 'deployed: 413,775 mini-apps from six super apps via an extended MiniCrawler; KeyMagnet (JAW-based) + UI Automator'],
  ['USENIX/2024/demystifying-the-security-implications-in-iot-device-rental-services', 'IN', 'section', 'deployed: 75 WeChat mini-programs among 92 IoT rental apps; BurpSuite black-box API testing'],
  ['IEEE-SP/2025/hey-your-secrets-leaked-detecting-and-characterizing-secret-leakage-in-the-wild', 'IN', 'section', 'deployed: 41,719 WeChat mini-programs (March-October 2023) as one of three corpora; KeySentinel'],
  ['CCS/2024/riotfuzzer-companion-app-assisted-remote-fuzzing-for-detecting-vulnerabilities-i', 'IN', 'section', 'host framework: the mini-app-to-Java bridge in four IoT all-in-one apps, 27 purchased devices; Frida + ChatGPT filtering (judgement call: reader IN-CORE; the object is the IoT device, the mini-app layer is what the fuzzer instruments around)'],
  // ---- decided from sentence contexts (out/mp/triage_ctx.md)
  ["CCS/2017/poster-neural-network-based-graph-embedding-for-malicious-accounts-detection", "OUT", "operator-data", "Alipay account-device graph supplied and deployed by the operator (Ant Financial authors); no mini-programs"],
  ["WWW/2017/appholmes-detecting-and-characterizing-app-collusion-among-third-party-android-m", "OUT", "host-sdk", "the WeChat/QQ SDK inside third-party Android apps (app collusion); predates mini-programs"],
  ["CCS/2019/detecting-fake-accounts-in-online-social-networks-at-the-time-of-registrations", "OUT", "operator-data", "WeChat registration data supplied by the operator; Sybil detection deployed by WeChat"],
  ["CCS/2019/tokenscope-automatically-detecting-inconsistent-behaviors-of-cryptocurrency-toke", "OUT", "homograph", "\"MiniApps\" is the name of an Ethereum token"],
  ["IEEE-SP/2019/stealthy-porn-understanding-real-world-adversarial-images-for-illicit-online-pro", "OUT", "messenger", "WeChat IDs and QR codes as contact points in promotional images"],
  ["NDSS/2019/understanding-open-ports-in-android-applications-discovery-diagnosis-and-security-assessment", "OUT", "host-app-analysis", "WeChat as one Android app with an open port"],
  ["USENIX/2019/devils-in-the-guidance-predicting-logic-vulnerabilities-in-payment-syndication-s", "OUT", "messenger", "Alipay/WeChat Pay as payment services in syndication SDKs"],
  ["WWW/2019/no-more-than-what-i-post-preventing-linkage-attacks-on-check-in-services", "OUT", "operator-data", "WeChat check-in dataset"],
  ["CCS/2020/lies-in-the-air-characterizing-fake-base-station-spam-ecosystem-in-china", "OUT", "messenger", "WeChat accounts as contact points in SMS spam"],
  ["CCS/2021/hello-its-me-deep-learning-based-speech-synthesis-attacks-in-the-real-world", "OUT", "messenger", "WeChat voiceprint login as an attack target"],
  ["CCS/2023/password-stealing-without-hacking-wi-fi-enabled-practical-keystroke-eavesdroppin", "OUT", "messenger", "WeChat Pay password as an attack target"],
  ["CCS/2023/poster-longitudinal-measurement-of-the-adoption-dynamics-in-apples-privacy-label", "OUT", "homograph", "\"In-App\" purchases, matched by the app-in-app probe"],
  ["CCS/2023/understanding-and-detecting-abused-image-hosting-modules-as-malicious-services", "CONTEXT", "source-repos", "GitHub repositories of mini-programs are one of three categories of abusers of image-hosting upload APIs; deployed mini-programs not analysed"],
  ["USENIX/2023/a-study-of-chinas-censorship-and-its-evasion-through-the-lens-of-online-gaming", "OUT", "recruitment", "survey posted to QQ/WeChat \"micro-applications\" and groups"],
  ["NDSS/2024/leaking-the-privacy-of-groups-and-more-understanding-privacy-risks-of-cross-app-content-sharing-in-mobile-ecosystem", "OUT", "host-sdk", "cross-app sharing into WeChat via its SDK; no mini-programs"],
  ["NDSS/2024/maginot-line-assessing-a-new-cross-app-threat-to-pii-as-factor-authentication-in-chinese-mobile-apps", "OUT", "host-app-analysis", "Alipay/WeChat as native apps with PII-as-factor authentication"],
  ["CCS/2024/a-first-look-at-security-and-privacy-risks-in-the-rapidapi-ecosystem", "OUT", "mention", "related work (app-in-app, AppSecret leaks)"],
  ["PETS/2024/honesty-is-the-best-policy-on-the-accuracy-of-apple-privacy-labels-compared-to-a", "OUT", "homograph", "\"In-App\" purchases, matched by the app-in-app probe"],
  ["IMC/2024/whatcha-lookin-at-investigating-third-party-web-content-in-popular-android-apps", "OUT", "mention", "related work (identity confusion; SaTS workshop paper)"],
  ["USENIX/2024/can-i-hear-your-face-pervasive-attack-on-voice-authentication-systems-with-a-sin", "OUT", "messenger", "WeChat voiceprint as an attack target"],
  ["USENIX/2024/famos-robust-privacy-preserving-authentication-on-payment-apps-via-federated-mul", "OUT", "operator-data", "Alipay behavioural dataset for an authentication system"],
  ["WWW/2024/phishinwebview-analysis-of-anti-phishing-entities-in-mobile-apps-with-webview-ta", "CONTEXT", "host-feature", "WeChat's in-app browser does not use Google Safe Browsing in China; the host app's web container, not its mini-programs"],
  ["CCS/2025/digital-safety-for-children-with-intellectual-disabilities-when-using-mobile-dev", "OUT", "recruitment", "WeChat public accounts and groups used to recruit"],
  ["IEEE-SP/2025/from-one-stolen-utterance-assessing-the-risks-of-voice-cloning-in-the-aigc-era", "OUT", "messenger", "WeChat voice login as an attack target"],
  ["CCS/2025/deep-dive-into-in-app-browsers-uncovering-hidden-pitfalls-in-certificate-validat", "CONTEXT", "host-feature", "in-app browsers incl. WeChat 8.0.51; WeChat would not launch on Android 11/15 emulators and was run on BlueStacks \u2014 a host-app instrumentation pitfall"],
  ["IEEE-SP/2025/prevalence-overshadows-concerns-understanding-chinese-users-privacy-awareness-an", "OUT", "mention", "mini-programs named as a channel to LLM services"],
  ["IEEE-SP/2025/gptracker-a-large-scale-measurement-of-misused-gpts", "OUT", "mention", "related work (WeChat Mini-App stores)"],
  ["IEEE-SP/2025/born-with-a-silver-spoon-on-the-in-security-of-native-granted-app-privileges-in", "OUT", "host-app-analysis", "custom ROMs granting privileges to super apps by package name; no mini-programs"],
  ["NDSS/2025/eagleye-exposing-hidden-web-interfaces-in-iot-devices-via-routing-analysis", "OUT", "mention", "related work (APIScope)"],
  ["PETS/2025/can-social-media-privacy-and-safety-features-protect-targets-of-interpersonal-at", "OUT", "host-app-analysis", "WeChat's privacy/safety features as one of ten social apps"],
  ["USENIX/2025/privacy-law-enforcement-under-centralized-governance-a-qualitative-analysis-of-f", "OUT", "recruitment", "five WeChat groups used to recruit interviewees about rectification campaigns"],
  ["NDSS/2026/characterizing-the-implementation-of-censorship-policies-in-chinese-llm-services", "OUT", "mention", "WeChat censorship as prior work"],
  ["NDSS/2026/chameleoscan-demystifying-and-detecting-ios-chameleon-apps-via-llm-powered-ui-exploration", "OUT", "messenger", "WeChat official accounts as a source of chameleon-app promotions"],
  ["USENIX/2026/cracks-in-the-walled-garden-dissecting-the-gray-market-of-unauthorized-ios-app-d", "OUT", "host-app-analysis", "140 modified WeChat IPA variants and their injected dylibs; the host binary, not mini-programs"],
  ["WWW/2025/detecting-and-understanding-the-promotion-of-illicit-goods-and-services-on-twitt", "OUT", "messenger", "WeChat accounts as contact points in tweets"],
  ["CCS/2025/leaky-apps-large-scale-analysis-of-secrets-distributed-in-android-and-ios-apps", "OUT", "mention", "contrasts itself with super-app secret-leak work; native apps only"],
  ["IEEE-SP/2025/the-file-that-contained-the-keys-has-been-removed-an-empirical-analysis-of-secre", "OUT", "mention", "related work (mini-program secret leaks)"],
  ["PETS/2026/designing-reflective-thinking-based-contextual-privacy-policy-for-mobile-applica", "OUT", "mention", "a participant quote and future work on WeChat microapps"],
  ["PETS/2026/personal-data-flows-and-privacy-policy-traceability-in-third-party-llm-apps-in-t", "OUT", "homograph", "\"Miniapp\" is miniapps.ai, an LLM-app platform"],
  ["USENIX/2026/turn-your-face-into-an-attack-surface-screen-attack-using-facial-reflections-in", "OUT", "messenger", "WeChat video calls as a conferencing platform"],
  ["IEEE-SP/2023/understanding-the-in-security-of-cross-side-face-verification-systems-in-mobile", "CONTEXT", "host-feature", "the host app's face-verification service, which sub-apps can invoke; mini-programs themselves not analysed"],
];
 
export const READ_IN_FULL = VERDICTS.slice(0, 16).map((v) => v[0]);
 
// ---------------------------------------------------------------- hand codes for the IN papers
// From each paper's own text via the reading notes. Values are small closed vocabularies so the
// report can tally them; "not-stated" means the paper does not say, never that it was not done.
//   unit            deployed (a population of deployed mini-programs) | host-framework (the host's APIs
//                   or binary, probed with the authors' own test mini-programs) | host-traffic (what the
//                   host itself sends) | usage-logs (the operator's logs of users' mini-program use)
//   acquisition     MiniCrawler | search-api (the host's internal search endpoint, reverse-engineered) |
//                   client-cache (packages the desktop client writes to disk, GUI-automated) |
//                   prior-crawl (reuses an earlier paper's collection method) | rankings | keyword-search
//                   (by hand in the app) | qr-codes | store-revisit | audit-set | official-docs |
//                   own-test-miniapps | app-store-crawl | purchased-devices | operator-logs
//   analysis        static | dynamic | manual | ml
//   hostVersion     pinned (a host-app version is stated) | os-only | not-stated | n/a
//   accounts        own-test-accounts | non-Chinese-phone | real-name | operator | not-stated
//   keywords        Chinese (the paper's OWN seed terms or NLP on Chinese text are described; a cited prior method does not
//                   count) | not-stated | n/a. Tallied only over deployed-population papers.
//   validation      manual-sample | manual-all | api-oracle | ground-truth-set | held-out | vendor-confirmation | manual-coding
//   denominator     yes (the paper says its sample is not the population, or why) | no
//   artifacts       code | code+data | gated-data | declined (discussed and withheld) | none-stated. Two readers missed a release URL printed only in
//                   the reference list (KeySentinel, RIoTFuzzer); the report's schema cross-check caught both.
export const HAND_FIELDS = ['unit', 'hosts', 'side', 'acquisition', 'collected', 'analysed', 'analysis', 'tools', 'instrumentation', 'hostVersion', 'accounts', 'keywords', 'validation', 'llm', 'irb', 'disclosure', 'artifacts', 'denominator'];
 
export const HAND = {
  'CCS/2020/demystifying-resource-management-risks-in-emerging-mobile-app-in-app-ecosystems': {
    unit: 'host-framework', hosts: ['WeChat', 'QQ', 'Baidu', 'TikTok/Douyin', 'other'], side: ['vulnerability'],
    acquisition: ['own-test-miniapps', 'official-docs'], collected: '11 hosts', analysed: '927 sub-app APIs',
    analysis: ['dynamic'], tools: ['Xposed', 'own test sub-apps'], instrumentation: ['Xposed'], hostVersion: 'os-only',
    accounts: 'not-stated', keywords: 'n/a', validation: 'manual-all', llm: 'no', irb: 'none-mentioned',
    disclosure: ['host-vendor', 'bug-bounty'], artifacts: 'code', denominator: 'no',
  },
  'CCS/2022/cross-miniapp-request-forgery-root-causes-attacks-and-vulnerability-detection': {
    unit: 'deployed', hosts: ['WeChat', 'Baidu'], side: ['vulnerability'],
    acquisition: ['MiniCrawler'], collected: '2,571,490 WeChat + 148,512 Baidu', analysed: 'all (static)',
    analysis: ['static'], tools: ['MiniCrawler', 'DoubleX', 'CmrfScanner'], instrumentation: [], hostVersion: 'not-stated',
    accounts: 'own-test-accounts', keywords: 'not-stated', validation: 'manual-sample', llm: 'no', irb: 'none-mentioned',
    disclosure: ['host-vendor'], artifacts: 'code', denominator: 'no',
  },
  'USENIX/2022/identity-confusion-in-webview-based-mobile-app-in-app-ecosystems': {
    unit: 'host-framework', hosts: ['WeChat', 'Alipay', 'Baidu', 'TikTok/Douyin', 'other'], side: ['vulnerability'],
    acquisition: ['app-store-crawl', 'own-test-miniapps'], collected: '6,000 Android apps', analysed: '47 super apps',
    analysis: ['static', 'dynamic'], tools: ['Xposed'], instrumentation: ['Xposed'], hostVersion: 'not-stated',
    accounts: 'own-test-accounts', keywords: 'n/a', validation: 'manual-sample', llm: 'no', irb: 'none-mentioned',
    disclosure: ['host-vendor', 'bug-bounty'], artifacts: 'none-stated', denominator: 'no',
  },
  'CCS/2023/dont-leak-your-keys-understanding-measuring-and-exploiting-the-appsecret-leaks-i': {
    unit: 'deployed', hosts: ['WeChat', 'Baidu'], side: ['vulnerability'],
    acquisition: ['search-api'], collected: '3,450,586 WeChat + 171,989 Baidu', analysed: 'all (static)',
    analysis: ['static'], tools: ['regex + key-validation API'], instrumentation: [], hostVersion: 'not-stated',
    accounts: 'own-test-accounts', keywords: 'Chinese', validation: 'api-oracle', llm: 'no', irb: 'none-mentioned',
    disclosure: ['host-vendor', 'bug-bounty'], artifacts: 'none-stated', denominator: 'yes',
  },
  'CCS/2023/uncovering-and-exploiting-hidden-apis-in-mobile-super-apps': {
    unit: 'host-framework', hosts: ['WeChat', 'QQ', 'Baidu', 'TikTok/Douyin', 'other'], side: ['vulnerability', 'host-API'],
    acquisition: ['official-docs', 'MiniCrawler'], collected: '5 host APKs; 267,359 WeChat miniapps', analysed: '1,829 API candidates',
    analysis: ['static', 'dynamic'], tools: ['Soot', 'Frida', 'MiniCrawler'], instrumentation: ['Frida'], hostVersion: 'pinned',
    accounts: 'own-test-accounts', keywords: 'not-stated', validation: 'manual-all', llm: 'no', irb: 'none-mentioned',
    disclosure: ['host-vendor', 'bug-bounty'], artifacts: 'none-stated', denominator: 'yes',
  },
  'USENIX/2023/one-size-does-not-fit-all-uncovering-and-exploiting-cross-platform-discrepant-ap': {
    unit: 'host-framework', hosts: ['WeChat'], side: ['vulnerability', 'host-API', 'privacy'],
    acquisition: ['official-docs', 'own-test-miniapps'], collected: '1,031 documented APIs', analysed: 'all, on Windows/Android/iOS',
    analysis: ['dynamic'], tools: ['Frida', 'JEB', 'IDA Pro'], instrumentation: ['Frida'], hostVersion: 'os-only',
    accounts: 'own-test-accounts', keywords: 'n/a', validation: 'manual-all', llm: 'no', irb: 'none-mentioned',
    disclosure: ['host-vendor', 'bug-bounty'], artifacts: 'code', denominator: 'no',
  },
  'CCS/2024/minicat-understanding-and-detecting-cross-page-request-forgery-vulnerabilities-i': {
    unit: 'deployed', hosts: ['WeChat'], side: ['vulnerability'],
    acquisition: ['client-cache'], collected: '44,273', analysed: '41,726',
    analysis: ['static'], tools: ['pywinauto', 'jieba', 'wxappUnpacker', 'CodeQL'], instrumentation: [], hostVersion: 'not-stated',
    accounts: 'own-test-accounts', keywords: 'Chinese', validation: 'manual-sample', llm: 'no', irb: 'none-mentioned',
    disclosure: ['developers', 'CERT'], artifacts: 'code', denominator: 'yes',
  },
  'NDSS/2026/better-safe-than-sorry-uncovering-the-insecure-resource-management-in-app-in-app-cloud-services': {
    unit: 'deployed', hosts: ['WeChat', 'TikTok/Douyin', 'Alipay', 'Baidu', 'other'], side: ['vulnerability', 'privacy'],
    acquisition: ['prior-crawl'], collected: '1,248,815', analysed: '22,695 using cloud services',
    analysis: ['static', 'dynamic', 'manual'], tools: ['ICREMiner', 'Gemini'], instrumentation: [], hostVersion: 'not-stated',
    accounts: 'own-test-accounts', keywords: 'not-stated', validation: 'manual-all', llm: 'yes', irb: 'approval',
    disclosure: ['host-vendor', 'developers'], artifacts: 'code+data', denominator: 'no',
  },
  'PETS/2025/what-wechat-knows-pervasive-first-party-tracking-in-a-billion-user-super-app-eco': {
    unit: 'host-traffic', hosts: ['WeChat'], side: ['privacy'],
    acquisition: ['rankings', 'keyword-search'], collected: '170', analysed: '104',
    analysis: ['dynamic', 'manual'], tools: ['Frida', 'Jadx', 'Ghidra', 'IDA Pro'], instrumentation: ['Frida'], hostVersion: 'pinned',
    accounts: 'non-Chinese-phone', keywords: 'not-stated', validation: 'manual-coding', llm: 'no', irb: 'none-mentioned',
    disclosure: ['host-vendor'], artifacts: 'code+data', denominator: 'yes',
  },
  'USENIX/2025/i-can-tell-your-secrets-inferring-privacy-attributes-from-mini-app-interaction-h': {
    unit: 'usage-logs', hosts: ['Alipay'], side: ['privacy'],
    acquisition: ['operator-logs'], collected: '288,895 users', analysed: '219,826 users',
    analysis: ['ml'], tools: ['THEFT (DNN)'], instrumentation: [], hostVersion: 'n/a',
    accounts: 'operator', keywords: 'n/a', validation: 'held-out', llm: 'no', irb: 'approval',
    disclosure: ['host-vendor'], artifacts: 'declined', denominator: 'yes',
  },
  'USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games': {
    unit: 'deployed', hosts: ['WeChat', 'other'], side: ['advertising'],
    acquisition: ['MiniCrawler', 'audit-set'], collected: '6,769', analysed: '2,076 Cocos',
    analysis: ['static'], tools: ['MiniCrawler', 'Esprima', 'swc', 'WALA', 'XLM-RoBERTa'], instrumentation: [], hostVersion: 'not-stated',
    accounts: 'not-stated', keywords: 'not-stated', validation: 'ground-truth-set', llm: 'no', irb: 'none-mentioned',
    disclosure: ['host-vendor'], artifacts: 'code', denominator: 'yes',
  },
  'NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization': {
    unit: 'deployed', hosts: ['WeChat'], side: ['malware'],
    acquisition: ['MiniCrawler', 'store-revisit'], collected: '4,595,680', analysed: 'all (static)',
    analysis: ['static'], tools: ['MiniCrawler'], instrumentation: [], hostVersion: 'not-stated',
    accounts: 'not-stated', keywords: 'not-stated', validation: 'manual-sample', llm: 'no', irb: 'none-mentioned',
    disclosure: ['host-vendor'], artifacts: 'gated-data', denominator: 'no',
  },
  'NDSS/2025/the-skeleton-keys-a-large-scale-analysis-of-credential-leakage-in-mini-apps': {
    unit: 'deployed', hosts: ['WeChat', 'Baidu', 'Alipay', 'TikTok/Douyin', 'other'], side: ['vulnerability'],
    acquisition: ['MiniCrawler'], collected: '413,775', analysed: '402,527',
    analysis: ['static', 'dynamic'], tools: ['MiniCrawler', 'JAW', 'UI Automator'], instrumentation: [], hostVersion: 'not-stated',
    accounts: 'own-test-accounts', keywords: 'not-stated', validation: 'manual-sample', llm: 'no', irb: 'approval',
    disclosure: ['host-vendor', 'developers', 'CVE'], artifacts: 'code', denominator: 'no',
  },
  'USENIX/2024/demystifying-the-security-implications-in-iot-device-rental-services': {
    unit: 'deployed', hosts: ['WeChat'], side: ['vulnerability'],
    acquisition: ['keyword-search', 'qr-codes'], collected: '75 mini-programs', analysed: '75',
    analysis: ['dynamic', 'manual'], tools: ['BurpSuite'], instrumentation: [], hostVersion: 'not-stated',
    accounts: 'real-name', keywords: 'not-stated', validation: 'manual-all', llm: 'no', irb: 'none-mentioned',
    disclosure: ['developers', 'CVE'], artifacts: 'none-stated', denominator: 'yes',
  },
  'IEEE-SP/2025/hey-your-secrets-leaked-detecting-and-characterizing-secret-leakage-in-the-wild': {
    unit: 'deployed', hosts: ['WeChat'], side: ['vulnerability'],
    acquisition: ['prior-crawl'], collected: '41,719', analysed: '41,719',
    analysis: ['static'], tools: ['KeySentinel'], instrumentation: [], hostVersion: 'not-stated',
    accounts: 'not-stated', keywords: 'not-stated', validation: 'manual-sample', llm: 'no', irb: 'none-mentioned',
    disclosure: ['developers'], artifacts: 'code', denominator: 'no',
  },
  'CCS/2024/riotfuzzer-companion-app-assisted-remote-fuzzing-for-detecting-vulnerabilities-i': {
    unit: 'host-framework', hosts: ['IoT host'], side: ['vulnerability'],
    acquisition: ['purchased-devices'], collected: '27 devices', analysed: '27',
    analysis: ['static', 'dynamic'], tools: ['Apktool', 'Androguard', 'Frida', 'mitmproxy', 'ChatGPT'], instrumentation: ['Frida'], hostVersion: 'pinned',
    accounts: 'own-test-accounts', keywords: 'n/a', validation: 'vendor-confirmation', llm: 'yes', irb: 'none-mentioned',
    disclosure: ['host-vendor', 'CVE'], artifacts: 'code', denominator: 'no',
  },
};
 
// Mini-program papers in the venue index (corpus2/.meta) that never reached the extraction. The
// report finds them by title and throws if one appears that has no entry here.
export const OUTSIDE_EXTRACTION = [
  { slug: 'raising-the-flag-detecting-missing-permission-controls-in-mini-program-apis', why: 'selected by screening; PDF not retrieved (empty fulltext directory)', use: 'read in full from the USENIX PDF (out/mp/ext/wei_zhiao.pdf); cited, not counted' },
  { slug: 'convenience-at-a-cost-the-security-risks-of-template-based-development-in-the-ap', why: 'IEEE S&P 2026: no abstract in OpenAlex, so never screened (under-selected by construction)', use: 'abstract only (Semantic Scholar); cited, not counted' },
  { slug: 'real-or-rogue-detecting-malicious-miniapps-with-deceptive-reporting-interface', why: 'TheWebConf 2026: no abstract in OpenAlex, so never screened (under-selected by construction)', use: 'abstract only (Semantic Scholar; ACM DL is Cloudflare-walled); cited, not counted' },
  { slug: 'sats23-the-1st-acm-workshop-on-secure-and-trustworthy-superapps', why: 'workshop front matter, screened out (correctly)', use: 'evidence that a dedicated CCS workshop exists' },
  { slug: 'sats-24-the-2nd-acm-workshop-on-secure-and-trustworthy-superapps', why: 'workshop front matter, screened out (correctly)', use: 'evidence that a dedicated CCS workshop exists' },
  { slug: 'sats-25-the-3rd-acm-workshop-on-security-and-privacy-of-ai-empowered-mobile-supe', why: 'workshop front matter, screened out (correctly)', use: 'evidence that a dedicated CCS workshop exists' },
];

The figure and quote verifier

verify_mini_programs_figures.mjs
// verify_mini_programs_figures.mjs — every quoted span and per-paper figure on
// design:mobile_and_app_measurement:mini_programs, checked against its source.
//
//   node scripts/verify_mini_programs_figures.mjs > scripts/verify_mini_programs_figures-output.txt
//
// Two passes, so that neither a new quote nor a new figure can reach the page unchecked:
//   A. SPANS — every //"…"// span is pulled out of the page source automatically. Each must be
//      (i) verbatim in the paper of a citekey on the same line (nearest by character offset first),
//      (ii) listed in EXTERNAL with the label of a check in external_checks_mini_programs.sh
//      that printed OK, or (iii) listed in NOT_A_QUOTE with a reason. Anything else is FAIL.
//   B. NEEDLES — per-paper figures (numbers the page attributes to a paper) with a needle long
//      enough to be specific. Needles under 20 characters print WEAK.
// Lookup order: paper.cols.txt, paper.norm.txt, paper.txt, then a pypdf re-extraction of paper.pdf;
// whitespace collapsed, ligatures, dashes and curly quotes folded. The winning rendering is printed.
// wei2026_raising is outside the extraction: its text is the publisher PDF extracted to
// out/mp/ext/wei_zhiao.txt by external_checks_mini_programs.sh. Mutated needles must NOT be found.
// Exits 1 on any failure.
import fs from 'node:fs';
import path from 'node:path';
import { execFileSync } from 'node:child_process';
import { dataRoot } from './lib.mjs';
 
const PAGE = 'pages/design_mobile_and_app_measurement_mini_programs.txt';
const EXT_OUT = 'scripts/external_checks_mini_programs-output.txt';
const ROOT = path.join(dataRoot(), 'fulltext');
const fold = (s) => s
  .replace(/ff/g, 'ff').replace(/fi/g, 'fi').replace(/fl/g, 'fl').replace(/ffi/g, 'ffi').replace(/ffl/g, 'ffl')
  .replace(/[‐-―−]/g, '-').replace(/[‘’]/g, "'").replace(/[“”]/g, '"').replace(/­/g, '')
  .replace(/\u0000/g, '').replace(/(\w)- (\w)/g, '$1$2').replace(/\s+/g, ' ')
  // An unmapped ligature glyph survives as U+FFFD (or its mojibake) in both .cols and pypdf; mark it.
  .replace(/\uFFFD|�/g, '§');
const esc = (x) => x.replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
// Needle variants tried in order: exact; ligature letters allowed to match the § marker; a hyphen
// between letters dropped (the text-side fold joins "self- constructed" into "selfconstructed").
function variants(n) {
  const lig = new RegExp(esc(n).replace(/ffi|ffl|ff|fi|fl/g, (m) => `(?:${m}|§)`));
  const dehy = n.replace(/(\w)-(\w)/g, '$1$2');
  return [(t) => t.includes(n), (t) => lig.test(t), (t) => t.includes(dehy)];
}
 
// citekey -> corpus key <venue>/<year>/<slug>, or EXT:<path> for a paper read outside the extraction
const KEYMAP = {
  lu2020_demystifying: 'CCS/2020/demystifying-resource-management-risks-in-emerging-mobile-app-in-app-ecosystems',
  yang2022_cross: 'CCS/2022/cross-miniapp-request-forgery-root-causes-attacks-and-vulnerability-detection',
  zhang2022_identity: 'USENIX/2022/identity-confusion-in-webview-based-mobile-app-in-app-ecosystems',
  zhang2023_leak: 'CCS/2023/dont-leak-your-keys-understanding-measuring-and-exploiting-the-appsecret-leaks-i',
  wang2023_uncovering: 'CCS/2023/uncovering-and-exploiting-hidden-apis-in-mobile-super-apps',
  wang2023_size: 'USENIX/2023/one-size-does-not-fit-all-uncovering-and-exploiting-cross-platform-discrepant-ap',
  zhang2024_minicat: 'CCS/2024/minicat-understanding-and-detecting-cross-page-request-forgery-vulnerabilities-i',
  shi2026_better: 'NDSS/2026/better-safe-than-sorry-uncovering-the-insecure-resource-management-in-app-in-app-cloud-services',
  wang2025_wechat: 'PETS/2025/what-wechat-knows-pervasive-first-party-tracking-in-a-billion-user-super-app-eco',
  cai2025_tell: 'USENIX/2025/i-can-tell-your-secrets-inferring-privacy-attributes-from-mini-app-interaction-h',
  chen2026_minigames: 'USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games',
  yang2025_miniapp: 'NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization',
  shi2025_skeleton: 'NDSS/2025/the-skeleton-keys-a-large-scale-analysis-of-credential-leakage-in-mini-apps',
  he2024_demystifying: 'USENIX/2024/demystifying-the-security-implications-in-iot-device-rental-services',
  zhou2025_secrets: 'IEEE-SP/2025/hey-your-secrets-leaked-detecting-and-characterizing-secret-leakage-in-the-wild',
  liu2024_riotfuzzer: 'CCS/2024/riotfuzzer-companion-app-assisted-remote-fuzzing-for-detecting-vulnerabilities-i',
  lee2025_deep: 'CCS/2025/deep-dive-into-in-app-browsers-uncovering-hidden-pitfalls-in-certificate-validat',
  wei2026_raising: 'EXT:out/mp/ext/wei_zhiao.txt',
  zhang2021_measurement: 'EXT:out/mp/ext/zhang2021_abstract.txt',
};
 
// Spans quoted from non-corpus sources: span -> label of the external check that must print OK.
const EXTERNAL = {
  '微信私有协议': 'WeChat network: cloud hosting via callContainer over the private protocol',
  '接口应在服务器端调用,不可在前端(小程序、网页、APP等)直接调用': 'code2Session: server-side only',
  'WeChat 7.0.19': 'MiniCrawler README pins WeChat 7.0.19 (metadata)',
  '7.0.20': 'MiniCrawler README pins WeChat 7.0.20 (download)',
  'requires additional consent and agreement': 'MiniSec datasets: AppSecret set requires additional consent and agreement',
  'due to potential legal implications': 'TaintMini README: unpacker not provided due to potential legal implications',
  '使用插件、外挂或非经腾讯授权的第三方工具': 'WeChat licence 8.2.1.4: no plug-ins or unauthorised third-party tools',
  'Agentic/LLM-based techniques': 'SaTS 26 topics include Agentic/LLM-based techniques',
  'launched right inside Telegram': 'Telegram Mini Apps: launched right inside Telegram',
  'can completely replace any website': 'Telegram Mini Apps: can completely replace any website',
  'should not be trusted': 'Telegram Mini Apps: initDataUnsafe should not be trusted',
  'selects an available version and obtains its code package': 'TikTok mini games: app selects a version and obtains its code package',
};
// Spans that are not quotations of any source.
const NOT_A_QUOTE = {
};
 
// Per-paper figures: [citekey, needle, what the page uses it for]
const N = [
  ['wang2023_uncovering', 'WeChat has 39 hidden unchecked APIs (7.77%) that invoke Android APIs protected by permissions', 'hidden unchecked APIs in all five hosts'],
  ['wang2023_uncovering', 'Tiktok has 32 (26.23%), and QQ has 38 (12.88%) such APIs', 'all five hosts'],
  ['wang2023_uncovering', 'Total 78,974 267,359 29.54 Table 6: The 3rd party WeChat miniapps that have used the undocumented APIs', '78,974 of 267,359 (29.54%) invoke a hidden API'],
  ['wang2023_uncovering', 'We collected 267, 359 miniapps using Mini-Crawler', '267,359 miniapps via Mini-Crawler'],
  ['wang2023_uncovering', 'APIScope recognized in total 1,829 API candidates', '1,829 API candidates'],
  ['wang2023_uncovering', 'Google Pixel 4 running Android 11', 'devices (host version pinned)'],
  ['lu2020_demystifying', 'Apinat found 39 System Resource Exposure flaws', '39 SRE flaws'],
  ['lu2020_demystifying', 'generated test cases for all 927 sub-app APIs', '927 sub-app APIs'],
  ['lu2020_demystifying', 'on a Google Pixel 2 phone and iPhone 8 against 11 host apps', '11 hosts'],
  ['zhang2023_leak', 'resulting in 14,020 keywords that retrieved 3,450,586 mini-programs', '14,020 keywords, 3,450,586'],
  ['zhang2023_leak', 'We employed 1,000 commonly used Chinese characters and 1,000 commonly used English words as seed keywords', 'seeds'],
  ['zhang2023_leak', 'we maintained a maximum of six requests per minute, taking over six months to collect them all', '6 per minute, six months'],
  ['zhang2023_leak', 'we identified 40,880 mini-programs that leaked their MKs', '40,880'],
  ['zhang2023_leak', 'the WeChat market has about 4 million mini-programs', 'about 4 million'],
  ['zhang2023_leak', '171,989 Baidu mini-programs', 'Baidu 171,989'],
  ['zhang2023_leak', 'ensure no false positives by validating the correctness of keys using API', 'API oracle'],
  ['wang2023_size', 'more than 4.3 million', 'more than 4.3 million'],
  ['zhang2024_minicat', 'We collected and successfully unpacked 44,273 WeChat mini-programs using Mini-Program Crawler', '44,273'],
  ['zhang2024_minicat', 'MiniCAT successfully analyzed 41,726/44,273 (94.2%), identifying 13,349/41,726', '41,726 and 13,349'],
  ['zhang2024_minicat', '32.0% (13,349/41,726) of analyzable mini-programs are potentially', '32.0%'],
  ['zhang2024_minicat', 'These mini-programs consisted of 2,264,377 pages', '2,264,377 pages'],
  ['zhang2024_minicat', 'enabling an average collection of 2,687 mini-programs daily', '2,687 per day'],
  ['zhang2024_minicat', '20 requests per minute', '20 per minute'],
  ['zhang2024_minicat', 'user_file/Applet/AppID', 'Applet cache directory'],
  ['zhang2024_minicat', 'we used pywinauto', 'pywinauto'],
  ['zhang2024_minicat', 'jieba', 'jieba'],
  ['zhang2024_minicat', 'we used wxappUnpacker (a public WeChat mini-program unpacking tool)', 'wxappUnpacker'],
  ['zhang2024_minicat', 'we did not detect any FPs', 'validation 100/100'],
  ['zhang2024_minicat', 'we discovered 3 FN cases', '3 FN'],
  ['zhou2025_secrets', 'Based on prior research [81], we developed a MP crawler, collecting (March-October 2023) 41,719 projects', '41,719, built on MiniCAT'],
  ['zhou2025_secrets', 'achieving a precision of 91.33% on GitHub, 91.67% on PyPI, and 96.33% on WeChat MPs', '300-sample validation'],
  ['zhou2025_secrets', 'we randomly selected and labeled 300 samples', '300 samples'],
  ['yang2022_cross', 'We obtained 2,571,490 WeChat miniapps in total', '2,571,490'],
  ['yang2022_cross', 'WeChat miniapps and 148,512 Baidu miniapps', '148,512'],
  ['yang2022_cross', 'we have identified 50,281 (95.97%) vulnerable miniapps out of 52,394 miniapps from the WeChat market', '50,281 of 52,394'],
  ['yang2022_cross', 'sampled 100 miniapps from the miniapps identified as not involving extra data to verify false negatives', 'validation 100/100'],
  ['yang2022_cross', 'DoubleX', 'DoubleX'],
  ['shi2025_skeleton', 'Finally, we totally collect 413,775 mini-apps as our database', '413,775'],
  ['shi2025_skeleton', 'We successfully analyze 402,527 mini-apps', '402,527'],
  ['shi2025_skeleton', '3,984 Line mini-apps and 425 VK miniapps', 'LINE, VK'],
  ['shi2025_skeleton', 'randomly selecting 500 mini-apps identified as vulnerable from each super-app', '500 flagged'],
  ['shi2025_skeleton', 'Android UI Automator', 'UI Automator'],
  ['shi2025_skeleton', 'minimal risk', 'IRB minimal risk'],
  ['shi2026_better', 'apply it to 1,248,815 real-world mini-apps', '1,248,815'],
  ['shi2026_better', 'we identified 22,695 mini-apps', '22,695'],
  ['shi2026_better', '2,815 mini-apps (12.40%) are affected by the insecure resource management', '2,815 (12.40%)'],
  ['shi2026_better', 'Gemini', 'Gemini'],
  ['shi2026_better', 'we collected emails for 1,869 mini-apps', '1,869 emailed'],
  ['shi2026_better', '893 mini-app developers have fixed the issues or taken down their mini-apps', '893 fixed'],
  ['shi2026_better', 'this study is considered as "minimal risk"', 'IRB minimal risk'],
  ['shi2026_better', '4,324 Line mini-apps and 862 VK mini-apps', 'LINE, VK'],
  ['yang2025_miniapp', 'amassing 4,595,680 miniapps', '4,595,680'],
  ['yang2025_miniapp', '360,467 miniapps had been removed by the time when finishing our second collection', '360,467 delisted'],
  ['yang2025_miniapp', 'completed the first round of miniapp collection in June 2022', 'June 2022'],
  ['yang2025_miniapp', 'carried out from March 2020 to December 2022', 'March 2020'],
  ['yang2025_miniapp', 'in 19, 905 malware samples', '19,905'],
  ['yang2025_miniapp', 'we sampled a total of 500 miniapps out of the', '500 sampled'],
  ['chen2026_minigames', 'This process yielded 6,769 mini-games in total, from which we further extracted 2,076 Cocos-based', '6,769 and 2,076'],
  ['chen2026_minigames', 'using the open-source tool mini-Crawler', 'miniCrawler'],
  ['chen2026_minigames', '457 / 915 (49.95%)', '457 of 915'],
  ['chen2026_minigames', 'anonymous official audit department', 'audit cases'],
  ['wang2025_wechat', 'We collected and attempted to analyze 170 different popular Mini Programs', '170'],
  ['wang2025_wechat', 'We were able to successfully collect analytics data from 104 Mini Programs', '104'],
  ['wang2025_wechat', 'We used both a rooted Pixel 6 device and an emulator, both running Android 14', 'Pixel 6 + emulator'],
  ['wang2025_wechat', 'We analyzed WeChat app version 8.0.23 and WeChat 8.0.49', 'host version pinned'],
  ['wang2025_wechat', 'We used Frida, a dynamic instrumentation toolkit', 'Frida'],
  ['wang2025_wechat', 'Jadx, a popular Android decompiler', 'Jadx'],
  ['wang2025_wechat', 'We also used Ghidra and IDA Pro', 'Ghidra, IDA'],
  ['wang2025_wechat', "both Baidu's Smart Program ecosystem [57] and Alipay's Mini Program ecosystem [58] advertise similar default tracking features", 'Alipay and Baidu advertise similar'],
  ['wang2025_wechat', 'The account associated with our research was registered to a US phone number', 'US phone number'],
  ['cai2025_tell', 'of which 288,895 users agreed to share their data', '288,895 users'],
  ['cai2025_tell', 'we construct the dataset with 1,099,130 data samples from 219,826 users', '219,826'],
  ['cai2025_tell', 'with IRB approval from AliPay', 'AliPay IRB'],
  ['cai2025_tell', 'We contacted the vendors of all 31 super-apps', '31 super apps'],
  ['he2024_demystifying', 'we get apps for a total of 81 Chinese rentable products (including 75 WeChat mini-programs and 6 Android apps)', '75 of 81'],
  ['he2024_demystifying', 'realname registration and a deposit', 'real-name + deposit'],
  ['he2024_demystifying', 'Burpsuite proxy', 'BurpSuite'],
  ['liu2024_riotfuzzer', 'We apply RIoTFuzzer to 27 IoT devices', '27 devices'],
  ['liu2024_riotfuzzer', 'Xiaomi, Jingdong, Huawei, and Tuya', 'four IoT platforms'],
  ['liu2024_riotfuzzer', 'ChatGPT', 'ChatGPT filtering'],
  ['zhang2022_identity', 'we randomly crawl 6,000 popular Android apps', '6,000 apps'],
  ['zhang2022_identity', 'The first step gives us 47 super-apps', '47 super apps'],
  ['zhang2022_identity', 'Xposed', 'Xposed'],
  ['wang2023_size', 'two Windows-11 desktops, and four smartphones (two with Android-13, and two with iOS-16)', 'platform builds'],
  ['wei2026_raising', 'ranging from 9.5% in Baidu to 19.5% in Alipay', 'context (not on page)'],
  ['wei2026_raising', 'We disclosed all identified vulnerabilities to the Security Response Centers of Tencent, Baidu, and Alipay', 'disclosure'],
  ['lee2025_deep', 'we conducted WeChat-specific experiments on BlueStacks', 'BlueStacks'],
  ['wang2023_size', 'we have 1,031 APIs in total', '1,031 documented APIs'],
  ['yang2022_cross', 'consume 6.29 TB disk storage. We also extended M', '6.29 TB'],
  ['zhang2024_minicat', 'which occupied a storage space of 126.38 GB', '126.38 GB'],
  ['zhang2024_minicat', 'introduced a 5-minute timeout for CodeQL queries', 'five-minute timeout'],
  ['shi2025_skeleton', 'The average precision of KeyMagnet is 95.04% and the recall is 85.56%', 'recall 85.56%'],
  ['chen2026_minigames', 'with a recall of 83.55%', 'recall 83.55%'],
  ['zhou2025_secrets', 'WeChat 902 89.91% 83.38%', 'WeChat benchmark recall 83.38% (Table 14)'],
  ['chen2026_minigames', 'This process produced 371 labeled Ad-behaviors', '371 labelled behaviours'],
  ['zhang2023_leak', 'Api.weixin.qq.com/cgi-bin/token', 'the oracle is the token endpoint'],
  ['chen2026_minigames', 'All test data was', 'crawl compliance assertion'],
  ['shi2026_better', 'we ensure compliance with all relevant laws and regulations', 'compliance with laws'],
  ['he2024_demystifying', "comply with the vendor's bug bounty plan", 'bug-bounty plan'],
  ['wang2023_uncovering', 'Our experiments were conducted primarily in 2021', 'the 975/570 count is from 2021'],
  ['liu2024_riotfuzzer', 'reported in [45], most mini-apps employ obfuscation techniques', 'citing zhang2021 (as is spliced off in .cols)'],
];
const CONTROLS = [
  ['zhang2023_leak', 'we identified 40,881 mini-programs that leaked their MKs'],
  ['zhang2024_minicat', 'MiniCAT successfully analyzed 41,727/44,273 (94.2%)'],
  ['wang2025_wechat', 'we collected and attempted to analyze 171 different popular Mini Programs'],
];
 
const cache = new Map();
function renderings(ck) {
  const k = KEYMAP[ck];
  if (!k) throw new Error(`no KEYMAP entry for ${ck}`);
  if (cache.has(k)) return cache.get(k);
  if (k.startsWith('EXT:')) {
    const f = k.slice(4);
    if (!fs.existsSync(f)) throw new Error(`missing external text ${f} (run external_checks_mini_programs.sh)`);
    const out = [['publisher-pdf', fold(fs.readFileSync(f, 'utf8'))]];
    cache.set(k, out);
    return out;
  }
  const [v, y, slug] = k.split('/');
  const dir = path.join(ROOT, y, v, slug);
  if (!fs.existsSync(dir)) throw new Error(`missing fulltext directory ${dir}`);
  const out = [];
  for (const f of ['paper.cols.txt', 'paper.norm.txt', 'paper.txt']) {
    const p = path.join(dir, f);
    if (fs.existsSync(p)) out.push([f, fold(fs.readFileSync(p, 'utf8'))]);
  }
  out.push(['pypdf', null, k]);
  cache.set(k, out);
  return out;
}
const pdfCache = new Map();
function pypdf(k) {
  if (!pdfCache.has(k)) pdfCache.set(k, fold(execFileSync('python3', ['scripts/pdftext.py', k], { encoding: 'utf8', maxBuffer: 1 << 28, stdio: ['ignore', 'pipe', 'ignore'] })));
  return pdfCache.get(k);
}
function locate(ck, needle) {
  const n = fold(needle);
  const vs = variants(n);
  for (const [name, text, k] of renderings(ck)) {
    const t = text === null ? pypdf(k) : text;
    for (let i = 0; i < vs.length; i++) if (vs[i](t)) return i === 0 ? name : `${name}${i === 1 ? '+lig' : '+dehyph'}`;
  }
  return null;
}
 
let fail = 0;
for (const [ck, needle] of CONTROLS) {
  const hit = locate(ck, needle);
  console.log(`${hit ? 'CONTROL FAILED (mutated needle located)' : 'control ok (mutated needle not located)'}  ${ck}  "${needle}"`);
  if (hit) fail += 1;
}
 
// ---------------------------------------------------------------- A. spans
const extOut = fs.readFileSync(EXT_OUT, 'utf8');
const okLabels = new Set(extOut.split('\n').filter((l) => l.startsWith('OK ')).map((l) => l.replace(/^OK\s+/, '').trim()));
const page = fs.readFileSync(PAGE, 'utf8');
const lines = page.split('\n');
let spans = 0;
const sroutes = {};
console.log('\n== A. Every //"…"// span on the page ==');
lines.forEach((line, i) => {
  const cks = [...line.matchAll(/\{\[([a-z0-9_]+)\]\}/g)].map((m) => [m[1], m.index]);
  for (const m of line.matchAll(/"\/\/(.+?)\/\/"/g)) {
    spans += 1;
    const span = m[1];
    let route = null, where = '';
    if (span in NOT_A_QUOTE) { route = 'NOT-A-QUOTE'; where = NOT_A_QUOTE[span]; }
    else if (span in EXTERNAL) {
      const lab = EXTERNAL[span];
      route = okLabels.has(lab) ? 'EXTERNAL' : null;
      where = lab + (okLabels.has(lab) ? '' : '  <- no OK line in external check output');
    } else {
      const order = [...cks].sort((a, b) => Math.abs(a[1] - m.index) - Math.abs(b[1] - m.index));
      for (const [ck] of order) { const r = locate(ck, span); if (r) { route = r; where = ck; break; } }
      if (!route) where = `not in ${order.map((x) => x[0]).join(', ') || '(no citekey on the line)'}`;
      else if (order[0][0] !== where) where += '  (WARN: located in a citekey other than the nearest)';
    }
    sroutes[route || 'FAIL'] = (sroutes[route || 'FAIL'] || 0) + 1;
    if (!route) fail += 1;
    console.log(`${route ? 'OK  ' : 'FAIL'} ${(route || '-').padEnd(14)} line ${String(i + 1).padStart(3)}  ${where}\n       "${span.slice(0, 160)}"`);
  }
});
console.log(`spans: ${spans}; by route: ${Object.entries(sroutes).map(([k, v]) => `${k} ${v}`).join(', ')}`);
for (const [span] of Object.entries(EXTERNAL)) if (!page.includes(`"//${span}//"`)) { console.log(`FAIL EXTERNAL entry no longer on the page: "${span}"`); fail += 1; }
 
// ---------------------------------------------------------------- B. needles
console.log('\n== B. Per-paper figures ==');
let weak = 0;
const nroutes = {};
for (const [ck, needle, what] of N) {
  const r = locate(ck, needle);
  if (fold(needle).length < 20) weak += 1;
  nroutes[r || 'FAIL'] = (nroutes[r || 'FAIL'] || 0) + 1;
  if (!r) fail += 1;
  console.log(`${r ? 'OK  ' : 'FAIL'} ${(r || '-').padEnd(14)} ${fold(needle).length < 20 ? 'WEAK ' : ''}${ck}  | ${what}\n       "${needle}"`);
}
console.log(`needles: ${N.length}; by route: ${Object.entries(nroutes).map(([k, v]) => `${k} ${v}`).join(', ')}; weak (<20 chars): ${weak}`);
 
console.log(`
=== EXTERNAL FIGURES (non-corpus; each re-fetched by external_checks_mini_programs.sh — see its labels) ===
wxappUnpacker archived, last commit 2020-04-18; wux1an/wxapkg v2.0.0 released 2026-04-16; wedecode tag v0.10.6, last commit 2026-08-27
MiniCrawler last commit 2021-06-29; README pins WeChat 7.0.19 and 7.0.20 with Xposed
WeChat privacy guideline enforcement announced for 2023-09-15, postponed to 2023-10-17; getUserProfile DevTools base libraries 2.10.4-2.16.1
code2Session error 40226 (high-risk users blocked)
MIIT notice 工信部信管〔2023〕105号 dated 2023-07-21, published 2023-08-04; filing campaign 2023-2024
WeChat licence clauses 8.2.1.2 and 8.2.1.4; international ToS last modified 2025-11-18
Telegram Bot API 8.0 dated 2024-11-17
Tencent: combined Weixin and WeChat MAU 1,439 million (30 June 2026); no mini-program count in the 2025 annual, Q1 2026 or Q2 2026 releases
SaTS workshop: 2023, 2024, 2025 at CCS; SaTS 26 co-located with CCS 2026
wei2026_raising (publisher PDF): 2,067 APIs, 183 (8.85%), four super apps, about 55 USD of LLM calls
=== CORPUS FIGURES (report_mini_programs.mjs) are checked by check_page_numbers.mjs against its output ===
`);
console.log(`not located / failed: ${fail}`);
if (fail) process.exit(1);
verify_mini_programs_figures-output.txt
control ok (mutated needle not located)  zhang2023_leak  "we identified 40,881 mini-programs that leaked their MKs"
control ok (mutated needle not located)  zhang2024_minicat  "MiniCAT successfully analyzed 41,727/44,273 (94.2%)"
control ok (mutated needle not located)  wang2025_wechat  "we collected and attempted to analyze 171 different popular Mini Programs"
 
== A. Every //"…"// span on the page ==
OK   paper.cols.txt line  50  wang2023_uncovering
       "590 public APIs, 502 undocumented unchecked APIs, and 65 undocumented checked APIs"
OK   paper.cols.txt line  50  wang2023_uncovering
       "comprises 975 APIs [8], while the English version has only 570 APIs"
OK   publisher-pdf  line  51  wei2026_raising  (WARN: located in a citekey other than the nearest)
       "183 (8.85%) APIs are not properly protected at the mini-program API layer"
OK   EXTERNAL       line  52  WeChat network: cloud hosting via callContainer over the private protocol
       "微信私有协议"
OK   EXTERNAL       line  53  code2Session: server-side only
       "接口应在服务器端调用,不可在前端(小程序、网页、APP等)直接调用"
OK   paper.cols.txt line  54  wang2023_size
       "105 APIs that have existence discrepancies between Windows and Android, 40 APIs between Windows and iOS, 69 APIs between Android and iOS"
OK   EXTERNAL       line  65  MiniCrawler README pins WeChat 7.0.19 (metadata)
       "WeChat 7.0.19"
OK   EXTERNAL       line  65  MiniCrawler README pins WeChat 7.0.20 (download)
       "7.0.20"
OK   pypdf          line  65  zhang2024_minicat
       "batch querying of AppIDs has become impossible due to Tencent's restriction on related API access"
OK   paper.cols.txt line  67  zhang2024_minicat  (WARN: located in a citekey other than the nearest)
       "our crawler cannot successfully collect some mini-programs due to the compatibility limitations of the WeChat Windows client"
OK   EXTERNAL       line  70  MiniSec datasets: AppSecret set requires additional consent and agreement
       "requires additional consent and agreement"
OK   paper.cols.txt line  72  shi2026_better
       "adhered to the methods established in previous studies"
OK   paper.cols.txt line  78  zhang2024_minicat
       "missing main packages"
OK   pypdf          line  78  wang2025_wechat
       "excluded "games" since WeChat Games are implemented differently from Mini Programs"
OK   paper.cols.txt line  83  zhang2024_minicat
       "absent codes from WeChat Cloud Development [28] or highly obfuscated source codes"
OK   paper.cols.txt line  83  zhang2024_minicat
       "14,920 out of 41,726 (35.8%) mini-programs were skipped due to timeouts"
OK   paper.cols.txt line  84  shi2025_skeleton
       "timeout or AST parsing errors"
OK   pypdf          line  86  wang2025_wechat
       "38.8% of them required ID verification or Chinese phone number verification"
OK   paper.cols.txt line  89  yang2022_cross
       "6.29 TB disk storage"
OK   paper.cols.txt line  98  zhang2024_minicat
       "due to their use of a newer version of the WeChat mini-program base library"
OK   pypdf          line 104  yang2025_miniapp
       "the malware may dynamically hide malicious contents without distributing them to the front-end by the time we tested the cases"
OK   EXTERNAL       line 106  TaintMini README: unpacker not provided due to potential legal implications
       "due to potential legal implications"
OK   EXTERNAL       line 106  WeChat licence 8.2.1.4: no plug-ins or unauthorised third-party tools
       "使用插件、外挂或非经腾讯授权的第三方工具"
OK   paper.cols.txt line 113  lee2025_deep
       "WeChat failed to launch even on Android 11 AVDs"
OK   paper.cols.txt line 113  lee2025_deep
       "a configuration publicly known to support the app reliably"
OK   paper.cols.txt line 116  wang2025_wechat
       "decided against further automation"
OK   EXTERNAL       line 118  SaTS 26 topics include Agentic/LLM-based techniques
       "Agentic/LLM-based techniques"
OK   paper.cols.txt line 124  wang2025_wechat
       "Of the 104 Mini Programs we coded, 51 had profile update flows, 85 had search flows, and 96 had browsing flows. 84.3%, 72.9%, and 76.0% of those flows were exfi"
OK   paper.cols.txt line 124  wang2025_wechat
       "we also identified browsing data in 89.7% of the traces we decrypted from 40 health-related Mini Programs"
OK   paper.cols.txt line 124  wang2025_wechat
       "reflect a baseline of WeChat analytics' data collection"
OK   paper.cols.txt line 125  cai2025_tell
       "only one superapp (WeChat) mentions that it collects Mini-H"
OK   paper.cols.txt line 135  shi2025_skeleton
       "84,491 credential leaks are detected, spanning over 54,728 mini-apps"
OK   paper.cols.txt line 138  zhou2025_secrets
       "30.08% of 41,719 WeChat MPs contain leaked secrets"
OK   paper.cols.txt line 140  chen2026_minigames
       "49.95% of ad-enabled mini-games"
OK   paper.cols.txt line 142  yang2022_cross
       "making the FN rate to 2%"
OK   paper.cols.txt line 146  wang2025_wechat
       "purchased Canadian and American phone numbers, which resulted in various restrictions and limitations to our study"
OK   paper.cols.txt line 146  wang2025_wechat  (WARN: located in a citekey other than the nearest)
       "evidence that WeChat behaves differently when registered to a Chinese phone number [55], which may bias our results"
OK   EXTERNAL       line 156  Telegram Mini Apps: launched right inside Telegram
       "launched right inside Telegram"
OK   EXTERNAL       line 156  Telegram Mini Apps: can completely replace any website
       "can completely replace any website"
OK   EXTERNAL       line 156  Telegram Mini Apps: initDataUnsafe should not be trusted
       "should not be trusted"
OK   EXTERNAL       line 157  TikTok mini games: app selects a version and obtains its code package
       "selects an available version and obtains its code package"
OK   pypdf          line 164  wang2023_uncovering
       "We have never uploaded our malicious miniapps onto the markets to harm other users"
OK   pypdf          line 165  yang2025_miniapp  (WARN: located in a citekey other than the nearest)
       "a few seconds per miniapp"
OK   paper.cols.txt line 166  shi2026_better
       "without accessing the cloud data"
OK   paper.cols.txt line 168  zhang2024_minicat
       "248/316 (78.5%)"
OK   paper.cols.txt line 170  chen2026_minigames
       "compliance with platform policies"
spans: 46; by route: paper.cols.txt 27, publisher-pdf 1, EXTERNAL 12, pypdf 6
 
== B. Per-paper figures ==
OK   pypdf          wang2023_uncovering  | hidden unchecked APIs in all five hosts
       "WeChat has 39 hidden unchecked APIs (7.77%) that invoke Android APIs protected by permissions"
OK   paper.cols.txt wang2023_uncovering  | all five hosts
       "Tiktok has 32 (26.23%), and QQ has 38 (12.88%) such APIs"
OK   paper.cols.txt wang2023_uncovering  | 78,974 of 267,359 (29.54%) invoke a hidden API
       "Total 78,974 267,359 29.54 Table 6: The 3rd party WeChat miniapps that have used the undocumented APIs"
OK   paper.cols.txt wang2023_uncovering  | 267,359 miniapps via Mini-Crawler
       "We collected 267, 359 miniapps using Mini-Crawler"
OK   pypdf          wang2023_uncovering  | 1,829 API candidates
       "APIScope recognized in total 1,829 API candidates"
OK   paper.cols.txt wang2023_uncovering  | devices (host version pinned)
       "Google Pixel 4 running Android 11"
OK   paper.cols.txt lu2020_demystifying  | 39 SRE flaws
       "Apinat found 39 System Resource Exposure flaws"
OK   paper.cols.txt lu2020_demystifying  | 927 sub-app APIs
       "generated test cases for all 927 sub-app APIs"
OK   paper.cols.txt lu2020_demystifying  | 11 hosts
       "on a Google Pixel 2 phone and iPhone 8 against 11 host apps"
OK   pypdf          zhang2023_leak  | 14,020 keywords, 3,450,586
       "resulting in 14,020 keywords that retrieved 3,450,586 mini-programs"
OK   pypdf          zhang2023_leak  | seeds
       "We employed 1,000 commonly used Chinese characters and 1,000 commonly used English words as seed keywords"
OK   paper.cols.txt zhang2023_leak  | 6 per minute, six months
       "we maintained a maximum of six requests per minute, taking over six months to collect them all"
OK   paper.cols.txt zhang2023_leak  | 40,880
       "we identified 40,880 mini-programs that leaked their MKs"
OK   paper.cols.txt zhang2023_leak  | about 4 million
       "the WeChat market has about 4 million mini-programs"
OK   paper.cols.txt zhang2023_leak  | Baidu 171,989
       "171,989 Baidu mini-programs"
OK   paper.cols.txt zhang2023_leak  | API oracle
       "ensure no false positives by validating the correctness of keys using API"
OK   paper.cols.txt wang2023_size  | more than 4.3 million
       "more than 4.3 million"
OK   paper.cols.txt zhang2024_minicat  | 44,273
       "We collected and successfully unpacked 44,273 WeChat mini-programs using Mini-Program Crawler"
OK   paper.cols.txt zhang2024_minicat  | 41,726 and 13,349
       "MiniCAT successfully analyzed 41,726/44,273 (94.2%), identifying 13,349/41,726"
OK   paper.cols.txt zhang2024_minicat  | 32.0%
       "32.0% (13,349/41,726) of analyzable mini-programs are potentially"
OK   paper.cols.txt zhang2024_minicat  | 2,264,377 pages
       "These mini-programs consisted of 2,264,377 pages"
OK   paper.cols.txt zhang2024_minicat  | 2,687 per day
       "enabling an average collection of 2,687 mini-programs daily"
OK   paper.cols.txt zhang2024_minicat  | 20 per minute
       "20 requests per minute"
OK   paper.cols.txt zhang2024_minicat  | Applet cache directory
       "user_file/Applet/AppID"
OK   paper.cols.txt WEAK zhang2024_minicat  | pywinauto
       "we used pywinauto"
OK   paper.cols.txt WEAK zhang2024_minicat  | jieba
       "jieba"
OK   paper.cols.txt zhang2024_minicat  | wxappUnpacker
       "we used wxappUnpacker (a public WeChat mini-program unpacking tool)"
OK   paper.cols.txt zhang2024_minicat  | validation 100/100
       "we did not detect any FPs"
OK   paper.cols.txt zhang2024_minicat  | 3 FN
       "we discovered 3 FN cases"
OK   paper.cols.txt zhou2025_secrets  | 41,719, built on MiniCAT
       "Based on prior research [81], we developed a MP crawler, collecting (March-October 2023) 41,719 projects"
OK   paper.cols.txt zhou2025_secrets  | 300-sample validation
       "achieving a precision of 91.33% on GitHub, 91.67% on PyPI, and 96.33% on WeChat MPs"
OK   paper.cols.txt zhou2025_secrets  | 300 samples
       "we randomly selected and labeled 300 samples"
OK   paper.cols.txt yang2022_cross  | 2,571,490
       "We obtained 2,571,490 WeChat miniapps in total"
OK   paper.cols.txt yang2022_cross  | 148,512
       "WeChat miniapps and 148,512 Baidu miniapps"
OK   paper.cols.txt yang2022_cross  | 50,281 of 52,394
       "we have identified 50,281 (95.97%) vulnerable miniapps out of 52,394 miniapps from the WeChat market"
OK   pypdf          yang2022_cross  | validation 100/100
       "sampled 100 miniapps from the miniapps identified as not involving extra data to verify false negatives"
OK   paper.cols.txt WEAK yang2022_cross  | DoubleX
       "DoubleX"
OK   pypdf+dehyph   shi2025_skeleton  | 413,775
       "Finally, we totally collect 413,775 mini-apps as our database"
OK   pypdf          shi2025_skeleton  | 402,527
       "We successfully analyze 402,527 mini-apps"
OK   paper.cols.txt shi2025_skeleton  | LINE, VK
       "3,984 Line mini-apps and 425 VK miniapps"
OK   pypdf          shi2025_skeleton  | 500 flagged
       "randomly selecting 500 mini-apps identified as vulnerable from each super-app"
OK   paper.cols.txt shi2025_skeleton  | UI Automator
       "Android UI Automator"
OK   paper.cols.txt WEAK shi2025_skeleton  | IRB minimal risk
       "minimal risk"
OK   paper.cols.txt shi2026_better  | 1,248,815
       "apply it to 1,248,815 real-world mini-apps"
OK   paper.cols.txt shi2026_better  | 22,695
       "we identified 22,695 mini-apps"
OK   paper.cols.txt shi2026_better  | 2,815 (12.40%)
       "2,815 mini-apps (12.40%) are affected by the insecure resource management"
OK   paper.cols.txt WEAK shi2026_better  | Gemini
       "Gemini"
OK   paper.cols.txt shi2026_better  | 1,869 emailed
       "we collected emails for 1,869 mini-apps"
OK   paper.cols.txt shi2026_better  | 893 fixed
       "893 mini-app developers have fixed the issues or taken down their mini-apps"
OK   paper.cols.txt shi2026_better  | IRB minimal risk
       "this study is considered as "minimal risk""
OK   pypdf          shi2026_better  | LINE, VK
       "4,324 Line mini-apps and 862 VK mini-apps"
OK   paper.cols.txt yang2025_miniapp  | 4,595,680
       "amassing 4,595,680 miniapps"
OK   paper.cols.txt yang2025_miniapp  | 360,467 delisted
       "360,467 miniapps had been removed by the time when finishing our second collection"
OK   paper.cols.txt yang2025_miniapp  | June 2022
       "completed the first round of miniapp collection in June 2022"
OK   paper.cols.txt yang2025_miniapp  | March 2020
       "carried out from March 2020 to December 2022"
OK   paper.cols.txt yang2025_miniapp  | 19,905
       "in 19, 905 malware samples"
OK   paper.cols.txt yang2025_miniapp  | 500 sampled
       "we sampled a total of 500 miniapps out of the"
OK   paper.cols.txt chen2026_minigames  | 6,769 and 2,076
       "This process yielded 6,769 mini-games in total, from which we further extracted 2,076 Cocos-based"
OK   paper.cols.txt chen2026_minigames  | miniCrawler
       "using the open-source tool mini-Crawler"
OK   paper.cols.txt WEAK chen2026_minigames  | 457 of 915
       "457 / 915 (49.95%)"
OK   paper.cols.txt chen2026_minigames  | audit cases
       "anonymous official audit department"
OK   paper.cols.txt wang2025_wechat  | 170
       "We collected and attempted to analyze 170 different popular Mini Programs"
OK   pypdf          wang2025_wechat  | 104
       "We were able to successfully collect analytics data from 104 Mini Programs"
OK   paper.cols.txt wang2025_wechat  | Pixel 6 + emulator
       "We used both a rooted Pixel 6 device and an emulator, both running Android 14"
OK   paper.cols.txt wang2025_wechat  | host version pinned
       "We analyzed WeChat app version 8.0.23 and WeChat 8.0.49"
OK   paper.cols.txt wang2025_wechat  | Frida
       "We used Frida, a dynamic instrumentation toolkit"
OK   paper.cols.txt wang2025_wechat  | Jadx
       "Jadx, a popular Android decompiler"
OK   paper.cols.txt wang2025_wechat  | Ghidra, IDA
       "We also used Ghidra and IDA Pro"
OK   paper.cols.txt wang2025_wechat  | Alipay and Baidu advertise similar
       "both Baidu's Smart Program ecosystem [57] and Alipay's Mini Program ecosystem [58] advertise similar default tracking features"
OK   paper.cols.txt wang2025_wechat  | US phone number
       "The account associated with our research was registered to a US phone number"
OK   paper.cols.txt cai2025_tell  | 288,895 users
       "of which 288,895 users agreed to share their data"
OK   paper.cols.txt cai2025_tell  | 219,826
       "we construct the dataset with 1,099,130 data samples from 219,826 users"
OK   paper.cols.txt cai2025_tell  | AliPay IRB
       "with IRB approval from AliPay"
OK   paper.cols.txt cai2025_tell  | 31 super apps
       "We contacted the vendors of all 31 super-apps"
OK   paper.cols.txt he2024_demystifying  | 75 of 81
       "we get apps for a total of 81 Chinese rentable products (including 75 WeChat mini-programs and 6 Android apps)"
OK   paper.cols.txt he2024_demystifying  | real-name + deposit
       "realname registration and a deposit"
OK   paper.cols.txt WEAK he2024_demystifying  | BurpSuite
       "Burpsuite proxy"
OK   paper.cols.txt liu2024_riotfuzzer  | 27 devices
       "We apply RIoTFuzzer to 27 IoT devices"
OK   paper.cols.txt liu2024_riotfuzzer  | four IoT platforms
       "Xiaomi, Jingdong, Huawei, and Tuya"
OK   paper.cols.txt WEAK liu2024_riotfuzzer  | ChatGPT filtering
       "ChatGPT"
OK   paper.cols.txt zhang2022_identity  | 6,000 apps
       "we randomly crawl 6,000 popular Android apps"
OK   paper.cols.txt zhang2022_identity  | 47 super apps
       "The first step gives us 47 super-apps"
OK   paper.cols.txt WEAK zhang2022_identity  | Xposed
       "Xposed"
OK   pypdf          wang2023_size  | platform builds
       "two Windows-11 desktops, and four smartphones (two with Android-13, and two with iOS-16)"
OK   publisher-pdf  wei2026_raising  | context (not on page)
       "ranging from 9.5% in Baidu to 19.5% in Alipay"
OK   publisher-pdf  wei2026_raising  | disclosure
       "We disclosed all identified vulnerabilities to the Security Response Centers of Tencent, Baidu, and Alipay"
OK   paper.cols.txt lee2025_deep  | BlueStacks
       "we conducted WeChat-specific experiments on BlueStacks"
OK   paper.cols.txt wang2023_size  | 1,031 documented APIs
       "we have 1,031 APIs in total"
OK   paper.cols.txt yang2022_cross  | 6.29 TB
       "consume 6.29 TB disk storage. We also extended M"
OK   paper.cols.txt zhang2024_minicat  | 126.38 GB
       "which occupied a storage space of 126.38 GB"
OK   paper.cols.txt zhang2024_minicat  | five-minute timeout
       "introduced a 5-minute timeout for CodeQL queries"
OK   pypdf          shi2025_skeleton  | recall 85.56%
       "The average precision of KeyMagnet is 95.04% and the recall is 85.56%"
OK   paper.cols.txt chen2026_minigames  | recall 83.55%
       "with a recall of 83.55%"
OK   paper.cols.txt zhou2025_secrets  | WeChat benchmark recall 83.38% (Table 14)
       "WeChat 902 89.91% 83.38%"
OK   paper.cols.txt chen2026_minigames  | 371 labelled behaviours
       "This process produced 371 labeled Ad-behaviors"
OK   paper.cols.txt zhang2023_leak  | the oracle is the token endpoint
       "Api.weixin.qq.com/cgi-bin/token"
OK   paper.cols.txt WEAK chen2026_minigames  | crawl compliance assertion
       "All test data was"
OK   paper.cols.txt shi2026_better  | compliance with laws
       "we ensure compliance with all relevant laws and regulations"
OK   paper.cols.txt he2024_demystifying  | bug-bounty plan
       "comply with the vendor's bug bounty plan"
OK   pypdf          wang2023_uncovering  | the 975/570 count is from 2021
       "Our experiments were conducted primarily in 2021"
OK   paper.cols.txt liu2024_riotfuzzer  | citing zhang2021 (as is spliced off in .cols)
       "reported in [45], most mini-apps employ obfuscation techniques"
needles: 101; by route: pypdf 12, paper.cols.txt 86, pypdf+dehyph 1, publisher-pdf 2; weak (<20 chars): 10
 
=== EXTERNAL FIGURES (non-corpus; each re-fetched by external_checks_mini_programs.sh — see its labels) ===
wxappUnpacker archived, last commit 2020-04-18; wux1an/wxapkg v2.0.0 released 2026-04-16; wedecode tag v0.10.6, last commit 2026-08-27
MiniCrawler last commit 2021-06-29; README pins WeChat 7.0.19 and 7.0.20 with Xposed
WeChat privacy guideline enforcement announced for 2023-09-15, postponed to 2023-10-17; getUserProfile DevTools base libraries 2.10.4-2.16.1
code2Session error 40226 (high-risk users blocked)
MIIT notice 工信部信管〔2023〕105号 dated 2023-07-21, published 2023-08-04; filing campaign 2023-2024
WeChat licence clauses 8.2.1.2 and 8.2.1.4; international ToS last modified 2025-11-18
Telegram Bot API 8.0 dated 2024-11-17
Tencent: combined Weixin and WeChat MAU 1,439 million (30 June 2026); no mini-program count in the 2025 annual, Q1 2026 or Q2 2026 releases
SaTS workshop: 2023, 2024, 2025 at CCS; SaTS 26 co-located with CCS 2026
wei2026_raising (publisher PDF): 2,067 APIs, 183 (8.85%), four super apps, about 55 USD of LLM calls
=== CORPUS FIGURES (report_mini_programs.mjs) are checked by check_page_numbers.mjs against its output ===
 
not located / failed: 0

The external checks

external_checks_mini_programs.sh
#!/usr/bin/env bash
# external_checks_mini_programs.sh — re-fetch every external fact design:mobile_and_app_measurement:mini_programs
# leans on and print the evidence. Each check prints OK or FAILED; the exit status is the number of
# FAILED checks (capped at 255), so a green run is a real assertion, not a transcript.
#   bash scripts/external_checks_mini_programs.sh > scripts/external_checks_mini_programs-output.txt
# Honours GH_TOKEN (never printed). Needs curl, python3 (pypdf), node + Playwright's headless shell.
set -u
cd "$(dirname "$0")/.."
UA='Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0 Safari/537.36'
FAILS=0
TMP=$(mktemp -d)
AUTH=()
if [ -n "${GH_TOKEN:+x}" ]; then AUTH=(-H "Authorization: Bearer $GH_TOKEN"); fi
export PLAYWRIGHT_BROWSERS_PATH=/workspace/.playwright
echo "run: $(date -u +%Y-%m-%dT%H:%MZ)"
 
# check <label> <file> <python-regex> [i] — whitespace-collapsed, tag-stripped search; 4th arg "i" = case-insensitive
check() {
  if python3 - "$2" "$3" "${4:-}" <<'PY'
import re, sys, html
t = open(sys.argv[1], encoding='utf-8', errors='replace').read()
t = html.unescape(re.sub(r'<[^>]+>', ' ', t))
t = re.sub(r'\s+', ' ', t).replace('’', "'").replace('‘', "'").replace('“', '"').replace('”', '"')
m = re.search(sys.argv[2], t, re.I if sys.argv[3] == 'i' else 0)
if not m: sys.exit(1)
s = max(0, m.start() - 60)
print('      > ' + t[s:m.end() + 60])
PY
  then echo "OK      $1"; else echo "FAILED  $1"; FAILS=$((FAILS+1)); fi
}
# get: fetch, print the final HTTP status, and FAIL on a non-2xx so a broken URL cannot pass on an accidental match
get() {
  local code; code=$(curl -sL -m 60 -A "$UA" -w '%{http_code}' "$1" -o "$2")
  echo "   GET $1 -> HTTP $code, $(wc -c < "$2") bytes"
  case "$code" in 2??) ;; *) echo "FAILED  HTTP $code for $1"; FAILS=$((FAILS+1));; esac
}
gh()  { curl -sL -m 40 "${AUTH[@]}" -H 'Accept: application/vnd.github+json' "https://api.github.com/repos/$1" -o "$2"; }
ghc() { curl -sL -m 40 "${AUTH[@]}" -H 'Accept: application/vnd.github+json' "https://api.github.com/repos/$1/commits?per_page=1" -o "$2"; }
# last default-branch commit date, printed for the record
lastc() { python3 -c "import json,sys;d=json.load(open('$2'));print('      $1 last commit', d[0]['commit']['committer']['date'], repr(d[0]['commit']['message'].splitlines()[0][:60]))"; }
 
echo; echo "== 1. Unpackers and research artefacts (GitHub API) =="
gh qwerty472123/wxappUnpacker $TMP/wxu.json; ghc qwerty472123/wxappUnpacker $TMP/wxu_c.json; lastc wxappUnpacker $TMP/wxu_c.json
check 'wxappUnpacker (qwerty472123) is archived' $TMP/wxu.json '"archived": ?true'
check 'wxappUnpacker last commit 2020-04-18' $TMP/wxu_c.json '"date": ?"2020-04-18'
curl -sL -m 40 "${AUTH[@]}" https://api.github.com/repos/wux1an/wxapkg/releases/latest -o $TMP/wux.json
check 'wux1an/wxapkg latest release v2.0.0 published 2026-04-16' $TMP/wux.json '"tag_name": ?"v2\.0\.0".*"published_at": ?"2026-04-16'
curl -sL -m 40 "${AUTH[@]}" "https://api.github.com/repos/biggerstar/wedecode/tags?per_page=100" -o $TMP/wed_tags.json
python3 -c "
import json,re
t=[x['name'] for x in json.load(open('$TMP/wed_tags.json'))]
v=sorted(t,key=lambda s:[int(x) for x in re.findall(r'\d+',s)])
print('      wedecode tags (unsorted API):',t,' semver max:',v[-1])
open('$TMP/wed_max.txt','w').write('max='+v[-1])"
check 'wedecode newest tag v0.10.6' $TMP/wed_max.txt 'max=v0\.10\.6$'
ghc biggerstar/wedecode $TMP/wed_c.json; lastc wedecode $TMP/wed_c.json
check 'wedecode last commit 2026-08-27' $TMP/wed_c.json '"date": ?"2026-08-27'
ghc OSUSecLab/MiniCrawler $TMP/mc_c.json; lastc MiniCrawler $TMP/mc_c.json
check 'MiniCrawler last commit 2021-06-29' $TMP/mc_c.json '"date": ?"2021-06-29'
get https://raw.githubusercontent.com/OSUSecLab/MiniCrawler/main/README.md $TMP/mc_readme.md
if [ "$(wc -c < $TMP/mc_readme.md)" -lt 200 ]; then get https://raw.githubusercontent.com/OSUSecLab/MiniCrawler/master/README.md $TMP/mc_readme.md; fi
check 'MiniCrawler README pins WeChat 7.0.19 (metadata)' $TMP/mc_readme.md 'Install Xposed and WeChat 7\.0\.19'
check 'MiniCrawler README pins WeChat 7.0.20 (download)' $TMP/mc_readme.md 'Install Xposed and WeChat 7\.0\.20'
for r in OSUSecLab/TaintMini kee1ongz/MiniCAT OSUSecLab/APIDiff; do
  n=$(echo $r | tr '/' '_')
  get https://raw.githubusercontent.com/$r/main/README.md $TMP/$n.md
  if [ "$(wc -c < $TMP/$n.md)" -lt 200 ]; then get https://raw.githubusercontent.com/$r/master/README.md $TMP/$n.md; fi
done
check 'TaintMini README: unpacker not provided due to potential legal implications' $TMP/OSUSecLab_TaintMini.md 'unable to provide such a tool directly due to potential legal implications'
check 'MiniCAT README: wxappUnpacker not provided due to potential legal implications' $TMP/kee1ongz_MiniCAT.md 'Due to potential legal implications, we DO NOT provide wxappUnpacker'
check 'APIDiff README: crawlers not provided due to potential legal implications' $TMP/OSUSecLab_APIDiff.md 'not able to provide such web crawlers due to potential legal implications'
curl -sL -m 40 "${AUTH[@]}" https://api.github.com/repos/Ackites/KillWxapkg/releases/latest -o $TMP/kw.json
ghc Ackites/KillWxapkg $TMP/kw_c.json; lastc KillWxapkg $TMP/kw_c.json
check 'KillWxapkg latest release v2.4.1' $TMP/kw.json '"tag_name": ?"v2\.4\.1"'
check 'KillWxapkg last commit 2024-09-20' $TMP/kw_c.json '"date": ?"2024-09-20'
curl -sL -m 40 "${AUTH[@]}" https://api.github.com/repos/EEEEEEcho/wemint/contents/wxappUnpacker -o $TMP/wemint_wxu.json
check 'WeMinT repository bundles wxappUnpacker (wuWxapkg.js present)' $TMP/wemint_wxu.json '"name": ?"wuWxapkg\.js"'
get https://minimalware.github.io/ $TMP/minimal.html
check 'MiniSec datasets: AppSecret set requires additional consent and agreement' $TMP/minimal.html 'requires additional consent and agreement'
check 'MiniSec datasets: institutional authentication and written consent' $TMP/minimal.html 'institutional authentication and written consent'
check 'MiniSec datasets: SIGMETRICS21 random samples' $TMP/minimal.html 'Randomly-selected miniapp samples.{0,80}SIGMETRICS21'
 
echo; echo "== 2. WeChat developer documentation =="
get https://developers.weixin.qq.com/miniprogram/dev/framework/ability/network.html $TMP/wx_net.html
check 'WeChat network: only whitelisted domains' $TMP/wx_net.html '只可以跟指定的域名进行网络通信'
check 'WeChat network: domains need ICP filing' $TMP/wx_net.html '域名必须经过 ?ICP ?备案'
check 'WeChat network: fixed Referer servicewechat.com/{appid}/{version}/page-frame.html' $TMP/wx_net.html 'https://servicewechat\.com/\{appid\}/\{version\}/page-frame\.html'
check 'WeChat network: cloud hosting via callContainer over the private protocol' $TMP/wx_net.html 'callContainer.{0,60}微信私有协议'
get https://developers.weixin.qq.com/miniprogram/dev/server/API/user-login/api_code2session.html $TMP/wx_c2s.html
check 'code2Session: server-side only' $TMP/wx_c2s.html '接口应在服务器端调用,不可在前端(小程序、网页、APP等)直接调用'
check 'code2Session: error 40226 blocks high-risk users' $TMP/wx_c2s.html '40226.{0,40}高风险等级用户'
get https://developers.weixin.qq.com/miniprogram/dev/framework/user-privacy/PrivacyAuthorize.html $TMP/wx_priv.html
check 'Privacy guideline: 2023-09-15 enforcement announced' $TMP/wx_priv.html '2023年9月15日之后'
check 'Privacy guideline: postponed to 2023-10-17' $TMP/wx_priv.html '延期至 ?2023年10月17日'
check 'Privacy guideline: undeclared APIs disabled' $TMP/wx_priv.html '若未声明,对应接口或组件将直接禁用'
get https://developers.weixin.qq.com/miniprogram/dev/api/open-api/user-info/wx.getUserProfile.html $TMP/wx_gup.html
check 'getUserProfile: DevTools 2.10.4-2.16.1 real data, devices anonymous' $TMP/wx_gup.html '开发者工具中 ?2\.10\.4 ?~ ?2\.16\.1 ?基础库版本.{0,80}真机上此区间会按照公告返回匿名数据'
get https://developers.weixin.qq.com/miniprogram/dev/devtools/remote-debug.html $TMP/wx_rd.html
check 'DevTools remote debugging packs and uploads the local code' $TMP/wx_rd.html '工具会将本地代码进行处理打包并上传'
get https://developers.weixin.qq.com/miniprogram/dev/api/base/debug/wx.setEnableDebug.html $TMP/wx_dbg.html
check 'wx.setEnableDebug: debug switch set by the mini-program, works on release builds' $TMP/wx_dbg.html '此开关对正式版也能生效'
get https://developers.weixin.qq.com/miniprogram/product/record/record_faq.html $TMP/wx_faq.html
check 'Filing FAQ: unfiled mini-programs become inaccessible or delisted' $TMP/wx_faq.html '无法访问或下架'
check 'Filing FAQ: MIIT system searchable by the 小程序 category' $TMP/wx_faq.html '选择小程序业务类别'
 
echo; echo "== 3. Regulator and terms =="
get https://www.miit.gov.cn/zwgk/zcwj/wjfb/tz/art/2023/art_920db564162e4312916a01bed6540ad8.html $TMP/miit.html
check 'MIIT notice 工信部信管〔2023〕105号' $TMP/miit.html '工信部信管〔2023〕105号'
check 'MIIT notice dated 2023-07-21, published 2023-08-04' $TMP/miit.html '成文日期: ?2023-07-21 ?发布日期: ?2023-08-04'
check 'MIIT notice covers distribution platforms incl. mini-programs' $TMP/miit.html '分发平台(含小程序、快应用等分发)'
get "https://weixin.qq.com/agreement?lang=zh_CN" $TMP/wx_lic.html
check 'WeChat licence 8.2.1.2: no reverse engineering' $TMP/wx_lic.html '8\.2\.1\.2 ?对本软件进行反向工程'
check 'WeChat licence 8.2.1.4: no plug-ins or unauthorised third-party tools' $TMP/wx_lic.html '8\.2\.1\.4.{0,200}使用插件、外挂或非经腾讯授权的第三方工具'
get https://www.wechat.com/en/service_terms.html $TMP/wx_tos.html
check 'WeChat licence 12.3: governed by mainland-Chinese law' $TMP/wx_lic.html '12\.3 ?本协议的成立、生效、履行、解释及纠纷解决,适用中华人民共和国大陆地区法律'
check 'WeChat international ToS: last modified 2025-11-18' $TMP/wx_tos.html 'Last modified: ?2025-11-18'
check 'WeChat international ToS: no reverse engineering of WeChat Software' $TMP/wx_tos.html 'reverse engineer or extract source codes from WeChat Software'
 
echo; echo "== 4. Other hosts =="
get https://core.telegram.org/bots/webapps $TMP/tg.html
check 'Telegram Mini Apps: launched right inside Telegram' $TMP/tg.html 'launched right inside Telegram'
check 'Telegram Mini Apps: can completely replace any website' $TMP/tg.html 'can completely replace any website'
check 'Telegram Mini Apps: initDataUnsafe should not be trusted' $TMP/tg.html 'Data from this field should not be trusted'
check 'Telegram Bot API 8.0 (2024-11-17): photo_url to all Mini Apps' $TMP/tg.html 'photo_url in the class WebAppUser is now available to all Mini Apps'
check 'Telegram Bot API 8.0 dated November 17, 2024' $TMP/tg.html 'November 17, 2024 Bot API 8\.0'
check 'Telegram: photo_url only if privacy settings allow' $TMP/tg.html 'if their privacy settings allow for it'
get https://developers.tiktok.com/docs/en/mini-games-overview $TMP/tt_over.html
if [ "$(wc -c < $TMP/tt_over.html)" -lt 5000 ]; then node scripts/pw_fetch_text.mjs https://developers.tiktok.com/docs/en/mini-games-overview $TMP/tt_over.html > /dev/null 2>&1; echo "   PW  tiktok overview -> $(wc -c < $TMP/tt_over.html) bytes"; fi
check 'TikTok mini games: launched markets (no EU country listed)' $TMP/tt_over.html 'Already launched in markets including the U\.S\., Japan, Indonesia, Turkey, Saudi Arabia, Thailand, Brazil, Malaysia, Philippines, and Vietnam'
get https://developers.tiktok.com/docs/en/mini-games-technical-overview $TMP/tt_tech.html
if [ "$(wc -c < $TMP/tt_tech.html)" -lt 5000 ]; then node scripts/pw_fetch_text.mjs https://developers.tiktok.com/docs/en/mini-games-technical-overview $TMP/tt_tech.html > /dev/null 2>&1; echo "   PW  tiktok technical -> $(wc -c < $TMP/tt_tech.html) bytes"; fi
check 'TikTok mini games: app selects a version and obtains its code package' $TMP/tt_tech.html 'selects an available version and obtains its code package'
 
echo; echo "== 5. Tencent's own results: no mini-program count =="
get https://www.prnewswire.com/apac/news-releases/tencent-announces-2026-second-quarter-results-302849608.html $TMP/tq2.html
check 'Tencent Q2 2026: combined Weixin and WeChat MAU 1,439 million' $TMP/tq2.html 'Combined MAU of Weixin and WeChat 1,439'
python3 - $TMP/tq2.html <<'PY'
import re, sys, html
t = re.sub(r'\s+', ' ', html.unescape(re.sub(r'<[^>]+>', ' ', open(sys.argv[1], encoding='utf-8', errors='replace').read())))
hits = [t[max(0, m.start() - 80):m.end() + 80] for m in re.finditer(r'[Mm]ini ?[Pp]rogram', t)]
print(f'      "Mini Program" mentions in the Q2 2026 release: {len(hits)}')
for h in hits: print('      > ' + h)
nums = [h for h in hits if re.search(r'\d[\d,.]* ?(million|billion|m\b|bn)', h)]
print(f'      of them next to a count: {len(nums)}')
open(sys.argv[1] + '.count', 'w').write(f'count_mentions={len(nums)}')
PY
check 'Tencent Q2 2026: no mini-program count next to any "Mini Program" mention' $TMP/tq2.html.count 'count_mentions=0'
# The same question over the two earlier releases (PDFs), so the claim covers 2025 annual/Q4, Q1 and Q2 2026.
for u in https://static.www.tencent.com/uploads/2026/03/18/e6a646796d0d869acc76271c9ee1a6a5.pdf https://static.www.tencent.com/uploads/2026/05/13/47382ae415a209fd161bc19a1f9b3704.pdf; do
  curl -sL -m 60 -A "$UA" "$u" -o $TMP/tq.pdf
  python3 - $TMP/tq.pdf "$u" <<'PY'
import re, sys, pypdf
t = re.sub(r'\s+', ' ', ' '.join((p.extract_text() or '') for p in pypdf.PdfReader(sys.argv[1]).pages))
mau = re.search(r'Combined MAU of Weixin and WeChat [\d,]+', t)
hits = [t[max(0, m.start() - 80):m.end() + 80] for m in re.finditer(r'[Mm]ini ?[Pp]rogram', t)]
nums = [h for h in hits if re.search(r'\d[\d,.]* ?(million|billion|m\b|bn)', h)]
print(f'      {sys.argv[2].rsplit("/",1)[1]}: positive control {mau.group(0) if mau else "MISSING"}; "Mini Program" mentions {len(hits)}, next to a count {len(nums)}')
for h in hits: print('      > ' + h)
open(sys.argv[1] + '.count', 'w').write(f'control={bool(mau)} count_mentions={len(nums)}')
PY
  check "Tencent results PDF $(basename $u): MAU line present, no mini-program count" $TMP/tq.pdf.count 'control=True count_mentions=0'
done
 
echo; echo "== 6. Venues, preprints and the workshop =="
get https://superappsec.github.io/ $TMP/sats.html
check 'SaTS 26 co-located with CCS 2026' $TMP/sats.html 'SaTS ’?.?26.{0,40}Co-located with ACM CCS 2026' i
check 'SaTS 26 topics include Agentic/LLM-based techniques' $TMP/sats.html 'Agentic/LLM-based techniques'
get https://arxiv.org/abs/2306.07495 $TMP/sok.html
check 'SoK super app: arXiv 2306.07495, Yang Wang Zhang Lin' $TMP/sok.html 'Decoding the Super App Enigma'
# arXiv's Journal-ref field is author-maintained and usually never filled, so "no venue" is asserted on OpenAlex's
# indexed locations instead: every location must be a repository (arXiv), none a conference or journal.
curl -sL -m 40 "https://api.openalex.org/works/doi:10.48550/arXiv.2306.07495" -o $TMP/sok_oa.json
python3 -c "
import json
d=json.load(open('$TMP/sok_oa.json'))
locs=[((l.get('source') or {}).get('display_name'),(l.get('source') or {}).get('type')) for l in d['locations']]
print('      OpenAlex locations:',locs)
open('$TMP/sok_oa.txt','w').write('nonrepo=%d total=%d' % (sum(1 for n,t in locs if t!='repository'), len(locs)))"
check 'SoK super app: OpenAlex lists only repository locations (no venue)' $TMP/sok_oa.txt 'nonrepo=0 total=[1-9]'
get https://arxiv.org/abs/2608.17538 $TMP/tenet.html
check 'TENET preprint: Telegram Mini App (in)security, plaintext tokens and mnemonics' $TMP/tenet.html 'store authentication materials---such as session tokens and wallet mnemonic phrases---in plaintext'
get https://arxiv.org/abs/2608.13390 $TMP/tgap.html
check 'TeleGapper preprint: privacy policies in Telegram Mini apps' $TMP/tgap.html 'On the \(un\)reliability of Privacy Policies in Telegram Mini apps'
curl -sL -m 40 "https://api.openalex.org/works/doi:10.1145/3460081" -o $TMP/mc_oa.json
python3 -c "
import json;d=json.load(open('$TMP/mc_oa.json'));ii=d.get('abstract_inverted_index') or {}
w=sorted((p,k) for k,v in ii.items() for p in v);open('$TMP/mc_abs.txt','w').write(' '.join(k for p,k in w))"
if [ ! -s $TMP/mc_abs.txt ]; then curl -sL -m 40 "https://api.semanticscholar.org/graph/v1/paper/DOI:10.1145/3460081?fields=abstract" -o $TMP/mc_abs.txt; fi
check 'SIGMETRICS 2021 abstract: measures obfuscation rate' $TMP/mc_abs.txt 'obfuscation rate'
get https://arxiv.org/abs/2607.08232 $TMP/oauth.html
check 'CCS 2026 OAuth-misuse preprint: accepted by ACM CCS 2026' $TMP/oauth.html 'Accepted by ACM CCS 2026'
for d in 10.1145/3460081 10.1145/3607199.3607236 10.1109/icse48619.2023.00086 10.1109/ase56229.2023.00151; do
  curl -sL -m 40 "https://api.crossref.org/works/$d" -o $TMP/cr.json
  python3 -c "import json;m=json.load(open('$TMP/cr.json'))['message'];print('      $d:',m['title'][0],'|',(m.get('container-title') or [''])[0],'|',m['issued']['date-parts'][0])"
  check "Crossref resolves $d" $TMP/cr.json '"status": ?"ok"'
done
for d in DOI:10.1145/3774904.3792470 DOI:10.1109/SP63933.2026.00074; do
  curl -sL -m 40 "https://api.semanticscholar.org/graph/v1/paper/$d?fields=title,abstract" -o $TMP/s2.json; sleep 2
  python3 -c "import json;d=json.load(open('$TMP/s2.json'));print('      $d:',d['title'],'|',d['abstract'][:160])"
done
 
echo; echo "== 7. Outside-the-extraction paper read from the publisher PDF =="
if [ ! -s out/mp/ext/wei_zhiao.txt ]; then
  curl -sL -m 60 -A "$UA" https://www.usenix.org/system/files/usenixsecurity26-wei-zhiao.pdf -o out/mp/ext/wei_zhiao.pdf
  python3 -c "import pypdf,re;r=pypdf.PdfReader('out/mp/ext/wei_zhiao.pdf');open('out/mp/ext/wei_zhiao.txt','w').write(re.sub(r'\s+',' ',' '.join(p.extract_text() or '' for p in r.pages)))" 2>/dev/null
fi
check 'Raising the Flag: 183 (8.85%) APIs not properly protected' out/mp/ext/wei_zhiao.txt '183 \(8\.85%\) APIs are not properly protected at the mini-program API layer'
check 'Raising the Flag: four super-apps WeChat, Alipay, Baidu, QQ' out/mp/ext/wei_zhiao.txt 'we analyze four popular super-apps: WeChat, Alipay, Baidu, and QQ'
check 'Raising the Flag: 2,067 analysed APIs' out/mp/ext/wei_zhiao.txt 'across 2,067 analyzed APIs'
check 'Raising the Flag: LLM test-case generation cost about 55 USD' out/mp/ext/wei_zhiao.txt 'approximately 55 USD'
check 'Raising the Flag: Xposed instruments the super-apps' out/mp/ext/wei_zhiao.txt 'we use Xposed \[ ?21\] to instru- ?ment super-apps'
 
echo; echo "FAILED checks: $FAILS"
[ $FAILS -gt 255 ] && FAILS=255
exit $FAILS
external_checks_mini_programs-output.txt
run: 2026-09-27T16:54Z
 
== 1. Unpackers and research artefacts (GitHub API) ==
      wxappUnpacker last commit 2020-04-18T15:00:46Z 'rm'
      > scussions": false, "forks_count": 2370, "mirror_url": null, "archived": true, "disabled": false, "open_issues_count": 54, "license": nul
OK      wxappUnpacker (qwerty472123) is archived
      > mitter": { "name": "4qwerty7", "email": "4qwerty7@163.com", "date": "2020-04-18T15:00:46Z" }, "message": "rm", "tree": { "sha": "a0381c1ee1
OK      wxappUnpacker last commit 2020-04-18
      > ", "site_admin": false }, "node_id": "RE_kwDOJmI-EM4SefKP", "tag_name": "v2.0.0", "target_commitish": "main", "name": "v2.0.0 全新跨平台 GUI 版本", "draft": false, "immutable": false, "prerelease": false, "created_at": "2026-04-16T17:55:53Z", "updated_at": "2026-04-16T18:03:16Z", "published_at": "2026-04-16T17:58:49Z", "assets": [ { "url": "https://api.github.com/re
OK      wux1an/wxapkg latest release v2.0.0 published 2026-04-16
      wedecode tags (unsorted API): ['v0.10.6', 'v0.10.5', 'v0.10.3', 'v0.10.1', 'v0.9.2', 'v0.3.0']  semver max: v0.10.6
      > max=v0.10.6
OK      wedecode newest tag v0.10.6
      wedecode last commit 2026-08-27T06:49:22Z 'Merge pull request #89 from Jah-yee/fix/typo-unkown'
      > ", "email": "58533654+biggerstar@users.noreply.github.com", "date": "2026-08-27T06:49:22Z" }, "committer": { "name": "GitHub", "email": "no
OK      wedecode last commit 2026-08-27
      MiniCrawler last commit 2021-06-29T20:03:45Z 'Update README.md'
      > b", "email": "35103616+OSUSecLab@users.noreply.github.com", "date": "2021-06-29T20:03:45Z" }, "committer": { "name": "GitHub", "email": "no
OK      MiniCrawler last commit 2021-06-29
   GET https://raw.githubusercontent.com/OSUSecLab/MiniCrawler/main/README.md -> HTTP 200, 2103 bytes
      > p metadata. ## How to run it **Crawl Mini-app Metadata** 1. Install Xposed and WeChat 7.0.19 on your phone. 2. Compile and install XposedPlugin, and ena
OK      MiniCrawler README pins WeChat 7.0.19 (metadata)
      > a created database file `data.db` **Download Mini-apps** 1. Install Xposed and WeChat 7.0.20 on your phone. 2. Compile and install XposedPlugin, and ena
OK      MiniCrawler README pins WeChat 7.0.20 (download)
   GET https://raw.githubusercontent.com/OSUSecLab/TaintMini/main/README.md -> HTTP 200, 5048 bytes
   GET https://raw.githubusercontent.com/kee1ongz/MiniCAT/main/README.md -> HTTP 200, 4777 bytes
   GET https://raw.githubusercontent.com/OSUSecLab/APIDiff/main/README.md -> HTTP 200, 5306 bytes
      > -Program unpacking tool in advance. Please note that we are unable to provide such a tool directly due to potential legal implications. We recommend seeking it out on external websites. ## Usage
OK      TaintMini README: unpacker not provided due to potential legal implications
      > nt MiniCAT relies on wxappUnpacker to unpack mini-programs. Due to potential legal implications, we DO NOT provide wxappUnpacker in our repository. Users should copy it to the wxappUnpacke
OK      MiniCAT README: wxappUnpacker not provided due to potential legal implications
      > the API document for `apitest-gen`. Please note that we are not able to provide such web crawlers due to potential legal implications. We offer the `typescript` object description for structure
OK      APIDiff README: crawlers not provided due to potential legal implications
      KillWxapkg last commit 2024-09-20T04:12:44Z 'style: v2.4.1'
      > ", "site_admin": false }, "node_id": "RE_kwDOMapcSc4KffDl", "tag_name": "v2.4.1", "target_commitish": "master", "name": "v2.4.1", "draft": f
OK      KillWxapkg latest release v2.4.1
      > "name": "Antkites", "email": "mayizhuifengzheng@gmail.com", "date": "2024-09-20T04:12:44Z" }, "committer": { "name": "Antkites", "email": "
OK      KillWxapkg last commit 2024-09-20
      > Echo/wemint/blob/master/wxappUnpacker/wuRestoreZ.js" } }, { "name": "wuWxapkg.js", "path": "wxappUnpacker/wuWxapkg.js", "sha": "32212530231cc
OK      WeMinT repository bundles wxappUnpacker (wuWxapkg.js present)
   GET https://minimalware.github.io/ -> HTTP 200, 8354 bytes
      > ] Dataset for Miniapps with AppSecret Leakage [CCS23] (This requires additional consent and agreement, contact me for details) Evasive miniapp malware [NDSS25] R
OK      MiniSec datasets: AppSecret set requires additional consent and agreement
      > arch in this field. However, to avoid misuse, we do require institutional authentication and written consent to not misuse the released dataset. If you would like to ac
OK      MiniSec datasets: institutional authentication and written consent
      > t, contact me for details) Evasive miniapp malware [NDSS25] Randomly-selected miniapp samples to facilitate your preliminary research [SIGMETRICS21] Analysis tools for CMRF vulnerability discovery, AppSecret
OK      MiniSec datasets: SIGMETRICS21 random samples
 
== 2. WeChat developer documentation ==
   GET https://developers.weixin.qq.com/miniprogram/dev/framework/ability/network.html -> HTTP 200, 146563 bytes
      > API 时,需要注意下列问题,请开发者提前了解。 # 1. 服务器域名配置 每个微信小程序需要事先设置通讯域名,小程序 只可以跟指定的域名进行网络通信 。包括普通 HTTPS 请求( wx.request )、上传文件( wx.uploadFile )、下载文件( wx
OK      WeChat network: only whitelisted domains
      > //myserver.com:443 请求则会失败。 对于 wss 域名,无需配置端口,默认允许请求该域名下所有端口。 域名必须经过 ICP 备案; 出于安全考虑, api.weixin.qq.com 不能被配置为服务器域名,相关 API 也不能在小程序内调用。 开
OK      WeChat network: domains need ICP filing
      > 优先级高于 app.json 中的配置 # 使用限制 网络请求的 referer header 不可设置。其格式固定为 https://servicewechat.com/{appid}/{version}/page-frame.html ,其中 {appid} 为小程序的 appid, {version} 为小程序的版本号,版本号为 0 表示为开发版、体
OK      WeChat network: fixed Referer servicewechat.com/{appid}/{version}/page-frame.html
      > 内的非本机 IP 以及配置过的服务器域名通信。 如使用 微信云托管 作为后端服务,则可无需配置通讯域名(在小程序内通过 callContainer 和 connectContainer 通过微信私有协议向云托管服务发起 HTTPS 调用和 WebSocket 通信)。 # 配置流程 服务器域名请在 「小程序后台-开发-开
OK      WeChat network: cloud hosting via callContainer over the private protocol
   GET https://developers.weixin.qq.com/miniprogram/dev/server/API/user-login/api_code2session.html -> HTTP 200, 192004 bytes
      > 管控原因查询 广告 回传广告数据 广告数据源查询 广告创建数据源 广告数据源报表查询 # 小程序登录凭证校验 调试诊断 接口应在服务器端调用,不可在前端(小程序、网页、APP等)直接调用,具体可参考 接口调用指南 。 接口英文名:code2Session 登录凭证校验。通过 wx.login 接口获得临时
OK      code2Session: server-side only
      > 决方案 -1 system error 系统繁忙,此时请开发者稍候再试 40029 code 无效 js_code无效 40226 code blocked 高风险等级用户,小程序登录拦截 。风险等级详见 用户安全解方案 45011 api minute-quota reach limit 
OK      code2Session: error 40226 blocks high-risk users
   GET https://developers.weixin.qq.com/miniprogram/dev/framework/user-privacy/PrivacyAuthorize.html -> HTTP 200, 170648 bytes
      > rivacyCheck__: true 后,会启用隐私相关功能,如果不配置或者配置为 false 则不会启用。 2)在 2023年9月15日之后,不论 app.json 中是否有配置 __usePrivacyCheck__ ,隐私相关功能都会启用。 接口用法可参考
OK      Privacy guideline: 2023-09-15 enforcement announced
      >  ,隐私相关功能都会启用。 接口用法可参考下方 完整示例demo 2023.09.14 更新: 1)隐私相关功能启用时间延期至 2023年10月17日。在 2023年10月17日之前,在 app.json 中配置 __usePrivacyCheck__: true 后,
OK      Privacy guideline: postponed to 2023-10-17
      > 见: 用户隐私保护指引填写说明 。 需要注意的是,仅有在指引中声明所处理的用户信息,才可以调用平台提供的对应接口或组件。若未声明,对应接口或组件将直接禁用。 隐私接口与对应的处理的信息关系可见: 小程序用户隐私保护指引内容介绍 。 配置完成后,对于每个使用小程序的用户,开发
OK      Privacy guideline: undeclared APIs disabled
   GET https://developers.weixin.qq.com/miniprogram/dev/api/open-api/user-info/wx.getUserProfile.html -> HTTP 200, 464408 bytes
      > p : wx.getUserProfile 返回的加密数据中不包含 openId 和 unionId 字段。 bug :开发者工具中 2.10.4 ~ 2.16.1 基础库版本通过 <button open-type="getUserInfo"> 会返回真实数据,真机上此区间会按照公告返回匿名数据。 < view class = " container " > < view class = " userinfo "
OK      getUserProfile: DevTools 2.10.4-2.16.1 real data, devices anonymous
   GET https://developers.weixin.qq.com/miniprogram/dev/devtools/remote-debug.html -> HTTP 200, 59902 bytes
      > .0 进行调试。 # 调试流程 要发起一个真机远程调试流程,需要先点击开发者工具的工具栏上 "真机调试" 按钮。 此时,工具会将本地代码进行处理打包并上传,就绪之后,使用手机客户端扫描二维码即可弹出调试窗口,开始远程调试。 # 远程调试窗口 使用手机扫描此二维码,即可开始远
OK      DevTools remote debugging packs and uploads the local code
   GET https://developers.weixin.qq.com/miniprogram/dev/api/base/debug/wx.setEnableDebug.html -> HTTP 200, 449919 bytes
      > Windows 版 :支持 微信 Mac 版 :支持 微信 鸿蒙 OS 版 :支持 # 功能描述 设置是否打开调试开关。此开关对正式版也能生效。 # 参数 # Object object 属性 类型 默认值 必填 说明 enableDebug boolean 是
OK      wx.setEnableDebug: debug switch set by the mini-program, works on release builds
   GET https://developers.weixin.qq.com/miniprogram/product/record/record_faq.html -> HTTP 200, 67116 bytes
      > 【注销主体】。需要注意,注销主体后该主体下的全部备案信息均会被注销,主体下的小程序、APP、网站、快应用等将因未备案导致无法访问或下架,且提交注销申请后无法撤回,建议谨慎操作。 # 2、什么情况下选择【注销小程序】? 当你确认多平台运营的小程序备案不再使
OK      Filing FAQ: unfiled mini-programs become inaccessible or delisted
      > 小程序通过备案,无法根据短信通知确认小程序对应的备案号时,可登录 工信部备案管理系统 进行查询。在搜索框内输入备案号后,选择小程序业务类别,点击搜索即可检索到相关信息。 # 2、如何查询主体下所有的ICP备案信息? 可登录 工信部备案管理系统 进行查询,在搜
OK      Filing FAQ: MIIT system searchable by the 小程序 category
 
== 3. Regulator and terms ==
   GET https://www.miit.gov.cn/zwgk/zcwj/wjfb/tz/art/2023/art_920db564162e4312916a01bed6540ad8.html -> HTTP 200, 19674 bytes
      > 布 > 通知 发文机关: 工业和信息化部 标 题: 工业和信息化部关于开展移动互联网应用程序备案工作的通知 发文字号: 工信部信管〔2023〕105号 成文日期: 2023-07-21 发布日期: 2023-08-04 发布机构: 信息通信管理局 分 类: 信息通信管理
OK      MIIT notice 工信部信管〔2023〕105号
      > 信息化部 标 题: 工业和信息化部关于开展移动互联网应用程序备案工作的通知 发文字号: 工信部信管〔2023〕105号 成文日期: 2023-07-21 发布日期: 2023-08-04 发布机构: 信息通信管理局 分 类: 信息通信管理 工业和信息化部关于开展移动互联网应用程序备案工作的通知 发布时间:
OK      MIIT notice dated 2023-07-21, published 2023-08-04
      > 基础电信企业,公益性互联单位、互联网接入服务提供者、互联网数据中心服务提供者、内容分发网络服务提供者,移动互联网应用程序分发平台(含小程序、快应用等分发)、智能终端生产企业、互联网信息服务提供者: 为落实《中华人民共和国反电信网络诈骗法》《互联网信息服务管理办法》(国务院令
OK      MIIT notice covers distribution platforms incl. mini-programs
   GET https://weixin.qq.com/agreement?lang=zh_CN -> HTTP 200, 47391 bytes
      > 法律允许或腾讯书面许可,你在使用本软件过程中不得从事下列行为: 8.2.1.1 删除本软件及其副本上关于著作权的信息; 8.2.1.2 对本软件进行反向工程、反向汇编、反向编译,或以其他方式尝试发现本软件的源代码; 8.2.1.3 对腾讯拥有知识产权的内容进行使用、出租、出借
OK      WeChat licence 8.2.1.2: no reverse engineering
      > .2.1.3 对腾讯拥有知识产权的内容进行使用、出租、出借、复制、修改、链接、转载、汇编、发表、出版、建立镜像站点等; 8.2.1.4 对本软件或本软件运行过程中释放到任何终端内存中的数据、软件运行过程中客户端与服务器端的交互数据,以及本软件运行所必需的系统数据,进行复制、修改、增加、删除、挂接运行或创作任何衍生作品,形式包括但不限于使用插件、外挂或非经腾讯授权的第三方工具/服务接入本软件和相关系统; 8.2.1.5 通过修改或伪造软件运行中的指令、数据,增加、删减、变动软件的功能或运行效果
OK      WeChat licence 8.2.1.4: no plug-ins or unauthorised third-party tools
   GET https://www.wechat.com/en/service_terms.html -> HTTP 200, 183449 bytes
      > 改后的协议。如果你不接受修改后的协议,应当停止使用本软件。 12.2 本协议签订地为中华人民共和国广东省深圳市南山区。 12.3 本协议的成立、生效、履行、解释及纠纷解决,适用中华人民共和国大陆地区法律(不包括冲突法)。 12.4 若你和腾讯之间发生任何纠纷或争议,首先应友好协商解决;协商不成的,你同意将纠纷或争议提交本
OK      WeChat licence 12.3: governed by mainland-Chinese law
      > der .logo a{display:inline-block} WECHAT – TERMS OF SERVICE Last modified: 2025-11-18 TABLE OF CONTENTS INTRODUCTION ADDITIONAL TERMS AND POLICIE
OK      WeChat international ToS: last modified 2025-11-18
      > not copy, modify, create derivative works, reverse compile, reverse engineer or extract source codes from WeChat Software, and you may not sell, distribute, redistribute or sublicen
OK      WeChat international ToS: no reverse engineering of WeChat Software
 
== 4. Other hosts ==
   GET https://core.telegram.org/bots/webapps -> HTTP 200, 193421 bytes
      > Script to create infinitely flexible interfaces that can be launched right inside Telegram — and can completely replace any website . Like bots, Mini 
OK      Telegram Mini Apps: launched right inside Telegram
      > interfaces that can be launched right inside Telegram — and can completely replace any website . Like bots, Mini Apps support seamless authorization , pay
OK      Telegram Mini Apps: can completely replace any website
      > bject with input data transferred to the Mini App. WARNING: Data from this field should not be trusted. You should only use data from initData on the bot's server
OK      Telegram Mini Apps: initDataUnsafe should not be trusted
      > the device's model and performance class. General The field photo_url in the class WebAppUser is now available to all Mini Apps, allowing them to access a user's profile photo if their pr
OK      Telegram Bot API 8.0 (2024-11-17): photo_url to all Mini Apps
      > cure local storage on the user's device for sensitive data. November 17, 2024 Bot API 8.0 This is the largest update in the history of Telegram mini 
OK      Telegram Bot API 8.0 dated November 17, 2024
      > l Mini Apps, allowing them to access a user's profile photo if their privacy settings allow for it. Third parties (e.g., Mini App builders, external SDKs etc.
OK      Telegram: photo_url only if privacy settings allow
   GET https://developers.tiktok.com/docs/en/mini-games-overview -> HTTP 200, 190714 bytes
      > ect search access, and more) to achieve seamless conversion Already launched in markets including the U.S., Japan, Indonesia, Turkey, Saudi Arabia, Thailand, Brazil, Malaysia, Philippines, and Vietnam, with new markets coming soon Leverages TikTok's global mon
OK      TikTok mini games: launched markets (no EU country listed)
   GET https://developers.tiktok.com/docs/en/mini-games-technical-overview -> HTTP 200, 258376 bytes
      >  it begins from a specific user entry point. The TikTok app selects an available version and obtains its code package. The runtime creates a game instance and executes JavaScrip
OK      TikTok mini games: app selects a version and obtains its code package
 
== 5. Tencent's own results: no mini-program count ==
   GET https://www.prnewswire.com/apac/news-releases/tencent-announces-2026-second-quarter-results-302849608.html -> HTTP 200, 332415 bytes
      >  Quarter- on-quarter change (in millions, unless specified) Combined MAU of Weixin and WeChat 1,439 1,411 2 % 1,432 0.5 % Mobile device MAU of QQ 520 532 -2 % 
OK      Tencent Q2 2026: combined Weixin and WeChat MAU 1,439 million
      "Mini Program" mentions in the Q2 2026 release: 0
      of them next to a count: 0
      > count_mentions=0
OK      Tencent Q2 2026: no mini-program count next to any "Mini Program" mention
      e6a646796d0d869acc76271c9ee1a6a5.pdf: positive control Combined MAU of Weixin and WeChat 1,418; "Mini Program" mentions 2, next to a count 0
      > ds (where the user clicks through to native transactional experiences, su ch as Mini Programs, Mini Shops, or Mini Games). Impression growth benefitted primarily from great
      >  We grew user engagement with Mini Shops, Mini Game s and other content-related Mini Programs at rapid year-on-year rates’, by strengthening Weixin’s commerce experience an
      > control=True count_mentions=0
OK      Tencent results PDF e6a646796d0d869acc76271c9ee1a6a5.pdf: MAU line present, no mini-program count
      47382ae415a209fd161bc19a1f9b3704.pdf: positive control Combined MAU of Weixin and WeChat 1,432; "Mini Program" mentions 0, next to a count 0
      > control=True count_mentions=0
OK      Tencent results PDF 47382ae415a209fd161bc19a1f9b3704.pdf: MAU line present, no mini-program count
 
== 6. Venues, preprints and the workshop ==
   GET https://superappsec.github.io/ -> HTTP 200, 15133 bytes
      > p on Security and Safety of AI-Empowered Mobile Super Apps (SaTS '26) Co-located with ACM CCS 2026 » November 19th, 2026 cfp anchor In response to adapt to po
OK      SaTS 26 co-located with CCS 2026
      > of Agentic frameworks and assistants on Mobile (Super) Apps Agentic/LLM-based techniques for security analysis in Mobile (Super) Apps User-centric a
OK      SaTS 26 topics include Agentic/LLM-based techniques
   GET https://arxiv.org/abs/2306.07495 -> HTTP 200, 41522 bytes
      > ment.documentElement.classList.add('js'); [2306.07495] SoK: Decoding the Super App Enigma: The Security Mechanisms, Threats, and Trade-offs in OS-ali
OK      SoK super app: arXiv 2306.07495, Yang Wang Zhang Lin
      OpenAlex locations: [('arXiv (Cornell University)', 'repository'), ('arXiv (Cornell University)', 'repository')]
      > nonrepo=0 total=2
OK      SoK super app: OpenAlex lists only repository locations (no venue)
   GET https://arxiv.org/abs/2608.17538 -> HTTP 200, 42857 bytes
      > s notable security risks. As we demonstrate, many Mini Apps store authentication materials---such as session tokens and wallet mnemonic phrases---in plaintext on client devices, exposing users to unauthorized access, i
OK      TENET preprint: Telegram Mini App (in)security, plaintext tokens and mnemonics
   GET https://arxiv.org/abs/2608.13390 -> HTTP 200, 43093 bytes
      > cumentElement.classList.add('js'); [2608.13390] TeleGapper: On the (un)reliability of Privacy Policies in Telegram Mini apps Skip to main content Search Submit Donate Log in Search arX
OK      TeleGapper preprint: privacy policies in Telegram Mini apps
      > asure their resource consumption, API usage, library usage, obfuscation rate, app categorization, and app ratings at an aggregated level
OK      SIGMETRICS 2021 abstract: measures obfuscation rate
   GET https://arxiv.org/abs/2607.08232 -> HTTP 200, 43541 bytes
      >  guidance and enhanced platform-level safeguards. Comments: Accepted by ACM CCS 2026 Subjects: Cryptography and Security (cs.CR) Cite as: arXiv:
OK      CCS 2026 OAuth-misuse preprint: accepted by ACM CCS 2026
      10.1145/3460081: A Measurement Study of Wechat Mini-Apps | Proceedings of the ACM on Measurement and Analysis of Computing Systems | [2021, 6]
      > {"status":"ok","message-type":"work","message-version":"1.0.0","message":{
OK      Crossref resolves 10.1145/3460081
      10.1145/3607199.3607236: Measuring the Leakage and Exploitability of Authentication Secrets in Super-apps: The WeChat Case | Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses | [2023, 10, 16]
      > {"status":"ok","message-type":"work","message-version":"1.0.0","message":{
OK      Crossref resolves 10.1145/3607199.3607236
      10.1109/icse48619.2023.00086: Taintmini: Detecting Flow of Sensitive Data in Mini-Programs with Static Taint Analysis | 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) | [2023, 5]
      > {"status":"ok","message-type":"work","message-version":"1.0.0","message":{
OK      Crossref resolves 10.1109/icse48619.2023.00086
      10.1109/ase56229.2023.00151: Wemint:Tainting Sensitive Data Leaks in WeChat Mini-Programs | 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE) | [2023, 9, 11]
      > {"status":"ok","message-type":"work","message-version":"1.0.0","message":{
OK      Crossref resolves 10.1109/ase56229.2023.00151
      DOI:10.1145/3774904.3792470: Real or Rogue? Detecting Malicious Miniapps with Deceptive Reporting Interface | Today, mobile super apps such as WeChat offer a wide array of services through integrated miniapps. While the miniapps provide self-contained services via JavaS
      DOI:10.1109/SP63933.2026.00074: Convenience at a Cost: the Security Risks of Template-Based Development in the App-in-App Ecosystem | Recently, many popular mobile applications, such as WeChat and Alipay, have evolved into super-apps, with numerous merchants tending to develop their own mini-a
 
== 7. Outside-the-extraction paper read from the publisher PDF ==
      > r empirical analysis of four major super- apps reveals that 183 (8.85%) APIs are not properly protected at the mini-program API layer. We discuss the challenges un- derlying this issue and advo
OK      Raising the Flag: 183 (8.85%) APIs not properly protected
      > scope for this permission. 3.4 Analysis Scope In this work, we analyze four popular super-apps: WeChat, Alipay, Baidu, and QQ. Unlike prior works that mainly per- form cross-platform an
OK      Raising the Flag: four super-apps WeChat, Alipay, Baidu, QQ
      > ntirely unguarded). Specifically, we identify 183 such APIs across 2,067 analyzed APIs, with WeChat exhibiting the highest number (96, 9.36%), fol
OK      Raising the Flag: 2,067 analysed APIs
      > e total monetary cost of LLM-based test-case generation was approximately 55 USD. Answer to RQ1: PERM SCOPE reliably generates valid test ca
OK      Raising the Flag: LLM test-case generation cost about 55 USD
      >  tasks forwarded via Binder IPC. To address this challenge, we use Xposed [ 21] to instru- ment super-apps and reconstruct mini-program API execution flows. Our appro
OK      Raising the Flag: Xposed instruments the super-apps
 
FAILED checks: 0

The reader brief, and the machine check of the readers' notes

mp_reader_brief.md
# Reader brief — design:mobile_and_app_measurement:mini_programs (hand codes)
 
You are reading research papers for a wiki page about measuring **mini-programs** (also called
mini-apps, miniapps, app-in-app, sub-apps, smart programs, mini-games): third-party programs that run
INSIDE a host "super app" (WeChat, Alipay, Baidu, TikTok/Douyin, QQ, Taobao, IoT "all-in-one" apps...)
and reach the device and user data only through APIs the HOST provides. The page is for a PhD student
about to run such a measurement. What they need from each paper: which host(s); how the population of
mini-programs was obtained when there is no public store listing; what counts as one unit; static vs
dynamic analysis and the exact tooling; the host-app version and device setup; accounts, location and
language (these ecosystems are China-centric); validation; ethics and disclosure; artefacts; and the
paper's measured results with their own denominators.
 
## Where the text is
Full text: `/workspace/publications_dataset/data/fulltext/<year>/<venue>/<slug>/paper.cols.txt`
(key format below is `<venue>/<year>/<slug>`; the directory order is year/venue/slug).
Grep it whitespace-collapsed: `tr -s '[:space:]' ' ' < paper.cols.txt | grep -oiE ".{0,300}mini-program.{0,300}"`.
Two-column PDFs are sometimes spliced in `.cols`; if a sentence looks broken, also try
`python3 /workspace/artifacts/wiki/scripts/pdftext.py <venue>/<year>/<slug> | grep -oiE ".{0,300}PATTERN.{0,300}"`.
Read the methodology / dataset / implementation / evaluation / ethics / discussion sections fully, not
just grep hits. The data is read-only.
 
## What to produce
Write ONE markdown file (path given in your task) with one section per paper, in this exact shape.
Every field that makes a factual claim carries a VERBATIM quote from the paper (copy-paste exact
characters from the text; no paraphrase inside quotation marks; never join two separate sentences with
"..."; if you must shorten, quote a shorter contiguous span instead). If the paper does not say, write
`not stated` — do not infer. Write the file as you go (append after each paper), so partial work survives.
 
```
### <key>
- verdict: IN-CORE | IN-SECTION | CONTEXT-<code> | OUT-<code>     (rule below)
- hosts: which host apps (WeChat / Alipay / Baidu / TikTok-Douyin / QQ / Taobao / IoT host / other), host
  versions and OS if stated  > "quote"
- side: vulnerability | privacy-tracking | malware | advertising/dark-patterns | host-API | other (say)
- object: one line — what the paper measures
- acquisition: how the mini-programs were obtained — one or more of: crawl-host-search(<endpoint/keywords>),
  crawl-directory(<which>), device-cache(packages pulled from the host app's storage on a phone),
  host-download-api, third-party-dataset(<which, e.g. MiniCrawler>), github-source, manual-selection,
  QR-codes, not stated  > "quote" per claim
- unit: what one sample is (appid / package file / version / subpackage), and dedup  > "quote" or not stated
- scale: candidates found vs downloaded vs successfully analysed, with dates of collection — each number
  with its quote  > "quote" (one per number)
- static: unpacker/decompiler (wxappUnpacker, unveilr, own...), analysis (taint, AST, call graph, regex,
  CodeQL, TaintMini, LLM...), language (JS/WXML) and failure rate  > "quote" per claim
- dynamic: device or emulator, OS, host-app version, instrumentation (Frida, WeChat DevTools, remote debug,
  proxy/mitmproxy, vConsole), UI automation (Appium/UIAutomator/own), coverage  > "quote" per claim
- host-api-knowledge: where the API inventory came from (official docs, reverse-engineered host code,
  hidden/undocumented APIs), and whether results are pinned to a host version  > "quote"
- accounts-vantage-language: real-name accounts, Chinese phone numbers, location of the measurement,
  Chinese-language handling (keywords, translation, NLP on Chinese text), region restrictions  > "quote"
- validation: how precision/recall of the detector was checked (manual sample size, ground truth)  > "quote"
- ethics-disclosure: IRB; testing on own accounts only; disclosure to host vendor (Tencent, Alibaba,
  ByteDance, Baidu...) and to mini-program developers; bug bounty/CVE; vendor response  > "quote" per claim
- artifacts: code/dataset released? URL as printed  > "quote" or not stated
- denominator: does the paper acknowledge its sample is not the population (search bias, popularity
  bias, unreachable/removed mini-programs, region)?  > "quote" or not stated
- headline-figures: up to 4 measured results with the paper's own denominator  > "quote" each
- notes: anything a methods-page author should know (surprises, limitations the paper states, tools named)
```
 
## Verdict rule (fixed; apply it, do not reinterpret)
IN-CORE    — mini-programs (or the host app's mini-program framework/APIs) are the paper's main object of
             measurement or analysis.
IN-SECTION — mini-programs are one analysed population among several (e.g. one of three platforms).
CONTEXT-<code> — adjacent: e.g. `source-repos` (mini-program source on GitHub, not deployed ones),
             `host-feature` (a super app's own feature, not its mini-programs), `user-study`.
OUT-<code> — WeChat/Alipay only as a messenger, payment method, dataset owner or recruitment channel;
             related-work mention; homograph.
For CONTEXT/OUT papers you may shorten to: verdict, one deciding quote, and any field useful to the page.
 
Your context may not be exhaustive; if a paper seems to be missing text, say so rather than guessing.
When done, reply with a one-line summary per paper (key: verdict).
mp_notes_quotecheck-output.txt
mp_papers_A.md: 64/96 quotes located
mp_papers_B.md: 79/119 quotes located
mp_papers_C.md: 61/123 quotes located
mp_papers_D.md: 69/90 quotes located
TOTAL: 273/428 located (63.8%); by first rendering: {'paper.cols.txt': 225, 'pypdf': 48}
MISS: 155; word-position coverage of the misses: {'>=90% (splice or agent insertion)': 109, '60-90%': 35, '<60% (paraphrase or fabrication)': 11}
  mp_papers_A.md CCS/2020/demystifying-resource-management-risks-in-emerging-mobile-app-in-app-ecosystems cov=0.87
    "All test cases were then used by a sub-app we built to invoke subapp APIs."
  mp_papers_A.md CCS/2022/cross-miniapp-request-forgery-root-causes-attacks-and-vulnerability-detection cov=1.00
    "A miniapp is a full-fledged app that is executed inside a mobile super app such as WeChat or SnapChat."
  mp_papers_A.md CCS/2022/cross-miniapp-request-forgery-root-causes-attacks-and-vulnerability-detection cov=1.00
    "We have tested CmrfScanner with 2,571,490 WeChat miniapps and 148,512 Baidu miniapps"
  mp_papers_A.md CCS/2022/cross-miniapp-request-forgery-root-causes-attacks-and-vulnerability-detection cov=1.00
    "To crawl the miniapp for the testing by CmrfScanner, we used our open source MiniCrawler [48] to download the miniapps from WeChat app store. We obtained 2,571,490 WeChat miniapps in total, which consume 6.29 TB disk sto"
  mp_papers_A.md CCS/2022/cross-miniapp-request-forgery-root-causes-attacks-and-vulnerability-detection cov=0.96
    "We also extended MiniCralwer to allow it to download miniapps from Baidu market, with which we also collected 148,512 Baidu miniapps (consuming 81 GB disk storage)."
  mp_papers_A.md CCS/2022/cross-miniapp-request-forgery-root-causes-attacks-and-vulnerability-detection cov=1.00
    "We have tested CmrfScanner with 2,571,490 WeChat miniapps and 148,512 Baidu miniapps, and identified 52,394 (2.04%) WeChat miniapps and 494 (0.33%) Baidu miniapps that involve cross-communication. Among them, CmrfScanner"
  mp_papers_A.md CCS/2022/cross-miniapp-request-forgery-root-causes-attacks-and-vulnerability-detection cov=1.00
    "we used DoubleX due to its easy AST traversal APIs and value analysis components. We also modified DoubleX to enable the parsing of JS files in miniapp and used domain knowledge for function entry identification, appID c"
  mp_papers_A.md CCS/2022/cross-miniapp-request-forgery-root-causes-attacks-and-vulnerability-detection cov=0.86
    "With the knowledge obtained from the reverse engineering of the victim miniapps, we then created the corresponding fake request and injected the message through navigateToMiniProgram (CMRF-DM), or obtained the transmitte"
  mp_papers_A.md CCS/2022/cross-miniapp-request-forgery-root-causes-attacks-and-vulnerability-detection cov=1.00
    "a receiver miniapp must fetch the sender miniapp's appId from the object referrerInfo.appId, which is set by the WeChat framework"
  mp_papers_A.md CCS/2022/cross-miniapp-request-forgery-root-causes-attacks-and-vulnerability-detection cov=0.93
    "Given that there is no ground truth... we sampled 100 miniapps from the miniapps identified as vulnerable and checked to verify the false positives, and sampled 100 miniapps from the miniapps identified as not involving "
  mp_papers_A.md CCS/2022/cross-miniapp-request-forgery-root-causes-attacks-and-vulnerability-detection cov=1.00
    "Among the 100 miniapps that are not identified vulnerable, we found 2 false negatives, making the FN rate to 2%."
  mp_papers_A.md CCS/2022/cross-miniapp-request-forgery-root-causes-attacks-and-vulnerability-detection cov=0.91
    "Our CmrfScanner, whose source code has been made available at https://github.com/OSUSecLab/CMRFScanner."
  mp_papers_A.md CCS/2022/cross-miniapp-request-forgery-root-causes-attacks-and-vulnerability-detection cov=1.00
    "We plan to release a set of the vulnerable miniapps including these manually labeled 200 miniapps to support open science"
  mp_papers_A.md CCS/2022/cross-miniapp-request-forgery-root-causes-attacks-and-vulnerability-detection cov=1.00
    "50,281 (95.97%) of WeChat miniapps, and 493 (99.80%) of Baidu miniapps lack the appID checks"
  mp_papers_A.md CCS/2022/cross-miniapp-request-forgery-root-causes-attacks-and-vulnerability-detection cov=0.89
    "620 and 4 miniapps from WeChat and Baidu, respectively"
  mp_papers_A.md USENIX/2022/identity-confusion-in-webview-based-mobile-app-in-app-ecosystems cov=0.98
    "Table 1 shows a list of the top 15 popular super-apps ranked by total downloads according to our survey study. These super-apps are diversified, which ranges from communication and social to finance and business and span"
  mp_papers_A.md USENIX/2022/identity-confusion-in-webview-based-mobile-app-in-app-ecosystems cov=1.00
    "The first step gives us 47 super-apps: The number of remaining apps after applying each filter is shown in Table 4."
  mp_papers_A.md USENIX/2022/identity-confusion-in-webview-based-mobile-app-in-app-ecosystems cov=0.97
    "we use dynamic instrumentation to discover indirect hidden API calls. Specifically, we hook statically-identified container objects (e.g., via Xposed [13]) and then generate test cases to trigger documented public runtim"
  mp_papers_A.md USENIX/2022/identity-confusion-in-webview-based-mobile-app-in-app-ecosystems cov=0.98
    "We had informed all the 47 super-apps of their vulnerabilities. Currently, 29 super-apps have confirmed their vulnerabilities, and 19 have already fixed them. Take Alipay, for an example. We had regular monthly meetings "
  mp_papers_A.md USENIX/2022/identity-confusion-in-webview-based-mobile-app-in-app-ecosystems cov=0.90
    "they (both the Android and iOS versions) are all vulnerable to at least one type of identity confusion attack... Nine super-apps adopt no identity checks at all... all 38 super-apps with AppID checks are vulnerable; all "
  mp_papers_A.md USENIX/2022/identity-confusion-in-webview-based-mobile-app-in-app-ecosystems cov=0.80
    "over 1.3 million WeChat mini-apps."
  mp_papers_A.md CCS/2023/dont-leak-your-keys-understanding-measuring-and-exploiting-the-appsecret-leaks-i cov=0.93
    "We used the innersearch API (obtained through reverse engineering of WeChat) to search for and download mini-programs, and the waVerifyInfo API to collect developer information... We employed 1,000 commonly used Chinese "
  mp_papers_A.md CCS/2023/dont-leak-your-keys-understanding-measuring-and-exploiting-the-appsecret-leaks-i cov=1.00
    "Note that the WeChat market has about 4 million mini-programs [5], making our dataset likely to cover the majority of them."
  mp_papers_A.md CCS/2023/dont-leak-your-keys-understanding-measuring-and-exploiting-the-appsecret-leaks-i cov=0.95
    "We checked if the mini-programs had leaked their MK(s) by identifying all 256-bit hexadecimal digit strings and pruning them based on the MK validation API... We used a regular expression with a 32-byte hex-string format"
  mp_papers_A.md CCS/2023/dont-leak-your-keys-understanding-measuring-and-exploiting-the-appsecret-leaks-i cov=0.83
    "obtained through reverse engineering of WeChat"
  mp_papers_A.md CCS/2023/dont-leak-your-keys-understanding-measuring-and-exploiting-the-appsecret-leaks-i cov=0.94
    "We employed 1,000 commonly used Chinese characters and 1,000 commonly used English words as seed keywords"
  mp_papers_A.md CCS/2023/dont-leak-your-keys-understanding-measuring-and-exploiting-the-appsecret-leaks-i cov=0.96
    "We respected rate limits defined by WeChat and Baidu servers throughout our study... Our testing was conducted within controlled parameters, solely involving our accounts, devices, and servers, avoiding attacks on third-"
  mp_papers_A.md CCS/2023/dont-leak-your-keys-understanding-measuring-and-exploiting-the-appsecret-leaks-i cov=0.91
    "we choose not to disclose their names due to ethics concerns."
  mp_papers_A.md CCS/2023/dont-leak-your-keys-understanding-measuring-and-exploiting-the-appsecret-leaks-i cov=0.88
    "the list of 40,880 mini-programs that leaked MKs"
  mp_papers_A.md CCS/2023/dont-leak-your-keys-understanding-measuring-and-exploiting-the-appsecret-leaks-i cov=0.00
    "not currently exploitable"
  mp_papers_A.md CCS/2023/dont-leak-your-keys-understanding-measuring-and-exploiting-the-appsecret-leaks-i cov=1.00
    "Note that the WeChat market has about 4 million mini-programs [5], making our dataset likely to cover the majority of them."
  mp_papers_A.md CCS/2023/dont-leak-your-keys-understanding-measuring-and-exploiting-the-appsecret-leaks-i cov=0.00
    "denominator acknowledged"
  mp_papers_B.md CCS/2023/uncovering-and-exploiting-hidden-apis-in-mobile-super-apps cov=1.00
    "We excluded other super apps such as Alipay and Snapchat particularly because they do not build on the V8 engine (making our tool unsuitable for them at this moment)."
  mp_papers_B.md CCS/2023/uncovering-and-exploiting-hidden-apis-in-mobile-super-apps cov=0.80
    "we found this hidden API was still functional... Our experiments were conducted primarily in 2021"
  mp_papers_B.md CCS/2023/uncovering-and-exploiting-hidden-apis-in-mobile-super-apps cov=0.94
    "Fortunately, we can use Frida [15], an Android hooking tool, to dynamically instrument the V8 Engine to invoke startProfiling of Profiler and let it start profiling, and collect the function traces."
  mp_papers_B.md CCS/2023/uncovering-and-exploiting-hidden-apis-in-mobile-super-apps cov=1.00
    "we have disclosed the vulnerabilities and our attacks against WeChat to Tencent in September 2021, and the other four super apps in November 2021. They have all acknowledged and confirmed our findings, and so far among t"
  mp_papers_B.md CCS/2023/uncovering-and-exploiting-hidden-apis-in-mobile-super-apps cov=0.40
    "8 APIs (7.08%) in Baidu... 32 APIs (26.67%) in Tiktok"
  mp_papers_B.md CCS/2023/uncovering-and-exploiting-hidden-apis-in-mobile-super-apps cov=0.93
    "During static API recognition, APIScope recognized in total 1,829 API candidates for these super apps."
  mp_papers_B.md USENIX/2023/one-size-does-not-fit-all-uncovering-and-exploiting-cross-platform-discrepant-ap cov=0.96
    "we run our A PI D IFF on six devices, two Windows-11 desktops, and four smartphones (two with Android-13, and two with iOS-16)."
  mp_papers_B.md USENIX/2023/one-size-does-not-fit-all-uncovering-and-exploiting-cross-platform-discrepant-ap cov=0.96
    "we focus exclusively on WeChat for four reasons. First, WeChat has the largest number of users (with 1.2 billion monthly active users)... Second, WeChat pioneered the concept of miniapp paradigm, and so far it has more t"
  mp_papers_B.md USENIX/2023/one-size-does-not-fit-all-uncovering-and-exploiting-cross-platform-discrepant-ap cov=0.84
    "For each API, A PI D IFF needs to create a test case, with the corresponding parameters properly initialized."
  mp_papers_B.md USENIX/2023/one-size-does-not-fit-all-uncovering-and-exploiting-cross-platform-discrepant-ap cov=0.85
    "some results from the experiments require us to reverse engineer WeChat, and therefore, we used JEB [5] and IDA Pro [23] to inspect the decompiled code statically"
  mp_papers_B.md USENIX/2023/one-size-does-not-fit-all-uncovering-and-exploiting-cross-platform-discrepant-ap cov=0.97
    "programmed Frida [25] scripts to dynamically verify our findings. Again, we run our A PI D IFF on six devices, two Windows-11 desktops, and four smartphones (two with Android-13, and two with iOS-16)."
  mp_papers_B.md USENIX/2023/one-size-does-not-fit-all-uncovering-and-exploiting-cross-platform-discrepant-ap cov=0.97
    "manual efforts were only required at the beginning to investigate the workflow, including investigating the possible error codes that may be observed for each API. After the tool is built, no further manual analysis is r"
  mp_papers_B.md USENIX/2023/one-size-does-not-fit-all-uncovering-and-exploiting-cross-platform-discrepant-ap cov=0.95
    "we propose first extracting the parameter type from the documentation and then initializing them based on the domain knowledge."
  mp_papers_B.md USENIX/2023/one-size-does-not-fit-all-uncovering-and-exploiting-cross-platform-discrepant-ap cov=0.94
    "we closed the super app first and then re-launched it again to run the experiment a second time. This was done to ensure that any observed differences were not caused by the running environment."
  mp_papers_B.md USENIX/2023/one-size-does-not-fit-all-uncovering-and-exploiting-cross-platform-discrepant-ap cov=0.98
    "We reported our findings to Tencent and they acknowledged them by awarding us bug bounties. Tencent's security engineers have actively worked with us over the past year, meeting online multiple times to discuss vulnerabi"
  mp_papers_B.md USENIX/2023/one-size-does-not-fit-all-uncovering-and-exploiting-cross-platform-discrepant-ap cov=0.94
    "we have identified 22 APIs that have output discrepancies, falling into 8 categories including UI, media, and device."
  mp_papers_B.md USENIX/2023/one-size-does-not-fit-all-uncovering-and-exploiting-cross-platform-discrepant-ap cov=0.00
    "unit = documented API"
  mp_papers_B.md USENIX/2023/one-size-does-not-fit-all-uncovering-and-exploiting-cross-platform-discrepant-ap cov=0.00
    "unit = deployed miniapp,"
  mp_papers_B.md CCS/2024/minicat-understanding-and-detecting-cross-page-request-forgery-vulnerabilities-i cov=0.96
    "Crawling WeChat mini-programs is challenging since there are no official or third-party markets similar to Google Play [14] or Apkpure [4] for Android apps."
  mp_papers_B.md CCS/2024/minicat-understanding-and-detecting-cross-page-request-forgery-vulnerabilities-i cov=1.00
    "Furthermore, even without routing paths and parameters (notably in Baidu and TikTok), attackers can exploit mini-programs with the same name on other platforms and then perform attacks."
  mp_papers_B.md CCS/2024/minicat-understanding-and-detecting-cross-page-request-forgery-vulnerabilities-i cov=0.84
    "we discovered that when a user accesses a mini-program on the WeChat Windows client, the client creates a directory under the user profile folder to store the mini-program, located at user_file/Applet/AppID."
  mp_papers_B.md CCS/2024/minicat-understanding-and-detecting-cross-page-request-forgery-vulnerabilities-i cov=0.90
    "batch querying of AppIDs has become impossible due to Tencent's restriction on related API access... we developed an automated crawler that simulates user actions on the WeChat Windows client... Our crawler, called Mini-"
  mp_papers_B.md CCS/2024/minicat-understanding-and-detecting-cross-page-request-forgery-vulnerabilities-i cov=1.00
    "MiniCAT successfully analyzed 41,726/44,273 (94.2%), identifying 13,349/41,726 (32.0%) as potentially vulnerable with a cumulative 119,471 risky pages."
  mp_papers_B.md CCS/2024/minicat-understanding-and-detecting-cross-page-request-forgery-vulnerabilities-i cov=0.97
    "For those non-unpackable mini-programs, we found that they were incompatible with the wxappUnpacker tool due to their use of a newer version of the WeChat mini-program base library. Additionally, failures in unpacking we"
  mp_papers_B.md CCS/2024/minicat-understanding-and-detecting-cross-page-request-forgery-vulnerabilities-i cov=0.94
    "According to the WeChat developer documentation [37], page routing APIs of WeChat mini-programs are called in the logic layer of the mini-program (i.e., in JavaScript files). We focus on three routing APIs: wx.navigateTo"
  mp_papers_B.md CCS/2024/minicat-understanding-and-detecting-cross-page-request-forgery-vulnerabilities-i cov=0.94
    "We have also contacted CNCERT/CC [8], the Chinese vulnerability coordination organization, and disclosed our findings with CNVD [7]... three cases have been confirmed (CNVD-2023-75836, CNVD-2023-75837, and CNVD-2024-0552"
  mp_papers_B.md CCS/2024/minicat-understanding-and-detecting-cross-page-request-forgery-vulnerabilities-i cov=0.94
    "Although WeChat has no official mini-program ranking, previous studies [70, 73] used ratings to assess their popularity."
  mp_papers_B.md CCS/2024/minicat-understanding-and-detecting-cross-page-request-forgery-vulnerabilities-i cov=0.00
    "batch AppID query blocked by Tencent"
  mp_papers_B.md NDSS/2026/better-safe-than-sorry-uncovering-the-insecure-resource-management-in-app-in-app-cloud-services cov=0.97
    "for the mini-app dataset collection, we adhered to the methods established in previous studies [26], [34]. To avoid overloading the super-app's servers, we also limited the download speed to a few seconds per mini-app."
  mp_papers_B.md NDSS/2026/better-safe-than-sorry-uncovering-the-insecure-resource-management-in-app-in-app-cloud-services cov=0.83
    "In total, we analyzed 1,943 mini-apps"
  mp_papers_B.md NDSS/2026/better-safe-than-sorry-uncovering-the-insecure-resource-management-in-app-in-app-cloud-services cov=0.96
    "attackers can leverage the code caching mechanism of super-apps to achieve this on their own devices without the need to modify network traffic."
  mp_papers_B.md NDSS/2026/better-safe-than-sorry-uncovering-the-insecure-resource-management-in-app-in-app-cloud-services cov=0.93
    "We analyze the developer documentation provided by super-app platforms to model their resource management workflows."
  mp_papers_B.md NDSS/2026/better-safe-than-sorry-uncovering-the-insecure-resource-management-in-app-in-app-cloud-services cov=0.88
    "Our work was reviewed by our institution's IRB and this study is considered as 'minimal risk'."
  mp_papers_B.md NDSS/2026/better-safe-than-sorry-uncovering-the-insecure-resource-management-in-app-in-app-cloud-services cov=0.94
    "we did not perform any write operations and conducted security inferences without accessing the cloud data... we did not conduct fuzz testing on cloud services. Instead, we limited the detection rate to simulate manual a"
  mp_papers_B.md NDSS/2026/better-safe-than-sorry-uncovering-the-insecure-resource-management-in-app-in-app-cloud-services cov=0.89
    "The super-app platforms have recognized this vulnerability. Besides, we actively worked together with them to fix these problems."
  mp_papers_B.md NDSS/2026/better-safe-than-sorry-uncovering-the-insecure-resource-management-in-app-in-app-cloud-services cov=0.94
    "the source code of the artifact, along with the required scripts, can be downloaded from: https://doi.org/10.5281/zenodo.16946146."
  mp_papers_B.md NDSS/2026/better-safe-than-sorry-uncovering-the-insecure-resource-management-in-app-in-app-cloud-services cov=0.70
    "We provide the test dataset in the directory: './StaticAnalysis/componentAnalyze/test dataset'."
  mp_papers_B.md NDSS/2026/better-safe-than-sorry-uncovering-the-insecure-resource-management-in-app-in-app-cloud-services cov=0.97
    "our current work focuses on the analysis in four prominent super-apps. The assessment can demonstrate the scalability of ICREM INER, and we also analyze the security issues in other super-apps, such as Line and VK."
  mp_papers_B.md NDSS/2026/better-safe-than-sorry-uncovering-the-insecure-resource-management-in-app-in-app-cloud-services cov=1.00
    "our method may introduce some false positives, primarily because some data are not privacy-sensitive. ICREM INER determines whether data is privacy-sensitive based on variable names and data dependencies."
  mp_papers_B.md NDSS/2026/better-safe-than-sorry-uncovering-the-insecure-resource-management-in-app-in-app-cloud-services cov=0.83
    "115 vulnerable mini-apps each have over 100,000 users, and the cumulative number of affected users exceeds 70 million"
  mp_papers_C.md PETS/2025/what-wechat-knows-pervasive-first-party-tracking-in-a-billion-user-super-app-eco cov=0.91
    "The current state-of-the-art in Mini Program measurement is to search common Chinese keywords and download the Mini Programs from the search results [8]. We used similar methods to identify and install popular Mini Progr"
  mp_papers_C.md PETS/2025/what-wechat-knows-pervasive-first-party-tracking-in-a-billion-user-super-app-eco cov=0.73
    ", and collected 45 apps in total from the search results."
  mp_papers_C.md PETS/2025/what-wechat-knows-pervasive-first-party-tracking-in-a-billion-user-super-app-eco cov=0.56
    "since WeChat Games are implemented differently from Mini Programs."
  mp_papers_C.md PETS/2025/what-wechat-knows-pervasive-first-party-tracking-in-a-billion-user-super-app-eco cov=0.92
    "Unfortunately, 38.8% of them required ID verification or Chinese phone number verification."
  mp_papers_C.md PETS/2025/what-wechat-knows-pervasive-first-party-tracking-in-a-billion-user-super-app-eco cov=0.82
    "we identified fine-grained browsing data in 89.7% of the traces we decrypted from 40 health-related Mini Programs."
  mp_papers_C.md PETS/2025/what-wechat-knows-pervasive-first-party-tracking-in-a-billion-user-super-app-eco cov=1.00
    "We used Frida, a dynamic instrumentation toolkit, to hook into the app's functions and manipulate application memory [39]. We also captured and studied network traffic using Wireshark and tcpdump to study the network pro"
  mp_papers_C.md PETS/2025/what-wechat-knows-pervasive-first-party-tracking-in-a-billion-user-super-app-eco cov=0.68
    "WeChat sends requests to three different internal APIs during regular Mini Program operation ... JsApiOperateRealtimeReport ... AppBrandIDKeyBatchReport ... JsApiOperateWXData."
  mp_papers_C.md PETS/2025/what-wechat-knows-pervasive-first-party-tracking-in-a-billion-user-super-app-eco cov=0.83
    "make this tooling public at https://github.com/citizenlab/wechat-security-report."
  mp_papers_C.md PETS/2025/what-wechat-knows-pervasive-first-party-tracking-in-a-billion-user-super-app-eco cov=0.77
    "we identified fine-grained browsing data in 76.0% of the network traces we decrypted."
  mp_papers_C.md PETS/2025/what-wechat-knows-pervasive-first-party-tracking-in-a-billion-user-super-app-eco cov=0.97
    "Of the 104 Mini Programs we coded, 51 had profile update flows, 85 had search flows, and 96 had browsing flows. 84.3%, 72.9%, and 76.0% of those flows were exfiltrated to WeChat."
  mp_papers_C.md USENIX/2025/i-can-tell-your-secrets-inferring-privacy-attributes-from-mini-app-interaction-h cov=0.91
    "we pushed the AliPay-dev version to their devices ... The only difference is that it displays an additional user agreement ... Note that we do not add extra instrumentation or probes in AliPay-dev."
  mp_papers_C.md USENIX/2025/i-can-tell-your-secrets-inferring-privacy-attributes-from-mini-app-interaction-h cov=0.90
    "internal engineers extracted their mini-app interaction history from AliPay's servers."
  mp_papers_C.md USENIX/2025/i-can-tell-your-secrets-inferring-privacy-attributes-from-mini-app-interaction-h cov=0.80
    "28 types following prior literature."
  mp_papers_C.md USENIX/2025/i-can-tell-your-secrets-inferring-privacy-attributes-from-mini-app-interaction-h cov=0.97
    "we record 1) the unique id miid of the mini-app (maintained by the AliPay backend), 2) the mini-app category code mic, and 3) the number of access times mif over the last 30 days."
  mp_papers_C.md USENIX/2025/i-can-tell-your-secrets-inferring-privacy-attributes-from-mini-app-interaction-h cov=0.91
    "For location, we give three labels: Tier-1 cities, Tier-2 cities, and Tier-3 cities... Tier-1 cities include municipalities, well-developed provincial capitals, and economic centers."
  mp_papers_C.md USENIX/2025/i-can-tell-your-secrets-inferring-privacy-attributes-from-mini-app-interaction-h cov=0.96
    "we conducted a chi-square test comparing these distributions with data from China's National Bureau of Statistics [7]. The test confirmed no significant statistical difference."
  mp_papers_C.md USENIX/2025/i-can-tell-your-secrets-inferring-privacy-attributes-from-mini-app-interaction-h cov=0.91
    "The average inference accuracy of THEFT is 65.2%... the p value is 0.02. Therefore, we reject the null hypothesis (p < 0.05)."
  mp_papers_C.md USENIX/2025/i-can-tell-your-secrets-inferring-privacy-attributes-from-mini-app-interaction-h cov=0.93
    "We contacted the vendors of all 31 super-apps in Table 2 and the standards association... Among the 31 vendors, eight acknowledged our report, and four acknowledged and provided feedback."
  mp_papers_C.md USENIX/2025/i-can-tell-your-secrets-inferring-privacy-attributes-from-mini-app-interaction-h cov=0.97
    "due to AliPay's IRB and business regulations, and strict user privacy concerns, all code and data are processed on supervised servers. Therefore, we cannot publicly release certain artifacts, including code and comprehen"
  mp_papers_C.md USENIX/2025/i-can-tell-your-secrets-inferring-privacy-attributes-from-mini-app-interaction-h cov=0.96
    "THEFT can achieve more than 95.5% accuracy in inferring privacy attributes of over 16.1% of more than 219K users with the training data from only 200 users."
  mp_papers_C.md USENIX/2025/i-can-tell-your-secrets-inferring-privacy-attributes-from-mini-app-interaction-h cov=0.80
    "only one super-app (WeChat) mentions that it collects Mini-H in the privacy policies and terms"
  mp_papers_C.md USENIX/2025/i-can-tell-your-secrets-inferring-privacy-attributes-from-mini-app-interaction-h cov=1.00
    "none of the existing academic papers has discussed the privacy issues of the mini-app interaction history"
  mp_papers_C.md USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games cov=0.92
    "we selected WeChat Mini Game, Facebook Instant Games, and QuickGame, as they host the largest mini-game ecosystems with each exceeding 100 million monthly active users (MAU)."
  mp_papers_C.md USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games cov=0.95
    "We then crawled the Game category of these platforms using the open-source tool miniCrawler [74]. This process yielded 6,769 mini-games in total, from which we further extracted 2,076 Cocos-based mini-games (1,593 from W"
  mp_papers_C.md USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games cov=0.96
    "we obtained 100 real-world mini-games from an anonymous official audit department. These cases originated from user complaints between January and December 2023 and had been validated by auditors."
  mp_papers_C.md USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games cov=0.96
    "This process yielded 6,769 mini-games in total, from which we further extracted 2,076 Cocos-based mini-games (1,593 from WeChat, 312 from Facebook, and 171 from QuickGame)."
  mp_papers_C.md USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games cov=0.83
    "This process produced 371 labeled Ad-behaviors."
  mp_papers_C.md USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games cov=0.94
    "Stage I unpacks the mini-game, then applies module-, function-, and polyfill-level pruning to strip away engine code."
  mp_papers_C.md USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games cov=0.97
    "we use Esprima [27] to convert JavaScript code to the AST. In order to unify the ES6+ syntax in the code into ES5, we use swc [48] to downgrade ES6+ JavaScript code. We chose WALA [8] for generating call graphs."
  mp_papers_C.md USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games cov=0.92
    "310/371 labeled behaviors as true positive outputs, with a recall of 83.55%."
  mp_papers_C.md USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games cov=0.98
    "we conducted manual behavior labeling: we interacted with the game for five minutes to trigger advertisements on physical devices, applied dynamic instrumentation to ad APIs to capture runtime logs and stack traces, and "
  mp_papers_C.md USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games cov=0.91
    "By reverse-engineering how the Cocos engine parses and links resource files, we gain insights into the conventions and rules governing cross-language references."
  mp_papers_C.md USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games cov=0.94
    "producing 371 labeled Ad-behaviors, which were independently cross-checked by three researchers to ensure consistency and accuracy."
  mp_papers_C.md USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games cov=0.00
    "Overall 854/903 (94.57%) [precision] ... 310/371 (83.55%) [recall]."
  mp_papers_C.md USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games cov=0.95
    "We conducted responsible disclosure through official channels prior to publication, notifying platforms of risky parameter settings and default behaviors... Both the WeChat and QuickGame teams confirmed the issues and ex"
  mp_papers_C.md USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games cov=0.89
    "All artifacts are available at the following link: https://doi.org/10.5281/zenodo.18227703."
  mp_papers_C.md USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games cov=0.83
    "For platform-specific rule-matching components, we provide a packaged decision system rather than exposing raw patterns... Sensitive or proprietary content (e.g., raw user traffic, unmodified platform submissions) has be"
  mp_papers_C.md USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games cov=0.95
    "MAAD is bounded by Cocos-centric design... Limited by our sample size and the fact that dynamic interaction cannot exhaustively trigger all behaviors, some aggressive behaviors, particularly deceptive advertising, have l"
  mp_papers_C.md USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games cov=0.93
    "67.39% of aggressive advertising mini-games configured pop-up intervals shorter than the 30s required by Facebook."
  mp_papers_C.md USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games cov=1.00
    "94.09% of bypass cases relying on dynamic, network-dependent triggers."
  mp_papers_C.md USENIX/2026/when-fun-turns-toxic-a-first-look-at-aggressive-advertising-in-mini-games cov=0.94
    "We identified 18 groups comprising 42 games with highly similar or even identical code, spanning 31 companies."
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=0.97
    "as popular super apps including ALIPAY, TIKTOK, and BAIDU commonly adopt JavaScript as the language for implementing miniapps, the concern of dynamic code execution and content rendering is applicable across platforms."
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=0.98
    "we revisited the miniapp store, querying the miniapp market whether the appIDs of the miniapps collected in the first round still exist. By the end of 2022, we noticed that a significant number of miniapps, 360,467 in to"
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=0.75
    "we collected over 4.5 million miniapps, identifying a subset (19,905) as malicious."
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=0.83
    "360,467 in total, had been delisted"
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=0.00
    "resulting in 19,905 malware samples"
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=1.00
    "carried out from March 2020 to December 2022."
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=0.92
    "829,288 miniapps by the end of the revisiting process in December 2022,"
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=1.00
    "Using static control and data flow techniques, we examined function calls, layout components, and content displayed by the miniapps."
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=0.98
    "In total, 34 of the 500 sampled miniapps are associated with content vetting evasion, whereas the rest 466 sampled miniapps are identified as code vetting evasion. Among all these miniapps, 487 miniapps are correctly ide"
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=0.98
    "our approach is based on downloaded miniapp packages, which are at the front-end. However, due to the evasive nature of these malware, the malware may dynamically hide malicious contents without distributing them to the "
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=0.84
    "we first examined through the official announcements from WECHAT about libraries deemed as capable for miniapps to implement hot update to evade code vetting... since July 2022, WECHAT prohibits the usage of a list of Ja"
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=0.86
    "we extracted signatures from publicized malware examples [9], as well as examples we obtained during our interaction with WECHAT security teams."
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=0.92
    "we sampled a total of 500 miniapps out of the 19,905 malware, and manually checked whether the miniapp involves evasive signatures... Among all these miniapps, 487 miniapps are correctly identified, whilst 13 miniapp wer"
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=0.90
    "we involved three security researchers to evaluate the security impacts"
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=0.96
    "we downloaded over 4 million miniapps. To prevent imposing a burden on Tencent's servers, we deliberately limited the download speed to a few seconds per miniapp."
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=0.90
    "we have reached out to Tencent and shared our findings... We are now working with Tencent security teams for improved countermeasures."
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=0.90
    "The instructions of requesting the dataset is published at https://minimalware.github.io/."
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=0.98
    "our approach is based on downloaded miniapp packages, which are at the front-end. However, due to the evasive nature of these malware, the malware may dynamically hide malicious contents without distributing them to the "
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=0.88
    "focuses solely on one super app, namely WECHAT."
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=0.75
    "we collected over 4.5 million miniapps, identifying a subset (19,905) as malicious."
  mp_papers_C.md NDSS/2025/understanding-miniapp-malware-identification-dissection-and-characterization cov=0.88
    "P4 Rogue Malware Web Earning 4,105 41 20.63%."
  mp_papers_D.md NDSS/2025/the-skeleton-keys-a-large-scale-analysis-of-credential-leakage-in-mini-apps cov=0.93
    "Finally, we totally collect 413,775 mini-apps as our database, including 214,602 WeChat mini-apps, 86,570 Baidu mini-apps, 93,130 Alipay mini-apps, 15,064 TikTok mini-apps, 3,984 Line mini-apps and 425 VK mini-apps."
  mp_papers_D.md NDSS/2025/the-skeleton-keys-a-large-scale-analysis-of-credential-leakage-in-mini-apps cov=0.00
    "should-hold-credential"
  mp_papers_D.md NDSS/2025/the-skeleton-keys-a-large-scale-analysis-of-credential-leakage-in-mini-apps cov=0.93
    "Then, we extend the opensource tool MiniCrawler [24] to support mini-app crawling of different super-apps."
  mp_papers_D.md NDSS/2025/the-skeleton-keys-a-large-scale-analysis-of-credential-leakage-in-mini-apps cov=0.94
    "we develop our own testing miniapps for the in-depth analysis and restrict the vulnerability exploitation verification to internal tests using our own accounts, with the explicit consent of all participants involved."
  mp_papers_D.md USENIX/2024/demystifying-the-security-implications-in-iot-device-rental-services cov=1.00
    "In China, most devices prefer WeChat mini-programs over native Android or iOS apps, as the convenience of accessing these mini-programs directly within the WeChat platform, which eliminates the need for users to install "
  mp_papers_D.md USENIX/2024/demystifying-the-security-implications-in-iot-device-rental-services cov=0.95
    "testing is limited to our private devices or app accounts, ensuring that both attacker and victim scenarios remain within our private environments."
  mp_papers_D.md USENIX/2024/demystifying-the-security-implications-in-iot-device-rental-services cov=0.95
    "we show that 84% of the existing vulnerabilities can be exploited to attack all devices or users of several vendors."
  mp_papers_D.md USENIX/2024/demystifying-the-security-implications-in-iot-device-rental-services cov=1.00
    "testing is limited to our private devices or app accounts, ensuring that both attacker and victim scenarios remain within our private environments, avoiding disruption to vendors or users."
  mp_papers_D.md USENIX/2024/demystifying-the-security-implications-in-iot-device-rental-services cov=0.95
    "not all of these apps (or mini-programs) can be selected for research since we cannot find valid QR codes for them."
  mp_papers_D.md IEEE-SP/2025/hey-your-secrets-leaked-detecting-and-characterizing-secret-leakage-in-the-wild cov=0.95
    "As a purely static tool, KEYSENTINEL struggles to detect secrets in encrypted or heavily obfuscated text, code files, and binary files."
  mp_papers_D.md IEEE-SP/2025/hey-your-secrets-leaked-detecting-and-characterizing-secret-leakage-in-the-wild cov=0.95
    "As a purely static tool, KEYSENTINEL struggles to detect secrets in encrypted or heavily obfuscated text, code files, and binary files."
  mp_papers_D.md IEEE-SP/2025/hey-your-secrets-leaked-detecting-and-characterizing-secret-leakage-in-the-wild cov=0.73
    "all analyses were conducted locally, and no actual attacks were used to validate detected secrets"
  mp_papers_D.md IEEE-SP/2025/hey-your-secrets-leaked-detecting-and-characterizing-secret-leakage-in-the-wild cov=0.00
    "near perfect agreement"
  mp_papers_D.md IEEE-SP/2025/hey-your-secrets-leaked-detecting-and-characterizing-secret-leakage-in-the-wild cov=0.88
    "The labeling process was double-blind... achieved a Cohen's Kappa value of 0.91 [59], indicating "near perfect agreement""
  mp_papers_D.md IEEE-SP/2025/hey-your-secrets-leaked-detecting-and-characterizing-secret-leakage-in-the-wild cov=0.95
    "Following ethical guidelines [65], all analyses were conducted locally, and no actual attacks were used to validate detected secrets."
  mp_papers_D.md IEEE-SP/2025/hey-your-secrets-leaked-detecting-and-characterizing-secret-leakage-in-the-wild cov=0.95
    "we used public contact information to reach project owners, sending 3,906 disclosure emails total and receiving 205 responses as the paper writing... secrets in 12 projects were still valid and posed active security risk"
  mp_papers_D.md IEEE-SP/2025/hey-your-secrets-leaked-detecting-and-characterizing-secret-leakage-in-the-wild cov=0.93
    "Obfuscated secrets fall outside our metrics, as none of the tools we employ perform deobfuscation"
  mp_papers_D.md IEEE-SP/2025/hey-your-secrets-leaked-detecting-and-characterizing-secret-leakage-in-the-wild cov=0.92
    "Since WeChat MPs are powered by JavaScript, .js files account for the majority of leakage instances... their proportion of total project files is relatively low (0.37%)."
  mp_papers_D.md CCS/2024/riotfuzzer-companion-app-assisted-remote-fuzzing-for-detecting-vulnerabilities-i cov=0.90
    "If the mutated key-value pair control command appears in the network payload, a mutation point is found, enabling further side-channel-guided fuzzing."
  mp_papers_D.md CCS/2024/riotfuzzer-companion-app-assisted-remote-fuzzing-for-detecting-vulnerabilities-i cov=0.64
    "All of them have been acknowledged by the corresponding vendors... 8 have been confirmed."
  mp_papers_D.md USENIX/2023/medusa-attack-exploring-security-hazards-of-in-app-qr-code-scanning cov=1.00
    "We conducted an empirical study on 800 very popular Android and iOS apps with billions of users in the two largest mobile ecosystems, the US and mainland China mobile markets, to investigate the prevalence and severity o"

Bibliography entries added

bib_additions_mini_programs.bib
@inproceedings{lu2020_demystifying,
  author        = {Lu, Haoran and Xing, Luyi and Xiao, Yue and Zhang, Yifan and Liao, Xiaojing and Wang, XiaoFeng and Wang, Xueqiang},
  title         = {Demystifying Resource Management Risks in Emerging Mobile App-in-App Ecosystems},
  booktitle     = {Proceedings of the ACM SIGSAC Conference on Computer and Communications Security},
  year          = {2020},
  series        = {CCS 2020},
  doi           = {10.1145/3372297.3417255},
}
 
@inproceedings{yang2022_cross,
  author        = {Yang, Yuqing and Zhang, Yue and Lin, Zhiqiang},
  title         = {Cross Miniapp Request Forgery: Root Causes, Attacks, and Vulnerability Detection},
  booktitle     = {Proceedings of the ACM SIGSAC Conference on Computer and Communications Security},
  year          = {2022},
  series        = {CCS 2022},
  doi           = {10.1145/3548606.3560597},
}
 
@inproceedings{zhang2022_identity,
  author        = {Zhang, Lei and Zhang, Zhibo and Liu, Ancong and Cao, Yinzhi and Zhang, Xiaohan and Chen, Yanjun and Zhang, Yuan and Yang, Guangliang and Yang, Min},
  title         = {Identity Confusion in WebView-based Mobile App-in-app Ecosystems},
  booktitle     = {Proceedings of the USENIX Security Symposium},
  year          = {2022},
  series        = {USENIX Security 2022},
  url           = {https://www.usenix.org/conference/usenixsecurity22/presentation/zhang-lei},
}
 
@inproceedings{zhang2023_leak,
  author        = {Zhang, Yue and Yang, Yuqing and Lin, Zhiqiang},
  title         = {Don't Leak Your Keys: Understanding, Measuring, and Exploiting the AppSecret Leaks in Mini-Programs},
  booktitle     = {Proceedings of the ACM SIGSAC Conference on Computer and Communications Security},
  year          = {2023},
  series        = {CCS 2023},
  doi           = {10.1145/3576915.3616591},
}
 
@inproceedings{wang2023_uncovering,
  author        = {Wang, Chao and Zhang, Yue and Lin, Zhiqiang},
  title         = {Uncovering and Exploiting Hidden APIs in Mobile Super Apps},
  booktitle     = {Proceedings of the ACM SIGSAC Conference on Computer and Communications Security},
  year          = {2023},
  series        = {CCS 2023},
  doi           = {10.1145/3576915.3616676},
}
 
@inproceedings{wang2023_size,
  author        = {Wang, Chao and Zhang, Yue and Lin, Zhiqiang},
  title         = {One Size Does Not Fit All: Uncovering and Exploiting Cross Platform Discrepant APIs in WeChat},
  booktitle     = {Proceedings of the USENIX Security Symposium},
  year          = {2023},
  series        = {USENIX Security 2023},
  url           = {https://www.usenix.org/conference/usenixsecurity23/presentation/wang-chao},
}
 
@inproceedings{zhang2024_minicat,
  author        = {Zhang, Zidong and Hou, Qinsheng and Ying, Lingyun and Diao, Wenrui and Gu, Yacong and Li, Rui and Guo, Shanqing and Duan, Haixin},
  title         = {MiniCAT: Understanding and Detecting Cross-Page Request Forgery Vulnerabilities in Mini-Programs},
  booktitle     = {Proceedings of the ACM SIGSAC Conference on Computer and Communications Security},
  year          = {2024},
  series        = {CCS 2024},
  doi           = {10.1145/3658644.3670294},
}
 
@inproceedings{shi2026_better,
  author        = {Shi, Yizhe and Yang, Zhemin and Liu, Dingyi and Zhong, Kangwei and Dai, Jiarun and Yang, Min},
  title         = {Better Safe than Sorry: Uncovering the Insecure Resource Management in App-in-App Cloud Services},
  booktitle     = {Proceedings of the Network and Distributed System Security Symposium},
  year          = {2026},
  series        = {NDSS 2026},
  url           = {https://www.ndss-symposium.org/ndss-paper/better-safe-than-sorry-uncovering-the-insecure-resource-management-in-app-in-app-cloud-services/},
}
 
@inproceedings{cai2025_tell,
  author        = {Cai, Yifeng and Zhang, Ziqi and Yao, Mengyu and Liu, Junlin and Zhao, Xiaoke and Fu, Xinyi and Li, Ruoyu and Liu, Zhe and Chen, Xiangqun and Guo, Yao and Li, Ding},
  title         = {I Can Tell Your Secrets: Inferring Privacy Attributes from Mini-app Interaction History in Super-apps},
  booktitle     = {Proceedings of the USENIX Security Symposium},
  year          = {2025},
  series        = {USENIX Security 2025},
  url           = {https://www.usenix.org/conference/usenixsecurity25/presentation/cai-yifeng},
}
 
@inproceedings{chen2026_minigames,
  author        = {Chen, Pei and Hong, Geng and Qin, Yicheng and Wang, Huazhe and Wu, Mengying and Yang, Min and Zhao, Ziru and Zhu, Yuanpeng and Su, Tao},
  title         = {When Fun Turns Toxic: A First Look at Aggressive Advertising in Mini-games},
  booktitle     = {Proceedings of the USENIX Security Symposium},
  year          = {2026},
  series        = {USENIX Security 2026},
  url           = {https://www.usenix.org/conference/usenixsecurity26/presentation/chen-pei},
}
 
@inproceedings{yang2025_miniapp,
  author        = {Yang, Yuqing and Zhang, Yue and Lin, Zhiqiang},
  title         = {Understanding Miniapp Malware: Identification, Dissection, and Characterization},
  booktitle     = {Proceedings of the Network and Distributed System Security Symposium},
  year          = {2025},
  series        = {NDSS 2025},
  url           = {https://www.ndss-symposium.org/ndss-paper/understanding-miniapp-malware-identification-dissection-and-characterization/},
}
 
@inproceedings{shi2025_skeleton,
  author        = {Shi, Yizhe and Yang, Zhemin and Zhong, Kangwei and Yang, Guangliang and Yang, Yifan and Zhang, Xiaohan and Yang, Min},
  title         = {The Skeleton Keys: A Large Scale Analysis of Credential Leakage in Mini-apps},
  booktitle     = {Proceedings of the Network and Distributed System Security Symposium},
  year          = {2025},
  series        = {NDSS 2025},
  url           = {https://www.ndss-symposium.org/ndss-paper/the-skeleton-keys-a-large-scale-analysis-of-credential-leakage-in-mini-apps/},
}
 
@inproceedings{he2024_demystifying,
  author        = {He, Yi and Guan, Yunchao and Lun, Ruoyu and Song, Shangru and Guo, Zhihao and Zhuge, Jianwei and Chen, Jianjun and Wei, Qiang and Wu, Zehui and Yu, Miao and Shi, Hetian and Li, Qi},
  title         = {Demystifying the Security Implications in IoT Device Rental Services},
  booktitle     = {Proceedings of the USENIX Security Symposium},
  year          = {2024},
  series        = {USENIX Security 2024},
  url           = {https://www.usenix.org/conference/usenixsecurity24/presentation/he-yi},
}
 
@inproceedings{liu2024_riotfuzzer,
  author        = {Liu, Kaizheng and Yang, Ming and Ling, Zhen and Zhang, Yue and Lei, Chongqing and Luo, Junzhou and Fu, Xinwen},
  title         = {RIoTFuzzer: Companion App Assisted Remote Fuzzing for Detecting Vulnerabilities in IoT Devices},
  booktitle     = {Proceedings of the ACM SIGSAC Conference on Computer and Communications Security},
  year          = {2024},
  series        = {CCS 2024},
  doi           = {10.1145/3658644.3670342},
}
 
@inproceedings{lee2025_deep,
  author        = {Lee, Woonghee and Hur, Junbeom and Kwon, Hyunsoo},
  title         = {Deep Dive into In-app Browsers: Uncovering Hidden Pitfalls in Certificate Validation},
  booktitle     = {Proceedings of the ACM SIGSAC Conference on Computer and Communications Security},
  year          = {2025},
  series        = {CCS 2025},
  doi           = {10.1145/3719027.3765215},
}
 
@inproceedings{wei2026_raising,
  author        = {Wei, Zhiao and Wang, Chao and Faheem, Haseeb-Ur-Rehman and Xing, Luyi and Aafer, Yousra and Lin, Zhiqiang},
  title         = {Raising the Flag: Detecting Missing Permission Controls in Mini-Program APIs},
  booktitle     = {Proceedings of the USENIX Security Symposium},
  year          = {2026},
  series        = {USENIX Security 2026},
  url           = {https://www.usenix.org/conference/usenixsecurity26/presentation/wei-zhiao},
}
 
@inproceedings{shi2026_convenience,
  author        = {Shi, Yizhe and Yang, Zhemin and Yang, Yifan and Yang, Yunteng and Yang, Min},
  title         = {Convenience at a Cost: the Security Risks of Template-Based Development in the App-in-App Ecosystem},
  booktitle     = {Proceedings of the IEEE Symposium on Security and Privacy},
  year          = {2026},
  series        = {IEEE S\&P 2026},
  doi           = {10.1109/sp63933.2026.00074},
}
 
@inproceedings{yang2026_real,
  author        = {Yang, Yuqing and Lin, Zhiqiang},
  title         = {Real or Rogue? Detecting Malicious Miniapps with Deceptive Reporting Interface},
  booktitle     = {Proceedings of the ACM Web Conference},
  year          = {2026},
  series        = {TheWebConf 2026},
  doi           = {10.1145/3774904.3792470},
}
 
@inproceedings{wang2025_wechat,
  author        = {Wang, Mona and Lin, Pellaeon and Knockel, Jeffrey and Greenberg, Will and Mayer, Jonathan and Mittal, Prateek},
  title         = {What WeChat Knows: Pervasive First-Party Tracking in a Billion-User Super-App Ecosystem},
  booktitle     = {Proceedings on Privacy Enhancing Technologies},
  year          = {2025},
  series        = {PoPETs 2025},
  doi           = {10.56553/popets-2025-0163},
}
 
@article{zhang2021_measurement,
  author        = {Zhang, Yue and Turkistani, Bayan and Yang, Allen Yuqing and Zuo, Chaoshun and Lin, Zhiqiang},
  title         = {A Measurement Study of {WeChat} Mini-Apps},
  journal       = {Proceedings of the ACM on Measurement and Analysis of Computing Systems},
  volume        = {5},
  number        = {2},
  year          = {2021},
  series        = {SIGMETRICS 2021},
  doi           = {10.1145/3460081},
}
 
@inproceedings{baskaran2023_measuring,
  author        = {Baskaran, Supraja and Zhao, Lianying and Mannan, Mohammad and Youssef, Amr},
  title         = {Measuring the Leakage and Exploitability of Authentication Secrets in Super-apps: The {WeChat} Case},
  booktitle     = {Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses},
  year          = {2023},
  series        = {RAID 2023},
  doi           = {10.1145/3607199.3607236},
}
 
@inproceedings{wang2023_taintmini,
  author        = {Wang, Chao and Ko, Ronny and Zhang, Yue and Yang, Yuqing and Lin, Zhiqiang},
  title         = {{TaintMini}: Detecting Flow of Sensitive Data in Mini-Programs with Static Taint Analysis},
  booktitle     = {Proceedings of the 45th IEEE/ACM International Conference on Software Engineering},
  year          = {2023},
  series        = {ICSE 2023},
  doi           = {10.1109/ICSE48619.2023.00086},
}
 
@inproceedings{meng2023_wemint,
  author        = {Meng, Shi and Wang, Liu and Wang, Shenao and Wang, Kailong and Xiao, Xusheng and Bai, Guangdong and Wang, Haoyu},
  title         = {{WeMinT}: Tainting Sensitive Data Leaks in {WeChat} Mini-Programs},
  booktitle     = {Proceedings of the 38th IEEE/ACM International Conference on Automated Software Engineering},
  year          = {2023},
  series        = {ASE 2023},
  doi           = {10.1109/ASE56229.2023.00151},
}
 
@misc{yang2023_sok,
  author        = {Yang, Yuqing and Wang, Chao and Zhang, Yue and Lin, Zhiqiang},
  title         = {{SoK}: Decoding the Super App Enigma: The Security Mechanisms, Threats, and Trade-offs in {OS}-alike Apps},
  year          = {2023},
  howpublished  = {arXiv:2306.07495},
  url           = {https://arxiv.org/abs/2306.07495},
}
 
@misc{zhang2026_oauth,
  author        = {Zhang, Zidong and Xie, Zhentao and Ying, Lingyun and Hou, Qinsheng and Gu, Yacong and Diao, Wenrui and Wu, Jianliang},
  title         = {Mini-Programs, Mega-Problems: Unveiling {OAuth}-based Authentication Misuses in Mini-Programs via Dynamic Analysis},
  year          = {2026},
  howpublished  = {arXiv:2607.08232, accepted at ACM CCS 2026},
  url           = {https://arxiv.org/abs/2607.08232},
}
 
@misc{ciccotelli2026_tenet,
  author        = {Ciccotelli, Andrea and Zappone, Federico and Di Pietro, Roberto},
  title         = {{TENET}: Telegram Mini App (in)security},
  year          = {2026},
  howpublished  = {arXiv:2608.17538},
  url           = {https://arxiv.org/abs/2608.17538},
}
 
@misc{ferrari2026_telegapper,
  author        = {Ferrari, Luca and Ceccato, Mariano and Verderame, Luca},
  title         = {{TeleGapper}: On the (un)reliability of Privacy Policies in Telegram Mini apps},
  year          = {2026},
  howpublished  = {arXiv:2608.13390},
  url           = {https://arxiv.org/abs/2608.13390},
}

References

[1]
Shi, Yizhe; Yang, Zhemin; Liu, Dingyi; Zhong, Kangwei; Dai, Jiarun; Yang, Min (2026): "Better Safe than Sorry: Uncovering the Insecure Resource Management in App-in-App Cloud Services", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[2]
Yang, Yuqing; Zhang, Yue; Lin, Zhiqiang (2025): "Understanding Miniapp Malware: Identification, Dissection, and Characterization", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[3]
Wang, Mona; Lin, Pellaeon; Knockel, Jeffrey; Greenberg, Will; Mayer, Jonathan; Mittal, Prateek (2025): "What WeChat Knows: Pervasive First-Party Tracking in a Billion-User Super-App Ecosystem", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[4]
Cai, Yifeng; Zhang, Ziqi; Yao, Mengyu; Liu, Junlin; Zhao, Xiaoke; Fu, Xinyi; Li, Ruoyu; Liu, Zhe; Chen, Xiangqun; Guo, Yao; Li, Ding (2025): "I Can Tell Your Secrets: Inferring Privacy Attributes from Mini-app Interaction History in Super-apps", in: Proceedings of the USENIX Security Symposium. (Link)
[5]
Liu, Kaizheng; Yang, Ming; Ling, Zhen; Zhang, Yue; Lei, Chongqing; Luo, Junzhou; Fu, Xinwen (2024): "RIoTFuzzer: Companion App Assisted Remote Fuzzing for Detecting Vulnerabilities in IoT Devices", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[6]
Zhang, Yue; Turkistani, Bayan; Yang, Allen Yuqing; Zuo, Chaoshun; Lin, Zhiqiang (2021): "A Measurement Study of WeChat Mini-Apps", Proceedings of the ACM on Measurement and Analysis of Computing Systems 5(2). (DOI)
[7]
Zhou, Jiawei; Zhang, Zidong; Ying, Lingyun; Chai, Huajun; Cao, Jiuxin; Duan, Haixin (2025): "Hey, Your Secrets Leaked! Detecting and Characterizing Secret Leakage in the Wild", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[8]
Lu, Haoran; Xing, Luyi; Xiao, Yue; Zhang, Yifan; Liao, Xiaojing; Wang, XiaoFeng; Wang, Xueqiang (2020): "Demystifying Resource Management Risks in Emerging Mobile App-in-App Ecosystems", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[9]
Zhang, Lei; Zhang, Zhibo; Liu, Ancong; Cao, Yinzhi; Zhang, Xiaohan; Chen, Yanjun; Zhang, Yuan; Yang, Guangliang; Yang, Min (2022): "Identity Confusion in WebView-based Mobile App-in-app Ecosystems", in: Proceedings of the USENIX Security Symposium. (Link)
[10]
Wei, Zhiao; Wang, Chao; Faheem, Haseeb-Ur-Rehman; Xing, Luyi; Aafer, Yousra; Lin, Zhiqiang (2026): "Raising the Flag: Detecting Missing Permission Controls in Mini-Program APIs", in: Proceedings of the USENIX Security Symposium. (Link)
[11]
Shi, Yizhe; Yang, Zhemin; Yang, Yifan; Yang, Yunteng; Yang, Min (2026): "Convenience at a Cost: the Security Risks of Template-Based Development in the App-in-App Ecosystem", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[12]
Yang, Yuqing; Lin, Zhiqiang (2026): "Real or Rogue? Detecting Malicious Miniapps with Deceptive Reporting Interface", in: Proceedings of the ACM Web Conference. (DOI)
[13]
Chen, Pei; Hong, Geng; Qin, Yicheng; Wang, Huazhe; Wu, Mengying; Yang, Min; Zhao, Ziru; Zhu, Yuanpeng; Su, Tao (2026): "When Fun Turns Toxic: A First Look at Aggressive Advertising in Mini-games", in: Proceedings of the USENIX Security Symposium. (Link)
[14]
He, Yi; Guan, Yunchao; Lun, Ruoyu; Song, Shangru; Guo, Zhihao; Zhuge, Jianwei; Chen, Jianjun; Wei, Qiang; Wu, Zehui; Yu, Miao; Shi, Hetian; Li, Qi (2024): "Demystifying the Security Implications in IoT Device Rental Services", in: Proceedings of the USENIX Security Symposium. (Link)
provenance/design/mobile_and_app_measurement/mini_programs.txt · Last modified: by karel.kubicek.claude