User Tools

Site Tools


design:mobile_and_app_measurement:mini_programs

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Next revision
Previous revision
design:mobile_and_app_measurement:mini_programs [2026/09/27 16:21] – New page: measuring super-app mini-programs (16 papers hand-verdicted from 57 candidates; routes, tools, host as first party). Authored by Claude karel.kubicek.claudedesign:mobile_and_app_measurement:mini_programs [2026/09/27 17:08] (current) – Review round 3 fix: two papers report recall on a labelled benchmark. Authored by Claude karel.kubicek.claude
Line 5: Line 5:
 Four things change enough to break an app-measurement intuition: Four things change enough to break an app-measurement intuition:
  
-  - **There is no public store listing.** No Google Play, no AndroZoo, no Tranco of mini-programs. Your population is whatever your crawler could search, scan or pull out of a client's cache, and every large corpus in these venues descends from one of three crawling routes, the oldest of which the host has since closed. See [[#Getting a Population]].+  - **There is no public store listing.** No Google Play, no AndroZoo, no Tranco of mini-programs. Your population is whatever your crawler could search, scan or pull out of a client's cache, and every large corpus in these venues descends from one of three crawling routes; one 2024 paper reports the oldest blocked for batch appid queries, while two later papers (2025, 2026) still name it or its extension for crawls they do not date, and a third (2026) reuses that extended crawler's method. See [[#Getting a Population]].
   - **There are two permission systems stacked on each other, and a third party watching both.** The OS grants permissions to the host; the host grants scopes to the mini-program; and the host itself sees every page view. The unit of study can be the deployed mini-programs, the host's own API surface, or the host's own telemetry — three designs with three different denominators. See [[#Which Population Is Yours]].   - **There are two permission systems stacked on each other, and a third party watching both.** The OS grants permissions to the host; the host grants scopes to the mini-program; and the host itself sees every page view. The unit of study can be the deployed mini-programs, the host's own API surface, or the host's own telemetry — three designs with three different denominators. See [[#Which Population Is Yours]].
   - **The package is only the front end.** Server logic, cloud-hosted functions and content fetched at run time are not in it, so a static result is a lower bound on what the mini-program does, and malware can hide from it by construction. See [[#Static and Dynamic Analysis]].   - **The package is only the front end.** Server logic, cloud-hosted functions and content fetched at run time are not in it, so a static result is a lower bound on what the mini-program does, and malware can hide from it by construction. See [[#Static and Dynamic Analysis]].
-  - **The ecosystem is Chinese, and so is its documentation.** WeChat's Chinese developer documentation lists 975 APIs where the English one lists 570 {[wang2023_uncovering]}; research accounts registered to non-Chinese phone numbers are locked out of a large share of popular mini-programs {[wang2025_wechat]}; and since the regulator's 2023–2024 filing campaign every mini-program operating in mainland China has had to be filed through its host. See [[#Accounts, Vantage and Language]].+  - **The ecosystem is Chinese, and so is its documentation.** WeChat's Chinese developer documentation listed 975 APIs where the English one listed 570 when {[wang2023_uncovering]} counted them in 2021; research accounts registered to non-Chinese phone numbers are locked out of a large share of popular mini-programs {[wang2025_wechat]}; and since the regulator's 2023–2024 filing campaign every mini-program operating in mainland China has had to be filed through its host. See [[#Accounts, Vantage and Language]].
  
 <WRAP important> <WRAP important>
-**The question a mini-program measurement has to answer that an app measurement mostly does not is "what fraction of the ecosystem could your crawler reach, and how would you know?"** The host publishes no list, and Tencent's quarterly results for 2025–2026 give no count of mini-programs at all, only the combined Weixin and WeChat user base. The ecosystem sizes the papers quote as denominators — "about 4 million", "more than 4.3 million" — are citations of secondary sources, not censuses, and the largest corpus in these venues (4,595,680 WeChat miniapps {[yang2025_miniapp]}) was collected between March 2020 and June 2022. Report what you searched for, how, when, how many packages you got and how many you could analyse, and say plainly that the rest is unknown. Six of the nine papers here that measure deployed mini-programs acknowledge that their sample is not the population; none can say how large the population is.+**The question a mini-program measurement has to answer that an app measurement mostly does not is "what fraction of the ecosystem could your crawler reach, and how would you know?"** The host publishes no list, and Tencent's 2025 annual results and its first- and second-quarter 2026 releases give no count of mini-programs at all, only the combined Weixin and WeChat user base. The ecosystem sizes the papers quote as denominators — "about 4 million", "more than 4.3 million" — are citations of secondary sources, not censuses, and the largest corpus in these venues (4,595,680 WeChat miniapps {[yang2025_miniapp]}) was collected between March 2020 and June 2022. Report what you searched for, how, when, how many packages you got and how many you could analyse, and say plainly that the rest is unknown. Four of the nine papers here that measure deployed mini-programs acknowledge that their sample is not the population; none can say how large the population is.
 </WRAP> </WRAP>
  
Line 20: Line 20:
 Six papers, each here for a methodological reason: Six papers, each here for a methodological reason:
  
-  - {[zhang2021_measurement]} (SIGMETRICS 2021, outside the seven venues) — **the crawler every early corpus descends from.** MiniCrawler and the first ecosystem-scale measurement of WeChat mini-apps. Read it for how the population was built, then read {[zhang2024_minicat]} for why that route no longer works as published.+  - {[zhang2021_measurement]} (SIGMETRICS 2021, outside the seven venues) — **the crawler every early corpus descends from.** MiniCrawler and the first ecosystem-scale measurement of WeChat mini-apps. Read it for how the population was built, then read {[zhang2024_minicat]} for its 2024 report that the route's batch appid queries are blocked.
   - {[zhang2023_leak]} — **acquisition, rate limit and validation done explicitly.** 3,450,586 WeChat mini-programs found through a reverse-engineered search endpoint seeded with 14,020 keywords, six requests per minute for over six months, and every detected secret validated against the host's own key-check API so the count has no false positives by construction.   - {[zhang2023_leak]} — **acquisition, rate limit and validation done explicitly.** 3,450,586 WeChat mini-programs found through a reverse-engineered search endpoint seeded with 14,020 keywords, six requests per minute for over six months, and every detected secret validated against the host's own key-check API so the count has no false positives by construction.
   - {[zhang2024_minicat]} — **the current route, and attrition reported at every stage.** Mini-programs pulled from the Windows client's on-disk cache by GUI automation; 44,273 collected, 41,726 analysable, 14,920 of those timed out. Copy the attrition reporting.   - {[zhang2024_minicat]} — **the current route, and attrition reported at every stage.** Mini-programs pulled from the Windows client's on-disk cache by GUI automation; 44,273 collected, 41,726 analysable, 14,920 of those timed out. Copy the attrition reporting.
Line 68: Line 68:
 | **Rankings and manual search** | marketing-firm top-100 lists and hand searches for keywords, installed and exercised one by one | 1: 2025 {[wang2025_wechat]} | Always works; a popularity sample of about a hundred. | | **Rankings and manual search** | marketing-firm top-100 lists and hand searches for keywords, installed and exercised one by one | 1: 2025 {[wang2025_wechat]} | Always works; a popularity sample of about a hundred. |
 | **Operator logs** | the host's own records, with its co-authors and its IRB | 1: 2025 {[cai2025_tell]} | Requires the host as a partner; not reproducible by anyone else. | | **Operator logs** | the host's own records, with its co-authors and its IRB | 1: 2025 {[cai2025_tell]} | Requires the host as a partner; not reproducible by anyone else. |
-| **Gated research datasets** | the OSU group behind MiniCrawler serves its CCS 2022, CCS 2023 and NDSS 2025 corpora and random SIGMETRICS 2021 samples on request, with institutional authentication and written consent | reused by later papers | Live (minimalware.github.io, fetched 2026-09-27); the AppSecret-leak set "//requires additional consent and agreement//". |+| **Gated research datasets** | the OSU group behind MiniCrawler serves datasets from its CCS 2022, CCS 2023 and NDSS 2025 papers and random samples of its SIGMETRICS 2021 crawl on request, with institutional authentication and written consent | none in the corpus (the group's own later papers re-crawl) | Live (minimalware.github.io, fetched 2026-09-27); the AppSecret-leak set "//requires additional consent and agreement//". The samples are from a 2020–2022 crawl. |
  
-The other routes in the corpus are specific to a vertical — QR codes on rentable scooters and chargers, where 75 of 81 Chinese products were reachable only as a WeChat mini-program {[he2024_demystifying]}; the control mini-app shipped with a purchased IoT device {[liu2024_riotfuzzer]}; a platform audit department's complaint cases as ground truth {[chen2026_minigames]} — or reuse an earlier paper's method: {[shi2026_better]} "//adhered to the methods established in previous studies//", citing {[zhang2023_leak]} and {[shi2025_skeleton]}.+The other routes in the corpus are specific to a vertical — QR codes on rentable scooters and chargers, where for 75 of 81 Chinese products the authors found a WeChat mini-program and no native app {[he2024_demystifying]}; the control mini-app shipped with a purchased IoT device {[liu2024_riotfuzzer]}; a platform audit department's complaint cases as ground truth {[chen2026_minigames]} — or reuse an earlier paper's method: {[shi2026_better]} "//adhered to the methods established in previous studies//", citing {[zhang2023_leak]} and {[shi2025_skeleton]}.
  
-One more frame exists outside the research literature. Since the regulator's 2023 filing notice, mini-programs operating in mainland China must be filed through their distribution platform, and WeChat's own FAQ says an unfiled mini-program will become inaccessible or be delisted; the filing system is searchable by the business category "小程序"((MIIT notice 工信部信管〔2023〕105号, dated 2023-07-21, published 2023-08-04, https://www.miit.gov.cn/zwgk/zcwj/wjfb/tz/art/2023/art_920db564162e4312916a01bed6540ad8.html; WeChat filing FAQ, https://developers.weixin.qq.com/miniprogram/product/record/record_faq.html; both fetched 2026-09-27.)). **No paper in these venues uses it**, and we did not establish whether it can be queried at scale or whether a record carries the appid. It is the only candidate for a population frame that is not "what our search returned".+One more frame exists outside the research literature. Since the regulator's 2023 filing notice, mini-programs operating in mainland China must be filed through their distribution platform, and WeChat's own FAQ says an unfiled mini-program will become inaccessible or be delisted; the filing system can be searched by filing number or mini-program name, with "小程序" as a category filter((MIIT notice 工信部信管〔2023〕105号, dated 2023-07-21, published 2023-08-04, https://www.miit.gov.cn/zwgk/zcwj/wjfb/tz/art/2023/art_920db564162e4312916a01bed6540ad8.html; WeChat filing FAQ, https://developers.weixin.qq.com/miniprogram/product/record/record_faq.html; both fetched 2026-09-27.)). **No paper in these venues uses it**, and we did not establish whether it can be enumerated rather than looked up, queried at scale, or whether a record carries the appid. If it can be enumerated, it is the only candidate for a population frame outside the host.
  
 ==== What counts as one mini-program ==== ==== What counts as one mini-program ====
Line 87: Line 87:
 | {[shi2026_better]} | 1,248,815 (four hosts) | 22,695 that use cloud services | out of the question's scope | | {[shi2026_better]} | 1,248,815 (four hosts) | 22,695 that use cloud services | out of the question's scope |
  
-Reported that way, a reader can see that MiniCAT's headline 32.0% is over the 41,726 analysable, not the 44,273 collected (30.2% of those — our arithmetic), and that it is silent about the 14,920 that timed out.+Reported that way, a reader can see that MiniCAT's headline 32.0% is over the 41,726 analysable, not the 44,273 collected (30.2% of those), and that the 41,726 include 14,920 on which the detector's five-minute CodeQL timeout fired — so the 13,349 flagged are 49.8% of the 26,806 on which it completed (both our arithmetic; the paper does not say whether a flag can come from a query that timed out). Storage is part of the budget too: 2,571,490 WeChat packages took "//6.29 TB disk storage//" {[yang2022_cross]} (about 2.45 MB each), and MiniCAT's 44,273 unpacked mini-programs 126.38 GB (about 2.85 MB each; our arithmetic).
  
 ===== Static and Dynamic Analysis ===== ===== Static and Dynamic Analysis =====
Line 97: Line 97:
 ^ Tool ^ What it is ^ Used by ^ State as of 2026-09-27 ^ ^ Tool ^ What it is ^ Used by ^ State as of 2026-09-27 ^
 | **wxappUnpacker** (qwerty472123) | the unpacker the corpus names | {[zhang2024_minicat]} | **Archived; last commit 2020-04-18.** MiniCAT reports packages it could not unpack "//due to their use of a newer version of the WeChat mini-program base library//". Superseded. | | **wxappUnpacker** (qwerty472123) | the unpacker the corpus names | {[zhang2024_minicat]} | **Archived; last commit 2020-04-18.** MiniCAT reports packages it could not unpack "//due to their use of a newer version of the WeChat mini-program base library//". Superseded. |
-| **wux1an/wxapkg**, **biggerstar/wedecode** | maintained unpackers | none in the corpus | v2.0.0 released 2026-04-16; tag v0.10.6, last commit 2026-08-27. Neither is named in any corpus paper — validate on your own packages before trusting its output. |+| **wux1an/wxapkg**, **biggerstar/wedecode** | maintained unpackers | none in the corpus | wxapkg v2.0.0, released 2026-04-16; wedecode tag v0.10.6, last commit 2026-08-27. Neither is named in any corpus paper — validate on your own packages before trusting its output. | 
 +| **Ackites/KillWxapkg** | the most-starred unpacker on GitHub (decrypts, unpacks, restores the project tree) | none in the corpus | v2.4.1, 2024-09-20, no commit since. Stalling. |
 | **DoubleX**, **JAW**, **CodeQL**, **Esprima + WALA** | general JS analysers with a mini-program model added | {[yang2022_cross]}, {[shi2025_skeleton]}, {[zhang2024_minicat]}, {[chen2026_minigames]} | General-purpose and maintained upstream; the mini-program model is each paper's own. | | **DoubleX**, **JAW**, **CodeQL**, **Esprima + WALA** | general JS analysers with a mini-program model added | {[yang2022_cross]}, {[shi2025_skeleton]}, {[zhang2024_minicat]}, {[chen2026_minigames]} | General-purpose and maintained upstream; the mini-program model is each paper's own. |
-| **TaintMini**, **WeMinT** | mini-program taint analysers | none in the corpus (ICSE 2023 {[wang2023_taintmini]}; ASE 2023 {[meng2023_wemint]}) | Code released; TaintMini's README declines to ship an unpacker. |+| **TaintMini**, **WeMinT** | mini-program taint analysers | none in the corpus (ICSE 2023 {[wang2023_taintmini]}; ASE 2023 {[meng2023_wemint]}) | Code released; TaintMini's README declines to ship an unpacker, WeMinT's bundles one. |
  
-Two things a static result cannot see. **The package is the front end only**: {[yang2025_miniapp]} warns that "//the malware may dynamically hide malicious contents without distributing them to the front-end by the time we tested the cases//", and MiniCAT could not analyse mini-programs whose logic lived in WeChat Cloud Development. And **packages are obfuscated**: {[liu2024_riotfuzzer]} chose not to analyse the mini-app at all, because "//most mini-apps employ obfuscation techniques//", and instrumented the host's Java bridge instead.+Two things a static result cannot see. **The package is the front end only**: {[yang2025_miniapp]} warns that "//the malware may dynamically hide malicious contents without distributing them to the front-end by the time we tested the cases//", and MiniCAT could not analyse mini-programs whose logic lived in WeChat Cloud Development. And **packages are obfuscated**: {[zhang2021_measurement]} measured obfuscation rates across its crawl, and {[liu2024_riotfuzzer]}, citing it — "as //reported in [45], most mini-apps employ obfuscation techniques//" — chose not to analyse the mini-app at all and instrumented the host's Java bridge instead.
  
-The research groups that publish these tools **do not ship unpackers**: TaintMini's and MiniCAT's READMEs both say so "//due to potential legal implications//", and APIDiff's says the same of its crawlers. The reason is in the host's user licence rather than its developer terms: WeChat's Chinese licence forbids reverse-engineering the client (clause 8.2.1.2) and copying, modifying or hooking the data it holds in memory or exchanges with the server, "//使用插件、外挂或非经腾讯授权的第三方工具//" — using plug-ins or third-party tools Tencent has not authorised (clause 8.2.1.4); the international terms (last modified 2025-11-18) forbid reverse-engineering "WeChat Software"((腾讯微信软件许可及服务协议, https://weixin.qq.com/agreement?lang=zh_CN; WeChat Terms of Service, https://www.wechat.com/en/service_terms.html; both fetched 2026-09-27.)). Both clauses are about the client, not third-party packages, and the mini-program developer terms contain no reverse-engineering clause; but Xposed or Frida hooking of the client, which every dynamic study here does, is exactly what 8.2.1.4 names. None of the 16 papers discusses it.+Most research groups that publish these tools **do not ship unpackers**: TaintMini's and MiniCAT's READMEs both say so "//due to potential legal implications//", and APIDiff's says the same of its crawlers. WeMinT is the exception: its repository bundles a copy of wxappUnpacker without comment. The reason is in the host's user licence rather than its developer terms: WeChat's Chinese licence forbids reverse-engineering the client (clause 8.2.1.2) and copying, modifying or hooking the data it holds in memory or exchanges with the server, "//使用插件、外挂或非经腾讯授权的第三方工具//" — using plug-ins or third-party tools Tencent has not authorised (clause 8.2.1.4); the international terms (last modified 2025-11-18) forbid reverse-engineering "WeChat Software"((腾讯微信软件许可及服务协议, https://weixin.qq.com/agreement?lang=zh_CN; WeChat Terms of Service, https://www.wechat.com/en/service_terms.html; both fetched 2026-09-27.)). Both clauses are about the client, not third-party packages, and the mini-program developer terms contain no reverse-engineering clause; but on our reading, hooking the client with Frida or Xposed, which most dynamic studies here do (Frida in 4 papers, Xposed in 2), is what 8.2.1.4 names. The licence is governed by mainland-Chinese law (clause 12.3), which is the practical fact for a researcher elsewhere. None of the 16 papers quotes or analyses the licence.
  
 ==== Dynamic: instrument the host ==== ==== Dynamic: instrument the host ====
Line 109: Line 110:
 Dynamic analysis of a mini-program is dynamic analysis of the **host app** with the mini-program running inside it, so everything on [[Design:Mobile and app measurement]] about rooted handsets, Frida and pinning applies to WeChat, not to the mini-program. What is specific here: Dynamic analysis of a mini-program is dynamic analysis of the **host app** with the mini-program running inside it, so everything on [[Design:Mobile and app measurement]] about rooted handsets, Frida and pinning applies to WeChat, not to the mini-program. What is specific here:
  
-  * **Hook the host's JavaScript bridge.** Frida is named by 4 of the 16 papers (2023–2025). Xposed — which the parent page dates as superseded for general app hooking — is named by 2 (2020, 2022), is what MiniCrawler requires, and is what {[wei2026_raising]} used in 2026 to hook WeChat's, Alipay's, Baidu's and QQ's bridge classes. +  * **Hook the host's JavaScript bridge.** Frida is named by 4 of the 16 papers (2023–2025). Xposed — which the parent page rates dead for general app hooking, with its successor LSPosed stalled — is named by 2 (2020, 2022), is what MiniCrawler requires, and is what {[wei2026_raising]} used in 2026 to hook WeChat's, Alipay's, Baidu's and QQ's bridge classes. 
-  * **WeChat resists emulators.** A 2025 study of in-app browsers found "//WeChat failed to launch even on Android 11 AVDs//" and ran it on BlueStacks instead {[lee2025_deep]}; {[wang2025_wechat]} used both a rooted Pixel 6 and an emulator on Android 14. Budget for a physical device. +  * **The stock emulator images do not run WeChat; others did.** A 2025 study of in-app browsers found "//WeChat failed to launch even on Android 11 AVDs//" and ran it on BlueStacks, "//a configuration publicly known to support the app reliably//" {[lee2025_deep]}; {[wang2025_wechat]} measured on an Android 14 emulator alongside a rooted Pixel 6. Say which emulator you used; a physical device is the safe budget. 
-  * **The developer tools are not the device.** WeChat's own documentation for ''wx.getUserProfile'' notes that base libraries 2.10.4–2.16.1 return real user data in DevTools but anonymous data on a real phone((WeChat API documentation, wx.getUserProfile, https://developers.weixin.qq.com/miniprogram/dev/api/open-api/user-info/wx.getUserProfile.html, fetched 2026-09-27.)). A measurement run in DevTools needs its device-vs-DevTools difference checked.+  * **The developer tools are not the device.** WeChat's own documentation for ''wx.getUserProfile'' notes that base libraries 2.10.4–2.16.1 return real user data in DevTools but anonymous data on a real phone((WeChat API documentation, wx.getUserProfile, https://developers.weixin.qq.com/miniprogram/dev/api/open-api/user-info/wx.getUserProfile.html, fetched 2026-09-27.)). That is a note about 2021-era base libraries, but the point stands, and a larger one with it: DevTools debugs **your own** project — its remote debugging packs and uploads the local code, and the release-build debug switch belongs to the mini-program's developer((WeChat DevTools documentation, https://developers.weixin.qq.com/miniprogram/dev/devtools/remote-debug.html, fetched 2026-09-27.)) — so there is no documented way to attach it to a third-party mini-program. Running unpacked third-party source as a local project under your own AppID is what the tool READMEs imply, not what WeChat documents.
   * **Two encryption layers.** Mini-program requests to a developer's server are HTTPS, so a trusted root on the device lets a proxy read them ({[he2024_demystifying]} captured and replayed rental mini-programs' requests with BurpSuite). What **WeChat itself** sends about mini-program use travels over MMTLS, WeChat's own protocol, which {[wang2025_wechat]} had to reverse-engineer with Frida, Jadx, Ghidra and IDA Pro before those flows could be read.   * **Two encryption layers.** Mini-program requests to a developer's server are HTTPS, so a trusted root on the device lets a proxy read them ({[he2024_demystifying]} captured and replayed rental mini-programs' requests with BurpSuite). What **WeChat itself** sends about mini-program use travels over MMTLS, WeChat's own protocol, which {[wang2025_wechat]} had to reverse-engineer with Frida, Jadx, Ghidra and IDA Pro before those flows could be read.
   * **Driving the UI** is mostly manual: {[wang2025_wechat]} "//decided against further automation//" to avoid account bans; {[shi2025_skeleton]} drives WebView-based mini-apps with Android UI Automator; automated exploration tools such as MiniScope exist in the software-engineering literature, outside these venues.   * **Driving the UI** is mostly manual: {[wang2025_wechat]} "//decided against further automation//" to avoid account bans; {[shi2025_skeleton]} drives WebView-based mini-apps with Android UI Automator; automated exploration tools such as MiniScope exist in the software-engineering literature, outside these venues.
Line 139: Line 140:
 | Aggressive advertising in mini-games {[chen2026_minigames]} | "//49.95% of ad-enabled mini-games//" (457 of 915) | — | | Aggressive advertising in mini-games {[chen2026_minigames]} | "//49.95% of ad-enabled mini-games//" (457 of 915) | — |
  
-Precision is measured almost everywhere and recall almost nowhere. The papers sample flagged and unflagged cases by hand — 100 and 100 {[yang2022_cross]}, {[zhang2024_minicat]}; up to 500 and 500 per host {[shi2025_skeleton]}; 500 flagged {[yang2025_miniapp]}; 300 {[zhou2025_secrets]} — or, better, validate against the host itself: {[zhang2023_leak]} checks every candidate key with the host's key-validation API, which is why it can claim no false positives. Report both directions of the sample, as the first two do, and say that a sample of 100 unflagged mini-programs bounds recall only loosely.+Precision is measured everywhere; recall less often and more weakly. Three deployed-population papers estimate it from a hand sample of unflagged cases — 100 {[yang2022_cross]} ("//making the FN rate to 2%//"), 100 {[zhang2024_minicat]}, up to 500 per host {[shi2025_skeleton]} (recall 85.56%) — and two against a labelled ground-truth set ({[chen2026_minigames]}, recall 83.55% on 371 audited behaviours; {[zhou2025_secrets]}, 83.38% on the WeChat part of its benchmark, while its 300-detection field check is precision-only); {[yang2025_miniapp]} samples 500 flagged cases only. {[zhang2023_leak]} validates every candidate key against the host, which is why it can claim no false positives — but the oracle is the host's token endpoint, so a "valid" result is a live access token for someone else's backend; the paper states no ethics review, and {[shi2025_skeleton]} ran the same kind of check under an IRB "minimal risk" finding. Say which you did. Report both directions of the sample, and say that 100 unflagged mini-programs bound recall only loosely.
  
 ===== Accounts, Vantage and Language ===== ===== Accounts, Vantage and Language =====
Line 147: Line 148:
 **Vantage.** The extraction records no measurement location for 15 of the 16 papers (and has no vantage tuple for the sixteenth), and our reading agrees: beyond the phone-number country {[wang2025_wechat]} reports, none of the 16 says where its crawler or test devices ran, although the host, the filing regime and the content are all mainland-Chinese. Choosing and reporting the vantage point is [[Design:Crawling location]]; what a non-Chinese vantage may not be shown is [[Design:Blocking and geodifference]]. **Vantage.** The extraction records no measurement location for 15 of the 16 papers (and has no vantage tuple for the sixteenth), and our reading agrees: beyond the phone-number country {[wang2025_wechat]} reports, none of the 16 says where its crawler or test devices ran, although the host, the filing regime and the content are all mainland-Chinese. Choosing and reporting the vantage point is [[Design:Crawling location]]; what a non-Chinese vantage may not be shown is [[Design:Blocking and geodifference]].
  
-**Language.** Seed keywords, mini-program names, UI text and the host's documentation are Chinese. Only 2 of the 9 deployed-population papers describe how they handled it — Chinese seed characters {[zhang2023_leak]} and ''jieba'' word segmentation {[zhang2024_minicat]} — and the English documentation gap above (570 of 975 APIs) means an English-only reading of the host's API surface misses two in five. Tooling for detecting and translating page language is on [[Programming:Multilingual support]].+**Language.** Seed keywords, mini-program names, UI text and the host's documentation are Chinese. Only 2 of the 9 deployed-population papers describe how they handled it — Chinese seed characters {[zhang2023_leak]} and ''jieba'' word segmentation {[zhang2024_minicat]} — and the English documentation gap above (570 of 975 APIs in 2021) meant an English-only reading of the host's API surface missed two in five; re-count before relying on the English documentation. Tooling for detecting and translating page language is on [[Programming:Multilingual support]].
  
 ===== Beyond WeChat ===== ===== Beyond WeChat =====
Line 153: Line 154:
 14 of the 16 papers study WeChat; Baidu appears in 7, Douyin/TikTok in 5, Alipay in 4, QQ in 2. Seven include a host outside those five — LINE and VK ({[shi2025_skeleton]}: 3,984 and 425 mini-apps; {[shi2026_better]}), Facebook Instant Games and QuickGame ({[chen2026_minigames]}), IoT all-in-one apps from Xiaomi, JD, Huawei and Tuya ({[liu2024_riotfuzzer]}), WeCom ({[wang2023_uncovering]}), and the non-Chinese super apps and browsers among the hosts of {[zhang2022_identity]} and {[lu2020_demystifying]}. Non-Chinese hosts are thinly measured: 14 of the 16 papers study WeChat; Baidu appears in 7, Douyin/TikTok in 5, Alipay in 4, QQ in 2. Seven include a host outside those five — LINE and VK ({[shi2025_skeleton]}: 3,984 and 425 mini-apps; {[shi2026_better]}), Facebook Instant Games and QuickGame ({[chen2026_minigames]}), IoT all-in-one apps from Xiaomi, JD, Huawei and Tuya ({[liu2024_riotfuzzer]}), WeCom ({[wang2023_uncovering]}), and the non-Chinese super apps and browsers among the hosts of {[zhang2022_identity]} and {[lu2020_demystifying]}. Non-Chinese hosts are thinly measured:
  
-  * **Telegram Mini Apps** are not packages at all: Telegram's documentation describes them as JavaScript interfaces "//launched right inside Telegram//" that "//can completely replace any website//" — a developer-hosted web page in a WebView, so the methods of [[Programming:Crawler]] apply and there is nothing to unpack. Every Mini App receives the user's id, name and, since Bot API 8.0 (2024-11-17), profile photo; the documentation warns that the client-side copy of that data "//should not be trusted//" and must be validated server-side by signature((Telegram Bot API, Mini Apps, https://core.telegram.org/bots/webapps, fetched 2026-09-27.)) — the same client-trust failure class as WeChat's AppSecret leaks. **No paper in these venues measures them.**+  * **Telegram Mini Apps** are not packages at all: Telegram's documentation describes them as JavaScript interfaces "//launched right inside Telegram//" that "//can completely replace any website//" — a developer-hosted web page in a WebView, so the methods of [[Programming:Crawler]] apply and there is nothing to unpack. Every Mini App receives the user's id and first name, and — since Bot API 8.0 (2024-11-17) — the profile photo URL if the user's privacy settings allow it; the documentation warns that the client-side copy of that data "//should not be trusted//" and must be validated server-side by signature((Telegram Bot API, Mini Apps, https://core.telegram.org/bots/webapps, fetched 2026-09-27.)) — a server that trusts the client copy is the Telegram analogue of the client-side trust failures the WeChat literature measured. **No paper in these venues measures them**; two arXiv preprints from August 2026, not yet in any venue, have started: plaintext session tokens and wallet phrases in Mini App client storage {[ciccotelli2026_tenet]}, and Mini App behaviour against their privacy policies {[ferrari2026_telegapper]}.
   * **TikTok mini games** are package-based like WeChat's — TikTok's documentation says the app "//selects an available version and obtains its code package//" and executes JavaScript or WebAssembly — and are offered in a list of markets that includes the United States, Japan, Brazil and Southeast Asia but no EU country((TikTok for Developers, https://developers.tiktok.com/docs/en/mini-games-overview and https://developers.tiktok.com/docs/en/mini-games-technical-overview, fetched 2026-09-27.)). The corpus has no measurement of them; {[chen2026_minigames]} covers WeChat, Facebook Instant Games and QuickGame.   * **TikTok mini games** are package-based like WeChat's — TikTok's documentation says the app "//selects an available version and obtains its code package//" and executes JavaScript or WebAssembly — and are offered in a list of markets that includes the United States, Japan, Brazil and Southeast Asia but no EU country((TikTok for Developers, https://developers.tiktok.com/docs/en/mini-games-overview and https://developers.tiktok.com/docs/en/mini-games-technical-overview, fetched 2026-09-27.)). The corpus has no measurement of them; {[chen2026_minigames]} covers WeChat, Facebook Instant Games and QuickGame.
   * Other hosts with mini-program platforms — Snapchat, LINE, Grab and others — are named in the papers' introductions and the SaTS workshop's call; beyond the two LINE/VK corpora above, the corpus has no measurement of them.   * Other hosts with mini-program platforms — Snapchat, LINE, Grab and others — are named in the papers' introductions and the SaTS workshop's call; beyond the two LINE/VK corpora above, the corpus has no measurement of them.
Line 163: Line 164:
   * **Attack only your own accounts, devices and test mini-programs**; never publish a malicious mini-program to the store ({[wang2023_uncovering]}: "//We have never uploaded our malicious miniapps onto the markets to harm other users//").   * **Attack only your own accounts, devices and test mini-programs**; never publish a malicious mini-program to the store ({[wang2023_uncovering]}: "//We have never uploaded our malicious miniapps onto the markets to harm other users//").
   * **Rate-limit the host**: six requests per minute {[zhang2023_leak]}, twenty {[zhang2024_minicat]}, "//a few seconds per miniapp//" {[yang2025_miniapp]}.   * **Rate-limit the host**: six requests per minute {[zhang2023_leak]}, twenty {[zhang2024_minicat]}, "//a few seconds per miniapp//" {[yang2025_miniapp]}.
-  * **Probe without reading**: {[shi2026_better]} inferred which cloud resources were exposed "//without accessing the cloud data//".+  * **Probe without reading**: {[shi2026_better]} inferred which cloud resources were exposed "//without accessing the cloud data//". Validating a leaked secret against the host is the opposite — it uses the secret — and needs saying (see [[#The Vulnerability Side]]).
   * **Gate what you release**: malware and leaked-secret corpora are served on request, with identity verification, not posted.   * **Gate what you release**: malware and leaked-secret corpora are served on request, with identity verification, not posted.
   * **Tell the developers, not only the host**: {[zhang2024_minicat]} found contacts for "//248/316 (78.5%)//" of its confirmed cases and emailed them; {[shi2026_better]} emailed 1,869 and reports that 893 fixed the issue or took the mini-program down.   * **Tell the developers, not only the host**: {[zhang2024_minicat]} found contacts for "//248/316 (78.5%)//" of its confirmed cases and emailed them; {[shi2026_better]} emailed 1,869 and reports that 893 fixed the issue or took the mini-program down.
  
-What the papers do not discuss is the host's terms. Crawling the host's search endpoint, downloading third-party packages at scale and hooking the client are what every large corpus did, and WeChat's user licence forbids hooking the client with unauthorised tools (see [[#Static and Dynamic Analysis]]); none of the 16 papers mentions it. See [[Practices:Ethics]] and [[Practices:Notifying websites]].+What the papers do not analyse is the host's terms. Searching the host at scale, downloading or caching third-party packages, and (in most dynamic studies) hooking the client are how these corpora were built, and WeChat's user licence forbids hooking the client with unauthorised tools (see [[#Static and Dynamic Analysis]]). None of the 16 quotes or analyses the licence; {[chen2026_minigames]} asserts that its public-source crawl ensured "//compliance with platform policies//" without saying which, and two others assert compliance with laws or a vendor's bug-bounty plan. See [[Practices:Ethics]] and [[Practices:Notifying websites]].
  
 ===== Use in Publications ===== ===== Use in Publications =====
Line 177: Line 178:
 A paper counts if **mini-programs, or the host's mini-program framework and its APIs, are measured or analysed** — core if they are the main object, section if they are one analysed population among several. WeChat or Alipay as a messenger, payment method, attack target, dataset owner or recruitment channel does not count, nor do mentions in related work. A paper counts if **mini-programs, or the host's mini-program framework and its APIs, are measured or analysed** — core if they are the main object, section if they are one analysed population among several. WeChat or Alipay as a messenger, payment method, attack target, dataset owner or recruitment channel does not count, nor do mentions in related work.
  
-The candidate set is the union of five probes over whitespace-collapsed full text and the extraction schema: the 2026-09-22 gap pass's rule (''mini-?programs?|WeChat'', case-insensitive, at least ten hits); the ecosystem vocabulary (mini-program, mini-app, mini-game, super app, app-in-app) at least three times; the WeChat package and markup names (''wxapkg'', WXML, WXSS, ''wx.*'' APIs) once; Telegram/Snap/LINE mini-app product names once; or the vocabulary or a host name in a paper's title, population sources, measured phenomena, tools or classifier resources — **57** papers. The 16 that looked like mini-program studies were read in full; the other 41 were decided from the sentence around every hit. A recall probe outside the set (applet, light app, quick app, H5, host app, sub-app, JSBridge, with a host name) surfaced one more paper, which was read and is not about mini-programs.+The candidate set is the union of five probes over whitespace-collapsed full text and the extraction schema: the 2026-09-22 gap pass's rule (''mini-?programs?|WeChat'', case-insensitive, at least ten hits); the ecosystem vocabulary (mini-program, mini-app, mini-game, super app, app-in-app) at least three times; the WeChat package and markup names (''wxapkg'', WXML, WXSS, ''wx.*'' APIs) once; Telegram/Snap/LINE mini-app product names once; or the vocabulary or a host name in a paper's title, population sources, measured phenomena, tools or classifier resources — **57** papers. The 16 that looked like mini-program studies were read in full; the other 41 were decided from the sentence around every hit. A recall probe outside the set (applet, light app, quick app, H5, host app, sub-app, JSBridge, with a host name) surfaced six papers: five use "host app" for an app embedding an SDK and were dismissed from context; one was read in full and is not about mini-programs.
  
 ^ Verdict ^ Papers ^ ^ Verdict ^ Papers ^
Line 199: Line 200:
 | this page's candidate set | 57 | 16 | 28.1% | 100% | | this page's candidate set | 57 | 16 | 28.1% | 100% |
  
-The gap rule misses one paper, {[liu2024_riotfuzzer]}, whose IoT host apps are not WeChat and which says "mini-app". The ecosystem vocabulary alone finds all 16 at twice the gap rule's precision; the recall of the candidate set is 100% **by construction** against the population it produced.+The gap rule misses one paper, {[liu2024_riotfuzzer]}, whose IoT host apps are not WeChat and which says "mini-app". The ecosystem vocabulary alone finds all 16, at 53.3% precision against the gap rule's 40.5%; the recall of the candidate set is 100% **by construction** against the population it produced.
  
 <WRAP tip> <WRAP tip>
Line 210: Line 211:
 | 2020 | CCS | {[lu2020_demystifying]} | host framework | 11 hosts incl. WeChat, QQ, Baidu, Douyin | 927 sub-app APIs | own test sub-apps, Xposed | | 2020 | CCS | {[lu2020_demystifying]} | host framework | 11 hosts incl. WeChat, QQ, Baidu, Douyin | 927 sub-app APIs | own test sub-apps, Xposed |
 | 2022 | CCS | {[yang2022_cross]} | deployed | WeChat, Baidu | 2,571,490 + 148,512 | MiniCrawler | | 2022 | CCS | {[yang2022_cross]} | deployed | WeChat, Baidu | 2,571,490 + 148,512 | MiniCrawler |
-| 2022 | USENIX Sec | {[zhang2022_identity]} | host framework | 47 super apps | 6,000 Android apps → 47 | app-store crawl, Xposed |+| 2022 | USENIX Sec | {[zhang2022_identity]} | host framework | 47 super apps incl. WeChat, Alipay, Baidu, TikTok | 6,000 Android apps → 47 | app-store crawl, Xposed |
 | 2023 | CCS | {[zhang2023_leak]} | deployed | WeChat, Baidu | 3,450,586 + 171,989 | reverse-engineered search API | | 2023 | CCS | {[zhang2023_leak]} | deployed | WeChat, Baidu | 3,450,586 + 171,989 | reverse-engineered search API |
 | 2023 | CCS | {[wang2023_uncovering]} | host framework | WeChat, WeCom, QQ, Baidu, TikTok | 1,829 API candidates; 267,359 miniapps | documentation, Soot, Frida, MiniCrawler | | 2023 | CCS | {[wang2023_uncovering]} | host framework | WeChat, WeCom, QQ, Baidu, TikTok | 1,829 API candidates; 267,359 miniapps | documentation, Soot, Frida, MiniCrawler |
-| 2023 | USENIX Sec | {[wang2023_size]} | host framework | WeChat (Windows, Android, iOS) | documented APIs | documentation, own test miniapps, Frida |+| 2023 | USENIX Sec | {[wang2023_size]} | host framework | WeChat (Windows, Android, iOS) | 1,031 documented APIs | documentation, own test miniapps, Frida |
 | 2024 | CCS | {[zhang2024_minicat]} | deployed | WeChat | 44,273 → 41,726 | desktop client cache | | 2024 | CCS | {[zhang2024_minicat]} | deployed | WeChat | 44,273 → 41,726 | desktop client cache |
 | 2024 | CCS | {[liu2024_riotfuzzer]} //(section)// | host framework | Xiaomi, JD, Huawei, Tuya | 27 devices | purchased devices, Frida | | 2024 | CCS | {[liu2024_riotfuzzer]} //(section)// | host framework | Xiaomi, JD, Huawei, Tuya | 27 devices | purchased devices, Frida |
Line 219: Line 220:
 | 2025 | IEEE S&P | {[zhou2025_secrets]} //(section)// | deployed | WeChat | 41,719 | built on the MiniCAT crawler | | 2025 | IEEE S&P | {[zhou2025_secrets]} //(section)// | deployed | WeChat | 41,719 | built on the MiniCAT crawler |
 | 2025 | NDSS | {[yang2025_miniapp]} | deployed | WeChat | 4,595,680 | MiniCrawler, store revisits | | 2025 | NDSS | {[yang2025_miniapp]} | deployed | WeChat | 4,595,680 | MiniCrawler, store revisits |
-| 2025 | NDSS | {[shi2025_skeleton]} | deployed | six hosts | 413,775 → 402,527 | extended MiniCrawler |+| 2025 | NDSS | {[shi2025_skeleton]} | deployed | WeChat, Baidu, Alipay, TikTok, LINE, VK | 413,775 → 402,527 | extended MiniCrawler |
 | 2025 | PETS | {[wang2025_wechat]} | host telemetry | WeChat | 170 → 104 | rankings, manual search | | 2025 | PETS | {[wang2025_wechat]} | host telemetry | WeChat | 170 → 104 | rankings, manual search |
 | 2025 | USENIX Sec | {[cai2025_tell]} | operator logs | Alipay | 288,895 users → 219,826 | the operator | | 2025 | USENIX Sec | {[cai2025_tell]} | operator logs | Alipay | 288,895 users → 219,826 | the operator |
Line 240: Line 241:
  
 ^ Practice ^ Status ^ Evidence ^ ^ Practice ^ Status ^ Evidence ^
-| Crawl WeChat with **MiniCrawler** | **Historical as published.** | Named in 2022, 2023, 2025, 2025, 2026; code last changed 2021-06-29 and pinned to WeChat 7.0.19/7.0.20 with Xposed; batch appid queries reported blocked in 2024. A 2026 paper still names it with no crawl date |+| Crawl WeChat with **MiniCrawler** | **Historical as published.** | Named in 2022, 2023, 2025, 2025, 2026; code last changed 2021-06-29 and pinned to WeChat 7.0.19/7.0.20 with Xposed; batch appid queries reported blocked in one 2024 paper; two later papers (2025, 2026) name it or its extension for undated crawls, a third (2026) reuses the method |
 | Crawl through the **desktop client cache** with GUI automation and Chinese keyword segmentation | **Current: the most recent route documented in these venues.** | 2024, and 2025 building on it; limited by the client's own compatibility | | Crawl through the **desktop client cache** with GUI automation and Chinese keyword segmentation | **Current: the most recent route documented in these venues.** | 2024, and 2025 building on it; limited by the client's own compatibility |
-| Use a **gated research corpus** instead of crawling | **Current, and the cheapest start.** | OSU MiniSec datasets live on 2026-09-27; institutional authentication required |+| Use a **gated research corpus** instead of crawling | **Available, untried in these venues.** | OSU MiniSec datasets live on 2026-09-27; institutional authentication required; samples from a 2020–2022 crawl, 7.8% of which was delisted by the end of 2022 |
 | Query the **regulator's filing system** as a population frame | **Untried.** | Mandatory filing since 2023–2024; no paper uses it; bulk access not established | | Query the **regulator's filing system** as a population frame | **Untried.** | Mandatory filing since 2023–2024; no paper uses it; bulk access not established |
 | Unpack with **wxappUnpacker** | **Superseded.** | Archived; last commit 2020-04-18; fails on newer base libraries (2024) | | Unpack with **wxappUnpacker** | **Superseded.** | Archived; last commit 2020-04-18; fails on newer base libraries (2024) |
-| Unpack with **wux1an/wxapkg** or **wedecode** | **Current, unvalidated in the literature.** | Releases/commits in 2026; no corpus paper names either |+| Unpack with **wux1an/wxapkg** or **wedecode** | **Current, unvalidated in the literature.** | Releases/commits in 2026; no corpus paper names either. KillWxapkg, the most-starred, has not changed since 2024-09-20 |
 | General JS analysers (**CodeQL**, **JAW**, **DoubleX**, **WALA**) with a mini-program model | **Current.** | 2022–2026; each paper builds its own model of the host APIs | | General JS analysers (**CodeQL**, **JAW**, **DoubleX**, **WALA**) with a mini-program model | **Current.** | 2022–2026; each paper builds its own model of the host APIs |
 | Hook the host with **Frida** | **Current.** | 2023–2025, 4 papers | | Hook the host with **Frida** | **Current.** | 2023–2025, 4 papers |
 | Hook the host's bridge with **Xposed** | **Still used here**, unlike general app hooking. | 2020, 2022, and {[wei2026_raising]} in 2026 | | Hook the host's bridge with **Xposed** | **Still used here**, unlike general app hooking. | 2020, 2022, and {[wei2026_raising]} in 2026 |
-| Run WeChat on an **emulator** | **Unreliable.** | WeChat would not launch on Android 11/15 AVDs (2025); use a rooted handset | +| Run WeChat on an **emulator** | **Mixed.** | Stock Android 11/15 AVDs would not launch it; BlueStacks and an Android 14 emulator were used (2025). Name the emulator | 
-| Treat **DevTools** as the device | **Wrong for user-data APIs.** | WeChat documents DevTools returning real data where devices return anonymous data | +| Treat **DevTools** as the device | **Documented to differ; and it cannot open a third-party mini-program.** | WeChat documents DevTools returning real data where devices returned anonymous data (2021-era base libraries); remote debugging uploads your own code | 
-| Validate with a **hand sample of flagged and unflagged cases** | **Current norm.** | 100 + 100 (2022, 2024), up to 500 + 500 (2025) | +| Validate with a **hand sample of flagged and unflagged cases** | **Current norm.** | 100 + 100 (2022, 2024), up to 500 + 500 (2025); a labelled ground-truth set twice (2025, 2026) | 
-| Validate against the **host's own oracle** (key-validation API) | **Best practice where it exists.** | 2023 |+| Validate against the **host's own oracle** (key-validation API) | **Best precision where it exists; the check is itself a use of the leaked credential.** | 2023, and a similar check in 2025 under IRB review |
 | **LLMs** in the pipeline | **New, not yet a practice.** | 2 of 16 (2024, 2026) as components; one more outside the extraction (2026); none as a classifier | | **LLMs** in the pipeline | **New, not yet a practice.** | 2 of 16 (2024, 2026) as components; one more outside the extraction (2026); none as a classifier |
-| Measure **non-Chinese hosts** (Telegram Mini Apps, Snap, LINE) | **Open.** | 2 corpora of LINE/VK (2025, 2026); none of Telegram Mini Apps |+| Measure **non-Chinese hosts** (Telegram Mini Apps, Snap, LINE) | **Open.** | 2 corpora of LINE/VK (2025, 2026); Telegram Mini Apps only in two August 2026 preprints |
  
 ===== What to Report ===== ===== What to Report =====
  
   - **The host, its version and build, the OS, and the date** — per platform build if you ran more than one. None of the 9 deployed-population papers states a host-app or base-library version.   - **The host, its version and build, the OS, and the date** — per platform build if you ran more than one. None of the 9 deployed-population papers states a host-app or base-library version.
-  - **The acquisition route**: tool, seed keywords (and their language), rate limit, dates of collection, and the account(s) used, with the phone number's country and verification state.+  - **The acquisition route**: tool, seed keywords (and their language), rate limit, dates of collection, storage, and the account(s) used, with the phone number's country and verification state.
   - **The unit**: appid, package version, sub-packages, pages, engine; and how you deduplicated.   - **The unit**: appid, package version, sub-packages, pages, engine; and how you deduplicated.
   - **Attrition, numerically**: found → downloaded → unpacked → analysable → timed out → analysed, with the reason at each step. Put every percentage over a named stage.   - **Attrition, numerically**: found → downloaded → unpacked → analysable → timed out → analysed, with the reason at each step. Put every percentage over a named stage.
Line 272: Line 273:
  
 <WRAP todo> <WRAP todo>
-  * **Nobody knows the size of any ecosystem.** Every denominator is "what we crawled"; the figures quoted for WeChat's total are citations of secondary sources, and Tencent's own 2025–2026 results publish none. The regulator's mandatory filing records are a candidate population frame that no paper in these venues has tried. +  * **No host publishes a count you can use as a denominator.** Every denominator is "what we crawled"; the figures quoted for WeChat's total are citations of secondary sources, and Tencent's 2025 annual and 2026 quarterly results we read publish none. The regulator's mandatory filing records are a candidate frame, if they can be enumerated, that no paper in these venues has tried. 
-  * **The crawling routes are closing.** MiniCrawler's batch appid route was reported blocked in 2024, and the desktop-cache route depends on one client's behaviour. A maintained, documented acquisition method would unblock the whole field. +  * **Nobody knows which crawling routes still work.** MiniCrawler's batch appid route was reported blocked in 2024 while later papers still name it, and the desktop-cache route depends on one client's behaviour. A maintained, documented acquisition method would unblock the whole field. 
-  * **What a host learns is measured once.** {[wang2025_wechat]} is the only first-party tracking measurement of a super app in these venues, on one host, over 104 popular mini-programs, with non-Chinese accounts. Alipay and Baidu advertise similar default analytics; nobody has measured them.+  * **What a host learns is measured once from the network side** ({[wang2025_wechat]}: one host, 104 popular mini-programs, non-Chinese accounts) **and once from the operator's own logs** ({[cai2025_tell]}, with the operator's cooperation). Alipay and Baidu advertise similar default analytics; nobody has measured them from outside.
   * **Declared versus actual privacy.** WeChat enforces a per-mini-program privacy declaration; comparing it with observed API use would be the mini-program analogue of the Data safety label studies on [[Design:Mobile and app measurement]].   * **Declared versus actual privacy.** WeChat enforces a per-mini-program privacy declaration; comparing it with observed API use would be the mini-program analogue of the Data safety label studies on [[Design:Mobile and app measurement]].
-  * **Telegram Mini Apps are unmeasured** in these venues, although their trust model reproduces the client-trust failures the WeChat literature measured. +  * **Telegram Mini Apps are unmeasured in these venues.** Two August 2026 preprints ({[ciccotelli2026_tenet]}, {[ferrari2026_telegapper]}) are the first measurements we found anywhere; a population method for a host whose mini-apps are ordinary web pages is still to be written. 
-  * **Recall.** Almost every detector here reports precision on a hand sample; recall is bounded, at best, by 100 unflagged mini-programs.+  * **Recall.** Every detector here reports precision; recall comes from hand samples of 100 to 500 unflagged mini-programs in three papers and from a labelled benchmark in two. None measures recall against an independent ground truth at the scale of its crawl.
 </WRAP> </WRAP>
  
design/mobile_and_app_measurement/mini_programs.1790526119.txt.gz · Last modified: by karel.kubicek.claude