User Tools

Site Tools


design:mobile_and_app_measurement:mini_programs

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
design:mobile_and_app_measurement:mini_programs [2026/09/27 16:55] – Review round 2 (generic) fixes: recall counts, four of nine, emulator evidence, MiniCrawler status both sides, DevTools, terms, storage, dated doc counts. Authored by Claude karel.kubicek.claudedesign:mobile_and_app_measurement:mini_programs [2026/09/27 17:08] (current) – Review round 3 fix: two papers report recall on a labelled benchmark. Authored by Claude karel.kubicek.claude
Line 140: Line 140:
 | Aggressive advertising in mini-games {[chen2026_minigames]} | "//49.95% of ad-enabled mini-games//" (457 of 915) | — | | Aggressive advertising in mini-games {[chen2026_minigames]} | "//49.95% of ad-enabled mini-games//" (457 of 915) | — |
  
-Precision is measured everywhere; recall less often and more weakly. Three deployed-population papers estimate it from a hand sample of unflagged cases — 100 {[yang2022_cross]} ("//making the FN rate to 2%//"), 100 {[zhang2024_minicat]}, up to 500 per host {[shi2025_skeleton]} (recall 85.56%) — and one against a labelled ground-truth set ({[chen2026_minigames]}, recall 83.55% on 371 audited behaviours); {[yang2025_miniapp]} samples 500 flagged cases only, and {[zhou2025_secrets]} 300 detections. {[zhang2023_leak]} validates every candidate key against the host, which is why it can claim no false positives — but the oracle is the host's token endpoint, so a "valid" result is a live access token for someone else's backend; the paper states no ethics review, and {[shi2025_skeleton]} ran the same kind of check under an IRB "minimal risk" finding. Say which you did. Report both directions of the sample, and say that 100 unflagged mini-programs bound recall only loosely.+Precision is measured everywhere; recall less often and more weakly. Three deployed-population papers estimate it from a hand sample of unflagged cases — 100 {[yang2022_cross]} ("//making the FN rate to 2%//"), 100 {[zhang2024_minicat]}, up to 500 per host {[shi2025_skeleton]} (recall 85.56%) — and two against a labelled ground-truth set ({[chen2026_minigames]}, recall 83.55% on 371 audited behaviours; {[zhou2025_secrets]}, 83.38% on the WeChat part of its benchmark, while its 300-detection field check is precision-only); {[yang2025_miniapp]} samples 500 flagged cases only. {[zhang2023_leak]} validates every candidate key against the host, which is why it can claim no false positives — but the oracle is the host's token endpoint, so a "valid" result is a live access token for someone else's backend; the paper states no ethics review, and {[shi2025_skeleton]} ran the same kind of check under an IRB "minimal risk" finding. Say which you did. Report both directions of the sample, and say that 100 unflagged mini-programs bound recall only loosely.
  
 ===== Accounts, Vantage and Language ===== ===== Accounts, Vantage and Language =====
Line 252: Line 252:
 | Run WeChat on an **emulator** | **Mixed.** | Stock Android 11/15 AVDs would not launch it; BlueStacks and an Android 14 emulator were used (2025). Name the emulator | | Run WeChat on an **emulator** | **Mixed.** | Stock Android 11/15 AVDs would not launch it; BlueStacks and an Android 14 emulator were used (2025). Name the emulator |
 | Treat **DevTools** as the device | **Documented to differ; and it cannot open a third-party mini-program.** | WeChat documents DevTools returning real data where devices returned anonymous data (2021-era base libraries); remote debugging uploads your own code | | Treat **DevTools** as the device | **Documented to differ; and it cannot open a third-party mini-program.** | WeChat documents DevTools returning real data where devices returned anonymous data (2021-era base libraries); remote debugging uploads your own code |
-| Validate with a **hand sample of flagged and unflagged cases** | **Current norm.** | 100 + 100 (2022, 2024), up to 500 + 500 (2025); a labelled ground-truth set once (2026) |+| Validate with a **hand sample of flagged and unflagged cases** | **Current norm.** | 100 + 100 (2022, 2024), up to 500 + 500 (2025); a labelled ground-truth set twice (2025, 2026) |
 | Validate against the **host's own oracle** (key-validation API) | **Best precision where it exists; the check is itself a use of the leaked credential.** | 2023, and a similar check in 2025 under IRB review | | Validate against the **host's own oracle** (key-validation API) | **Best precision where it exists; the check is itself a use of the leaked credential.** | 2023, and a similar check in 2025 under IRB review |
 | **LLMs** in the pipeline | **New, not yet a practice.** | 2 of 16 (2024, 2026) as components; one more outside the extraction (2026); none as a classifier | | **LLMs** in the pipeline | **New, not yet a practice.** | 2 of 16 (2024, 2026) as components; one more outside the extraction (2026); none as a classifier |
Line 278: Line 278:
   * **Declared versus actual privacy.** WeChat enforces a per-mini-program privacy declaration; comparing it with observed API use would be the mini-program analogue of the Data safety label studies on [[Design:Mobile and app measurement]].   * **Declared versus actual privacy.** WeChat enforces a per-mini-program privacy declaration; comparing it with observed API use would be the mini-program analogue of the Data safety label studies on [[Design:Mobile and app measurement]].
   * **Telegram Mini Apps are unmeasured in these venues.** Two August 2026 preprints ({[ciccotelli2026_tenet]}, {[ferrari2026_telegapper]}) are the first measurements we found anywhere; a population method for a host whose mini-apps are ordinary web pages is still to be written.   * **Telegram Mini Apps are unmeasured in these venues.** Two August 2026 preprints ({[ciccotelli2026_tenet]}, {[ferrari2026_telegapper]}) are the first measurements we found anywhere; a population method for a host whose mini-apps are ordinary web pages is still to be written.
-  * **Recall.** Every detector here reports precision; recall comes from hand samples of 100 to 500 unflagged mini-programs in three papers and from a labelled set in one. None measures recall against an independent ground truth at the scale of its crawl.+  * **Recall.** Every detector here reports precision; recall comes from hand samples of 100 to 500 unflagged mini-programs in three papers and from a labelled benchmark in two. None measures recall against an independent ground truth at the scale of its crawl.
 </WRAP> </WRAP>
  
design/mobile_and_app_measurement/mini_programs.1790528159.txt.gz · Last modified: by karel.kubicek.claude