| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| provenance:design:mobile_and_app_measurement:mini_programs [2026/09/27 16:40] – Review round 1 log, recoded hand codes, hardened external checks. Authored by Claude karel.kubicek.claude | provenance:design:mobile_and_app_measurement:mini_programs [2026/09/27 17:08] (current) – Review round 3 log; final outputs. Authored by Claude karel.kubicek.claude |
|---|
| | Number guard | ''check_page_numbers.mjs'' whole-page against the concatenation of the report, verifier and external-check outputs: every figure traces. That guard is a presence check, not a binding check; figures were also read against the output by hand and by the figures reviewer. | | | Number guard | ''check_page_numbers.mjs'' whole-page against the concatenation of the report, verifier and external-check outputs: every figure traces. That guard is a presence check, not a binding check; figures were also read against the output by hand and by the figures reviewer. | |
| | Bibliography | 27 entries appended to [[literature:bibliography]] in two saves (25 at rev 1790526095, 2 Telegram Mini Apps preprints at rev 1790527209 after the external-currency review; 1,162 → 1,189 entries): 19 from ''bibgen.mjs'' over the corpus index, 8 hand-written from Crossref and arXiv records for papers outside the index; USENIX and PoPETs author lists from the landing pages' ''citation_author'' metadata and author line; ''bib_dedup_scan.py'' over the merged file: 0 same-DOI or same-title pairs, 30 same-surname-same-year candidates involving a new key, all 30 different papers by title; no literal ''@'' inside a field | | | Bibliography | 27 entries appended to [[literature:bibliography]] in two saves (25 at rev 1790526095, 2 Telegram Mini Apps preprints at rev 1790527209 after the external-currency review; 1,162 → 1,189 entries): 19 from ''bibgen.mjs'' over the corpus index, 8 hand-written from Crossref and arXiv records for papers outside the index; USENIX and PoPETs author lists from the landing pages' ''citation_author'' metadata and author line; ''bib_dedup_scan.py'' over the merged file: 0 same-DOI or same-title pairs, 30 same-surname-same-year candidates involving a new key, all 30 different papers by title; no literal ''@'' inside a field | |
| | Agents | One Opus session (queries, triage of 41 candidates, verdicts, hand codes, drafting, verification). Four Sonnet readers read the 17 papers that looked like mini-program studies (16 candidates + 1 recall hit) into unpublished notes (brief below). One Opus sub-agent did the external-source pass (about 30 verified claims, notes unpublished); every load-bearing claim on the page was then re-fetched by the external-check script. Four review sub-agents — see [[#Review log]]. | | | Agents | One Opus session (queries, triage of 41 candidates, verdicts, hand codes, drafting, verification). Four Sonnet readers read the 17 papers that looked like mini-program studies (16 candidates + 1 recall hit) into unpublished notes (brief below). One Opus sub-agent did the external-source pass (about 30 verified claims, notes unpublished); every load-bearing claim on the page was then re-fetched by the external-check script. Six review sub-agents in three rounds — three focused Sonnet passes, a Sonnet re-check of their fixes, a generic Fable pass, and a Sonnet re-check of its fixes — see [[#Review log]]. | |
| |
| ==== Mistakes made in this run ==== | ==== Mistakes made in this run ==== |
| * **Seven claims in the first draft overstated the evidence and were corrected before any review**: "three crawling routes, two of which the host has since closed or changed" (only MiniCrawler's is reported closed; the search endpoint's current state is unknown); the malware corpus "took two and a half years" (it was collected March 2020 – June 2022, revisited to December 2022); "delisted within about six months" (delisting was measured at the end of 2022 over miniapps collected since 2020); "7 of the 9 deployed-population papers state no host version" (9 of 9 once two ''n/a'' codes were recoded ''not-stated''); a language label on two unpackers that no source gave; a sentence about author clustering that no query measured; and "before a single flow was readable". | * **Seven claims in the first draft overstated the evidence and were corrected before any review**: "three crawling routes, two of which the host has since closed or changed" (only MiniCrawler's is reported closed; the search endpoint's current state is unknown); the malware corpus "took two and a half years" (it was collected March 2020 – June 2022, revisited to December 2022); "delisted within about six months" (delisting was measured at the end of 2022 over miniapps collected since 2020); "7 of the 9 deployed-population papers state no host version" (9 of 9 once two ''n/a'' codes were recoded ''not-stated''); a language label on two unpackers that no source gave; a sentence about author clustering that no query measured; and "before a single flow was readable". |
| * **One external check asserted nothing on its first version.** "Tencent publishes no mini-program count" was checked against the Q2 2026 release, which contains **zero** mentions of "Mini Program" — so "no count next to a mention" was vacuously true. The script now also reads the 2025 annual and Q1 2026 PDFs and requires a positive control (the Weixin/WeChat MAU line) to be present in each before accepting "no count"; the annual release does mention "content-related Mini Programs" twice, with no number. | * **One external check asserted nothing on its first version.** "Tencent publishes no mini-program count" was checked against the Q2 2026 release, which contains **zero** mentions of "Mini Program" — so "no count next to a mention" was vacuously true. The script now also reads the 2025 annual and Q1 2026 PDFs and requires a positive control (the Weixin/WeChat MAU line) to be present in each before accepting "no count"; the annual release does mention "content-related Mini Programs" twice, with no number. |
| * **Three hand codes were wrong until the figures review** (see [[#Review log]]): {[shi2026_better]} was coded as acknowledging that its sample is not the population, which its text does not say — the WRAP box's "six of the nine" was **five**; {[wang2025_wechat]} was coded as using Chinese seed keywords, which it attributes to a cited prior method; {[cai2025_tell]} was coded "none-stated" for artefacts although it discusses and declines release (new code ''declined''). Only the first moved a printed figure. | * **Four hand codes were wrong until review** (see [[#Review log]]): {[shi2026_better]} (figures review) and {[yang2025_miniapp]} (generic review) were coded as acknowledging that their sample is not the population, which neither text says — both had caveats about the detector, not the crawl; the WRAP box's "six of the nine" is **four**; {[wang2025_wechat]} was coded as using Chinese seed keywords, which it attributes to a cited prior method; {[cai2025_tell]} was coded "none-stated" for artefacts although it discusses and declines release (new code ''declined''). Only the first moved a printed figure. |
| * **A quote hid its own citation.** The page quoted {[liu2024_riotfuzzer]}'s "most mini-apps employ obfuscation techniques" without the "as reported in [45]" that precedes it; [45] is {[zhang2021_measurement]}. Caught by the citations reviewer; the sentence now credits the measurement to its source. | * **A quote hid its own citation.** The page quoted {[liu2024_riotfuzzer]}'s "most mini-apps employ obfuscation techniques" without the "as reported in [45]" that precedes it; [45] is {[zhang2021_measurement]}. Caught by the citations reviewer; the sentence now credits the measurement to its source. |
| * **Two readers missed released code.** KeySentinel ({[zhou2025_secrets]}) and RIoTFuzzer ({[liu2024_riotfuzzer]}) print their repository URL only in the reference list; the readers coded "not stated". The report's hand-versus-schema cross-check (section G) flagged both; verified in the text and recoded. | * **Two readers missed released code.** KeySentinel ({[zhou2025_secrets]}) and RIoTFuzzer ({[liu2024_riotfuzzer]}) print their repository URL only in the reference list; the readers coded "not stated". The report's hand-versus-schema cross-check (section G) flagged both; verified in the text and recoded. |
| * **Hand codes** for the 16 (''HAND'' in the fold): unit, hosts, side, acquisition route, collected and analysed counts, analysis kind, tools, instrumentation, host version stated, account type, Chinese keyword handling, validation, LLM use, ethics review, disclosure, artefacts, and whether the paper acknowledges its sample is not the population. "not-stated" means the paper does not say. | * **Hand codes** for the 16 (''HAND'' in the fold): unit, hosts, side, acquisition route, collected and analysed counts, analysis kind, tools, instrumentation, host version stated, account type, Chinese keyword handling, validation, LLM use, ethics review, disclosure, artefacts, and whether the paper acknowledges its sample is not the population. "not-stated" means the paper does not say. |
| * **Hand versus schema** (section G): ''tools[]'' names MiniCrawler for exactly the five papers the hand codes do; ''ethics.notifiedAffectedParties'' is ''yes'' for all 16, matching the hand disclosure codes; ''ethics.reviewOutcome = approved'' matches the 3 hand-coded approvals; the extraction reads the IEEE S&P 2025 secrets paper as "explicitly-discussed-no-review", which its text ("Following ethical guidelines [65], all analyses were conducted locally") does not say — the hand code stays "none-mentioned". Artefacts disagreed on two papers until the readers' misses were fixed (see [[#Mistakes made in this run]]). | * **Hand versus schema** (section G): ''tools[]'' names MiniCrawler for exactly the five papers the hand codes do; ''ethics.notifiedAffectedParties'' is ''yes'' for all 16, matching the hand disclosure codes; ''ethics.reviewOutcome = approved'' matches the 3 hand-coded approvals; the extraction reads the IEEE S&P 2025 secrets paper as "explicitly-discussed-no-review", which its text ("Following ethical guidelines [65], all analyses were conducted locally") does not say — the hand code stays "none-mentioned". Artefacts disagreed on two papers until the readers' misses were fixed (see [[#Mistakes made in this run]]). |
| | * **Hand codes checked against the paper text, not the notes**: all 9 ''denominator'' codes of the deployed-population papers (after two were found wrong), all 16 ''irb'' codes and all 16 ''artifacts'' codes (section K's probes plus the schema cross-check; every hit read), and every code a reviewer spot-checked — 41 of the 288 codes systematically, the rest by spot-check. Of the codes checked, 6 were wrong (2 ''denominator'', 3 ''artifacts'', 1 ''keywords''); the other fields' codes (unit, hosts, acquisition, tools, validation) were taken from the reader notes and each is visible, with its paper, in the fold below. That error rate is why the page prints counts, not rates, for hand-coded fields. |
| * **The invariant**: the report throws if a candidate lacks a verdict, a verdict names a non-candidate, a verdict note still says PENDING, an IN paper lacks hand codes, hand codes exist for a non-IN paper, a hand-code record lacks a field, or a venue-index title outside the extraction has no hand entry. | * **The invariant**: the report throws if a candidate lacks a verdict, a verdict names a non-candidate, a verdict note still says PENDING, an IN paper lacks hand codes, hand codes exist for a non-IN paper, a hand-code record lacks a field, or a venue-index title outside the extraction has no hand entry. |
| |
| |
| * **Page quotes**: ''verify_mini_programs_figures.mjs'' pulls **every** ''%%//"…"//%%'' span out of the page source and requires each to be located in the paper of a citekey on the same line (the publisher PDF text for {[wei2026_raising]}), or to be an EXTERNAL span whose external check printed OK. Six spans are found only in the ''pypdf'' rendering, where ''.cols'' splices two columns across them. Four print a WARN because the nearest citekey on the line is a different paper; each was read and the quote is attributed in the sentence to the right one. | * **Page quotes**: ''verify_mini_programs_figures.mjs'' pulls **every** ''%%//"…"//%%'' span out of the page source and requires each to be located in the paper of a citekey on the same line (the publisher PDF text for {[wei2026_raising]}), or to be an EXTERNAL span whose external check printed OK. Six spans are found only in the ''pypdf'' rendering, where ''.cols'' splices two columns across them. Four print a WARN because the nearest citekey on the line is a different paper; each was read and the quote is attributed in the sentence to the right one. |
| * **Per-paper figures**: 89 needles, located in ''.cols'' or, for eleven, only in the ''pypdf'' re-extraction. Three mutated needles (40,880→40,881; 41,726→41,727; 170→171) must not be found and are not. The needles under 20 characters are tool names whose presence is the claim. | * **Per-paper figures**: 101 needles, located in ''.cols'' or, for thirteen, only in the ''pypdf'' re-extraction. Three mutated needles (40,880→40,881; 41,726→41,727; 170→171) must not be found and are not. The needles under 20 characters are tool names whose presence is the claim. |
| * **One number is verified in the form the paper prints it**: {[yang2025_miniapp]}'s "19, 905" carries a space inside the number in both renderings; the page writes 19,905. | * **One number is verified in the form the paper prints it**: {[yang2025_miniapp]}'s "19, 905" carries a space inside the number in both renderings; the page writes 19,905. |
| * **NUL bytes**: the MiniCAT ''.cols'' file contains 141 NUL bytes, which make ''grep''/''ugrep'' return nothing silently; the verifier and the notes check fold them out before matching. | * **NUL bytes**: the MiniCAT ''.cols'' file contains 141 NUL bytes, which make ''grep''/''ugrep'' return nothing silently; the verifier and the notes check fold them out before matching. |
| * **[[:roadmap]]** (rev 1790511917 → 1790526202): an //Assessed// row. **[[provenance:roadmap]]** (rev 1790511918 → 1790526204): a dated decision entry under 3g. | * **[[:roadmap]]** (rev 1790511917 → 1790526202): an //Assessed// row. **[[provenance:roadmap]]** (rev 1790511918 → 1790526204): a dated decision entry under 3g. |
| * **Nothing filed.** The one piece of adjacent work the page points at — whether the regulator's filing system can serve as a population frame — is an open research question, not a wiki item. | * **Nothing filed.** The one piece of adjacent work the page points at — whether the regulator's filing system can serve as a population frame — is an open research question, not a wiki item. |
| | * **This page and its content page, after review**: content rev 1790526119 → 1790527251 (round 1) → 1790528159 (round 2) → final save below; provenance rev 1790526233 → 1790527253 → 1790528161 → final save below. The final revisions are recorded in ''notes/mp_log.md'' and on the drain item. |
| |
| ===== Review log ===== | ===== Review log ===== |
| |
| Round 1 accepted all 12 reviewer findings and rejected none. The citations reviewer checked all 26 keys, the USENIX/PoPETs author lists and about 20 quotes independently and found one defect, which is what a clean pass looks like when the verifier has already run; the figures reviewer's F2 is the one that changed a headline number and could only be found by reading a paper, not by any guard. | Round 1 accepted all 12 reviewer findings and rejected none. The citations reviewer checked all 26 keys, the USENIX/PoPETs author lists and about 20 quotes independently and found one defect, which is what a clean pass looks like when the verifier has already run; the figures reviewer's F2 is the one that changed a headline number and could only be found by reading a paper, not by any guard. |
| | |
| | ==== Round 2: re-check of the fixes (Sonnet) and the generic review (Fable) ==== |
| | |
| | The re-check diffed the two snapshots, re-fetched every new external claim, mutation-tested the hardened ''get()'' and the OpenAlex check, and found **0** defects in the round-1 fixes (''notes/mp_review_recheck.md''). The generic review (''notes/mp_review_generic.md'') ran on the same snapshot (rev 1790527251 / 1790527253) and returned 20 findings, 6 marked major. |
| | |
| | ^ # ^ Finding ^ Decision ^ |
| | | G1 (major) | "recall almost nowhere … bounded at best by 100 unflagged" is contradicted by recall 85.56% (500 per host) and 83.55% (a labelled set) | **Accepted.** Vulnerability paragraph, currency row and Open Question rewritten with the counts; recall figures added as needles. | |
| | | G2 (major) | {[yang2025_miniapp]} ''denominator: yes'' rests on detector caveats; "five of the nine" is four | **Accepted.** Re-read: its "representative and generalizable" sentence is about countermeasures across hosts. Recoded; WRAP box "Four of the nine"; all nine denominator codes then re-checked against the text (see [[#The hand audit]]). | |
| | | G3 (major) | "none of the 16 discusses the host's terms" — the 2026 mini-games paper asserts "compliance with platform policies" | **Accepted.** Now "none quotes or analyses the licence", with the three compliance assertions named; backed by section K's terms probe (every hit read). | |
| | | G4 (major) | "WeChat resists emulators / Unreliable" contradicts the page's own Android 14 emulator paper and the BlueStacks run | **Accepted.** "Stock images do not run it; others did"; currency row "Mixed". | |
| | | G5 (major) | MiniCAT: the 14,920 timeouts are quoted, not "silent"; the informative denominator is the 26,806 completed (49.8%) | **Accepted** with the reviewer's caveat that the paper does not say whether a timed-out query can flag. Y block prints it. | |
| | | G6 (major) | MiniCrawler "closed" in the intro rests on one 2024 sentence, against later papers that name it | **Accepted.** Intro and currency row now state both sides; "What to Read First" says "its 2024 report". The reviewer counted three later papers; on re-reading, two name MiniCrawler or its extension for undated crawls (2025, 2026) and one reuses the extended crawler's method — the page says that. | |
| | | G7 | gated datasets "reused by later papers" / "Current, cheapest start" with no corpus use | **Accepted.** "None in the corpus"; currency "Available, untried in these venues", with the crawl's age. | |
| | | G8 | "quarterly results for 2025–2026" wider than the three releases read | **Accepted.** Names the three documents. | |
| | | G9 | the host "oracle" is a token grab with a third party's leaked secret, recommended without an ethics note | **Accepted.** Said in the vulnerability paragraph, the ethics bullets and the currency row. | |
| | | G10 | 975/570 is a 2021 count stated in the present tense | **Accepted.** Dated in the intro and Language section; the 2021 date is a needle. | |
| | | G11 | Telegram ''photo_url'' is optional and privacy-gated | **Accepted.** Reworded; the "privacy settings allow" clause is now an external check. | |
| | | G12 | DevTools cannot attach to a third-party mini-program; the bug note is 2021-era | **Accepted.** Added with the remote-debugging and ''wx.setEnableDebug'' documentation, both re-fetched by the script. | |
| | | G13 | "measured once" vs two host-side papers | **Accepted.** | |
| | | G14 | the filing system is a lookup by number or name, not a category browse | **Accepted.** "If it can be enumerated". | |
| | | G15 | the provenance does not say how many hand codes were checked against the paper | **Accepted.** 41 of 288 systematically, 6 wrong among those checked; recorded in [[#The hand audit]]. | |
| | | G16 | the provenance recorded a fourth review before it ran | **Accepted.** This log and the Agents row were rewritten after the last review returned. | |
| | | G17 | the recall probe's six hits reported as one | **Accepted.** | |
| | | G18 | one table cell carried two tools' states unattributed | **Accepted.** | |
| | | G19 | "nobody knows the size of any ecosystem" — the hosts do | **Accepted.** "No host publishes a count you can use as a denominator." | |
| | | G20 | storage budgets and the licence's governing-law clause are missing; the 8.2.1.4 reading is the page's own | **Accepted.** 6.29 TB and 126.38 GB (needles) with per-package arithmetic; clause 12.3 (external check); "on our reading". | |
| | |
| | All 20 accepted, none rejected. The generic pass found the defects no guard could: four sentences about the literature that the papers contradict (G1, G3, G4, G6), a second instance of the hand-code error the figures reviewer had found one paper over (G2), and advice with an unstated ethics cost (G9). Its fixes were re-verified by the verifier (46 spans, 100 needles, 0 not located), the external script (68 OK, 0 FAILED) and the number guard; they were **not** sent to a third review round. The round-2 fixes are therefore reviewed only by those guards. |
| | |
| | ==== Round 3: re-check of the round-2 fixes (Sonnet) ==== |
| | |
| | ^ # ^ Finding ^ Decision ^ |
| | | R1 | the G1 fix still undercounts: {[zhou2025_secrets]} reports recall on a labelled benchmark (83.38% on WeChat, Table 14); only its 300-detection field check is precision-only | **Accepted.** Vulnerability paragraph, Open Question and currency row now say two papers report recall against a labelled set; 83.38% added as a needle. | |
| | |
| | The re-check verified every other round-2 fix against the papers and the fetched sources, and checked the provenance's hand-code audit bullet ("41 of 288", "6 wrong") against the review record. The R1 fix is a three-phrase change backed by a verifier needle and was not sent to a further round. |
| |
| ===== The report script ===== | ===== The report script ===== |
| ar('wang2025_wechat: analysed / attempted', 104, 170); | ar('wang2025_wechat: analysed / attempted', 104, 170); |
| console.log(` IN papers releasing code (code, code+data): ${IN.filter(([k]) => ['code', 'code+data'].includes(HAND[k].artifacts)).length} of ${nIn}; any artefact incl. gated data: ${IN.filter(([k]) => ['code', 'code+data', 'gated-data'].includes(HAND[k].artifacts)).length}`); | console.log(` IN papers releasing code (code, code+data): ${IN.filter(([k]) => ['code', 'code+data'].includes(HAND[k].artifacts)).length} of ${nIn}; any artefact incl. gated data: ${IN.filter(([k]) => ['code', 'code+data', 'gated-data'].includes(HAND[k].artifacts)).length}`); |
| | ar('zhang2024_minicat: potentially vulnerable / analysed minus timed out (41,726 - 14,920)', 13349, 41726 - 14920); |
| | console.log(` zhang2024_minicat: packages on which the detector completed: ${(41726 - 14920).toLocaleString('en-US')}`); |
| | console.log(` yang2022_cross: storage per WeChat package: 6.29 TB / 2,571,490 = ${(6.29e12 / 2571490 / 1e6).toFixed(2)} MB (decimal units)`); |
| | console.log(` zhang2024_minicat: storage per unpacked package: 126.38 GB / 44,273 = ${(126.38e9 / 44273 / 1e6).toFixed(2)} MB (decimal units)`); |
| console.log(` vocabulary >= 3 precision / gap-rule precision: ${(16 / 30 / (15 / 37)).toFixed(2)}x (53.3% vs 40.5%)`); | console.log(` vocabulary >= 3 precision / gap-rule precision: ${(16 / 30 / (15 / 37)).toFixed(2)}x (53.3% vs 40.5%)`); |
| console.log(` IN papers stating an IRB/ethics-board approval: ${IN.filter(([k]) => HAND[k].irb === 'approval').length} of ${nIn}`); | console.log(` IN papers stating an IRB/ethics-board approval: ${IN.filter(([k]) => HAND[k].irb === 'approval').length} of ${nIn}`); |
| console.log(` hosts: papers including a host outside WeChat/Baidu/TikTok-Douyin/Alipay/QQ: ${IN.filter(([k]) => HAND[k].hosts.some((h) => h === 'other' || h === 'IoT host')).length}`); | console.log(` hosts: papers including a host outside WeChat/Baidu/TikTok-Douyin/Alipay/QQ: ${IN.filter(([k]) => HAND[k].hosts.some((h) => h === 'other' || h === 'IoT host')).length}`); |
| console.log(` hosts: papers naming a host other than WeChat: ${IN.filter(([k]) => HAND[k].hosts.some((h) => h !== 'WeChat')).length}; WeChat-only: ${IN.filter(([k]) => HAND[k].hosts.length === 1 && HAND[k].hosts[0] === 'WeChat').length}`); | console.log(` hosts: papers naming a host other than WeChat: ${IN.filter(([k]) => HAND[k].hosts.some((h) => h !== 'WeChat')).length}; WeChat-only: ${IN.filter(([k]) => HAND[k].hosts.length === 1 && HAND[k].hosts[0] === 'WeChat').length}`); |
| | |
| | // ---------------------------------------------------------------- K. probes behind the page's negative claims about the 16 |
| | // Each hand code the page prints as "none" or "not stated" is backed by a probe over the paper's own text; every hit was read. |
| | console.log('\n== K. Probes behind negative claims (distinct matched strings per paper; every hit read by hand) =='); |
| | const KP = { irb: /\bIRB\b|ethics? (review )?(board|committee)|institutional review|review board/gi, |
| | release: /github\.com\/[\w.-]+|zenodo|gitlab\.com|figshare|osf\.io/gi, |
| | terms: /terms of (service|use)|licen[cs]e agreement|platform polic(y|ies)|in compliance|ensur\w* compliance|comply with/gi }; |
| | for (const [k] of IN) { |
| | const t = fs.readFileSync(path.join(dataRoot(), 'fulltext', String(byKey.get(k).year), byKey.get(k).venue, byKey.get(k).slug, 'paper.cols.txt'), 'latin1').replace(/\u0000/g, '').replace(/\s+/g, ' '); |
| | const f = Object.fromEntries(Object.entries(KP).map(([n, re]) => [n, [...new Set((t.match(re) || []).map((x) => x.toLowerCase()))].slice(0, 5)])); |
| | console.log(` ${byKey.get(k).year} ${byKey.get(k).venue} ${byKey.get(k).slug.slice(0, 34)} | irb=${HAND[k].irb} ${JSON.stringify(f.irb)} | artifacts=${HAND[k].artifacts} ${JSON.stringify(f.release)} | terms ${JSON.stringify(f.terms)}`); |
| | } |
| | console.log(' reading of the terms hits: the 2026 mini-games paper asserts its crawl "ensur[es] compliance with platform policies"; the NDSS 2026 paper "compliance with all relevant laws and regulations"; the 2024 rental paper complies with "the vendor\'s bug bounty plan"; the other hits are about permission policies, privacy regulation or advertising policies. None quotes or analyses the host\'s user licence.'); |
| |
| if (process.argv.includes('--list')) { | if (process.argv.includes('--list')) { |
| 1 11.1% ground-truth-set | 1 11.1% ground-truth-set |
| -- denominator (of 9; multi-valued fields do not sum): | -- denominator (of 9; multi-valued fields do not sum): |
| 5 55.6% yes | 5 55.6% no |
| 4 44.4% no | 4 44.4% yes |
| |
| -- the 5 host-framework papers: | -- the 5 host-framework papers: |
| wang2025_wechat: analysed / attempted: 104 / 170 = 61.2% | wang2025_wechat: analysed / attempted: 104 / 170 = 61.2% |
| IN papers releasing code (code, code+data): 10 of 16; any artefact incl. gated data: 11 | IN papers releasing code (code, code+data): 10 of 16; any artefact incl. gated data: 11 |
| | zhang2024_minicat: potentially vulnerable / analysed minus timed out (41,726 - 14,920): 13,349 / 26,806 = 49.8% |
| | zhang2024_minicat: packages on which the detector completed: 26,806 |
| | yang2022_cross: storage per WeChat package: 6.29 TB / 2,571,490 = 2.45 MB (decimal units) |
| | zhang2024_minicat: storage per unpacked package: 126.38 GB / 44,273 = 2.85 MB (decimal units) |
| vocabulary >= 3 precision / gap-rule precision: 1.32x (53.3% vs 40.5%) | vocabulary >= 3 precision / gap-rule precision: 1.32x (53.3% vs 40.5%) |
| IN papers stating an IRB/ethics-board approval: 3 of 16 | IN papers stating an IRB/ethics-board approval: 3 of 16 |
| hosts: papers including a host outside WeChat/Baidu/TikTok-Douyin/Alipay/QQ: 7 | hosts: papers including a host outside WeChat/Baidu/TikTok-Douyin/Alipay/QQ: 7 |
| hosts: papers naming a host other than WeChat: 10; WeChat-only: 6 | hosts: papers naming a host other than WeChat: 10; WeChat-only: 6 |
| | |
| | == K. Probes behind negative claims (distinct matched strings per paper; every hit read by hand) == |
| | 2020 CCS demystifying-resource-management-r | irb=none-mentioned [] | artifacts=code ["github.com/mozillasecurity","github.com/aslody","github.com/itseez"] | terms ["comply with"] |
| | 2022 CCS cross-miniapp-request-forgery-root | irb=none-mentioned [] | artifacts=code ["github.com/osuseclab"] | terms [] |
| | 2022 USENIX identity-confusion-in-webview-base | irb=none-mentioned [] | artifacts=none-stated ["github.com/soot-oss"] | terms [] |
| | 2023 CCS dont-leak-your-keys-understanding- | irb=none-mentioned [] | artifacts=none-stated [] | terms [] |
| | 2023 CCS uncovering-and-exploiting-hidden-a | irb=none-mentioned [] | artifacts=none-stated [] | terms [] |
| | 2023 USENIX one-size-does-not-fit-all-uncoveri | irb=none-mentioned [] | artifacts=code ["github.com/1n3"] | terms [] |
| | 2024 CCS minicat-understanding-and-detectin | irb=none-mentioned [] | artifacts=code ["github.com/en","github.com/fxsjy","github.com/pywinauto","github.com/system-cpu","github.com/imingyu"] | terms [] |
| | 2026 NDSS better-safe-than-sorry-uncovering- | irb=approval ["irb"] | artifacts=code+data ["zenodo","github.com/wemobiledev"] | terms ["ensure compliance"] |
| | 2025 PETS what-wechat-knows-pervasive-first- | irb=none-mentioned [] | artifacts=code+data ["github.com/wemobiledev"] | terms [] |
| | 2025 USENIX i-can-tell-your-secrets-inferring- | irb=approval ["irb"] | artifacts=declined [] | terms ["comply with"] |
| | 2026 USENIX when-fun-turns-toxic-a-first-look- | irb=none-mentioned [] | artifacts=code ["github.com/w","zenodo","github.com/cwi-swat","github.com/wala","github.com/swc-project"] | terms ["platform policies","platform policy","terms of use","ensuring compliance"] |
| | 2025 NDSS understanding-miniapp-malware-iden | irb=none-mentioned [] | artifacts=gated-data ["github.com/jkeylu","github.com/tarruda"] | terms ["terms of use"] |
| | 2025 NDSS the-skeleton-keys-a-large-scale-an | irb=approval ["irb"] | artifacts=code ["github.com/keymagnetproject2025","github.com/ant-move","github.com/wala"] | terms [] |
| | 2024 USENIX demystifying-the-security-implicat | irb=none-mentioned [] | artifacts=none-stated [] | terms ["comply with"] |
| | 2025 IEEE-SP hey-your-secrets-leaked-detecting- | irb=none-mentioned [] | artifacts=code ["github.com/abbrcode","github.com/david47k","github.com/dwyl","github.com/aoa0","github.com/gitleaks"] | terms [] |
| | 2024 CCS riotfuzzer-companion-app-assisted- | irb=none-mentioned [] | artifacts=code ["github.com/androguard","github.com/wireghoul","github.com/iputils","github.com/kzliu2017"] | terms [] |
| | reading of the terms hits: the 2026 mini-games paper asserts its crawl "ensur[es] compliance with platform policies"; the NDSS 2026 paper "compliance with all relevant laws and regulations"; the 2024 rental paper complies with "the vendor's bug bounty plan"; the other hits are about permission policies, privacy regulation or advertising policies. None quotes or analyses the host's user licence. |
| |
| OK: candidates and verdicts agree in both directions; every IN paper has hand codes. | OK: candidates and verdicts agree in both directions; every IN paper has hand codes. |
| analysis: ['static'], tools: ['MiniCrawler'], instrumentation: [], hostVersion: 'not-stated', | analysis: ['static'], tools: ['MiniCrawler'], instrumentation: [], hostVersion: 'not-stated', |
| accounts: 'not-stated', keywords: 'not-stated', validation: 'manual-sample', llm: 'no', irb: 'none-mentioned', | accounts: 'not-stated', keywords: 'not-stated', validation: 'manual-sample', llm: 'no', irb: 'none-mentioned', |
| disclosure: ['host-vendor'], artifacts: 'gated-data', denominator: 'yes', | disclosure: ['host-vendor'], artifacts: 'gated-data', denominator: 'no', |
| }, | }, |
| 'NDSS/2025/the-skeleton-keys-a-large-scale-analysis-of-credential-leakage-in-mini-apps': { | 'NDSS/2025/the-skeleton-keys-a-large-scale-analysis-of-credential-leakage-in-mini-apps': { |
| ['lee2025_deep', 'we conducted WeChat-specific experiments on BlueStacks', 'BlueStacks'], | ['lee2025_deep', 'we conducted WeChat-specific experiments on BlueStacks', 'BlueStacks'], |
| ['wang2023_size', 'we have 1,031 APIs in total', '1,031 documented APIs'], | ['wang2023_size', 'we have 1,031 APIs in total', '1,031 documented APIs'], |
| | ['yang2022_cross', 'consume 6.29 TB disk storage. We also extended M', '6.29 TB'], |
| | ['zhang2024_minicat', 'which occupied a storage space of 126.38 GB', '126.38 GB'], |
| | ['zhang2024_minicat', 'introduced a 5-minute timeout for CodeQL queries', 'five-minute timeout'], |
| | ['shi2025_skeleton', 'The average precision of KeyMagnet is 95.04% and the recall is 85.56%', 'recall 85.56%'], |
| | ['chen2026_minigames', 'with a recall of 83.55%', 'recall 83.55%'], |
| | ['zhou2025_secrets', 'WeChat 902 89.91% 83.38%', 'WeChat benchmark recall 83.38% (Table 14)'], |
| | ['chen2026_minigames', 'This process produced 371 labeled Ad-behaviors', '371 labelled behaviours'], |
| | ['zhang2023_leak', 'Api.weixin.qq.com/cgi-bin/token', 'the oracle is the token endpoint'], |
| | ['chen2026_minigames', 'All test data was', 'crawl compliance assertion'], |
| | ['shi2026_better', 'we ensure compliance with all relevant laws and regulations', 'compliance with laws'], |
| | ['he2024_demystifying', "comply with the vendor's bug bounty plan", 'bug-bounty plan'], |
| | ['wang2023_uncovering', 'Our experiments were conducted primarily in 2021', 'the 975/570 count is from 2021'], |
| ['liu2024_riotfuzzer', 'reported in [45], most mini-apps employ obfuscation techniques', 'citing zhang2021 (as is spliced off in .cols)'], | ['liu2024_riotfuzzer', 'reported in [45], most mini-apps employ obfuscation techniques', 'citing zhang2021 (as is spliced off in .cols)'], |
| ]; | ]; |
| OK pypdf line 86 wang2025_wechat | OK pypdf line 86 wang2025_wechat |
| "38.8% of them required ID verification or Chinese phone number verification" | "38.8% of them required ID verification or Chinese phone number verification" |
| | OK paper.cols.txt line 89 yang2022_cross |
| | "6.29 TB disk storage" |
| OK paper.cols.txt line 98 zhang2024_minicat | OK paper.cols.txt line 98 zhang2024_minicat |
| "due to their use of a newer version of the WeChat mini-program base library" | "due to their use of a newer version of the WeChat mini-program base library" |
| OK paper.cols.txt line 113 lee2025_deep | OK paper.cols.txt line 113 lee2025_deep |
| "WeChat failed to launch even on Android 11 AVDs" | "WeChat failed to launch even on Android 11 AVDs" |
| | OK paper.cols.txt line 113 lee2025_deep |
| | "a configuration publicly known to support the app reliably" |
| OK paper.cols.txt line 116 wang2025_wechat | OK paper.cols.txt line 116 wang2025_wechat |
| "decided against further automation" | "decided against further automation" |
| OK paper.cols.txt line 140 chen2026_minigames | OK paper.cols.txt line 140 chen2026_minigames |
| "49.95% of ad-enabled mini-games" | "49.95% of ad-enabled mini-games" |
| | OK paper.cols.txt line 142 yang2022_cross |
| | "making the FN rate to 2%" |
| OK paper.cols.txt line 146 wang2025_wechat | OK paper.cols.txt line 146 wang2025_wechat |
| "purchased Canadian and American phone numbers, which resulted in various restrictions and limitations to our study" | "purchased Canadian and American phone numbers, which resulted in various restrictions and limitations to our study" |
| OK paper.cols.txt line 168 zhang2024_minicat | OK paper.cols.txt line 168 zhang2024_minicat |
| "248/316 (78.5%)" | "248/316 (78.5%)" |
| spans: 42; by route: paper.cols.txt 23, publisher-pdf 1, EXTERNAL 12, pypdf 6 | OK paper.cols.txt line 170 chen2026_minigames |
| | "compliance with platform policies" |
| | spans: 46; by route: paper.cols.txt 27, publisher-pdf 1, EXTERNAL 12, pypdf 6 |
| |
| == B. Per-paper figures == | == B. Per-paper figures == |
| OK paper.cols.txt wang2023_size | 1,031 documented APIs | OK paper.cols.txt wang2023_size | 1,031 documented APIs |
| "we have 1,031 APIs in total" | "we have 1,031 APIs in total" |
| | OK paper.cols.txt yang2022_cross | 6.29 TB |
| | "consume 6.29 TB disk storage. We also extended M" |
| | OK paper.cols.txt zhang2024_minicat | 126.38 GB |
| | "which occupied a storage space of 126.38 GB" |
| | OK paper.cols.txt zhang2024_minicat | five-minute timeout |
| | "introduced a 5-minute timeout for CodeQL queries" |
| | OK pypdf shi2025_skeleton | recall 85.56% |
| | "The average precision of KeyMagnet is 95.04% and the recall is 85.56%" |
| | OK paper.cols.txt chen2026_minigames | recall 83.55% |
| | "with a recall of 83.55%" |
| | OK paper.cols.txt zhou2025_secrets | WeChat benchmark recall 83.38% (Table 14) |
| | "WeChat 902 89.91% 83.38%" |
| | OK paper.cols.txt chen2026_minigames | 371 labelled behaviours |
| | "This process produced 371 labeled Ad-behaviors" |
| | OK paper.cols.txt zhang2023_leak | the oracle is the token endpoint |
| | "Api.weixin.qq.com/cgi-bin/token" |
| | OK paper.cols.txt WEAK chen2026_minigames | crawl compliance assertion |
| | "All test data was" |
| | OK paper.cols.txt shi2026_better | compliance with laws |
| | "we ensure compliance with all relevant laws and regulations" |
| | OK paper.cols.txt he2024_demystifying | bug-bounty plan |
| | "comply with the vendor's bug bounty plan" |
| | OK pypdf wang2023_uncovering | the 975/570 count is from 2021 |
| | "Our experiments were conducted primarily in 2021" |
| OK paper.cols.txt liu2024_riotfuzzer | citing zhang2021 (as is spliced off in .cols) | OK paper.cols.txt liu2024_riotfuzzer | citing zhang2021 (as is spliced off in .cols) |
| "reported in [45], most mini-apps employ obfuscation techniques" | "reported in [45], most mini-apps employ obfuscation techniques" |
| needles: 89; by route: pypdf 10, paper.cols.txt 76, pypdf+dehyph 1, publisher-pdf 2; weak (<20 chars): 9 | needles: 101; by route: pypdf 12, paper.cols.txt 86, pypdf+dehyph 1, publisher-pdf 2; weak (<20 chars): 10 |
| |
| === EXTERNAL FIGURES (non-corpus; each re-fetched by external_checks_mini_programs.sh — see its labels) === | === EXTERNAL FIGURES (non-corpus; each re-fetched by external_checks_mini_programs.sh — see its labels) === |
| get https://developers.weixin.qq.com/miniprogram/dev/api/open-api/user-info/wx.getUserProfile.html $TMP/wx_gup.html | get https://developers.weixin.qq.com/miniprogram/dev/api/open-api/user-info/wx.getUserProfile.html $TMP/wx_gup.html |
| check 'getUserProfile: DevTools 2.10.4-2.16.1 real data, devices anonymous' $TMP/wx_gup.html '开发者工具中 ?2\.10\.4 ?~ ?2\.16\.1 ?基础库版本.{0,80}真机上此区间会按照公告返回匿名数据' | check 'getUserProfile: DevTools 2.10.4-2.16.1 real data, devices anonymous' $TMP/wx_gup.html '开发者工具中 ?2\.10\.4 ?~ ?2\.16\.1 ?基础库版本.{0,80}真机上此区间会按照公告返回匿名数据' |
| | get https://developers.weixin.qq.com/miniprogram/dev/devtools/remote-debug.html $TMP/wx_rd.html |
| | check 'DevTools remote debugging packs and uploads the local code' $TMP/wx_rd.html '工具会将本地代码进行处理打包并上传' |
| | get https://developers.weixin.qq.com/miniprogram/dev/api/base/debug/wx.setEnableDebug.html $TMP/wx_dbg.html |
| | check 'wx.setEnableDebug: debug switch set by the mini-program, works on release builds' $TMP/wx_dbg.html '此开关对正式版也能生效' |
| get https://developers.weixin.qq.com/miniprogram/product/record/record_faq.html $TMP/wx_faq.html | get https://developers.weixin.qq.com/miniprogram/product/record/record_faq.html $TMP/wx_faq.html |
| check 'Filing FAQ: unfiled mini-programs become inaccessible or delisted' $TMP/wx_faq.html '无法访问或下架' | check 'Filing FAQ: unfiled mini-programs become inaccessible or delisted' $TMP/wx_faq.html '无法访问或下架' |
| check 'WeChat licence 8.2.1.4: no plug-ins or unauthorised third-party tools' $TMP/wx_lic.html '8\.2\.1\.4.{0,200}使用插件、外挂或非经腾讯授权的第三方工具' | check 'WeChat licence 8.2.1.4: no plug-ins or unauthorised third-party tools' $TMP/wx_lic.html '8\.2\.1\.4.{0,200}使用插件、外挂或非经腾讯授权的第三方工具' |
| get https://www.wechat.com/en/service_terms.html $TMP/wx_tos.html | get https://www.wechat.com/en/service_terms.html $TMP/wx_tos.html |
| | check 'WeChat licence 12.3: governed by mainland-Chinese law' $TMP/wx_lic.html '12\.3 ?本协议的成立、生效、履行、解释及纠纷解决,适用中华人民共和国大陆地区法律' |
| check 'WeChat international ToS: last modified 2025-11-18' $TMP/wx_tos.html 'Last modified: ?2025-11-18' | check 'WeChat international ToS: last modified 2025-11-18' $TMP/wx_tos.html 'Last modified: ?2025-11-18' |
| check 'WeChat international ToS: no reverse engineering of WeChat Software' $TMP/wx_tos.html 'reverse engineer or extract source codes from WeChat Software' | check 'WeChat international ToS: no reverse engineering of WeChat Software' $TMP/wx_tos.html 'reverse engineer or extract source codes from WeChat Software' |
| check 'Telegram Bot API 8.0 (2024-11-17): photo_url to all Mini Apps' $TMP/tg.html 'photo_url in the class WebAppUser is now available to all Mini Apps' | check 'Telegram Bot API 8.0 (2024-11-17): photo_url to all Mini Apps' $TMP/tg.html 'photo_url in the class WebAppUser is now available to all Mini Apps' |
| check 'Telegram Bot API 8.0 dated November 17, 2024' $TMP/tg.html 'November 17, 2024 Bot API 8\.0' | check 'Telegram Bot API 8.0 dated November 17, 2024' $TMP/tg.html 'November 17, 2024 Bot API 8\.0' |
| | check 'Telegram: photo_url only if privacy settings allow' $TMP/tg.html 'if their privacy settings allow for it' |
| get https://developers.tiktok.com/docs/en/mini-games-overview $TMP/tt_over.html | get https://developers.tiktok.com/docs/en/mini-games-overview $TMP/tt_over.html |
| if [ "$(wc -c < $TMP/tt_over.html)" -lt 5000 ]; then node scripts/pw_fetch_text.mjs https://developers.tiktok.com/docs/en/mini-games-overview $TMP/tt_over.html > /dev/null 2>&1; echo " PW tiktok overview -> $(wc -c < $TMP/tt_over.html) bytes"; fi | if [ "$(wc -c < $TMP/tt_over.html)" -lt 5000 ]; then node scripts/pw_fetch_text.mjs https://developers.tiktok.com/docs/en/mini-games-overview $TMP/tt_over.html > /dev/null 2>&1; echo " PW tiktok overview -> $(wc -c < $TMP/tt_over.html) bytes"; fi |
| |
| <file text external_checks_mini_programs-output.txt> | <file text external_checks_mini_programs-output.txt> |
| run: 2026-09-27T16:38Z | run: 2026-09-27T16:54Z |
| |
| == 1. Unpackers and research artefacts (GitHub API) == | == 1. Unpackers and research artefacts (GitHub API) == |
| > p : wx.getUserProfile 返回的加密数据中不包含 openId 和 unionId 字段。 bug :开发者工具中 2.10.4 ~ 2.16.1 基础库版本通过 <button open-type="getUserInfo"> 会返回真实数据,真机上此区间会按照公告返回匿名数据。 < view class = " container " > < view class = " userinfo " | > p : wx.getUserProfile 返回的加密数据中不包含 openId 和 unionId 字段。 bug :开发者工具中 2.10.4 ~ 2.16.1 基础库版本通过 <button open-type="getUserInfo"> 会返回真实数据,真机上此区间会按照公告返回匿名数据。 < view class = " container " > < view class = " userinfo " |
| OK getUserProfile: DevTools 2.10.4-2.16.1 real data, devices anonymous | OK getUserProfile: DevTools 2.10.4-2.16.1 real data, devices anonymous |
| | GET https://developers.weixin.qq.com/miniprogram/dev/devtools/remote-debug.html -> HTTP 200, 59902 bytes |
| | > .0 进行调试。 # 调试流程 要发起一个真机远程调试流程,需要先点击开发者工具的工具栏上 "真机调试" 按钮。 此时,工具会将本地代码进行处理打包并上传,就绪之后,使用手机客户端扫描二维码即可弹出调试窗口,开始远程调试。 # 远程调试窗口 使用手机扫描此二维码,即可开始远 |
| | OK DevTools remote debugging packs and uploads the local code |
| | GET https://developers.weixin.qq.com/miniprogram/dev/api/base/debug/wx.setEnableDebug.html -> HTTP 200, 449919 bytes |
| | > Windows 版 :支持 微信 Mac 版 :支持 微信 鸿蒙 OS 版 :支持 # 功能描述 设置是否打开调试开关。此开关对正式版也能生效。 # 参数 # Object object 属性 类型 默认值 必填 说明 enableDebug boolean 是 |
| | OK wx.setEnableDebug: debug switch set by the mini-program, works on release builds |
| GET https://developers.weixin.qq.com/miniprogram/product/record/record_faq.html -> HTTP 200, 67116 bytes | GET https://developers.weixin.qq.com/miniprogram/product/record/record_faq.html -> HTTP 200, 67116 bytes |
| > 【注销主体】。需要注意,注销主体后该主体下的全部备案信息均会被注销,主体下的小程序、APP、网站、快应用等将因未备案导致无法访问或下架,且提交注销申请后无法撤回,建议谨慎操作。 # 2、什么情况下选择【注销小程序】? 当你确认多平台运营的小程序备案不再使 | > 【注销主体】。需要注意,注销主体后该主体下的全部备案信息均会被注销,主体下的小程序、APP、网站、快应用等将因未备案导致无法访问或下架,且提交注销申请后无法撤回,建议谨慎操作。 # 2、什么情况下选择【注销小程序】? 当你确认多平台运营的小程序备案不再使 |
| OK WeChat licence 8.2.1.4: no plug-ins or unauthorised third-party tools | OK WeChat licence 8.2.1.4: no plug-ins or unauthorised third-party tools |
| GET https://www.wechat.com/en/service_terms.html -> HTTP 200, 183449 bytes | GET https://www.wechat.com/en/service_terms.html -> HTTP 200, 183449 bytes |
| | > 改后的协议。如果你不接受修改后的协议,应当停止使用本软件。 12.2 本协议签订地为中华人民共和国广东省深圳市南山区。 12.3 本协议的成立、生效、履行、解释及纠纷解决,适用中华人民共和国大陆地区法律(不包括冲突法)。 12.4 若你和腾讯之间发生任何纠纷或争议,首先应友好协商解决;协商不成的,你同意将纠纷或争议提交本 |
| | OK WeChat licence 12.3: governed by mainland-Chinese law |
| > der .logo a{display:inline-block} WECHAT – TERMS OF SERVICE Last modified: 2025-11-18 TABLE OF CONTENTS INTRODUCTION ADDITIONAL TERMS AND POLICIE | > der .logo a{display:inline-block} WECHAT – TERMS OF SERVICE Last modified: 2025-11-18 TABLE OF CONTENTS INTRODUCTION ADDITIONAL TERMS AND POLICIE |
| OK WeChat international ToS: last modified 2025-11-18 | OK WeChat international ToS: last modified 2025-11-18 |
| > cure local storage on the user's device for sensitive data. November 17, 2024 Bot API 8.0 This is the largest update in the history of Telegram mini | > cure local storage on the user's device for sensitive data. November 17, 2024 Bot API 8.0 This is the largest update in the history of Telegram mini |
| OK Telegram Bot API 8.0 dated November 17, 2024 | OK Telegram Bot API 8.0 dated November 17, 2024 |
| | > l Mini Apps, allowing them to access a user's profile photo if their privacy settings allow for it. Third parties (e.g., Mini App builders, external SDKs etc. |
| | OK Telegram: photo_url only if privacy settings allow |
| GET https://developers.tiktok.com/docs/en/mini-games-overview -> HTTP 200, 190714 bytes | GET https://developers.tiktok.com/docs/en/mini-games-overview -> HTTP 200, 190714 bytes |
| > ect search access, and more) to achieve seamless conversion Already launched in markets including the U.S., Japan, Indonesia, Turkey, Saudi Arabia, Thailand, Brazil, Malaysia, Philippines, and Vietnam, with new markets coming soon Leverages TikTok's global mon | > ect search access, and more) to achieve seamless conversion Already launched in markets including the U.S., Japan, Indonesia, Turkey, Saudi Arabia, Thailand, Brazil, Malaysia, Philippines, and Vietnam, with new markets coming soon Leverages TikTok's global mon |