User Tools

Site Tools


provenance:privacy:age_assurance

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
provenance:privacy:age_assurance [2026/09/15 17:14] – Generic review log, the mistyped-slug defect that produced a published falsehood, and the reclassification of GAN-Invert to MENTION (20/7). Authored by Claude karel.kubicek.claudeprovenance:privacy:age_assurance [2026/09/15 17:21] (current) – Second-round review log, the de-hyphenation-as-someone-else's-defect mistake, three stale figures corrected, and the completed neighbour backlinks. Authored by Claude karel.kubicek.claude
Line 12: Line 12:
 | 11 candidates from the title-and-summary probe in ''scripts/gap_probe_roadmap.mjs'' | The 11 are really **10** — the probe has no word boundary before ''age'' (below) — and **3** of them are in the derived population: //Easy As Child's Play//, //Tales from the Porn// and the CCS 2022 kids'-apps poster. **27.3% precision** against the 11. A full-text probe finds **38** candidates and **5** population papers, so the title probe's **recall is 3 of 5 (60%)**: it misses {[west2024_picture]} and {[alomar2022_developers]}, whose titles say nothing about age. | | 11 candidates from the title-and-summary probe in ''scripts/gap_probe_roadmap.mjs'' | The 11 are really **10** — the probe has no word boundary before ''age'' (below) — and **3** of them are in the derived population: //Easy As Child's Play//, //Tales from the Porn// and the CCS 2022 kids'-apps poster. **27.3% precision** against the 11. A full-text probe finds **38** candidates and **5** population papers, so the title probe's **recall is 3 of 5 (60%)**: it misses {[west2024_picture]} and {[alomar2022_developers]}, whose titles say nothing about age. |
 | "most are children's-privacy and COPPA-compliance work" | Confirmed, and sized: **36** papers name COPPA in ''legal[]'', against **5** in the age-assurance population, and exactly **1** paper is in both. | | "most are children's-privacy and COPPA-compliance work" | Confirmed, and sized: **36** papers name COPPA in ''legal[]'', against **5** in the age-assurance population, and exactly **1** paper is in both. |
-| "the one squarely on it is //Easy As Child's Play//" | Confirmed. It is the only ''OBJECT'' verdict in the audit, and it has 194 phrase matches against the runner-up's 22. |+| "the one squarely on it is //Easy As Child's Play//" | Confirmed. It is the only ''OBJECT'' verdict in the audit, and it dominates the probe: the matched-form counts in the script output below are driven by it. |
 | "the 2019 IMC porn-ecosystem paper is adjacent and useful for the web case" | Confirmed, and it turned out to be the **only** web-platform measurement of age gates in the corpus, and the source of the page's vantage-point argument. | | "the 2019 IMC porn-ecosystem paper is adjacent and useful for the web case" | Confirmed, and it turned out to be the **only** web-platform measurement of age gates in the corpus, and the source of the page's vantage-point argument. |
 | "the regulatory surface is moving years ahead of the published measurement" | Confirmed with nine dated primary sources, 2025-01-16 to 2026-09-11. | | "the regulatory surface is moving years ahead of the published measurement" | Confirmed with nine dated primary sources, 2025-01-16 to 2026-09-11. |
  
-One thing the queue did not anticipate and the page now leads on: **six papers hit age assurance as an obstacle to a measurement about something else**, and that is the most common way it appears in the corpusFour of those six are from 2025–2026.+One thing the queue did not anticipate and the page now leads on: **six papers hit age assurance as an obstacle to a measurement about something else** — more than measure itthough fewer than the 20 that merely mention it. Five of those six are from 2025–2026.
  
 ===== The population ===== ===== The population =====
Line 316: Line 316:
 console.log(table(['year', 'candidates', 'measures it', 'obstructed by it'], yearRows)); console.log(table(['year', 'candidates', 'measures it', 'obstructed by it'], yearRows));
 console.log('* 2025-2026 are provisional venue-years — see literature:corpus.'); console.log('* 2025-2026 are provisional venue-years — see literature:corpus.');
 +// A share of candidates means nothing without the corpus's own share for the
 +// same years: the corpus grew, so "more mentions lately" is partly arithmetic.
 +const recentCand = rows.filter((r) => r.p.year >= 2025).length;
 +const recentCorpus = P.filter((p) => p.year >= 2025).length;
 +console.log(`\n2025-2026: ${recentCand} of ${tightKeys.size} candidates (${pct(recentCand, tightKeys.size)})` +
 +  ` against ${recentCorpus} of ${P.length} papers in the corpus (${pct(recentCorpus, P.length)})`);
 +console.log(`  of those ${recentCand} candidates, ${rows.filter((r) => r.p.year >= 2025 && (r.v === 'OBJECT' || r.v === 'SECTION')).length} measure age assurance` +
 +  ` and ${rows.filter((r) => r.p.year >= 2025 && r.v === 'OBSTACLE').length} were obstructed by it`);
 const venueRows = [...new Set(P.map((p) => p.venue))] const venueRows = [...new Set(P.map((p) => p.venue))]
   .map((v) => [v, rows.filter((r) => r.p.venue === v).length,   .map((v) => [v, rows.filter((r) => r.p.venue === v).length,
Line 659: Line 667:
 2026 *  6                      1 2026 *  6                      1
 * 2025-2026 are provisional venue-years — see literature:corpus. * 2025-2026 are provisional venue-years — see literature:corpus.
 +
 +2025-2026: 17 of 38 candidates (44.7%) against 1185 of 5859 papers in the corpus (20.2%)
 +  of those 17 candidates, 1 measure age assurance and 5 were obstructed by it
  
 venue    candidates  measures it venue    candidates  measures it
Line 1051: Line 1062:
   'risk-based, flexible, tech-neutral and future-proof': 'aa/ico_joint.txt',   'risk-based, flexible, tech-neutral and future-proof': 'aa/ico_joint.txt',
   'laid before the end of the year, and the changes should be implemented in Spring 2027': 'aa/gov_factsheet.txt',   'laid before the end of the year, and the changes should be implemented in Spring 2027': 'aa/gov_factsheet.txt',
 +  'should be implemented in Spring 2027': 'aa/gov_factsheet.txt',
   'serious doubts': 'aa/ofcom2026.txt',   'serious doubts': 'aa/ofcom2026.txt',
   'Almost all analysed pornography services relied exclusively on third-party vendors, with only one analysed pornography service using an in-house solution': 'aa/ofcom2026.txt',   'Almost all analysed pornography services relied exclusively on third-party vendors, with only one analysed pornography service using an in-house solution': 'aa/ofcom2026.txt',
Line 1238: Line 1250:
 [ ok ] when creating an account, minors below 13 can enter a false birth date  — aa/ec_meta.txt [ ok ] when creating an account, minors below 13 can enter a false birth date  — aa/ec_meta.txt
 [ ok ] laid before the end of the year, and the changes should be implemented  — aa/gov_factsheet.txt [ ok ] laid before the end of the year, and the changes should be implemented  — aa/gov_factsheet.txt
 +[ ok ] should be implemented in Spring 2027  — aa/gov_factsheet.txt
 [ ok ] serious doubts  — aa/ofcom2026.txt [ ok ] serious doubts  — aa/ofcom2026.txt
  
-48 spans: 8 not-a-quote, 16 external, 24 paper; 0 failed, 2 located in a paper other than the nearest citekey+49 spans: 8 not-a-quote, 17 external, 24 paper; 0 failed, 2 located in a paper other than the nearest citekey
 control ok: a fabricated span is not located control ok: a fabricated span is not located
 control ok: a real span resolves to its own paper, not to the other key in the block control ok: a real span resolves to its own paper, not to the other key in the block
Line 1355: Line 1368:
   - **The roadmap's own committed probe has the same defect** and is the reason //F-BLEAU: Fast Black-Box Leakage Estimation// is one of its 11 age-assurance candidates.   - **The roadmap's own committed probe has the same defect** and is the reason //F-BLEAU: Fast Black-Box Leakage Estimation// is one of its 11 age-assurance candidates.
   - **The first quote check failed on a true quote.** {[west2024_picture]}'s contribution sentence is spliced across a column boundary in all three ''.txt'' renderings; ''pypdf'' has it verbatim. The checker was rewritten to try four renderings and to print which one located each needle, and a fabricated-needle control was added so the check cannot pass vacuously.   - **The first quote check failed on a true quote.** {[west2024_picture]}'s contribution sentence is spliced across a column boundary in all three ''.txt'' renderings; ''pypdf'' has it verbatim. The checker was rewritten to try four renderings and to print which one located each needle, and a fabricated-needle control was added so the check cannot pass vacuously.
-  - **The bibliography cache served a stale parse.** After appending 11 entries and saving the page, 11 of 20 references rendered as allocated-but-empty numbers while both sources looked perfect. Purging ''literature/bibliography?purge=true'' and then the page fixed it; the rendered reference count was then checked against the distinct marker count (20 = 20, 42 markers, 42 ''bibtex_citekey'' spans).+  - **The bibliography cache served a stale parse.** After appending 11 entries and saving the page, 11 of 20 references rendered as allocated-but-empty numbers while both sources looked perfect. Purging ''literature/bibliography?purge=true'' and then the page fixed it. The check that matters is **distinct citekeys against rendered references**: 20 = 20. The marker-to-span ratio is not a check — the plugin emits **two** ''bibtex_citekey'' spans per marker, so the page's 55 markers render as 110 spans, and an earlier note here that read "42 markers, 42 spans" was counting one of them with a regex that only matched the opening span.
   - **ofcom.org.uk is unreachable from this sandbox** (403 to curl, to WebFetch, and to a full headless-Chromium context). The January 2025 quotes come from an Internet Archive capture; the July 2026 report was obtained from the UK government's asset host, which is not blocked. The page says so in both footnotes rather than implying a direct read.   - **ofcom.org.uk is unreachable from this sandbox** (403 to curl, to WebFetch, and to a full headless-Chromium context). The January 2025 quotes come from an Internet Archive capture; the July 2026 report was obtained from the UK government's asset host, which is not blocked. The page says so in both footnotes rather than implying a direct read.
 +  - **A de-hyphenation artefact was published as a defect in someone else's paper.** The artefacts table asserted that the repository URL printed in {[moti2024_targeted]} 404s. It does not. The paper prints ''https://github.com/targeted-and-troublesome/'', broken across a line at ''targeted-and-''; ''paper.cols.txt'', ''paper.norm.txt'' and the extraction's ''artifacts.codeUrl'' all rejoin it as ''targeted-andtroublesome'', which is what 404s. Only ''paper.txt'', which preserves the line break, shows the truth. The generic reviewer found it by reading all three renderings. **The tooling's own artefact was published as their error**, which is the second time in this run that an ENOENT or a mangled string was read as a fact about someone else's work.
   - **A mistyped slug produced a published falsehood about the corpus itself.** An early context pull used ''…privacy-preserving-facial-transformation**s**'' where the directory is singular. The ENOENT was read as "this paper has no full text", and that became an ''ARTEFACT'' verdict, a script comment, a provenance bullet and a limitation on the content page saying one candidate had been judged from its title alone. The file is there, 102 KB; the probe read it; its five matches are //age estimation// as the name of a face-attribute ML task. The generic reviewer found it by checking the claim against the script's **own printed list** of the four papers without full text — GAN-Invert is not on it. The verdict is now ''MENTION'' (20/7) and the limitation is deleted. Nothing about the population changed, but four separate places had repeated the same unchecked inference.   - **A mistyped slug produced a published falsehood about the corpus itself.** An early context pull used ''…privacy-preserving-facial-transformation**s**'' where the directory is singular. The ENOENT was read as "this paper has no full text", and that became an ''ARTEFACT'' verdict, a script comment, a provenance bullet and a limitation on the content page saying one candidate had been judged from its title alone. The file is there, 102 KB; the probe read it; its five matches are //age estimation// as the name of a face-attribute ML task. The generic reviewer found it by checking the claim against the script's **own printed list** of the four papers without full text — GAN-Invert is not on it. The verdict is now ''MENTION'' (20/7) and the limitation is deleted. Nothing about the population changed, but four separate places had repeated the same unchecked inference.
   - **The page's most important figures were the ones a first draft got wrong, and neither was caught by a guard.** The 2019 age-verification percentages were published against the wrong denominator — the paper's 6,843-site corpus and "six vantage points", where the section itself is a hand check of at most fifty sites in four countries — because the quote-check located the sentence and nothing checked what the sentence was a share //of//. And the whole 2026 UK deployment picture was missing, because the page was written from the corpus and the corpus stops at seven academic venues. Both came from reviewers.   - **The page's most important figures were the ones a first draft got wrong, and neither was caught by a guard.** The 2019 age-verification percentages were published against the wrong denominator — the paper's 6,843-site corpus and "six vantage points", where the section itself is a hand check of at most fifty sites in four countries — because the quote-check located the sentence and nothing checked what the sentence was a share //of//. And the whole 2026 UK deployment picture was missing, because the page was written from the corpus and the corpus stops at seven academic venues. Both came from reviewers.
Line 1405: Line 1419:
 | The page claims //"a verbatim-quote check of every per-paper number"//; eleven numbers sit outside every needle and a mutation of 14% to 24% passed both guards. The reviewer hand-verified all eleven and they are correct. | **Accepted.** The methodology section now names the eleven figures that are hand-checked rather than guarded. | | The page claims //"a verbatim-quote check of every per-paper number"//; eleven numbers sit outside every needle and a mutation of 14% to 24% passed both guards. The reviewer hand-verified all eleven and they are correct. | **Accepted.** The methodology section now names the eleven figures that are hand-checked rather than guarded. |
 | Missing practical point in the ethics section: verification vendors expose sandbox modes. | **Rejected for now.** No primary source was reached for which vendors do, and the page does not name vendors at all; asserting it would be exactly the vendor-marketing claim the source policy on this page rejects. Recorded here so the next run can close it. | | Missing practical point in the ethics section: verification vendors expose sandbox modes. | **Rejected for now.** No primary source was reached for which vendors do, and the page does not name vendors at all; asserting it would be exactly the vendor-marketing claim the source policy on this page rejects. Recorded here so the next run can close it. |
-| No neighbouring page links back to this one. | **Accepted in principle, not done in this sitting.** Adding a row to four neighbours' cross-reference tables is a separate edit against four pages whose own figures are current; doing it here risks the stale-snapshot failure this run has already paid for onceFiled as follow-up work. |+| No neighbouring page links back to this one. | **Accepted and done**, as a separate edit after the page was frozen for the re-review: one line each in [[:design:blocking_and_geodifference]], [[:design:crawling_location]], [[:privacy:consent]] and [[:practices:ethics]]No figures on those pages were touched. |
 | Several phrases repeat (//at most fifty// five times, //one of its two case studies// three times). | **Partly accepted.** Two instances trimmed; the rest carry the caveat in places a reader may arrive at directly. | | Several phrases repeat (//at most fifty// five times, //one of its two case studies// three times). | **Partly accepted.** Two instances trimmed; the rest carry the caveat in places a reader may arrive at directly. |
 | Nits: {[west2024_picture]} uses Frida rather than the camera; the 5,855-apps / 5,855-papers coincidence; an unsupported causal claim about the six obstacle papers; the unused //"we did not find any instance of AgeID being deployed"//. | **All four accepted.** | | Nits: {[west2024_picture]} uses Frida rather than the camera; the 5,855-apps / 5,855-papers coincidence; an unsupported causal claim about the six obstacle papers; the unused //"we did not find any instance of AgeID being deployed"//. | **All four accepted.** |
 +
 +Both passes were re-run once more against the frozen pages. What they found the second time:
 +
 +^ Finding ^ Disposition ^
 +| **The artefacts section published a defect in someone else's paper that does not exist.** {[moti2024_targeted]} prints its repository URL correctly; ''paper.cols.txt'', ''paper.norm.txt'' and the extraction all de-hyphenate it across a line break into a URL that 404s. | **Accepted — the worst finding of the second round**, and recorded in the mistakes list above. The row now states the correct URL, its HTTP 200, and the artefact that caused the confusion. |
 +| The Zenodo record for GUARD is titled //Drexel-SePAL/AgeScope: v1.0.1-1//, so searching for "GUARD" will not find it. | **Accepted.** The row now says so. |
 +| Three unbounded negatives introduced by the new artefacts section — //"there is no dataset"//, //"nobody has published one"//, and an ordering sentence attributing a sort to Ofcom that Ofcom does not perform. | **Accepted, all three.** |
 +| The closing argument said a crawl against adult sites would measure //the smaller half//, which nothing on the page sizes, and asserted the Spring 2027 date the fact sheet hedges. | **Accepted.** Both softened to what the sources support. |
 +| The intro's verdict clause attached //"the name of an unrelated ML task"// to the 7 ''ARTEFACT'' papers, where it describes the 20th ''MENTION''. | **Accepted.** Reworded. |
 +| The methodology section's list of eleven hand-checked figures is incomplete — the artefact file sizes, Ofcom's 25%-to-43% and 64-of-100, the survey base of 50 and several paraphrased denominators are in the same position. | **Accepted, and the enumeration was abandoned.** It was incomplete twice; the page now states the class instead, which is both shorter and true. |
 +| Three stale numbers in this provenance page: the "most common way it appears in the corpus" phrasing, an unsourced "194 phrase matches", and a marker-to-span identity that was never consistent. | **Accepted, all three.** The last one is worth keeping in mind: the bibtex plugin emits **two** spans per marker, so marker count and span count are not a check. |
 +| Scripts reproduce their committed outputs byte-for-byte; the verdict split, 44.7%/20.2%, the four artefact URLs, 55,481 bytes, 2,004 distinct URLs and the 2025-02-04 push date all independently verified. | **Accepted as a pass.** |
 +| The two ''WARN'' spans are correctly attributed; the log did not name them. | **Accepted.** They are //"Despite being rated as 17+…"// (nearest key ''vallina2019_porn'', found in ''yao2025_easy'') and //"seven out of 20 apps…"// (nearest key ''ardi2023_prevalence'', found in ''moti2025_whispertest''). In both, the previous bullet's trailing citekey is closer in characters than the bullet's own key at its end. A third ''WARN'' would be a new thing to look at. |
  
  
provenance/privacy/age_assurance.1789492475.txt.gz · Last modified: by karel.kubicek.claude