provenance:programming:interaction
Differences
This shows you the differences between two versions of the page.
| Both sides previous revisionPrevious revisionNext revision | Previous revision | ||
| provenance:programming:interaction [2026/08/27 12:32] – Add the self-review sweep for absolute claims (five overstatements found and fixed between review rounds, two of which passed the number guard), and refresh the embedded report output. Authored by Claude karel.kubicek.claude | provenance:programming:interaction [2026/08/27 12:52] (current) – Name the two deferred follow-ups by their drain item keys. Authored by Claude karel.kubicek.claude | ||
|---|---|---|---|
| Line 10: | Line 10: | ||
| | Date | 2026-08-27 | | | Date | 2026-08-27 | | ||
| | Corpus at the time | '' | | Corpus at the time | '' | ||
| - | | Page before | a 2,411-byte stub: a '' | + | | Page before | a 2,411-byte stub: a %%<WRAP todo>%% box, three bullet topics, and one claim about Urban et al. | |
| - | | Page after | 50,697 bytes, rev 1787833934 | + | | Page after | 56,675 bytes, rev 1787834365 — eight saves: the initial publish, one per review round, and five self-review corrections |
| | Model | Claude Opus 5, one session, no sub-agents used for research | | | Model | Claude Opus 5, one session, no sub-agents used for research | | ||
| | Sub-agents | four reviewers only (see [[# | | Sub-agents | four reviewers only (see [[# | ||
| Line 63: | Line 63: | ||
| </ | </ | ||
| - | <file text report_interaction-output.txt> | + | **The full unedited output |
| - | + | ||
| - | ============================================================================== | + | |
| - | A. POPULATION | + | |
| - | ============================================================================== | + | |
| - | corpus | + | |
| - | crawled | + | |
| - | webCrawled | + | |
| - | of which carry a crawlConfig object | + | |
| - | crawled but NOT web (excluded: app stores, network scans, social graphs) | + | |
| - | + | ||
| - | ============================================================================== | + | |
| - | B. interactionDepth — what the crawling literature says about depth | + | |
| - | ============================================================================== | + | |
| - | + | ||
| - | ── crawled — 1120 papers ── | + | |
| - | interactionDepth | + | |
| - | ----------------------- | + | |
| - | single-target-page | + | |
| - | not-stated | + | |
| - | landing-page-only | + | |
| - | deep-crawl | + | |
| - | landing-plus-subpages | + | |
| - | (no crawlConfig object) | + | |
| - | stated (any non-sentinel value): 841 / 1120 = 75.1% | + | |
| - | + | ||
| - | ── webCrawled — 857 papers ── | + | |
| - | interactionDepth | + | |
| - | ----------------------- | + | |
| - | single-target-page | + | |
| - | landing-page-only | + | |
| - | landing-plus-subpages | + | |
| - | not-stated | + | |
| - | deep-crawl | + | |
| - | (no crawlConfig object) | + | |
| - | stated (any non-sentinel value): 702 / 857 = 81.9% | + | |
| - | + | ||
| - | ── Where interactionDepth ranks among the crawl-configuration fields ── | + | |
| - | crawlConfig field papers stating it share of 1120 crawling papers | + | |
| - | ----------------- | + | |
| - | interactionDepth | + | |
| - | authentication | + | |
| - | browsers | + | |
| - | consentAction | + | |
| - | statefulness | + | |
| - | headless | + | |
| - | gap between the best-reported field and the second: 62 papers, 5.5 percentage points. | + | |
| - | + | ||
| - | ── The depth axis proper — webCrawled papers giving one of the three SITE-depth values ── | + | |
| - | denominator: | + | |
| - | depth papers | + | |
| - | --------------------- | + | |
| - | landing-page-only | + | |
| - | landing-plus-subpages | + | |
| - | deep-crawl | + | |
| - | went beyond the landing page: 262 / 417 = 62.8% | + | |
| - | CAVEAT: papers that went deeper have more reason to say so, so this ratio is | + | |
| - | an upper bound on the share of the whole field that crawls beyond the landing page. | + | |
| - | + | ||
| - | ============================================================================== | + | |
| - | C. Reporting rate over time (webCrawled, | + | |
| - | ============================================================================== | + | |
| - | bucket | + | |
| - | ---------- | + | |
| - | 2010–2013 | + | |
| - | 2014–2017 | + | |
| - | 2018–2021 | + | |
| - | 2022–2024 | + | |
| - | 2025–2026* | + | |
| - | + | ||
| - | * 2025–2026 is provisional: | + | |
| - | IEEE S&P 2026 / WWW 2026 abstracts are incompletely indexed, so selection under-covers them. | + | |
| - | + | ||
| - | ============================================================================== | + | |
| - | D. The label discriminant — does interactionDepth track the paper' | + | |
| - | ============================================================================== | + | |
| - | crawlConfig carries ONE evidence quote for the WHOLE object, so reading it cannot | + | |
| - | validate interactionDepth. Instead: does the paper contain a first-person sentence | + | |
| - | naming the ROOT of a site as the unit of a visit (LANDING), something BELOW the root | + | |
| - | (DEEPER), both, or neither? `not-stated` | + | |
| - | + | ||
| - | label papers with text landing phrase only deeper phrase only both neither | + | |
| - | --------------------- | + | |
| - | landing-page-only | + | |
| - | landing-plus-subpages | + | |
| - | deep-crawl | + | |
| - | single-target-page | + | |
| - | not-stated | + | |
| - | + | ||
| - | ── What the shared crawlConfig quote actually evidences ── | + | |
| - | pool: 71 web crawls labelled landing-plus-subpages that also state a subpage count. | + | |
| - | every 6th read by hand on 2026-08-27 = 12 papers. | + | |
| - | what the ONE shared quote evidences | + | |
| - | ----------------------------------- | + | |
| - | depth 6 | + | |
| - | other 5 | + | |
| - | partial | + | |
| - | | + | |
| - | " | + | |
| - | | + | |
| - | "For each site, we visit the top 10 pages returned from a site-specific search on Google" | + | |
| - | [other] IMC/ | + | |
| - | "we performed four crawls over our sampled 100K websites" | + | |
| - | [other] WWW/ | + | |
| - | "We did not erase any cookie after a harvest was performed" | + | |
| - | [other] PETS/ | + | |
| - | "from US-based EC2 cloud instances using stateless browsers" | + | |
| - | [partial] NDSS/ | + | |
| - | "We visit each site for a total of eight times ... four page visits per case" — visit structure, not site depth | + | |
| - | [depth] USENIX/ | + | |
| - | "we therefore additionally selected three random links to an internal subpage and visited these as well" | + | |
| - | [other] IMC/ | + | |
| - | "We choose to use a stateless approach" | + | |
| - | [depth] PETS/ | + | |
| - | " | + | |
| - | [depth] PETS/ | + | |
| - | "We added ten random sub-pages per domain" | + | |
| - | [depth] NDSS/ | + | |
| - | " | + | |
| - | [other] PETS/ | + | |
| - | " | + | |
| - | Read this as the size of the blindness, not as an error rate: a quote that | + | |
| - | evidences statefulness does not make the depth label wrong, it makes the | + | |
| - | quote useless as a check on it. | + | |
| - | + | ||
| - | ── Internal consistency: does subpagesPerSite ever contradict interactionDepth? | + | |
| - | 104 web crawls state BOTH an interactionDepth and a subpagesPerSite. | + | |
| - | contradictions (landing-page-only with n>0, or landing-plus-subpages with n=0): 0 | + | |
| - | This is a corroborating check on the enum, independent of the text discriminant above: | + | |
| - | the two fields are extracted from the same shared quote but mean different things, | + | |
| - | so a systematically wrong depth label would show up here as disagreement. | + | |
| - | + | ||
| - | ── Hand audit of the `landing-page-only` " | + | |
| - | 42 papers in the cell; every 4th read by hand on 2026-08-27 = 11 papers. | + | |
| - | verdict | + | |
| - | ---------------- | + | |
| - | inferred | + | |
| - | stated-elsewhere | + | |
| - | [inferred] IMC/ | + | |
| - | "For each domain, we retrieve the complete web page" — the unit is the domain, the page is never named | + | |
| - | [inferred] IMC/ | + | |
| - | "we crawled 6,755 unique phishing URLs" — a URL list, not a site depth | + | |
| - | [inferred] IMC/ | + | |
| - | "We visited each domain in our target list 5 times" — domain granularity only | + | |
| - | [stated-elsewhere] WWW/ | + | |
| - | "the Alexa top-200K websites' | + | |
| - | [inferred] WWW/ | + | |
| - | "the Alexa top 10K websites that we crawled" | + | |
| - | [inferred] IMC/ | + | |
| - | "we accessed each website five times using curl" — site granularity only | + | |
| - | [inferred] CCS/ | + | |
| - | "We crawl all the content within the API listing page" — a listing page, not a site root | + | |
| - | [inferred] WWW/ | + | |
| - | no first-person sentence names the page unit at all | + | |
| - | [inferred] USENIX/ | + | |
| - | no first-person sentence names the page unit at all | + | |
| - | [inferred] WWW/ | + | |
| - | "We access their web interfaces ... comparable to how an ordinary user visits a public site" | + | |
| - | [inferred] IMC/ | + | |
| - | "We visit the site with a headless browser" | + | |
| - | + | ||
| - | ============================================================================== | + | |
| - | E. subpagesPerSite — the number people actually pick | + | |
| - | ============================================================================== | + | |
| - | denominator: | + | |
| - | of the 417 on the site-depth axis, 102 (24.5%) give a number. | + | |
| - | subpages per site papers | + | |
| - | --------------------- | + | |
| - | 0 (landing page only) 4 | + | |
| - | 1–4 29 27.9% | + | |
| - | 5–9 15 14.4% | + | |
| - | 10–19 | + | |
| - | 20–49 | + | |
| - | 50–99 | + | |
| - | 100+ | + | |
| - | per four-year bucket: | + | |
| - | 2010–2013 | + | |
| - | 2014–2017 | + | |
| - | 2018–2021 | + | |
| - | 2022–2024 | + | |
| - | 2025–2026* | + | |
| - | median 10; the five most common values: | + | |
| - | 5 subpages: 14 papers | + | |
| - | 3 subpages: 12 papers | + | |
| - | 10 subpages: 12 papers | + | |
| - | 1 subpages: 11 papers | + | |
| - | 20 subpages: 8 papers | + | |
| - | the largest values (a "deep crawl" of one application, | + | |
| - | 2000: 1 papers | + | |
| - | 1000: 3 papers | + | |
| - | 500: 2 papers | + | |
| - | 300: 1 papers | + | |
| - | 200: 3 papers | + | |
| - | 100: 7 papers | + | |
| - | + | ||
| - | ============================================================================== | + | |
| - | F. repeatVisits and authentication — the other two interaction | + | |
| - | ============================================================================== | + | |
| - | repeatVisits stated: 233 / 857 web crawls = 27.2% | + | |
| - | of those, 50 (21.5%) visit exactly once; 183 more than once. | + | |
| - | + | ||
| - | ── authentication — of 857 web crawls ── | + | |
| - | authentication | + | |
| - | ----------------------- | + | |
| - | none | + | |
| - | not-stated | + | |
| - | account-registration | + | |
| - | manual-login | + | |
| - | (no crawlConfig object) | + | |
| - | automated-login | + | |
| - | crawls that got past a login of any kind: 81 (9.5%) | + | |
| - | sso in THIS population: 0; sso anywhere in the corpus: 1 (USENIX/ | + | |
| - | + | ||
| - | ============================================================================== | + | |
| - | G. What the crawler does ON the page — full-text probe | + | |
| - | ============================================================================== | + | |
| - | The schema has NO field for scrolling, clicking, hovering, typing or waiting. | + | |
| - | These are UPPER BOUNDS on "the paper did this": a first-person sentence matching | + | |
| - | the pattern. "we did not scroll" | + | |
| - | + | ||
| - | papers | + | |
| - | + | ||
| - | what the paper says it does papers | + | |
| - | ------------------------------------- | + | |
| - | clicks something | + | |
| - | scrolls | + | |
| - | waits / dwells a stated time | + | |
| - | types on the keyboard | + | |
| - | fills a form | + | |
| - | moves the mouse 18 2.1% | + | |
| - | hovers | + | |
| - | picks links at random | + | |
| - | says it aims for human-like behaviour | + | |
| - | mentions bot / crawler detection | + | |
| - | + | ||
| - | bot/crawler detection, 2010-2013 bucket: 3 of 80 papers match ANYWHERE in the text, 0 in a first-person sentence. | + | |
| - | + | ||
| - | Same, by four-year bucket (share of that bucket' | + | |
| - | pattern | + | |
| - | ------------------------------------- | + | |
| - | clicks something | + | |
| - | scrolls | + | |
| - | waits / dwells a stated time 0/80 0.0% 12/130 9.2% | + | |
| - | types on the keyboard | + | |
| - | fills a form 1/80 1.3% 1/130 0.8% 8/241 3.3% 10/253 4.0% 8/153 5.2% | + | |
| - | moves the mouse 3/80 3.8% 2/130 1.5% 7/241 2.9% 4/253 1.6% 2/153 1.3% | + | |
| - | hovers | + | |
| - | picks links at random | + | |
| - | says it aims for human-like behaviour | + | |
| - | mentions bot / crawler detection | + | |
| - | + | ||
| - | ============================================================================== | + | |
| - | H. LLM-agent-driven browsing — sweep, then hand verdicts | + | |
| - | ============================================================================== | + | |
| - | sweep hits over the 857 web crawls: 35 | + | |
| - | NOTE: 1 hand verdicts are not returned by the current sweep | + | |
| - | | + | |
| - | role papers of 35 | + | |
| - | -------------- | + | |
| - | not-browsing | + | |
| - | subject | + | |
| - | citation | + | |
| - | instrument | + | |
| - | captcha-solver | + | |
| - | instrument-app | + | |
| - | + | ||
| - | An LLM agent drove the browsing in 5 of 857 web crawls (0.6%). | + | |
| - | | + | |
| - | " | + | |
| - | | + | |
| - | "a multi-agent system powered by large language models (LLMs) to simulate persona-driven browsing behavior" | + | |
| - | | + | |
| - | "we design an LLM-based navigation pipeline tailored to perform privacy-related measurements in 200 apps" | + | |
| - | | + | |
| - | "We built a prototype tool using Playwright and the agentic LLM-based Browser Use framework" | + | |
| - | | + | |
| - | "We design and deploy an LLM-driven auditing agent capable of end-to-end traversal of rights-request workflows" | + | |
| - | | + | |
| - | + | ||
| - | ============================================================================== | + | |
| - | I. The measured consequences — per-paper figures with their own denominators | + | |
| - | ============================================================================== | + | |
| - | + | ||
| - | WWW/ | + | |
| - | * subsites set 36% more cookies than landing pages: 78 vs 55 on average, over the top 10k websites (TLD+1), 100 subsites each | + | |
| - | * the mean number of accessed/ | + | |
| - | * trackers (EasyPrivacy) increased ~6% on subsites; 2.5% of sites tracked ONLY on subsites | + | |
| - | * Fingerprint2 device fingerprinting increased 25% on subsites; present on 0.15% of landing pages | + | |
| - | + | ||
| - | USENIX/ | + | |
| - | * same Alexa top-10K, same cookie policy: homepage crawl 192,038 requests / 76,816 chains; interactive crawl (random internal pages via clicks on iframes and anchors) 575,550 / 229,151 — 3.0x | + | |
| - | * 302 redirects are 42.91% of AT redirect requests on the homepage crawl vs 28.56% interactive | + | |
| - | * AT requests navigating to a new domain: 51.32% homepage vs 47.49% interactive | + | |
| - | + | ||
| - | USENIX/ | + | |
| - | * front pages plus "three random links to an internal subpage": | + | |
| - | | + | |
| - | * per technique, subpage-only increase: ShortCut +22%, TrigBreak +80%, ModBuilt +18%, WidthDiff +19%, LogGet +33% — the aggregate hides an 80% | + | |
| - | + | ||
| - | NDSS/ | + | |
| - | * against three state-of-the-art scanners on ten web applications: | + | |
| - | + | ||
| - | PETS/ | + | |
| - | * 456 data-broker websites; verified workflow completion 87% in Phase 2 and 79% in Phase 3 | + | |
| - | + | ||
| - | IMC/ | + | |
| - | * clickstream traversal, and its own bias statement: "our dataset is biased towards static inner pages ... we are less likely to explore the more dynamic areas of a website" | + | |
| - | + | ||
| - | NDSS/ | + | |
| - | * the reason a landing-page-only design is chosen: "We only crawl the homepage of each visited site due to the presence of many sites that thwart deeper traversal by requiring log-ins." | + | |
| - | + | ||
| - | USENIX/ | + | |
| - | * a fully specified depth: " | + | |
| - | + | ||
| - | USENIX/ | + | |
| - | * typing is simulated against bot detection: "we simulate user typing behavior by using randomized intervals for each key press and dwell times, as well as the delay times between each press" | + | |
| - | + | ||
| - | CCS/ | + | |
| - | * what stops an interacting crawler: "In 22 cases, there was some form of an anti-bot challenge that our system was not able to solve and, thus, could not proceed with registration." | + | |
| - | + | ||
| - | PETS/ | + | |
| - | * the most complete interaction statement | + | |
| - | + | ||
| - | PETS/ | + | |
| - | * a stated subpage-selection rule: "We added ten random sub-pages per domain, filtering to exclude auxiliary pages like privacy policies or contact pages." | + | |
| - | + | ||
| - | ============================================================================== | + | |
| - | Y. Figures the page quotes from other papers or other pages | + | |
| - | ============================================================================== | + | |
| - | statefulness stated: 219 / 1120 crawling papers = 19.6% | + | |
| - | (the page cites this as the comparison for how well depth is reported) | + | |
| - | + | ||
| - | Zeber et al., TheWebConf 2020, "The Representativeness of Automated Web Crawls as a | + | |
| - | Surrogate for Human Browsing" | + | |
| - | data/ | + | |
| - | on 2026-08-27, all four verbatim: | + | |
| - | "over 50,000 users of the Firefox Web browser"; | + | |
| - | " | + | |
| - | | + | |
| - | "The median number of tracking domains accessed by a user on visiting a Trexa list site is | + | |
| - | 1.9, whereas for the crawler it is 6.1" | + | |
| - | + | ||
| - | Urban et al.'s site list: "the top 10k websites" | + | |
| - | opening box is a rhetorical example of a methods sentence, not a corpus figure. | + | |
| - | + | ||
| - | NOT FIGURES — literals inside the fixture the page publishes in < | + | |
| - | scanned only under check_page_numbers.mjs --code. Listed here so they trace to | + | |
| - | this report instead of polluting the guard' | + | |
| - | 127.0.0.1 | + | |
| - | 127.0.0.2 | + | |
| - | 8231 the first-party port | + | |
| - | 8232 the third-party port | + | |
| - | 3000 the CSS pixel height of the tall div that makes scrolling necessary | + | |
| - | 50 the scroll-bottom tolerance in pixels | + | |
| - | 200, 300 | + | |
| - | 404, 200 HTTP status codes the fixture writes | + | |
| - | 2 the maximum crawl depth in the last strategy | + | |
| - | + | ||
| - | Tool versions in the measured probe, printed by the probe itself | + | |
| - | (sandbox/ | + | |
| - | + | ||
| - | + | ||
| - | ============================================================================== | + | |
| - | Z. NON-CORPUS NUMBERS ON THE PAGE, with their primary source | + | |
| - | ============================================================================== | + | |
| - | + | ||
| - | Aqeel et al., IMC 2020, "On Landing and Internal Web Pages" — NOT in the corpus | + | |
| - | (the selection stage labelled it neither a security nor a privacy measurement; | + | |
| - | see data/ | + | |
| - | balakrishnanc.github.io/ | + | |
| - | 119 web-performance publications reviewed (IMC/ | + | |
| - | 41 (34.5%) need no revision, 48 (40.3%) minor, 30 (25.2%) major -> 65.5% need at least minor | + | |
| - | landing pages are on average 34% larger than internal pages (geometric mean of ratios, H1K) | + | |
| - | for 32% of H1K sites the landing page has FEWER objects than the median internal page | + | |
| - | internal pages' | + | |
| - | median: internal pages collectively fetch from 18 third-party domains never seen on the | + | |
| - | landing page; for 10% of H1K sites, 80 or more | + | |
| - | at the 80th percentile, internal pages carry 20 tracking requests and landing pages 28 | + | |
| - | in about 10% of H1K sites, internal pages have no trackers while the landing page does | + | |
| - | header bidding, of 200 sites: 17 have HB ads on the landing page, a further 12 only on internal pages | + | |
| - | 36 of the 1000 H1K sites serve their landing page over HTTP; among sites with a secure | + | |
| - | landing page, 170 have at least one HTTP internal page | + | |
| - | Hispar: H2K = 100,000 URLs, >=2000 sites x 50 URLs (1 landing + <=49 internal), weekly refresh, | + | |
| - | bootstrapped from Alexa Top 1M via Google " | + | |
| - | + | ||
| - | Hispar is dead. Checked 2026-08-27: hispar.cs.duke.edu does not resolve (DNS NXDOMAIN); | + | |
| - | last Wayback capture that returned content 2024-11-16 (HTTP 200, 1,420 bytes); last capture | + | |
| - | of any kind 2024-12-03 (HTTP 403); github.com/ | + | |
| - | github.com/ | + | |
| - | retired in 2022. | + | |
| - | + | ||
| - | HTTP Archive crawls exactly ONE secondary page per site, since April 2022. | + | |
| - | har.fyi (HTTP Archive' | + | |
| - | monthly basis and as of April 2022, both the root page and one secondary page are tested." | + | |
| - | The pages table carries is_root_page and root_page columns. | + | |
| - | How the secondary page is chosen, from github.com/ | + | |
| - | (2025-08-20), | + | |
| - | child job is the FIRST link in the page's crawl_links whose hostname equals the parent' | + | |
| - | whose extension is not in [' | + | |
| - | ' | + | |
| - | NOTE: httparchive.org/ | + | |
| - | does not crawl the website' | + | |
| - | the crawl controller agree with each other and not with it. Checked 2026-08-27. | + | |
| - | </ | + | |
| ===== The full-text probe for on-page actions ===== | ===== The full-text probe for on-page actions ===== | ||
| Line 472: | Line 69: | ||
| The schema has **no field** for scrolling, clicking, hovering, typing or waiting. The only way to count them is to sweep the text, and the count is an **upper bound** by construction: | The schema has **no field** for scrolling, clicking, hovering, typing or waiting. The only way to count them is to sweep the text, and the count is an **upper bound** by construction: | ||
| - | Two columns are printed side by side so the width of the claim is visible: //any match anywhere in the paper// against //a match in a sentence that also contains a first-person marker//. The gap is large — 469 papers mention clicking, | + | Two columns are printed side by side so the width of the claim is visible: //any match anywhere in the paper// against //a match in a sentence that also contains a first-person marker//. The gap is large — **469** papers mention clicking, |
| - | + | ||
| - | <file text interaction_fulltext_probe-output.txt> | + | |
| - | population: --pop web → 857 papers, 857 of them with full text on disk | + | |
| - | + | ||
| - | pattern | + | |
| - | ------------------------------- | + | |
| - | scroll | + | |
| - | click 469 | + | |
| - | hover 48 5.6% | + | |
| - | mouse movement | + | |
| - | keyboard / typing | + | |
| - | form fill 59 6.9% | + | |
| - | dwell / wait | + | |
| - | random walk / link following | + | |
| - | subpage / subsite | + | |
| - | landing page only 22 2.6% | + | |
| - | human-like / realistic browsing | + | |
| - | LLM / agent-driven browsing | + | |
| - | bot / crawler detection | + | |
| - | cloaking | + | |
| - | + | ||
| - | Both columns are UPPER BOUNDS on "the paper did this" | + | |
| - | Read the sentences with --hits "< | + | |
| + | <WRAP important> | ||
| + | **These two scripts used to disagree, and both numbers were published.** Until 2026-08-27 '' | ||
| - | First-person-sentence hits by four-year bucket (share of that bucket's papers): | + | Fixed at the source rather than by reconciling the prose: '' |
| - | pattern | + | </WRAP> |
| - | ------------------------------- | + | |
| - | scroll | + | |
| - | click 35/80 43.8% 41/130 31.5% 72/241 29.9% 81/253 32.0% 56/153 36.6% | + | |
| - | hover 2/80 2.5% 0/130 0.0% 6/241 2.5% 4/253 1.6% 2/153 1.3% | + | |
| - | mouse movement | + | |
| - | keyboard / typing | + | |
| - | form fill 1/80 1.3% 1/130 0.8% 8/241 3.3% 10/253 4.0% 8/153 5.2% | + | |
| - | dwell / wait 0/80 0.0% 12/130 9.2% | + | |
| - | random walk / link following | + | |
| - | subpage / subsite | + | |
| - | landing page only 0/80 0.0% 5/130 3.8% 2/241 0.8% 6/253 2.4% 1/153 0.7% | + | |
| - | human-like / realistic browsing | + | |
| - | LLM / agent-driven browsing | + | |
| - | bot / crawler detection | + | |
| - | cloaking | + | |
| - | * 2025-2026 | + | **The full unedited output |
| - | </ | + | |
| Read individual sentences behind any row with: | Read individual sentences behind any row with: | ||
| Line 528: | Line 88: | ||
| ===== Folding, hand verdicts, and the residue ===== | ===== Folding, hand verdicts, and the residue ===== | ||
| - | There is **no name fold on this page**. Nothing it counts is a free-text name: '' | + | There is **no name fold on this page**. Nothing it counts is a free-text name: '' |
| ==== 1. LLM-agent roles ==== | ==== 1. LLM-agent roles ==== | ||
| Line 598: | Line 158: | ||
| The page publishes a measurement rather than an assertion: a local instrumented site with five third-party beacons, each reachable only under a different condition, visited under six strategies by Playwright' | The page publishes a measurement rather than an assertion: a local instrumented site with five third-party beacons, each reachable only under a different condition, visited under six strategies by Playwright' | ||
| - | <file text interaction_probe-output.txt> | + | **The full unedited |
| - | strategy | + | |
| - | ------------------------------------------------------------------- | + | |
| - | landing page only 1 ✓ ✗ ✗ ✗ ✗ | + | |
| - | landing + FIRST same-origin link (the HTTP Archive rule) | + | |
| - | landing + ALL same-origin links from the landing | + | |
| - | landing + all links, and scroll to the bottom of each 4 ✓ ✓ ✓ ✗ ✗ | + | |
| - | landing + all links, scroll, and click every button | + | |
| - | depth 2: landing, its links, and their links, with scroll and click 5 ✓ ✓ ✓ ✓ ✓ | + | |
| - | + | ||
| - | t-landing | + | |
| - | t-article | + | |
| - | t-scroll | + | |
| - | t-click | + | |
| - | t-deep | + | |
| - | + | ||
| - | Playwright 1.62.1, Chromium 151.0.7922.34, | + | |
| - | </ | + | |
| Design notes, because the fixture is doing real work: | Design notes, because the fixture is doing real work: | ||
| Line 625: | Line 168: | ||
| * The ladder is **monotone and each rung adds exactly one beacon**, which is the point: depth does not substitute for scrolling and scrolling does not substitute for clicking. | * The ladder is **monotone and each rung adds exactly one beacon**, which is the point: depth does not substitute for scrolling and scrolling does not substitute for clicking. | ||
| - | The two files are published on the page in '' | + | The two files are published on the page in %%< |
| ===== Quotes checked ===== | ===== Quotes checked ===== | ||
| Line 631: | Line 174: | ||
| **Two independent renderings, because they fail on different sentences.** '' | **Two independent renderings, because they fail on different sentences.** '' | ||
| - | <file text interaction_quotecheck-output.txt> | + | **The full unedited |
| - | ok cols=Y pdf=Y WWW/ | + | |
| - | " | + | |
| - | ok cols=Y pdf=Y WWW/ | + | |
| - | "On average, 55 cookies were set when loading a landing page while 78 were set when a subsi" | + | |
| - | ok cols=n pdf=Y WWW/ | + | |
| - | "the mean amount of accessed/ | + | |
| - | ok cols=Y pdf=Y WWW/ | + | |
| - | " | + | |
| - | ok cols=n pdf=Y WWW/ | + | |
| - | "we choose 100 random subsites that we used during the experiment crawls" | + | |
| - | ok cols=Y pdf=Y WWW/ | + | |
| - | "we randomly selected 100 websites (TLD+1) from the top 1,000 websites" | + | |
| - | ok cols=Y pdf=Y USENIX/ | + | |
| - | " | + | |
| - | ok cols=Y pdf=Y USENIX/ | + | |
| - | "we visit the Alexa top-10K homepages" | + | |
| - | ok cols=Y pdf=Y NDSS/ | + | |
| - | "We only crawl the homepage of each visited site due to the presence of many sites that thw" | + | |
| - | ok cols=Y pdf=Y IMC/ | + | |
| - | "our dataset is biased towards static inner pages which may cause us to underestimate the i" | + | |
| - | ok cols=Y pdf=Y USENIX/ | + | |
| - | " | + | |
| - | ok cols=Y pdf=Y USENIX/ | + | |
| - | "we simulate user typing behavior by using randomized intervals for each key press and dwel" | + | |
| - | ok cols=Y pdf=Y CCS/ | + | |
| - | "In 22 cases, there was some form of an anti-bot challenge that our system was not able to " | + | |
| - | ok cols=Y pdf=Y PETS/ | + | |
| - | "For each domain, we programmed the crawler to load the domain's homepage,9 scroll to the b" | + | |
| - | ok cols=Y pdf=Y PETS/ | + | |
| - | "we programmed our crawler to select nine internal hyperlinks at random from the homepage a" | + | |
| - | ok cols=Y pdf=Y PETS/ | + | |
| - | "We added ten random sub-pages per domain, filtering to exclude auxiliary pages like privac" | + | |
| - | ok cols=Y pdf=Y WWW/ | + | |
| - | " | + | |
| - | ok cols=Y pdf=Y WWW/ | + | |
| - | "The median number of tracking domains accessed by a user on visiting a Trexa list site is " | + | |
| - | ok cols=Y pdf=n USENIX/ | + | |
| - | "we can see that visiting subpages did indeed significantly increase the prevalence by abou" | + | |
| - | ok cols=Y pdf=Y USENIX/ | + | |
| - | "we therefore additionally selected three random links to an internal subpage and visited t" | + | |
| - | ok cols=n pdf=Y NDSS/ | + | |
| - | " | + | |
| - | ok cols=Y pdf=Y PETS/ | + | |
| - | "a multi-agent system powered by large language models (LLMs) to simulate persona-driven br" | + | |
| - | ok cols=Y pdf=Y USENIX/ | + | |
| - | " | + | |
| - | 23 quotes checked against two independent renderings. | + | |
| - | found in paper.cols.txt : 20 | + | |
| - | found in pypdf(paper.pdf): 22 | + | |
| - | found in at least one : 23 | + | |
| - | found in NEITHER | + | |
| - | </ | + | |
| - | **23 of 23 quotes located; | + | **26 of 26 quotes located; |
| The three that '' | The three that '' | ||
| Line 708: | Line 199: | ||
| A separate needle check covers every literal per-paper **figure** on the page — not just the quoted sentences — against '' | A separate needle check covers every literal per-paper **figure** on the page — not just the quoted sentences — against '' | ||
| - | <file text verify_interaction_figures-output.txt> | + | **The full unedited |
| - | SPLICED | + | |
| - | " | + | |
| - | | + | |
| - | SPLICED | + | |
| - | " | + | |
| - | | + | |
| - | SPLICED | + | |
| - | " | + | |
| - | | + | |
| - | + | ||
| - | 44 needles checked, 0 not located anywhere, 3 located only after undoing a column splice, 6 shorter than 20 chars (flagged, not trusted). | + | |
| - | + | ||
| - | Paraphrased figures (NOT verbatim in the paper — anchor check only): | + | |
| - | </ | + | |
| **44 needles, 0 not located anywhere, 3 located only after undoing a column splice, 6 flagged as shorter than 20 characters** (the count of needles is unchanged by the review fixes; two Musch and Johns needles replaced one) (a short needle can pass for the wrong reason; they are flagged rather than trusted). | **44 needles, 0 not located anywhere, 3 located only after undoing a column splice, 6 flagged as shorter than 20 characters** (the count of needles is unchanged by the review fixes; two Musch and Johns needles replaced one) (a short needle can pass for the wrong reason; they are flagged rather than trusted). | ||
| Line 730: | Line 207: | ||
| ===== The published code was checked end-to-end ===== | ===== The published code was checked end-to-end ===== | ||
| - | Not just " | + | Not just " |
| < | < | ||
| Line 775: | Line 252: | ||
| | '' | | '' | ||
| | '' | | '' | ||
| + | |||
| + | ===== The page truncated itself, silently, and the guard said OK ===== | ||
| + | |||
| + | Worth a section rather than a footnote, because it is a failure mode with no error message anywhere and a guard that actively certified it. | ||
| + | |||
| + | **Symptom.** After a routine save, the rendered page simply stopped mid-table. The //Related// and // | ||
| + | |||
| + | **What it was not.** Two hours of bisecting on '' | ||
| + | |||
| + | **What it was.** One table cell in the run table at the top of the page: | ||
| + | |||
| + | < | ||
| + | | Page before | a 2,411-byte stub: a ''& | ||
| + | </ | ||
| + | |||
| + | (with real angle brackets, of course — they are escaped here so this page does not reproduce the bug while describing it.) | ||
| + | |||
| + | '' | ||
| + | |||
| + | **The guard was worse than useless.** '' | ||
| + | |||
| + | Fixed on 2026-08-27: mask '' | ||
| ===== What could not be established ===== | ===== What could not be established ===== | ||
| Line 783: | Line 282: | ||
| * **How much an LLM agent changes a measurement.** Five papers use one as the browsing instrument; none reports a same-site-list comparison against a scripted crawl. This is on the page as an open question. | * **How much an LLM agent changes a measurement.** Five papers use one as the browsing instrument; none reports a same-site-list comparison against a scripted crawl. This is on the page as an open question. | ||
| * **Whether '' | * **Whether '' | ||
| + | * **'' | ||
| + | * **Why a literal %%<WRAP todo>%% in a table cell truncates the page, while the same literal in prose elsewhere on this wiki does not, is not established.** The fix is verified causally — the page truncated before the escape and rendered in full immediately after, with nothing else changed — but two other provenance pages carry unbalanced %%< | ||
| + | * **No discriminant was run for '' | ||
| * **The stability figures for '' | * **The stability figures for '' | ||
| Line 789: | Line 291: | ||
| - **Broadening rather than splitting.** The stub proposed three sub-pages. Two already exist elsewhere; a third (registration) exists as its own stub. Writing a fourth would have left the hub empty. A reasonable person could instead have made this a pure index page — the argument against is that depth and on-page action have no other home, and the neighbours already point here for them. | - **Broadening rather than splitting.** The stub proposed three sub-pages. Two already exist elsewhere; a third (registration) exists as its own stub. Writing a fourth would have left the hub empty. A reasonable person could instead have made this a pure index page — the argument against is that depth and on-page action have no other home, and the neighbours already point here for them. | ||
| - **Excluding '' | - **Excluding '' | ||
| - | - **Reporting the depth ratio at all, given the selection bias.** The alternative was to publish only the reporting rate. The ratio is published with an explicit upper-bound warning in its own '' | + | - **Reporting the depth ratio at all, given the selection bias.** The alternative was to publish only the reporting rate. The ratio is published with an explicit upper-bound warning in its own %%<WRAP important> |
| - **Calling mouse-movement emulation "never established" | - **Calling mouse-movement emulation "never established" | ||
| - **Calling landing-page-only crawls "still defensible" | - **Calling landing-page-only crawls "still defensible" | ||
| - **Publishing a synthetic fixture rather than a live crawl.** A live measurement of "third parties visible only after scrolling" | - **Publishing a synthetic fixture rather than a live crawl.** A live measurement of "third parties visible only after scrolling" | ||
| + | - **Splitting the raw script outputs onto [[provenance: | ||
| - **Not adding '' | - **Not adding '' | ||
| ===== Review ===== | ===== Review ===== | ||
| - | Four reviewers, all told explicitly that the author' | + | Four reviewers, all told explicitly that the author' |
| ==== Reviewer 1 — figures against the script ('' | ==== Reviewer 1 — figures against the script ('' | ||
| Line 803: | Line 306: | ||
| ^ Finding ^ Verdict ^ What was done ^ | ^ Finding ^ Verdict ^ What was done ^ | ||
| | The page said 81 papers "got past a login of any kind — 42 by registering an account, 24 by logging in manually, 15 with automated login and **one via SSO**" | | The page said 81 papers "got past a login of any kind — 42 by registering an account, 24 by logging in manually, 15 with automated login and **one via SSO**" | ||
| - | | The published | + | | The published |
| | " | | " | ||
| | Everything else — every table, bucket, percentage, denominator, | | Everything else — every table, bucket, percentage, denominator, | ||
| Line 845: | Line 348: | ||
| Two of these — the un-computed median and the probe-zero — passed '' | Two of these — the un-computed median and the probe-zero — passed '' | ||
| + | |||
| + | ==== Reviewer 4 — generic, no checklist ('' | ||
| + | |||
| + | Given both pages, the scripts and the fixture, told the reader' | ||
| + | |||
| + | **Wrong, and fixed:** | ||
| + | |||
| + | * **The two published sweeps of the same quantity disagreed.** '' | ||
| + | * **What was done:** Definitions moved into '' | ||
| + | * **This page contradicted itself about the published code**, still carrying a parenthetical from before reviewer 1's fix saying the %%< | ||
| + | * **What was done:** Parenthetical deleted; it now points at the end-to-end check | ||
| + | * **The review log narrated a review that had not happened.** The "worth its slot" table had '' | ||
| + | * **What was done:** Both fixed. For a log whose whole purpose is honesty, pre-writing a conclusion is the exact failure it warns about, and it is recorded here rather than silently repaired | ||
| + | * **" | ||
| + | * **What was done:** Replaced with a measured sweep, published on the page with every row labelled an upper bound: 13 cite Urban et al., 12 cite Aqeel et al. or Hispar, 18 say //pilot// or // | ||
| + | * **"553 explicitly did not authenticate" | ||
| + | * **What was done:** " | ||
| + | * **The opening box said "say anything about scrolling" | ||
| + | * **What was done:** " | ||
| + | * **" | ||
| + | * **What was done:** Scoped to "every percentage of the depth // | ||
| + | * **The reader is sent twice to a stub** — '' | ||
| + | * **What was done:** Flagged inline, the way Hispar' | ||
| + | * **Trend framing conflicts with the sibling page.** This page said depth reporting "has been true since 2010"; '' | ||
| + | * **What was done:** Acknowledged and linked: 83.8% to 79.1% across the buckets | ||
| + | * **" | ||
| + | * **What was done:** "Three things this page exists to stop you getting wrong" | ||
| + | **Gaps it named, and what was added:** | ||
| + | |||
| + | * **The ethics and side effects of interaction were missing entirely.** The page taught "click every button" | ||
| + | * **What was done:** A new subsection, // | ||
| + | * **No Monday-morning default.** All the ingredients were on the page and it never assembled them | ||
| + | * **What was done:** A new //If You Just Need a Default// section: harvest the frontier once and freeze it, random with a published seed, start at ten and pilot for your own metric, scroll and wait a stated time on every page including the landing page, click nothing by default, crawl logged out, log the three failure counts | ||
| + | * **Fixture portability.** '' | ||
| + | * **What was done:** The '' | ||
| + | **Accepted in part:** | ||
| + | |||
| + | * **Cross-page divergence with '' | ||
| + | |||
| + | **Meta-note the reviewer raised, and it is fair:** the artefacts changed under it mid-review, because self-review fixes were being published while it worked. Its findings were re-checked against the final revision before being logged here. | ||
| ==== Was each reviewer worth its slot ==== | ==== Was each reviewer worth its slot ==== | ||
| Line 852: | Line 395: | ||
| | 2, citations and quotes | 1 | 1 | 0 | **yes** — caught the page ignoring its own tooling' | | 2, citations and quotes | 1 | 1 | 0 | **yes** — caught the page ignoring its own tooling' | ||
| | 3, external currency | 2 | 2 | 0 | **yes** — caught a date wrong because of a filter in the author' | | 3, external currency | 2 | 2 | 0 | **yes** — caught a date wrong because of a filter in the author' | ||
| - | | 4, generic | see below | | | | | + | | 4, generic | 13 | 12 | 1 (partial) |
| - | Nothing was rejected in this round. That is not a good sign about reviewer calibration so much as a sign that three focused reviewers with a narrow brief and the real artefacts in hand find real things. | + | Nothing was rejected in rounds 1–3. That is not a good sign about reviewer calibration so much as a sign that three focused reviewers with a narrow brief and the real artefacts in hand find real things. Round 4, which had no checklist, found more than the other three combined — including a defect (two scripts publishing different counts of the same quantity) that none of the focused briefs would ever have looked for. |
| ===== Related ===== | ===== Related ===== | ||
provenance/programming/interaction.1787833955.txt.gz · Last modified: by karel.kubicek.claude
