User Tools

Site Tools


provenance:design:connected_tv

This is an old revision of the document!


Provenance: Connected TV

Working log for connected_tv. Every figure on that page is produced by one script, printed here with its denominator, and every quote it uses is machine-checked against both renderings of its source paper. Corpus-level caveats — the seven venues, the funnel, the provisional 2025–2026 years — are on corpus and are not restated.

Run date 2026-09-12. Corpus at the time: 5,859 papers with extracted full text, data/extract/run1/extractions.jsonl, seven venues, 2010–2026.

1. Why this page exists, and what it is not

roadmap queued design:connected_tv on 2026-09-07 with a 16-paper candidate set from scripts/gap_probe_roadmap.mjs (family ctv_streaming), and roadmap §5 recorded a condition on it: Connected TV will need its population derived from platform fields rather than the probe, because the probe's web-platform column is exactly the wrong filter for it.” That condition was honoured, and it turned out to understate the problem in one direction and overstate it in another.

  • The title probe's precision is 62.5% — 10 of its 16 are in the final population.
  • Its recall is much worse: it misses 25 of the 35. The largest TV app analysis in the corpus, [1Tileria, Marcos; Blasco, Jorge (2022): "Watch Over Your TV: A Security and Privacy Analysis of the Android TV Ecosystem", in: Proceedings on Privacy Enhancing Technologies. (DOI)] (4,745 Android TV APKs), is not in the 16 at all; neither is [2Rye, Erik C.; Levin, Dave (2024): "Surveilling the Masses with Wi-Fi-Based Positioning Systems", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)], in which Roku devices are two of the five most common vendor prefixes in a 490-million-row dataset.
  • “Derive from platform fields” cannot mean “filter on a platform field”. There is no tv value in the platforms enum, and the closest one, iot, holds 436 corpus papers of which the overwhelming majority are smart speakers, cameras, plugs and firmware. Platform fields are used here as evidence about the population, not as the selector: 26 of the 35 carry iot, 1 carries web, and only 6 are inside the corpus-wide crawled population. That is the page's thesis, measured.

The roadmap row also named one paper that is not in the population. It listed Endangered Privacy (USENIX 2025) [3Björklund, Martin; Duvignau, Romaric (2025): "Endangered Privacy: Large-Scale Monitoring of Video Streaming Services", in: Proceedings of the USENIX Security Symposium. (Link)] among the six-paper “spine”. Reading it, it identifies videos from encrypted MPEG-DASH traffic against Amazon Prime Video, Max and SVT Play; its populations are 242,364 video manifests, 900 sampled titles and the VNAT flow dataset. No television is measured anywhere in it. It is recorded as ADJ with that reason. The roadmap row is left as written, because it is a record of what was believed on 2026-09-07; this is the correction.

2. The inclusion rule, fixed before any figure

Written into scripts/ctv_fold.mjs before the first count was taken:

  • A television-class endpoint is a smart TV set, a TV operating system (Android TV / Google TV, tvOS, Tizen, webOS, Roku OS, Fire OS), a streaming stick, box or set-top box, an app running on one, or the broadcast path (HbbTV / DVB) into one.
  • A — the paper's central object of measurement is a television-class endpoint.
  • B — television-class devices are part of a broader measured population and the paper reports at least one result broken out for them.
  • ADJ — adjacent. Video streaming measured off a TV (browser DRM, piracy websites, encrypted-traffic video fingerprinting), or a TV used as apparatus rather than measured.
  • OUT — the TV name is a passing reference, a survey answer option, a related-work sentence, or a homonym.

Two boundary decisions are worth naming because a reasonable person would draw them differently.

  • A TV used as apparatus is OUT, even when it has its own results row. [4Mavroudis, Vasilios; Hao, Shuang; Fratantonio, Yanick; Maggi, Federico; Kruegel, Christopher; Vigna, Giovanni (2017): "On the Privacy and Security of the Ultrasound Ecosystem", in: Proceedings on Privacy Enhancing Technologies. (DOI)]'s ultrasound work and Void (USENIX 2020, a Samsung Smart TV used as a replay loudspeaker with its own 24,282-sample row) both fail the rule for the same reason: the television generates a stimulus, it is not measured. Under a purely mechanical “named device with a results row” rule, Void would be Tier B.
  • Survey and interview studies are OUT even when every participant owns a TV. This removes about twenty smart-home qualitative papers. They are real research about televisions; they are not measurements of one, and including them would have made the “15 of 35 are IoT device sets” finding meaningless.

3. Every query, with its population

# Question Population Answer
Q1 How many corpus papers have full text to probe? all 5,859 5,855 scanned; 4 have no paper.cols.txt
Q2 What does the roadmap's title+summary probe return? all 5,859 16; 1 carries web
Q3 Wide TV-vocabulary gate (gate 1) all 5,855 scanned 142
Q4 Audit set (gate 2) gate 1 103
Q5 Population after hand audit the 103 35 (A 13, B 22); ADJ 16, OUT 52; precision 34.0%
Q6 Population papers the roadmap probe misses the 35 25
Q7 Population by venue the 35 IMC 11, USENIX 9, NDSS 6, PETS 6, CCS 2, IEEE S&P 1, WWW 0
Q8 platforms[] distribution the 35 iot 26, other-online-service 12, mobile 6, offline 3, web 1
Q9 Inside the corpus crawled population the 35 6
Q10 population[].unit the 35 iot-devices 20, other 17, mobile-apps 5
Q11 Named instruments, alias-folded, used only the 35 Wireshark 10, tcpdump 9, mitmproxy 4, Frida 3, adb 3
Q12 Interception-evidence probes 13 Tier A / 35 A+B router/AP 11/30, mitm 10/22, DNS 5/12, HDMI 6/9, remote-control 12/16, undecryptable reported, narrow probe 1, wide probe 3 (see §6b)
Q13 crawlConfig fields stated 6 papers with a crawlConfig interactionDepth 5, consentAction 0, statefulness 0, browsers 0
Q14 ethics.reviewOutcome stated 34 with an ethics object 14 (41.2%) vs corpus-empirical 1,728 of 4,472 (38.6%)
Q15 artifacts.availability stated 34 with an artifacts object 27 (79.4%) vs 2,890 of 4,854 (59.5%)
Q16 temporal.spanStart stated 34 empirical 27 (79.4%) vs 2,882 of 5,118 (56.3%)
Q17 detection[].prevalence coverage 70 Tier A tuples 68 carry a prevalence (97.1%); corpus-wide 26,316 of 27,241 (96.6%)
Q18 population[].listVersion stated 105 population tuples in the 35 42
Q19 Vantage location stated 34 with a vantage tuple 18 (52.9%)

Denominators that are easy to get wrong here, spelled out. ethics and artifacts are nullable in the schema — one of the 35 papers has neither object — so Q14–Q16 divide by the papers that carry the object, never by 35 and never by the corpus. crawlConfig is null for 29 of the 35, so Q13 divides by 6, and its zeros are read on the page rather than reported bare. No TV figure on either page divides by 5,859. The one place that denominator appears is the corpus-share column of the platform table on the content page, which exists precisely to say what share of the whole corpus each platform value has — that column divides by 5,859 by design, and is labelled as doing so. An earlier version of this sentence said “nothing on either page divides by 5,859”, which was wrong.

4. The candidate pool, and why it has two gates

Full text is whitespace-collapsed (soft hyphens stripped, hyphen-newline joined, runs of whitespace reduced to one space) before any regex runs. Without that a phrase broken across a line silently fails to match.

Nine probes over every paper.cols.txt:

smarttv      /\bsmart[-\s]?TVs?\b/i
ctv          /\bconnected[-\s]TVs?\b|\bCTV\b/i
ott          /\bover[-\s]the[-\s]top\b|\bOTT\b/i
hbbtv        /\bHbbTV\b|\bhybrid broadcast broadband\b/i
acr          /\bautomatic content recognition\b|\bACR\b/i
platformdev  /\bRoku\b|\bFire ?TV\b|\bApple ?TV\b|\bChromecast\b|\bAndroid ?TV\b|\bGoogle ?TV\b|\btvOS\b|\bWebOS\b|\bTizen\b|\bset[-\s]?top box(es)?\b/i
streamsvc    /\bNetflix\b|\bHulu\b|\bDisney\+|\bAmazon Prime Video\b|\bYouTube ?TV\b|\bTwitch\b/i
tvapp        /\bTV app(s|lication)?\b|\btelevision app(s)?\b/i
iptv         /\bIPTV\b|\binternet protocol television\b/i

Plus a device-name probe used to find televisions inside broader IoT device sets, which is where 15 of the 35 came from:

DEV  /\b(Roku|Fire ?TV|Apple ?TV|Chromecast|Android ?TV|Google ?TV|tvOS|WebOS|Tizen|Vizio|Hisense|Bravia|Nvidia Shield|Samsung(?: Smart)? TV|LG(?: Smart)? TV|TCL|smart[- ]?TVs?|set[- ]?top box(?:es)?)\b/gi

Gate 1 (142 papers)core >= 2 || acr >= 2 || iptv >= 2 || hbbtv >= 1 || tvapp >= 1 || titleHit, where core sums smarttv + ctv + hbbtv + platformdev + tvapp.

Gate 2, the audit set (103 papers) — gate 1 narrowed by devN >= 4 || brands >= 3 || titleHit || hbbtv > 0 || acr >= 2 || iptv >= 2.

The 39 papers dropped between the gates all have a single-brand, low-count mention — a Tizen in a list of embedded platforms, one Apple TV in an enumeration of Apple hardware. That is a judgement, not a measurement: it was not hand-audited, and if a television study exists that names exactly one TV-class device three times or fewer, this page does not contain it. A wider audit would be the cheapest improvement to make here.

Homonyms found, and what they cost. ACR is the American College of Radiology, an authentication context reference, and an arbitrary abbreviation in a privacy-policy paper; IPTV appears in leaked-credential corpora and in X spam campaigns; Tizen is a smartwatch platform and an open-source project under fuzz testing; WebOS matches both the LG TV OS and unrelated prose. 52 of 103 audit-set papers are OUT, and the biggest single class is the smart-home survey, where “smart TV” is an answer option. The precision of the audit set is 34.0% — for comparison, the ad-archives row on roadmap recorded 14.5% and the authentication row 24.5%.

5. Verdict map — the 68 papers NOT in the population

Published in full, because the rejections are the only record of where the line was drawn.

0c. ADJACENT AND OUT — the audit trail for what was NOT counted
==============================================================================
V    Year  Venue    Title                                                           Reason
---  ----  -------  --------------------------------------------------------------  -----------------------------------------------------------------------------------------------------------------------------------------------
ADJ  2011  IMC      Measurement and analysis of a large scale commercial mobile in  "TV" delivered to mobile handsets, not to a television.
ADJ  2012  IMC      Watching videos from everywhere: a study of the PPTV mobile Vo  Mobile VoD; no television endpoint.
ADJ  2013  IMC      Analyzing the potential benefits of CDN augmentation strategie  CDN augmentation for video workloads; no TV endpoint.
ADJ  2013  IMC      Peer-assisted content distribution in Akamai netsession         Peer-assisted CDN; set-top-box mention is background.
ADJ  2016  IMC      Performance Characterization of a Commercial Video Streaming S  Streaming service performance from browser/CDN vantage; no TV-specific result.
ADJ  2016  IMC      Anatomy of a Personalized Livestreaming System                  Livestreaming (Periscope) system measurement; no TV endpoint.
ADJ  2017  PETS     On the Privacy and Security of the Ultrasound Ecosystem         Ultrasonic cross-device tracking (uXDT): beacons emitted by TV adverts and picked up by phone SDKs. The TV is the emitter, never measured.
ADJ  2019  WWW      Exploiting Diversity in Android TLS Implementations for Mobile  Android app traffic classification; TLS-fingerprint method later reused on TV apps.
ADJ  2020  USENIX   Void: A fast and light voice liveness detection system          A Samsung Smart TV is used as a replay LOUDSPEAKER; the TV is apparatus, not the measured object.
ADJ  2022  NDSS     A Lightweight IoT Cryptojacking Detection Mechanism in Heterog  Authors implement their own cryptojacking PoC on an LG webOS TV to test a detector; no deployed-TV population.
ADJ  2022  USENIX   OVRseen: Auditing Network Traffic and Privacy Policies in Ocul  VR headsets; smart TVs used as the comparison ecosystem. Same lab, same pipeline shape.
ADJ  2023  IMC      Tracking, Profiling, and Ad Targeting in the Alexa Echo Smart   Smart speaker study; TVs cited as the comparable prior ecosystem, not measured. The closest methodological sibling.
ADJ  2023  PETS     Your DRM Can Watch You Too: Exploring the Privacy Implications  Widevine EME in browsers and Android; TVs named as another Widevine host, not measured.
ADJ  2024  IMC      Cost-Saving Streaming: Unlocking the Potential of Alternative   Edge-node economics for streaming delivery; no TV endpoint measured.
ADJ  2025  PETS     Unmasking the Shadows: A Cross-Country Study of Online Trackin  Illegal movie streaming WEBSITES crawled with a browser; a web-tracking study, not a TV study.
ADJ  2025  USENIX   Endangered Privacy: Large-Scale Monitoring of Video Streaming   Video identification from encrypted MPEG-DASH traffic. The roadmap listed it as CTV spine; it measures the SERVICE and its traffic, never a TV.
OUT  2010  IMC      What happened in my network: mining network events from router  IPTV named as the service carried; router syslogs are the object.
OUT  2011  CCS      On the vulnerability of FPGA bitstream encryption against powe  Set-top box named as an FPGA application domain.
OUT  2011  IMC      Broadcast yourself: understanding YouTube uploaders             IPTV appears once in related work.
OUT  2014  CCS      (Nothing else) MATor(s): Monitoring the Anonymity of Tor's Pat  "ACR" homonym.
OUT  2015  USENIX   A Placement Vulnerability Study in Multi-Tenant Public Clouds   Title probe matched "streaming"; cloud VM placement.
OUT  2015  USENIX   Rocking Drones with Intentional Sound Noise on Gyroscopic Sens  Passing mention.
OUT  2016  CCS      SandScout: Automatic Detection of Flaws in iOS Sandbox Profile  Apple TV named as a device that runs iOS/tvOS; iOS sandbox is the object.
OUT  2016  IMC      Entropy/IP: Uncovering Structure in IPv6 Addresses              "ACR" homonym.
OUT  2016  USENIX   You Are Who You Know and How You Behave: Attribute Inference A  IPTV homonym.
OUT  2017  CCS      POSTER: Watch Out Your Smart Watch When Paired                  Tizen here is the smartwatch platform, not the TV one.
OUT  2017  PETS     Why can’t users choose their identity providers on the web?     "ACR" homonym.
OUT  2017  USENIX   Same-Origin Policy: Evaluation in Modern Browsers               Passing mention of TV browsers.
OUT  2017  WWW      FLOCK: Combating Astroturfing on Livestreaming Platforms        Title probe matched "streaming platform"; astroturfing detection on Twitch-like sites.
OUT  2018  CCS      Medical Devices are at Risk: Information Security on Diagnosti  "ACR" = American College of Radiology.
OUT  2019  IEEE-SP  Drones' Cryptanalysis - Smashing Cryptography with a Flicker    IPTV/ACR homonyms.
OUT  2019  NDSS     cleaning-up-the-internet-of-evil-things-real-world-evidence-on  One infected set-top box in a Mirai remediation table; no TV finding.
OUT  2019  NDSS     latex-gloves-protecting-browser-extensions-from-probing-and-re  The Chromecast browser EXTENSION, not the device.
OUT  2019  USENIX   A Billion Open Interfaces for Eve and Mallory: MitM, DoS, and   tvOS listed among Apple OSes; AWDL is the object.
OUT  2019  WWW      Snapshot-based Loading Acceleration of Web Apps with Nondeterm  Tizen/webOS named as embedded web-app platforms; benchmark is web apps.
OUT  2020  CCS      iDEA: Static Analysis on the Security of Apple Kernel Drivers   tvOS is one of four Apple OSes scanned; no TV-specific result.
OUT  2020  PETS     Smart Devices in Airbnbs: Considering Privacy and Security for  Survey; smart TV is a scenario option.
OUT  2022  IMC      Deep dive into the IoT backend ecosystem                        Backend infrastructure; TV mentions are motivation and a citation to FingerprinTV.
OUT  2022  PETS     A Multi-Region Investigation of the Perceptions and Use of Sma  Survey; smart TV is a related-work citation and an ownership option.
OUT  2022  PETS     Exploring the Privacy Concerns of Bystanders in Smart Homes fr  Survey; smart TV is an example in a prompt.
OUT  2023  CCS      IoTFlow: Inferring IoT Device Behavior at Scale through Static  Companion-app analysis; no TV breakout.
OUT  2023  IEEE-SP  Characterizing Everyday Misuse of Smart Home Devices            Survey of 483 people; smart TV is an ownership option, not a measured device.
OUT  2023  IEEE-SP  WebSpec: Towards Machine-Checked Analysis of Browser Security   "ACR" homonym.
OUT  2023  IEEE-SP  UTopia: Automatic Generation of Fuzz Driver using Unit Tests    Tizen as an open-source project under test; no TV device.
OUT  2023  PETS     No Privacy Among Spies: Assessing the Functionality and Insecu  Android stalkerware; "ACR" homonym.
OUT  2023  USENIX   Examining Consumer Reviews to Understand Security and Privacy   Review-text analysis; set-top box is a Mirai product category, no TV measurement.
OUT  2023  USENIX   Examining Power Dynamics and User Privacy in Smart Technology   Interview study; TVs are participant device inventories.
OUT  2023  USENIX   Internet Service Providers' and Individuals' Attitudes, Barrie  Interview and survey; TV is a device-ownership row.
OUT  2023  USENIX   "It's the Equivalent of Feeling Like You're in Jail”: Lessons   Interview study on IPV; TV is a reported abuse vector, not measured.
OUT  2023  USENIX   Measuring Up to (Reasonable) Consumer Expectations: Providing   Vignette survey; Vizio appears only in a news citation.
OUT  2023  USENIX   Abuse Vectors: A Framework for Conceptualizing IoT-Enabled Int  Qualitative framework; TV is an example abuse vector.
OUT  2023  USENIX   Exploring Tenants' Preferences of Privacy Negotiation in Airbn  Vignette survey; smart TV is a device-type option.
OUT  2023  WWW      SISSI: An Architecture for Semantic Interoperable Self-Soverei  "ACR" homonym (authentication context reference).
OUT  2024  IEEE-SP  SoK: Technical Implementation and Human Impact of Internet Pri  SoK; TV work cited, not measured.
OUT  2024  PETS     A Bilingual Longitudinal Analysis of Privacy Policies Measurin  ACR homonym: "ACR" is not automatic content recognition here.
OUT  2024  PETS     Contextualizing Interpersonal Data Sharing in Smart Homes       Vignette survey; "viewing history from your smart TV" is a question stem.
OUT  2024  PETS     "My Best Friend's Husband Sees and Knows Everything": A Cross-  Survey; smart TV is a free-text mention count.
OUT  2024  USENIX   Co-Designing a Mobile App for Bystander Privacy Protection in   Interview study; TV names are participant-reported device inventories.
OUT  2025  IEEE-SP  Analyzing the iOS Local Network Permission from a Technical an  Chromecast is one of four IoT devices used to trigger the permission; no TV result.
OUT  2025  IEEE-SP  Hey, Your Secrets Leaked! Detecting and Characterizing Secret   IPTV homonym in leaked-credential data.
OUT  2025  NDSS     Non-intrusive and Unconstrained Keystroke Inference in VR Plat  VR; smart TV appears only as a citation to HomeSpy.
OUT  2025  PETS     Help Me Help You: Privacy Considerations for Third Party IoT D  Vignette survey; TVs appear in a device-category prompt.
OUT  2025  PETS     Who Cares? Contextual Privacy Judgments from Owner and Bystand  Survey; smart TV is a device-category option.
OUT  2025  USENIX   Regulating Smart Device Support Periods: User Expectations and  Survey; Smart TV is a self-reported ownership category.
OUT  2026  IEEE-SP  Privacy Perspectives and Practices of Chinese Smart Home Produ  Interview study; smart TV is a company product-line row.
OUT  2026  NDSS     TBTrackerX: Fantastic Trigger Bots and Where to Find Malicious  IPTV spam homonym.
OUT  2026  PETS     Dead Domains, Living Data: A Privacy Risk Analysis of Domain L  Android apps; one expired-domain example happens to also ship on Roku.
OUT  2026  USENIX   PANGOLIN: Fuzzing Multilingual IoT Firmware with LLM-Driven Co  "SmartTVs" is a citation to the 2021 fuzzing paper, used as a baseline name.

==============================================================================

And the 35 that are in:

0b. THE POPULATION, PAPER BY PAPER
==============================================================================
Tier  Year  Venue    Topic                 Title
----  ----  -------  --------------------  ----------------------------------------------------------------------------------------------------------------
A     2011  IMC      delivery-performance  Understanding couch potatoes: measurement and modeling of interactive usage of IPTV at large scale
A     2011  IMC      delivery-performance  Q-score: proactive service quality assessment in a large IPTV system
A     2014  USENIX   broadcast             From the Aether to the Ethernet—Attacking the Internet using Broadcast Digital Television
A     2019  CCS      tracking              Watching You Watch: The Tracking Ecosystem of Over-the-Top TV Streaming Devices
A     2020  PETS     tracking              The TV is Smart and Full of Trackers: Measuring Smart TV Advertising and Tracking
A     2021  USENIX   vulnerability         Android SmartTVs Vulnerability Discovery via Log-Guided Fuzzing
A     2022  PETS     tracking              FingerprinTV: Fingerprinting Smart TV Apps
A     2022  PETS     app-analysis          Watch Over Your TV: A Security and Privacy Analysis of the Android TV Ecosystem
A     2023  NDSS     broadcast             I Still Know What You Watched Last Sunday: Privacy of the HbbTV Protocol in the European Smart TV Landscape
A     2023  USENIX   side-channel          HOMESPY: The Invisible Sniffer of Infrared Remote Control of Smart TVs
A     2024  IMC      acr                   Watching TV with the Second-Party: A First Look at Automatic Content Recognition Tracking in Smart TVs
A     2024  NDSS     side-channel          Acoustic Keystroke Leakage on Smart Televisions
A     2025  USENIX   piracy                Watch Out Your TV Box: Reversing and Blocking a P2P-based Illegal Streaming Ecosystem
B     2018  IMC      delivery-performance  Understanding Video Management Planes
B     2019  IMC      iot-device-set        Information Exposure From Consumer IoT Devices: A Multidimensional, Network-Informed Measurement Approach
B     2019  USENIX   iot-device-set        All Things Considered: An Analysis of IoT Devices on Home Networks
B     2020  IMC      iot-device-set        A Haystack Full of Needles: Scalable Detection of IoT Devices in the Wild
B     2020  NDSS     iot-device-set        Packet-Level Signatures for Smart Home Devices
B     2020  NDSS     iot-device-set        Et Tu Alexa? When Commodity WiFi Devices Turn into Adversarial Motion Sensors
B     2020  USENIX   iot-device-set        You Are What You Broadcast: Identification of Mobile and IoT Devices from (Public) WiFi
B     2021  IMC      iot-device-set        IoTLS: understanding TLS usage in consumer IoT devices
B     2021  PETS     iot-device-set        Blocking Without Breaking: Identification and Mitigation of Non-Essential IoT Traffic
B     2022  CCS      app-analysis          Understanding IoT Security from a Market-Scale Perspective
B     2022  PETS     iot-device-set        Analyzing the Feasibility and Generalizability of Fingerprinting Internet of Things Devices
B     2022  USENIX   iot-device-set        Lumos: Identifying and Localizing Diverse Hidden IoT Devices in an Unfamiliar Environment
B     2023  IMC      iot-device-set        Behind the Scenes: Uncovering TLS and Server Certificate Practice of IoT Device Vendors in the Wild
B     2023  IMC      iot-device-set        In the Room Where It Happens: Characterizing Local Communication and Threats in Smart Homes
B     2024  IEEE-SP  device-population     Surveilling the Masses with Wi-Fi-Based Positioning Systems
B     2024  IMC      iot-device-set        IoT Bricks Over v6: Understanding IPv6 Usage in Smart Homes
B     2024  IMC      delivery-performance  Characterizing User Platforms for Video Streaming in Broadband Networks
B     2024  PETS     iot-device-set        Connecting the Dots: Tracing Data Endpoints in IoT Devices
B     2025  NDSS     iot-device-set        Evaluating Machine Learning-Based IoT Device Identification Models for Security Applications
B     2025  USENIX   vulnerability         Tracking You from a Thousand Miles Away! Turning a Bluetooth Device into an Apple AirTag Without Root Privileges
B     2026  NDSS     vulnerability         BLERP: BLE Re-Pairing Attacks and Defenses
B     2026  USENIX   policy-compliance     Missing, Present and Conflicting: A Large Scale Analysis of IoT Update Information in the EU Market

==============================================================================

5b. The guard on this map, mutation-tested

A review pass on 2026-09-12 mutation-tested the map rather than reading it, and found the guard was weaker than it looked.

  • Deleting an entry — the script throws: audit set has 1 slug(s) with no verdict in ctv_fold.mjs. Working as documented.
  • Flipping a tier letter on a paper that stays in the candidate set — B to OUT on Tracking You from a Thousand Miles Awaythe script exited 0, silently recomputed the population as 34 instead of 35 and OUT as 53 instead of 52, and every percentage on the content page would have moved with it. No assertion fired, because the only check was a slug-set difference.

Fixed the same day, by pinning the split to what connected_tv publishes:

const PUBLISHED_SPLIT = { A: 13, B: 22, ADJ: 16, OUT: 52 };

And that fix was itself insufficient — found by the next review pass, on 2026-09-13. A count is not a membership. Mutating two entries at once, swapping a genuine Tier B paper out and a genuine OUT paper in, leaves all four counts identical and exits 0, while the population silently acquires a paper that is not about television at all. The reviewer demonstrated it with BLERP: BLE Re-Pairing Attacks and Defenses (B to OUT) against a bilingual privacy-policy paper whose only TV content is an ACR homonym (OUT to B): platform web went from 1 to 2, PETS 6 to 7, NDSS 6 to 5, and the iot share from 74.3% to 71.4% — every one of them a published figure, none of them guarded.

So membership itself is now pinned, per tier, as a digest of the sorted slug list — and the topic tag, which had no guard at all and drives the iot-device-set list on the content page, is pinned the same way:

const PUBLISHED_MEMBERS = {
  A: 'b7410e5f7a33a91e', B: '4a71c82990115cdf', ADJ: '1eb63ba64295cda8', OUT: '22530ed51f0ceb35',
};
const PUBLISHED_TOPICS = 'fdb92e2a4c3ab263';

Re-run on 2026-09-13, all four mutations now fail and the unmutated script exits 0:

Mutation Before 2026-09-13 Now
Delete a map entry throws throws (audit set has 1 slug(s) with no verdict)
Flip one tier letter throws (since 2026-09-12) throws (verdict split moved)
Swap two papers between tiers exit 0, population silently wrong throws (tier B membership changed)
Change a topic tag only exit 0, the iot-device-set list silently wrong throws (topic assignments … changed)

The lesson worth carrying: the round-1 fix was written by reading the failure the reviewer demonstrated, and it closed exactly that failure and nothing adjacent to it. Only a second mutation test, by a second reviewer with the same brief, found the hole next to it.

6. Folding, and the residue in full

Two folds are used, and only one aggregates anything.

Tool-name alias fold. report_connected_tv.mjs maps a lower-cased alphanumeric skeleton of each tools[].name onto a canonical display name, for 27 skeletons mapping onto 22 canonical names: mitmproxy / mitmdump / mitmweb to mitmproxy, wireshark / tshark to Wireshark, adb / androiddebugbridge to adb, charles / charlesproxy to Charles Proxy, plus one-to-one entries for tcpdump, Frida, Pi-hole, VirusTotal, apktool, jadx, FlowDroid, LibScout, EasyList, Scapy, Selenium, OpenWPM, Mercury, PingPong, Appium, Raspberry Pi, Monkey and UIAutomator. Only usedOrMentioned == “used” tuples are counted, and the unit is the paper.

The unmapped residue is 191 distinct raw tool names across the 35 papers. Printed in full, because a residue that lives only in a local file is a residue nobody reads — and because in this case reading it is how the broadcast-side instruments were found:

3× DBSCAN | 2× Censys | 2× dnsmasq | 2× Google voice synthesizer | 2× IoT Inspector | 2× nmap | 2× OpenSSL
2× random forest | 2× scikit-learn | 2× t-SNE | 2× TF-IDF | 2× WHOIS | 1× Adam | 1× adb_shell
1× Afatech AF9015 | 1× agglomerative clustering | 1× Anaconda | 1× Analysis Scripts | 1× Androguard
1× Android Debug Bridge (adb) | 1× Android Debug Bridge (ADB) | 1× Android Studio APK Analyzer
1× AntMonitor | 1× apk-mitm | 1× apksigner | 1× AppCensus | 1× Apple trust store
1× Apple Wi-Fi geolocation API | 1× Application Exerciser Monkey | 1× Apriori | 1× arecord | 1× ARKit
1× Avalpa OpenCaster | 1× BeautifulSoup | 1× BeEF Toolkit | 1× BERT | 1× BiLSTM | 1× Bing | 1× Bleak
1× Bumble | 1× Chapoly1305/FindMy | 1× ChatGPT (OpenAI's TextCompletion API) | 1× Chrome | 1× CICFlowmeter
1× CogniCrypt | 1× Common CA Database | 1× Conviva | 1× cosine distance | 1× Criminal IP | 1× crt.sh
1× cryptography/fernet | 1× CryptoGuard | 1× curl | 1× DekTec DTU-215 | 1× DekTec StreamXpress
1× DICE coefficient | 1× Dijkstra's algorithm | 1× DNSDB | 1× DPDK | 1× DroidBot | 1× fastText
1× FCC database of digital TV towers | 1× Flight Radar 24 | 1× Forward feature selection (FFS)
1× Fourier transform | 1× generic deep neural network | 1× GNU TLS | 1× Google Play API
1× Google Public DNS | 1× Google search | 1× Google Search | 1× Google Voice synthesizer | 1× GPS Tracks
1× Gradient Boosting Decision Tree | 1× GSDMM | 1× HDMI Video Capture Device | 1× HiDes UT-100c
1× Hurricane Electric IPv6-over-IPv4 tunnel | 1× IDA Pro | 1× IDAPython
1× IEEE Organizationally Unique Identifier registry | 1× IFTTT | 1× Intel RealSense Camera T265
1× InternalBlue | 1× IP2Location | 1× IPFIX | 1× iptables | 1× IRDB | 1× irgen | 1× IrScrutinizer | 1× Java
1× Keras | 1× Latent Dirichlet Allocation | 1× LightGBM | 1× logistic regression (custom) | 1× MakeHex
1× MAPS | 1× Maven Repository | 1× MaxMind | 1× MaxMind GeoLite2 | 1× MaxMind geolocation database
1× Mbed TLS | 1× MbedTLS | 1× McAfee | 1× median absolute deviation (MAD) | 1× Microsoft trust store
1× Mon(IoT)r | 1× Monkey Application Exerciser | 1× Monkey Application Exerciser for Android Studio
1× Monte Carlo sampling | 1× Mother of all Ad-Blocking | 1× Mozilla trust store | 1× Naïve Bayes
1× NASA SEDAC Metropolitan Statistical Areas dataset | 1× nDPI | 1× nearest-neighbor classifier | 1× Nessus
1× Netdisco | 1× NetFlow | 1× Netify | 1× Nexmon | 1× NFF-Go | 1× NimBLE
1× Non-Negative Matrix Factorization (NMF) | 1× NoxPlayer | 1× Objection | 1× OpenAI Text Completion API
1× OpenCaster | 1× OpenDNS | 1× OpenWRT | 1× OpenWrt/LEDE | 1× OPP-115 | 1× Oracle Java
1× passive network telescope | 1× Passport | 1× Pi-hole Default blocklist | 1× PostgreSQL | 1× PrivBERT
1× Prodigy | 1× ProVerif | 1× pyshark | 1× Python | 1× Python requests/2.31.0 | 1× Python TLS implementation
1× Radare2 | 1× Random Forest | 1× Randoop | 1× Raspberry Pi 3 | 1× Raspberry Pi 4 | 1× Redis
1× RedOrbit HbbTV Emulator | 1× Remote Central Forums | 1× RIPE IPmap | 1× Roku External Control Protocol
1× SciPy | 1× Secure Transport | 1× SHAP | 1× Similarweb | 1× Snorkel | 1× Softflowd | 1× SoSci Survey
1× spaCy | 1× spaCy en_core_web_lg | 1× StopAd smart TV blocklist | 1× TensorFlow
1× The Big Blocklist Collection (Firebog) | 1× TP-Link power plugs | 1× traceroute | 1× TrafficPassthrough
1× Trigger Scripts | 1× TSDuck | 1× Tuya Smart app | 1× TV Fool | 1× tvbus.exe | 1× Unity
1× Validation Scripts | 1× VLC Player | 1× VS1838B | 1× WALA | 1× WiFi Inspector | 1× WiGLE | 1× WiGLE API
1× WireShark/tshark | 1× wolfSSL | 1× WolfSSL | 1× word2vec | 1× XCUITest | 1× XGBoost | 1× YAF
1× Yersinia | 1× Zeek

Nothing in the residue was silently merged and nothing was dropped: the page's instrument table is exactly the 27-alias slice, and the residue is everything else. The residue is the interesting half here. Avalpa OpenCaster, TSDuck, DekTec DTU-215, HiDes UT-100c, Afatech AF9015 and RedOrbit HbbTV Emulator are DVB modulation and stream-authoring tools with no counterpart anywhere else on this wiki; IRDB, irgen, IrScrutinizer, MakeHex and VS1838B are infrared remote tooling; HDMI Video Capture Device, Roku External Control Protocol, tvbus.exe and NoxPlayer are TV-specific automation. A fold that had merged these into “other” would have hidden the page's most useful finding.

No fold is applied to vantage.locations or population.sourceList, and both are published unfolded on the content page as rankings only, never as percentages. This is deliberate: at n=35 the folding error that corpus measures on the vantage field (280 versus 498 for the United States, corpus-wide) is not worth introducing, and the raw strings — “Apartment 1”, “lab space”, “241 countries and territories”, “e-bike route” — are themselves informative about what a TV vantage point is.

7. Quotes: checked against both renderings

scripts/ctv_quotecheck.py matches every phrase either page quotes, or leans on for a figure, against both paper.cols.txt and the PDF text layer via pypdf. Both are needed, and this run proves why:

  • 28 of 33 located in both renderings.
  • 2 in .cols only — the FingerprinTV DBF sentence and the Roku ECP URL, which the PDF text layer scrambles.
  • 3 in the PDF only — de-columning splices. The clearest is [5Kumar, Deepak; Shen, Kelly; Case, Benton; Garg, Deepali; Alperovich, Galina; Kuznetsov, Dmitry; Gupta, Rajarshi; Durumeric, Zakir (2019): "All Things Considered: An Analysis of IoT Devices on Home Networks", in: Proceedings of the USENIX Security Symposium. (Link)]: the sentence “the most popular vendor, Roku, only accounts for 17.4% of media devices” is spliced in .cols into “nd the most poputions of IoT device types, except when a device type accounts lar vendor, Roku, only accounts for 17.4% of media devices for fewer than 1% of devices”. A .cols-only check would have reported a correct quote as NOTFOUND.
  • 0 in neither.

One needle was genuinely wrong, and the check caught it. The first draft asserted “decryption fails for 1 out of 5 (or fewer) TLS connections for 80% of all apps” against [6Varmarken, Janus; Le, Hieu; Shuba, Anastasia; Markopoulou, Athina; Shafiq, Zubair (2020): "The TV is Smart and Full of Trackers: Measuring Smart TV Advertising and Tracking", in: Proceedings on Privacy Enhancing Technologies. (DOI)]. The paper says “decryption fails for 1 out of 10 (or fewer) TLS connections for 55% of all apps; 1 out of 5 (or fewer) TLS connections for 80% of all apps” — two clauses, and the draft had glued the opening of the first to the end of the second, producing a sentence the paper does not contain. The needle was narrowed to the clause that is actually there. This is the reason bare-number needles are avoided in that script.

Unedited output:

cols+pdf  CCS 2019 watching-you-watch-the-tracking-ecosystem-of "present on 69% of Roku channels and 89% of Amazon Fire TV channels"
cols+pdf  CCS 2019 watching-you-watch-the-tracking-ecosystem-of "we were able to install our own cert on the device which allowed u"
cols+pdf  CCS 2019 watching-you-watch-the-tracking-ecosystem-of "that leaked the title of the video to a tracking domain"
PDF only  PETS 2020 the-tv-is-smart-and-full-of-trackers-measuri "1 out of 5 (or fewer) TLS connections for 80% of all apps"
cols+pdf  PETS 2020 the-tv-is-smart-and-full-of-trackers-measuri "314 ATS domains that are unique to the Roku dataset"
cols only PETS 2022 fingerprintv-fingerprinting-smart-tv-apps    "among 80 apps that are made available on all three smart TV platfo"
cols+pdf  PETS 2022 watch-over-your-tv-a-security-and-privacy-an "The analysis found at least one sensitive data flow in 78% of the "
cols+pdf  NDSS 2023 i-still-know-what-you-watched-last-sunday-pr "26 communicate with trackers before the user has expressed their c"
cols+pdf  NDSS 2023 i-still-know-what-you-watched-last-sunday-pr "only block at maximum 44% in 2021 and 81% in 2022"
cols+pdf  IMC 2024 watching-tv-with-the-second-party-a-first-lo "there is a complete absence of communication with any previously i"
cols+pdf  IMC 2024 watching-tv-with-the-second-party-a-first-lo "smart TVs in the UK and the US contact distinct ACR domains"
cols+pdf  IMC 2024 watching-tv-with-the-second-party-a-first-lo "ACR network traffic exists when watching linear TV and when using "
PDF only  USENIX 2019 all-things-considered-an-analysis-of-iot-dev "the most popular vendor, Roku, only accounts for 17.4% of media de"
cols+pdf  IMC 2018 understanding-video-management-planes        "streaming set-top boxes1 dominate by view-hours"
cols+pdf  IMC 2023 in-the-room-where-it-happens-characterizing- "the analysis of the Smart TV ecosystem is left for future work"
cols+pdf  IEEE-SP 2024 surveilling-the-masses-with-wi-fi-based-posi "belong to the streaming television equipment manufacturer Roku"
cols+pdf  IMC 2021 iotls-understanding-tls-usage-in-consumer-io "such as voice assistants, smart TVs and video doorbells"
cols+pdf  USENIX 2025 watch-out-your-tv-box-reversing-and-blocking "they are offered only to those who have purchased specific"
cols+pdf  USENIX 2014 from-the-aether-to-the-ethernet-attacking-th "which requires a minimal budget and infrastructure"
cols+pdf  IMC 2024 iot-bricks-over-v6-understanding-ipv6-usage- "only eight out of 93 devices remain functional"
cols+pdf  CCS 2019 watching-you-watch-the-tracking-ecosystem-of "On Roku, a total of 43 channels failed to properly verify the serv"
cols+pdf  CCS 2019 watching-you-watch-the-tracking-ecosystem-of "794 of the 1000 Roku channels sent at least one request in clearte"
cols+pdf  CCS 2019 watching-you-watch-the-tracking-ecosystem-of "We found 9 channels on Roku and 14 channels on the Fire TV"
cols only CCS 2019 watching-you-watch-the-tracking-ecosystem-of "an HTTP GET request to "http://ROKU_ DEVICE_IP_ADDRESS:8060/keydow"
cols+pdf  PETS 2020 the-tv-is-smart-and-full-of-trackers-measuri "697 Fire TV apps that expose advertising ID alongside serial numbe"
cols+pdf  PETS 2022 fingerprintv-fingerprinting-smart-tv-apps    "96% (N = 961) of the top"
PDF only  PETS 2022 watch-over-your-tv-a-security-and-privacy-an "75% of the apps contain analytics libraries and 77% contain advert"
cols+pdf  USENIX 2021 android-smarttvs-vulnerability-discovery-via "37 unique vulnerabilities, including 11 high-impact cyber threats,"
cols+pdf  USENIX 2023 homespy-the-invisible-sniffer-of-infrared-re "The accuracy increases to 70% for Top3 and 77% for Top5"
cols+pdf  NDSS 2024 acoustic-keystroke-leakage-on-smart-televisi "up to 60.19% of common passwords"
cols+pdf  IMC 2024 watching-tv-with-the-second-party-a-first-lo "the fact that we observe network traffic every 15 seconds suggests"
cols+pdf  IMC 2011 understanding-couch-potatoes-measurement-and "The average number of set-top boxes provisioned was approximately "
cols+pdf  USENIX 2019 all-things-considered-an-analysis-of-iot-dev "are the most common type of device in seven of the eleven regions"

33 quotes: 28 in both renderings, 2 in .cols only, 3 in the PDF only, 0 in neither.

8. External and industry sources

Every one fetched on 2026-09-12, and every one a primary source: a vendor's own developer documentation, a standards body, or a regulator's own press release. None of the figures on the content page comes from a news article, a vendor blog post or a comparison site.

Claim on the page Source How verified
HbbTV 2.0.5, published 2026-02-25, incremental over 2.0.4 (March 2023) HbbTV Association specifications page fetched; the version table lists 1.0 (2010) through 2.0.5 (2026-02-25)
Roku ECP is “a simple RESTful API accessed using HTTP on port 8060”, no authentication documented Roku developer docs, External Control API fetched; the port and the query/device-info endpoint quoted verbatim. Cross-checked against [7Moghaddam, Hooman Mohajeri; Acar, Gunes; Burgess, Ben; Mathur, Arunesh; Huang, Danny Yuxing; Feamster, Nick; Felten, Edward W.; Mittal, Prateek; Narayanan, Arvind (2019): "Watching You Watch: The Tracking Ecosystem of Over-the-Top TV Streaming Devices", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], which uses the same URL form
RIDA, GetRIDA() / IsRIDADisabled(), 30-day temporary ID under limit-ad-tracking Roku developer docs: integrating-roku-advertising-framework for GetRIDA() and the 30-day ID, ifDeviceInfo for IsRIDADisabled() (which is not on the RAF page) fetched; re-verified 2026-09-13
TIFA, getTIFA() / isLATEnabled(), resettable, “no connection to any PII … or DUID” Samsung Smart TV developer docs fetched from the unique-identifiers-for-smarttv guide
Fire TV Advertising ID, advertising_id / limit_ad_tracking, Fire OS 5.2.1.1+ on TV Amazon Developer Policy Center, Advertising ID Policy fetched
Wireless adb needs Android 13 (API 33) for TV, against Android 11 for phones Android developer docs, adb page fetched; the TV/WearOS threshold is stated separately from the phone one
FTC/NJ–VIZIO, $2.2m, 11 million televisions, second-by-second, delete pre-2016-03-01 data US FTC press release, 2017-02-06 fetched
Walmart completed the VIZIO acquisition 2024-12-03 Walmart corporate newsroom fetched
Texas sues Sony, Samsung, LG, Hisense, TCL, 2025-12-15; “every 500 milliseconds” Texas Attorney General press release fetched with curl and a browser User-Agent (WebFetch returns HTTP 402 on this host); date and quote read from the rendered page
Hisense TRO, 2025-12-17 Texas Attorney General press release fetched the same way
Samsung agreement, 2026-02-26; LG agreement, 2026-05-11 Texas Attorney General press releases fetched the same way; the Samsung URL is not the one a search result suggested and 404s under the guessed slug
ATSC 3.0 reaches “more than 76% of U.S. households”; broadcaster applications support profile-based personalisation ATSC deployments and NextGen TV pages fetched; the page's own deployment map is dated July 2026
Artifact repository currency (5 repositories, none touched in 2025–2026); mitmproxy v12.2.3, 2026-05-12 GitHub REST API repos/<owner>/<repo> for archived and pushed_at, plus the default branch's newest commit date, because pushed_at counts any branch

Rejected, and why. A search for recent ACR measurement returned several consumer-facing articles (a “how to disable ACR in 2026” listicle, a cybersecurity blog summarising the IMC paper, a compliance vendor's education page) which between them asserted the LG-15-seconds and Samsung-per-minute cadences, the Texas lawsuit and the Samsung settlement. None was used. The cadences were taken from [8Anselmi, Gianluca; Vekaria, Yash; D'Souza, Alexander; Callejo, Patricia; Mandalari, Anna Maria; Shafiq, Zubair (2024): "Watching TV with the Second-Party: A First Look at Automatic Content Recognition Tracking in Smart TVs", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]'s own text, and the enforcement dates from the Attorney General's releases. One of those articles also attributed the “every 500 milliseconds” figure to the research paper; it is the regulator's pleading, and the paper's own 500 ms figure is an estimate of Samsung's capture rate. The content page keeps those two apart deliberately.

One dead link, and it is in the corpus's own data. The artifact URL the extraction recovered for [9Gopalakrishnan, Vijay; Jana, Rittwik; Ramakrishnan, K. K.; Swayne, Deborah F.; Vaishampayan, Vinay A. (2011): "Understanding couch potatoes: measurement and modeling of interactive usage of IPTV at large scale", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] — www.research.att.com/~kkrama/papers/streamcontrol.pdf — returns HTTP 403 over both http and https and with the tilde encoded. It appears in this page's artifact listing because it is what the paper stated, not because anything on either page rests on it; it is left in place as a record of the paper's own claim. Every URL actually cited as evidence on connected_tv returns 200.

Two further checks worth recording. The Mon(IoT)r testbed software is live at github.com/djdubois/moniotr-core (last pushed 2024-08-09) but the lab's tools page does not publish a smart-TV dataset for download, so the page describes testbed captures as a route without promising a TV dataset exists to fetch. And the PETS landing pages were used to recover author lists for five entries the corpus index lacks; [10Ahmed, Dilawer; Das, Anupam; Zaffar, Fareed (2022): "Analyzing the Feasibility and Generalizability of Fingerprinting Internet of Things Devices", in: Proceedings on Privacy Enhancing Technologies. (DOI)]'s authors were additionally cross-checked against Crossref because its stored PDF has no usable text layer on the title page.

8b. Sources added or corrected in round 2 (2026-09-13)

Claim Primary source, and how it was checked Result
ATSC 3.0 broadcaster applications: advertisingId, filterCode, receiver cookies ATSC A/344:2026-04, ATSC 3.0 Interactive Content, 14 April 2026. PDF downloaded (HTTP 200, 2,363,014 bytes, 202 pages), text extracted with pypdf, each needle matched in whitespace-collapsed text All five quotes FOUND verbatim. data collection returns 0 hits in the whole standard — the phrase the withdrawn 2016 quote used
A/344 current revision the A/344 document page lists 2026-04 (14 April 2026) above 2026-02; the 2026-04 PDF's own title block reads A/344:2026-04 2026-04, not the 2026-02 the author first downloaded. One round-2 reviewer read the standards listing as showing 2025-07 as the latest approved; the PDF's own designation settles it
EU: joint Article 62 GDPR operation on smart TVs Autoriteit Persoonsgegevens landing page (HTTP 200) and the report PDF (HTTP 200, 1,005,269 bytes, 14 pages), both needing a browser User-Agent and a Referer header Real. NL/HU/IT/LI, published 23 September 2025. The off-state percentages (97.52 / 98.84 / 91.10) read out of the PDF's own table
UK: ICO connected-TV programme ICO news release, 11 June 2026 (HTTP 403 to a bare fetcher, HTTP 200 with a browser User-Agent) Real. Quote and attribution to William Malcolm verified in the fetched HTML
Walmart press release names Platform+ the release (HTTP 200, 180,202 bytes), searched after double HTML-entity decoding It does. The round-1 footnote said it did not. The plus sign is entity-encoded, so a tag-strip-only search misses it — this is why the check has to decode entities, not just strip tags. Inscape is genuinely absent
ahn2025_watch Zenodo access Zenodo REST API, records/15646588 (HTTP 200): access_right: open, CC-BY-4.0, title ends [Public Artifact] The paper has two records; the extraction's restricted is true of zenodo.15602938 only. Round 1 attached the label to the open DOI
FingerprinTV code release README fetched raw (HTTP 200) “The FingerprinTV dataset has already been released.” / “Once it is ready for release to the public, the code will be added to this repository. Please stay tuned.” Four years on
FCC Fifth FNPRM is still pending Checked for a Report and Order; none found as of 2026-09-13. The adopted item is FCC 25-72, adopted 2025-10-28, released 2025-10-29 The page's fact-sheet citation stands. One reviewer surfaced a trade-press claim of a May 2026 FCC vote mandating ATSC 3.0 tuners; it appears nowhere on fcc.gov and is rejected

Rejected in round 2. A blog URL for “ADB Wi-Fi 2.0” offered in round 1 returns 404 and stays rejected; the claim rests on developer.android.com/tools/adb instead. The trade-press “FCC voted 3–2 in May 2026” story is rejected as above. Samsung's briefly-vacated TRO is still judged a procedural detail and stays off the content page.

9. Bibliography

32 entries were added to bibliography in this sitting — 31 in the first pass and [11Zhu, Yanzi; Xiao, Zhujun; Chen, Yuxin; Li, Zhijing; Liu, Max; Zhao, Ben Y.; Zheng, Haitao (2020): "Et Tu Alexa? When Commodity WiFi Devices Turn into Adversarial Motion Sensors", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] added after review (see §15) — generated by scripts/bibgen.mjs from data/corpus2/.meta so authors, titles and DOIs are publisher metadata rather than recall. Checks run before appending:

  • Citekey collisions: none of the 31 keys exists in the live bibliography.
  • Duplicate-paper scan: scripts/bib_dedup_scan.py on the merged file reports 0 definite duplicates (rule A, same DOI; rule B, same squashed title). Three of the 89 pre-existing rule-D candidates involve a new key — hu2024_bricks against hu2024_unmasking, and wang2024_characterizing against two other wang2024 keys — and all three are different first authors, so none is a duplicate.
  • Literal @ in a field: none. A raw ASCII @ inside a BibTeX field makes the bibtex4dw plugin drop the entry and every marker to it, silently.
  • Authors: PETS and USENIX records carry no authors in the index. Five PETS entries were filled from the publisher's landing pages (Varmarken et al., Tileria and Blasco, Mandalari et al., Ahmed et al., Mavroudis et al.); eight USENIX entries were resolved by scripts/fetch_authors.py. First and last author of all 31 were then checked against the paper's own PDF text layer, which passed for 30; the one that failed, [10Ahmed, Dilawer; Das, Anupam; Zaffar, Fareed (2022): "Analyzing the Feasibility and Generalizability of Fingerprinting Internet of Things Devices", in: Proceedings on Privacy Enhancing Technologies. (DOI)], has no text layer on its front matter and was confirmed against Crossref instead.
  • Two hand corrections to bibgen's output: the citekey bjrklund2025_endangered was corrected to bjorklund2025_endangered (the generator drops the ö rather than transliterating it), and Jad Al Aaraj was re-split from Aaraj, Jad Al to Al Aaraj, Jad.
  • bibgen's stdout carries QA notes (“no DOI available”, “metadata source: venue-page”, citation counts). Only the @ entries were appended; the notes were stripped by a regex that extracts complete entries, because those notes have previously gone live on the public bibliography page.

This provenance page adds no bibliography entries of its own and uses only keys the content page already uses, plus [3Björklund, Martin; Duvignau, Romaric (2025): "Endangered Privacy: Large-Scale Monitoring of Video Streaming Services", in: Proceedings of the USENIX Security Symposium. (Link)] for the roadmap correction in §1.

10. What could not be established

  • How big the true population is. The 35 is a floor. The 39 papers dropped between gate 1 and gate 2 were not read, and no probe can reach a paper that measures a television without naming a TV-class term in its full text.
  • Whether the Tier B line is where somebody else would put it. “Reports a result broken out for a television” is a reading, not a field. Moving Void and the cryptojacking proof-of-concept in would make it 37; requiring a TV-specific privacy or security result rather than any result would make it roughly 28.
  • What ACR does now. The only measurement is 2024, on two sets, and both of those vendors have since agreed consent changes with a US regulator. The page says the 2024 figures are a pre-order baseline; it does not claim to know the current behaviour, and nothing in the corpus does.
  • Anything about ATSC 3.0. Not one corpus paper measures it. The deployment share and the personalisation capabilities on the page are ATSC's own statements about its own standard, labelled as such.
  • TV market share by installed base. Repeatedly useful and repeatedly unavailable from a primary source that is not a paid analyst report. The page therefore never says which platform is biggest, only what the literature measured.
  • Whether TLS interception has improved on closed platforms since 2019. No paper in the corpus revisits it. The 4.3% figure is quoted with its date attached rather than as a current state.
  • A figure for the TV share of household or video traffic that a privacy paper could use as a denominator. [12Wang, Yifan; Lyu, Minzhao; Sivaraman, Vijay (2024): "Characterizing User Platforms for Video Streaming in Broadband Networks", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] and [13Akhtar, Zahaib; Nam, Yun Seong; Chen, Jessica; Govindan, Ramesh; Katz-Bassett, Ethan; Rao, Sanjay G.; Zhan, Jibin; Zhang, Hui (2018): "Understanding Video Management Planes", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] come closest and neither gives one; this is written up as an open question rather than back-calculated.

11. Judgement calls

  • A new page rather than a section of mobile_and_app_measurement. Android TV apps genuinely are Android apps, and that page's store, static-analysis and pinning material transfers. But ACR, HbbTV and “you cannot install a certificate at all” have no mobile analogue, and 26 of the 35 papers carry iot rather than mobile. The content page points at the mobile page rather than restating it, in four places.
  • Not a section of platforms. A television is a device, not a platform whose API you negotiate access to. The overlap is the store chart, and that is one row.
  • The blocklist-coverage numbers are shared with filter_lists rather than moved. That page already carries a smart-TV row citing [6Varmarken, Janus; Le, Hieu; Shuba, Anastasia; Markopoulou, Athina; Shafiq, Zubair (2020): "The TV is Smart and Full of Trackers: Measuring Smart TV Advertising and Tracking", in: Proceedings on Privacy Enhancing Technologies. (DOI)] with the 22%/27% figures. The content page repeats them once, in the section explaining why filter lists do not transfer, and links there rather than re-deriving.
  • IPTV performance work is Tier A, not adjacent. [9Gopalakrishnan, Vijay; Jana, Rittwik; Ramakrishnan, K. K.; Swayne, Deborah F.; Vaishampayan, Vinay A. (2011): "Understanding couch potatoes: measurement and modeling of interactive usage of IPTV at large scale", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] and [14Song, Han Hee; Ge, Zihui; Mahimkar, Ajay; Wang, Jia; Yates, Jennifer; Zhang, Yin; Basso, Andrea; Chen, Min (2011): "Q-score: proactive service quality assessment in a large IPTV system", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] measure set-top boxes, which the rule calls television-class. They are tagged delivery-performance so the topic table shows that two of the four oldest papers in the population are not privacy work at all. A privacy-only page would have dropped them and reported a smaller, tidier, less honest literature.
  • The enforcement timeline is on the content page, not only here. A student planning an ACR measurement in 2026 who does not know about the Texas agreements will measure a consent flow and report it as a default. That is a methodological fact, so it belongs on the content page.
  • Per-year counts stop being a series. With 35 papers over 16 years, the page reports four-to-five-year windows and labels 2025–2026 provisional in the table itself, rather than drawing a trend.

12. The scripts

Three files, committed with their real output.

ctv_fold.mjs
// Population map for `design:connected_tv`.
//
// The roadmap queued this page against a 16-paper TITLE+SUMMARY candidate set
// (scripts/gap_probe_roadmap.mjs, family `ctv_streaming`). That probe is a
// floor, and it is also the wrong instrument twice over: it misses papers whose
// title never says "TV" (Watch Over Your TV is in it only by accident of the
// word "TV"; the Android TV ecosystem paper was NOT in the 16), and its `web`
// column — 1 of 16 — is not a filter anyone should apply here, because a
// television is not a web-platform measurement in the extraction's sense.
//
// So the population is derived instead from a full-text recall probe over all
// 5,859 `paper.cols.txt` files (scripts/_ctv_probe1.mjs / _ctv_probe2.mjs),
// gated into an audit set, and then HAND-AUDITED against a written rule.
//
// THE RULE, fixed before any figure was computed:
//
//   A "television-class endpoint" is a smart TV set, a TV operating system
//   (Android TV / Google TV, tvOS, Tizen, webOS, Roku OS, Fire OS), a streaming
//   stick / box / set-top box, an app running on one, or the broadcast path
//   (HbbTV / DVB) delivered into one.
//
//   A  the paper's central object of measurement is a television-class endpoint
//      (its traffic, apps, firmware, broadcast channel or user interaction).
//   B  television-class devices are part of a broader measured population AND
//      the paper reports at least one result broken out for them.
//   ADJ adjacent: cite where relevant, do not count. Video streaming measured
//      off a TV (browser DRM, piracy websites, encrypted-traffic video
//      fingerprinting), or a TV used as apparatus rather than measured.
//   OUT the TV name is a passing reference, a survey answer option, a related-
//      work sentence, or a homonym.
//
// `topic` is hand-assigned and only used for a ranking, never a percentage.
//
// Every slug in the audit set appears here. report_connected_tv.mjs throws if
// the audit set and this map disagree, so widening a probe breaks the report
// instead of silently moving the page's denominator.
 
export const MAP = {
  // ---------------------------------------------------------------- Tier A
  'from-the-aether-to-the-ethernet-attacking-the-internet-using-broadcast-digital-t':
    ['A', 'broadcast', 'Injects HbbTV/DVB payloads into smart TVs over the broadcast band; TVs and STBs are the target population.'],
  'watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices':
    ['A', 'tracking', '1,000 Roku channels and 1,000 Fire TV channels crawled on real devices with TLS interception.'],
  'the-tv-is-smart-and-full-of-trackers-measuring-smart-tv-advertising-and-tracking':
    ['A', 'tracking', 'Roku and Fire TV app traffic, testbed plus in-the-wild; the reference smart-TV tracking measurement.'],
  'android-smarttvs-vulnerability-discovery-via-log-guided-fuzzing':
    ['A', 'vulnerability', 'Log-guided fuzzing of 11 Android TV devices; the TV firmware is the object.'],
  'fingerprintv-fingerprinting-smart-tv-apps':
    ['A', 'tracking', 'Top-1000 apps on each of Apple TV, Fire TV and Roku; network fingerprints of TV apps.'],
  'watch-over-your-tv-a-security-and-privacy-analysis-of-the-android-tv-ecosystem':
    ['A', 'app-analysis', '4,745 Android TV APKs statically analysed plus 21 apps intercepted. NOT in the roadmap candidate set.'],
  'i-still-know-what-you-watched-last-sunday-privacy-of-the-hbbtv-protocol-in-the-european-smart-tv-landscape':
    ['A', 'broadcast', '36 European HbbTV channels on real TVs, plus a 174-respondent awareness survey.'],
  'homespy-the-invisible-sniffer-of-infrared-remote-control-of-smart-tvs':
    ['A', 'side-channel', 'IR remote-control signals of smart TVs sniffed by neighbouring IoT devices.'],
  'acoustic-keystroke-leakage-on-smart-televisions':
    ['A', 'side-channel', 'On-screen-keyboard keystrokes recovered from TV audio on Apple and Samsung TVs.'],
  'watching-tv-with-the-second-party-a-first-look-at-automatic-content-recognition':
    ['A', 'acr', 'Two smart TVs (LG, Samsung) in UK and US; ACR traffic across six viewing scenarios.'],
  'watch-out-your-tv-box-reversing-and-blocking-a-p2p-based-illegal-streaming-ecosy':
    ['A', 'piracy', 'Reverses the EVPAD illegal-streaming set-top box and its P2P ecosystem.'],
  'understanding-couch-potatoes-measurement-and-modeling-of-interactive-usage-of-ip':
    ['A', 'delivery-performance', 'Two years of interaction traces from ~3M IPTV set-top boxes.'],
  'q-score-proactive-service-quality-assessment-in-a-large-iptv-system':
    ['A', 'delivery-performance', 'IPTV service quality inferred for millions of set-top boxes from network measurements and STB logs.'],
 
  // ---------------------------------------------------------------- Tier B
  'information-exposure-from-consumer-iot-devices-a-multidimensional-network-inform':
    ['B', 'iot-device-set', '81-device lab; Apple TV, Fire TV, LG TV, Roku TV and Samsung TV are named devices with per-device destinations.'],
  'a-haystack-full-of-needles-scalable-detection-of-iot-devices-in-the-wild':
    ['B', 'iot-device-set', 'Video category = Apple TV, Fire TV, LG TV, Roku TV, Samsung TV; detected in IXP flow data.'],
  'iotls-understanding-tls-usage-in-consumer-iot-devices':
    ['B', 'iot-device-set', 'TV category n=5 (Fire TV, Samsung TV, LG TV, Roku TV, Apple TV) with per-device TLS results.'],
  'blocking-without-breaking-identification-and-mitigation-of-non-essential-iot-tra':
    ['B', 'iot-device-set', 'Fire TV and Roku TV in the 31-device set with per-device destination and breakage results.'],
  'packet-level-signatures-for-smart-home-devices':
    ['B', 'iot-device-set', 'Smart TVs among the devices from which packet-level signatures were extracted.'],
  'analyzing-the-feasibility-and-generalizability-of-fingerprinting-internet-of-thi':
    ['B', 'iot-device-set', 'Roku TV reported as its own row with per-device fingerprinting accuracy and a confusion analysis.'],
  'behind-the-scenes-uncovering-tls-and-server-certificate-practice-of-iot-device-v':
    ['B', 'iot-device-set', 'A dedicated "smart TV and local device" capture is analysed as a case study.'],
  'in-the-room-where-it-happens-characterizing-local-communication-and-threats-in-s':
    ['B', 'iot-device-set', 'Smart TVs are one of eight testbed device categories; the paper also flags TV apps as a local-network threat.'],
  'iot-bricks-over-v6-understanding-ipv6-usage-in-smart-homes':
    ['B', 'iot-device-set', '93 devices; three smart TVs among the eight that still work on IPv6-only, with per-device domains.'],
  'connecting-the-dots-tracing-data-endpoints-in-iot-devices':
    ['B', 'iot-device-set', 'Roku and Samsung Smart TV among the fingerprinted device set; User-Agent and OUI evidence quoted for TVs.'],
  'evaluating-machine-learning-based-iot-device-identification-models-for-security-applications':
    ['B', 'iot-device-set', 'Smart TV is a labelled class (Fire TV, Apple TV, LG webOS TV, Roku TV, Samsung SmartTV) with per-device idleness results.'],
  'lumos-identifying-and-localizing-diverse-hidden-iot-devices-in-an-unfamiliar-env':
    ['B', 'iot-device-set', 'TV class = Vizio, Panasonic, TCL in the 44-device set; unseen smart TVs discussed in the field test.'],
  'et-tu-alexa-when-commodity-wifi-devices-turn-into-adversarial-motion-sensors':
    ['B', 'iot-device-set', 'Chromecast, Apple TV and Roku form the "Smart TV (& Sticks)" row with its own packet-rate measurement.'],
  'all-things-considered-an-analysis-of-iot-devices-on-home-networks':
    ['B', 'iot-device-set', '15.5M homes; Media/TV is a device category and Roku is named with a 17.4% within-category share.'],
  'you-are-what-you-broadcast-identification-of-mobile-and-iot-devices-from-public':
    ['B', 'iot-device-set', 'Apple TV identified from mDNS/DHCP/SSDP views, with a false-positive case study on Apple TV.'],
  'characterizing-user-platforms-for-video-streaming-in-broadband-networks':
    ['B', 'delivery-performance', 'Smart TV is one of the classified device types for 100M+ video flows; gives the TV share of streaming.'],
  'understanding-video-management-planes':
    ['B', 'delivery-performance', 'Streaming set-top boxes (Roku, Fire TV, Apple TV) are a platform category and dominate by view-hours.'],
  'missing-present-and-conflicting-a-large-scale-analysis-of-iot-update-information':
    ['B', 'policy-compliance', 'Smart TVs are one of five device types crawled across 58 EU stores, with a TV-specific disclosure result.'],
  'understanding-iot-security-from-a-market-scale-perspective':
    ['B', 'app-analysis', 'Market-scale mobile-IoT app analysis; Fire TV and HiSense TV appear as identified products with findings.'],
  'blerp-ble-re-pairing-attacks-and-defenses':
    ['B', 'vulnerability', 'TCL 43P638 Android TV is row 12 of the tested-device table with its own attack outcome.'],
  'tracking-you-from-a-thousand-miles-away-turning-a-bluetooth-device-into-an-apple':
    ['B', 'vulnerability', 'Sony Bravia A80J (Android TV 10) is in the tested-device table with its own result.'],
 
  'surveilling-the-masses-with-wi-fi-based-positioning-systems':
    ['B', 'device-population', 'Roku streaming devices are two of the five most-observed BSSID OUIs in a 490M-entry Wi-Fi positioning dataset; Roku-specific shares reported.'],
 
  // ---------------------------------------------------------------- Adjacent
  'endangered-privacy-large-scale-monitoring-of-video-streaming-services':
    ['ADJ', 'streaming-service', 'Video identification from encrypted MPEG-DASH traffic. The roadmap listed it as CTV spine; it measures the SERVICE and its traffic, never a TV.'],
  'unmasking-the-shadows-a-cross-country-study-of-online-tracking-in-illegal-movie':
    ['ADJ', 'streaming-service', 'Illegal movie streaming WEBSITES crawled with a browser; a web-tracking study, not a TV study.'],
  'your-drm-can-watch-you-too-exploring-the-privacy-implications-of-browsers-mis-im':
    ['ADJ', 'streaming-service', 'Widevine EME in browsers and Android; TVs named as another Widevine host, not measured.'],
  'cost-saving-streaming-unlocking-the-potential-of-alternative-edge-node-resources':
    ['ADJ', 'streaming-service', 'Edge-node economics for streaming delivery; no TV endpoint measured.'],
  'measurement-and-analysis-of-a-large-scale-commercial-mobile-internet-tv-system':
    ['ADJ', 'streaming-service', '"TV" delivered to mobile handsets, not to a television.'],
  'watching-videos-from-everywhere-a-study-of-the-pptv-mobile-vod-system':
    ['ADJ', 'streaming-service', 'Mobile VoD; no television endpoint.'],
  'performance-characterization-of-a-commercial-video-streaming-service':
    ['ADJ', 'streaming-service', 'Streaming service performance from browser/CDN vantage; no TV-specific result.'],
  'analyzing-the-potential-benefits-of-cdn-augmentation-strategies-for-internet-vid':
    ['ADJ', 'streaming-service', 'CDN augmentation for video workloads; no TV endpoint.'],
  'anatomy-of-a-personalized-livestreaming-system':
    ['ADJ', 'streaming-service', 'Livestreaming (Periscope) system measurement; no TV endpoint.'],
  'peer-assisted-content-distribution-in-akamai-netsession':
    ['ADJ', 'streaming-service', 'Peer-assisted CDN; set-top-box mention is background.'],
  'auto-draft-196': // slug is a placeholder in the index; title = A Lightweight IoT Cryptojacking Detection Mechanism in Heterogeneous Smart Home Networks (NDSS 2022)
    ['ADJ', 'vulnerability', 'Authors implement their own cryptojacking PoC on an LG webOS TV to test a detector; no deployed-TV population.'],
  'void-a-fast-and-light-voice-liveness-detection-system':
    ['ADJ', 'apparatus', 'A Samsung Smart TV is used as a replay LOUDSPEAKER; the TV is apparatus, not the measured object.'],
  'tracking-profiling-and-ad-targeting-in-the-alexa-echo-smart-speaker-ecosystem':
    ['ADJ', 'iot-device-set', 'Smart speaker study; TVs cited as the comparable prior ecosystem, not measured. The closest methodological sibling.'],
  'ovrseen-auditing-network-traffic-and-privacy-policies-in-oculus-vr':
    ['ADJ', 'iot-device-set', 'VR headsets; smart TVs used as the comparison ecosystem. Same lab, same pipeline shape.'],
  'exploiting-diversity-in-android-tls-implementations-for-mobile-app-traffic-class':
    ['ADJ', 'app-analysis', 'Android app traffic classification; TLS-fingerprint method later reused on TV apps.'],
 
  'on-the-privacy-and-security-of-the-ultrasound-ecosystem':
    ['ADJ', 'acr', 'Ultrasonic cross-device tracking (uXDT): beacons emitted by TV adverts and picked up by phone SDKs. The TV is the emitter, never measured.'],
 
  // ---------------------------------------------------------------- Out
  'a-bilingual-longitudinal-analysis-of-privacy-policies-measuring-the-impacts-of-t': ['OUT','','ACR homonym: "ACR" is not automatic content recognition here.'],
  'a-billion-open-interfaces-for-eve-and-mallory-mitm-dos-and-tracking-attacks-on-i': ['OUT','','tvOS listed among Apple OSes; AWDL is the object.'],
  'a-multi-region-investigation-of-the-perceptions-and-use-of-smart-home-devices': ['OUT','','Survey; smart TV is a related-work citation and an ownership option.'],
  'a-placement-vulnerability-study-in-multi-tenant-public-clouds': ['OUT','','Title probe matched "streaming"; cloud VM placement.'],
  'analyzing-the-ios-local-network-permission-from-a-technical-and-user-perspective': ['OUT','','Chromecast is one of four IoT devices used to trigger the permission; no TV result.'],
  'broadcast-yourself-understanding-youtube-uploaders': ['OUT','','IPTV appears once in related work.'],
  'characterizing-everyday-misuse-of-smart-home-devices': ['OUT','','Survey of 483 people; smart TV is an ownership option, not a measured device.'],
  'co-designing-a-mobile-app-for-bystander-privacy-protection-in-jordanian-smart-ho': ['OUT','','Interview study; TV names are participant-reported device inventories.'],
  'contextualizing-interpersonal-data-sharing-in-smart-homes': ['OUT','','Vignette survey; "viewing history from your smart TV" is a question stem.'],
  'dead-domains-living-data-a-privacy-risk-analysis-of-domain-lifecycle-in-android': ['OUT','','Android apps; one expired-domain example happens to also ship on Roku.'],
  'deep-dive-into-the-iot-backend-ecosystem': ['OUT','','Backend infrastructure; TV mentions are motivation and a citation to FingerprinTV.'],
  'drones-cryptanalysis-smashing-cryptography-with-a-flicker': ['OUT','','IPTV/ACR homonyms.'],
  'entropy-ip-uncovering-structure-in-ipv6-addresses': ['OUT','','"ACR" homonym.'],
  'examining-consumer-reviews-to-understand-security-and-privacy-issues-in-the-mark': ['OUT','','Review-text analysis; set-top box is a Mirai product category, no TV measurement.'],
  'examining-power-dynamics-and-user-privacy-in-smart-technology-use-among-jordania': ['OUT','','Interview study; TVs are participant device inventories.'],
  'exploring-the-privacy-concerns-of-bystanders-in-smart-homes-from-the-perspective': ['OUT','','Survey; smart TV is an example in a prompt.'],
  'flock-combating-astroturfing-on-livestreaming-platforms': ['OUT','','Title probe matched "streaming platform"; astroturfing detection on Twitch-like sites.'],
  'help-me-help-you-privacy-considerations-for-third-party-iot-device-repair': ['OUT','','Vignette survey; TVs appear in a device-category prompt.'],
  'hey-your-secrets-leaked-detecting-and-characterizing-secret-leakage-in-the-wild': ['OUT','','IPTV homonym in leaked-credential data.'],
  'idea-static-analysis-on-the-security-of-apple-kernel-drivers': ['OUT','','tvOS is one of four Apple OSes scanned; no TV-specific result.'],
  'internet-service-providers-and-individuals-attitudes-barriers-and-incentives-to': ['OUT','','Interview and survey; TV is a device-ownership row.'],
  'iotflow-inferring-iot-device-behavior-at-scale-through-static-mobile-companion-a': ['OUT','','Companion-app analysis; no TV breakout.'],
  'its-the-equivalent-of-feeling-like-youre-in-jail-lessons-from-firsthand-and-seco': ['OUT','','Interview study on IPV; TV is a reported abuse vector, not measured.'],
  'measuring-up-to-reasonable-consumer-expectations-providing-an-empirical-basis-fo': ['OUT','','Vignette survey; Vizio appears only in a news citation.'],
  'medical-devices-are-at-risk-information-security-on-diagnostic-imaging-system': ['OUT','','"ACR" = American College of Radiology.'],
  'my-best-friends-husband-sees-and-knows-everything-a-cross-contextual-and-cross-c': ['OUT','','Survey; smart TV is a free-text mention count.'],
  'no-privacy-among-spies-assessing-the-functionality-and-insecurity-of-consumer-an': ['OUT','','Android stalkerware; "ACR" homonym.'],
  'non-intrusive-and-unconstrained-keystroke-inference-in-vr-platforms-via-infrared-side-channel': ['OUT','','VR; smart TV appears only as a citation to HomeSpy.'],
  'nothing-else-mator-s-monitoring-the-anonymity-of-tors-path-selection': ['OUT','','"ACR" homonym.'],
  'on-the-vulnerability-of-fpga-bitstream-encryption-against-power-analysis-attacks': ['OUT','','Set-top box named as an FPGA application domain.'],
  'privacy-perspectives-and-practices-of-chinese-smart-home-product-teams': ['OUT','','Interview study; smart TV is a company product-line row.'],
  'regulating-smart-device-support-periods-user-expectations-and-the-european-cyber': ['OUT','','Survey; Smart TV is a self-reported ownership category.'],
  'rocking-drones-with-intentional-sound-noise-on-gyroscopic-sensors': ['OUT','','Passing mention.'],
  'same-origin-policy-evaluation-in-modern-browsers': ['OUT','','Passing mention of TV browsers.'],
  'sandscout-automatic-detection-of-flaws-in-ios-sandbox-profiles': ['OUT','','Apple TV named as a device that runs iOS/tvOS; iOS sandbox is the object.'],
  'sissi-an-architecture-for-semantic-interoperable-self-sovereign-identity-based-a': ['OUT','','"ACR" homonym (authentication context reference).'],
  'smart-devices-in-airbnbs-considering-privacy-and-security-for-both-guests-and-ho': ['OUT','','Survey; smart TV is a scenario option.'],
  'snapshot-based-loading-acceleration-of-web-apps-with-nondeterministic-javascript': ['OUT','','Tizen/webOS named as embedded web-app platforms; benchmark is web apps.'],
  'sok-technical-implementation-and-human-impact-of-internet-privacy-regulations': ['OUT','','SoK; TV work cited, not measured.'],
  'tbtrackerx-fantastic-trigger-bots-and-where-to-find-malicious-campaigns-on-x': ['OUT','','IPTV spam homonym.'],
  'webspec-towards-machine-checked-analysis-of-browser-security-mechanisms': ['OUT','','"ACR" homonym.'],
  'what-happened-in-my-network-mining-network-events-from-router-syslogs': ['OUT','','IPTV named as the service carried; router syslogs are the object.'],
  'who-cares-contextual-privacy-judgments-from-owner-and-bystander-perspectives-in': ['OUT','','Survey; smart TV is a device-category option.'],
  'why-can-t-users-choose-their-identity-providers-on-the-web': ['OUT','','"ACR" homonym.'],
  'you-are-who-you-know-and-how-you-behave-attribute-inference-attacks-via-users-so': ['OUT','','IPTV homonym.'],
 
  'poster-watch-out-your-smart-watch-when-paired': ['OUT','','Tizen here is the smartwatch platform, not the TV one.'],
  'cleaning-up-the-internet-of-evil-things-real-world-evidence-on-isp-and-consumer-efforts-to-remove-mirai': ['OUT','','One infected set-top box in a Mirai remediation table; no TV finding.'],
  'latex-gloves-protecting-browser-extensions-from-probing-and-revelation-attacks': ['OUT','','The Chromecast browser EXTENSION, not the device.'],
  'abuse-vectors-a-framework-for-conceptualizing-iot-enabled-interpersonal-abuse': ['OUT','','Qualitative framework; TV is an example abuse vector.'],
  'exploring-tenants-preferences-of-privacy-negotiation-in-airbnb': ['OUT','','Vignette survey; smart TV is a device-type option.'],
  'pangolin-fuzzing-multilingual-iot-firmware-with-llm-driven-code-analysis': ['OUT','','"SmartTVs" is a citation to the 2021 fuzzing paper, used as a baseline name.'],
  'utopia-automatic-generation-of-fuzz-driver-using-unit-tests': ['OUT','','Tizen as an open-source project under test; no TV device.'],
};
 
export const tier = (t) => Object.entries(MAP).filter(([, v]) => v[0] === t).map(([k]) => k);
report_connected_tv.mjs
// Report script for `design:connected_tv`.
//
// Every figure on that page is printed here with its own denominator. The page
// carries no number this script cannot produce.
//
// Structure:
//   0. Rebuild the candidate pool from the raw corpus and assert it still
//      matches scripts/ctv_fold.mjs exactly, in both directions.
//   1. The population, by tier, venue, year and topic.
//   2. Why the roadmap's `web` column was the wrong filter: platform fields.
//   3. How these papers get at the traffic (tools, interception, vantage).
//   4. What they sample (population sources, n, units) — the no-Tranco problem.
//   5. Measured results, quoted from detection[].prevalence.
//   6. Where the field goes quiet, on the CTV population vs the corpus.
//   7. Quote verification against paper.cols.txt.
//
// Usage: node scripts/report_connected_tv.mjs
//        node scripts/report_connected_tv.mjs --format wiki
import fs from 'node:fs';
import path from 'node:path';
import { createHash } from 'node:crypto';
import { loadExtractions, dataRoot, isSentinel, pct, table, wikiTable } from './lib.mjs';
import { MAP } from './ctv_fold.mjs';
 
const ROOT = dataRoot();
const WIKI = process.argv.includes('--format') && process.argv[process.argv.indexOf('--format') + 1] === 'wiki';
const H = (s) => console.log('\n' + '='.repeat(78) + '\n' + s + '\n' + '='.repeat(78));
const T = (h, r) => console.log(WIKI ? wikiTable(h, r) : table(h, r));
 
const ALL = loadExtractions();
const ftPath = (p) => path.join(ROOT, 'fulltext', String(p.year), p.venue, p.slug, 'paper.cols.txt');
const collapse = (s) => s.replace(/­/g, '').replace(/-\n/g, '').replace(/\s+/g, ' ');
const ftCache = new Map();
const fulltext = (p) => {
  if (!ftCache.has(p.slug)) {
    const f = ftPath(p);
    ftCache.set(p.slug, fs.existsSync(f) ? collapse(fs.readFileSync(f, 'utf8')) : '');
  }
  return ftCache.get(p.slug);
};
 
// ---------------------------------------------------------------------------
// 0. Rebuild the candidate pool and check the hand map against it.
// ---------------------------------------------------------------------------
const PROBES = {
  smarttv: /\bsmart[-\s]?TVs?\b/i,
  ctv: /\bconnected[-\s]TVs?\b|\bCTV\b/i,
  ott: /\bover[-\s]the[-\s]top\b|\bOTT\b/i,
  hbbtv: /\bHbbTV\b|\bhybrid broadcast broadband\b/i,
  acr: /\bautomatic content recognition\b|\bACR\b/i,
  platformdev: /\bRoku\b|\bFire ?TV\b|\bApple ?TV\b|\bChromecast\b|\bAndroid ?TV\b|\bGoogle ?TV\b|\btvOS\b|\bWebOS\b|\bTizen\b|\bset[-\s]?top box(es)?\b/i,
  streamsvc: /\bNetflix\b|\bHulu\b|\bDisney\+|\bAmazon Prime Video\b|\bYouTube ?TV\b|\bTwitch\b/i,
  tvapp: /\bTV app(s|lication)?\b|\btelevision app(s)?\b/i,
  iptv: /\bIPTV\b|\binternet protocol television\b/i,
};
// The roadmap's own title+summary probe, reproduced verbatim from
// scripts/gap_probe_roadmap.mjs so the two can be compared.
const ROADMAP_TITLE = /smart ?TV|connected TV|\bCTV\b|roku|streaming (device|platform|service)|set-?top box/i;
// Device-name probe used to find TVs inside broader IoT device sets.
const DEV = /\b(Roku|Fire ?TV|Apple ?TV|Chromecast|Android ?TV|Google ?TV|tvOS|WebOS|Tizen|Vizio|Hisense|Bravia|Nvidia Shield|Samsung(?: Smart)? TV|LG(?: Smart)? TV|TCL|smart[- ]?TVs?|set[- ]?top box(?:es)?)\b/gi;
 
let noFulltext = 0;
const scored = [];
for (const p of ALL) {
  const txt = fulltext(p);
  if (!txt) { noFulltext += 1; continue; }
  const c = {};
  for (const [k, re] of Object.entries(PROBES)) c[k] = (txt.match(new RegExp(re.source, re.flags + 'g')) || []).length;
  const core = c.smarttv + c.ctv + c.hbbtv + c.platformdev + c.tvapp;
  const brands = new Set((txt.match(DEV) || []).map((x) => x.toLowerCase().replace(/\s+/g, '')));
  const devN = (txt.match(DEV) || []).length;
  const titleHit = ROADMAP_TITLE.test(`${p.title ?? ''} • ${p.summary ?? ''}`);
  scored.push({ p, c, core, brands, devN, titleHit });
}
 
// GATE 1 (wide): anything with a real amount of TV vocabulary anywhere.
const gate1 = scored.filter((s) => s.core >= 2 || s.c.acr >= 2 || s.c.iptv >= 2 || s.c.hbbtv >= 1 || s.c.tvapp >= 1 || s.titleHit);
// GATE 2 (audit set): gate 1 narrowed to what is worth reading in full.
const AUDIT = gate1.filter((s) => s.devN >= 4 || s.brands.size >= 3 || s.titleHit || s.c.hbbtv > 0 || s.c.acr >= 2 || s.c.iptv >= 2);
const auditSlugs = new Set(AUDIT.map((s) => s.p.slug));
 
const mapSlugs = new Set(Object.keys(MAP));
const missingFromMap = [...auditSlugs].filter((s) => !mapSlugs.has(s));
const missingFromAudit = [...mapSlugs].filter((s) => !auditSlugs.has(s));
if (missingFromMap.length) throw new Error(`audit set has ${missingFromMap.length} slug(s) with no verdict in ctv_fold.mjs:\n  ${missingFromMap.join('\n  ')}`);
if (missingFromAudit.length) throw new Error(`ctv_fold.mjs has ${missingFromAudit.length} slug(s) the audit set no longer selects:\n  ${missingFromAudit.join('\n  ')}`);
 
const bySlug = new Map(ALL.map((p) => [p.slug, p]));
const verdict = (t) => Object.entries(MAP).filter(([, v]) => v[0] === t).map(([s]) => ({ p: bySlug.get(s), v: MAP[s] }))
  .sort((a, b) => a.p.year - b.p.year || a.p.venue.localeCompare(b.p.venue));
const A = verdict('A'), B = verdict('B'), ADJ = verdict('ADJ'), OUT = verdict('OUT');
const POP = [...A, ...B];                       // the page's population
const nA = A.length, nB = B.length, nPOP = POP.length;
 
// The slug-set check above only catches a slug appearing or disappearing. A
// mutation test on 2026-09-12 showed that FLIPPING a tier letter — 'B' to 'OUT'
// on a paper that stays in the candidate set — passed every assertion and
// silently moved the population from 35 to 34, changing every percentage on the
// page. So the split itself is pinned to what design:connected_tv publishes.
// Changing the map deliberately means changing these four numbers too.
const PUBLISHED_SPLIT = { A: 13, B: 22, ADJ: 16, OUT: 52 };
const actualSplit = { A: nA, B: nB, ADJ: ADJ.length, OUT: OUT.length };
for (const k of Object.keys(PUBLISHED_SPLIT)) {
  if (actualSplit[k] !== PUBLISHED_SPLIT[k])
    throw new Error(
      `verdict split moved: ${k} is ${actualSplit[k]}, design:connected_tv publishes ${PUBLISHED_SPLIT[k]}. ` +
      `Full split now ${JSON.stringify(actualSplit)} against published ${JSON.stringify(PUBLISHED_SPLIT)}. ` +
      `If the map change is intended, update PUBLISHED_SPLIT and every figure on the page.`
    );
}
const splitTotal = nA + nB + ADJ.length + OUT.length;
if (splitTotal !== AUDIT.length)
  throw new Error(`verdicts sum to ${splitTotal} but the audit set is ${AUDIT.length}`);
 
// A second mutation test, on 2026-09-13, broke the four counts above: SWAPPING
// two papers between tiers — one genuine Tier B out, one genuine OUT in —
// leaves A/B/ADJ/OUT unchanged and exits 0, while the population silently gains
// a paper that is not about television (platform 'web' went 1 -> 2, the venue
// and year tables both moved). Counts are not membership. So the membership
// itself is pinned, per tier, as a digest of the sorted slug list.
const digest = (rows) => createHash('sha256')
  .update(rows.map((x) => x.p.slug).sort().join('\n')).digest('hex').slice(0, 16);
const PUBLISHED_MEMBERS = {
  A: 'b7410e5f7a33a91e', B: '4a71c82990115cdf', ADJ: '1eb63ba64295cda8', OUT: '22530ed51f0ceb35',
};
const actualMembers = { A: digest(A), B: digest(B), ADJ: digest(ADJ), OUT: digest(OUT) };
for (const k of Object.keys(PUBLISHED_MEMBERS)) {
  if (actualMembers[k] !== PUBLISHED_MEMBERS[k])
    throw new Error(
      `tier ${k} membership changed (count is still ${actualSplit[k]}, so the count check above could not see it). ` +
      `digest is ${actualMembers[k]}, design:connected_tv was published against ${PUBLISHED_MEMBERS[k]}. ` +
      `Full digests now ${JSON.stringify(actualMembers)}. If the map change is intended, update ` +
      `PUBLISHED_MEMBERS and re-check every per-venue, per-year, per-platform and per-topic figure on the page.`
    );
}
 
// The topic tag has no count to pin — it is hand-assigned and published as a
// ranking. It still drives the "Fifteen of the 35 papers are IoT studies" list,
// so its membership is pinned the same way.
const topicDigest = createHash('sha256').update(
  Object.entries(MAP).filter(([, v]) => v[0] === 'A' || v[0] === 'B')
    .map(([s, v]) => `${s}\t${v[1]}`).sort().join('\n')).digest('hex').slice(0, 16);
const PUBLISHED_TOPICS = 'fdb92e2a4c3ab263';
if (topicDigest !== PUBLISHED_TOPICS)
  throw new Error(
    `topic assignments for the 35-paper population changed: digest ${topicDigest}, published against ${PUBLISHED_TOPICS}. ` +
    `Update PUBLISHED_TOPICS and re-check the "What this literature measures" table and the iot-device-set list on the page.`
  );
 
H('0. CANDIDATE POOL AND AUDIT');
console.log(`corpus papers with full text scanned : ${scored.length}   (no paper.cols.txt: ${noFulltext})`);
console.log(`roadmap title+summary probe alone    : ${scored.filter((s) => s.titleHit).length}`);
console.log(`  ...of which carry platform 'web'   : ${scored.filter((s) => s.titleHit && s.p.platforms.includes('web')).length}`);
console.log(`gate 1, any TV vocabulary            : ${gate1.length}`);
console.log(`gate 2, the hand-audited set         : ${AUDIT.length}`);
console.log(`  A  television is the study object  : ${nA}`);
console.log(`  B  television inside a device set  : ${nB}`);
console.log(`  ADJ adjacent, cited not counted    : ${ADJ.length}`);
console.log(`  OUT off topic                      : ${OUT.length}`);
console.log(`POPULATION (A+B)                     : ${nPOP}`);
console.log(`precision of the audit set           : ${pct(nPOP, AUDIT.length)}`);
console.log(`precision of the roadmap's 16        : ${pct(AUDIT.filter((s) => s.titleHit && MAP[s.p.slug] && ['A','B'].includes(MAP[s.p.slug][0])).length, scored.filter((s) => s.titleHit).length)}`);
const titleSet = new Set(scored.filter((s) => s.titleHit).map((s) => s.p.slug));
console.log(`population papers the roadmap probe MISSES: ${POP.filter((x) => !titleSet.has(x.p.slug)).length} of ${nPOP}`);
 
H('0b. THE POPULATION, PAPER BY PAPER');
T(['Tier', 'Year', 'Venue', 'Topic', 'Title'], POP.map((x) => [x.v[0], x.p.year, x.p.venue, x.v[1], x.p.title.replace(/\.$/, '')]));
 
H('0c. ADJACENT AND OUT — the audit trail for what was NOT counted');
T(['V', 'Year', 'Venue', 'Title', 'Reason'], [...ADJ, ...OUT].map((x) => [x.v[0], x.p.year, x.p.venue, x.p.title.slice(0, 62).replace(/\.$/, ''), x.v[2]]));
 
// ---------------------------------------------------------------------------
// 1. Shape of the population
// ---------------------------------------------------------------------------
H('1. THE POPULATION BY VENUE, YEAR AND TOPIC');
const venues = [...new Set(ALL.map((p) => p.venue))].sort();
T(['Venue', 'A', 'B', 'A+B', 'Venue papers in corpus', 'Share of venue'],
  venues.map((v) => {
    const a = A.filter((x) => x.p.venue === v).length, b = B.filter((x) => x.p.venue === v).length;
    const tot = ALL.filter((p) => p.venue === v).length;
    return [v, a, b, a + b, tot, pct(a + b, tot)];
  }).sort((x, y) => y[3] - x[3]));
 
const YB = [['20102014', 2010, 2014], ['20152018', 2015, 2018], ['20192021', 2019, 2021], ['20222024', 2022, 2024], ['20252026*', 2025, 2026]];
T(['Window', 'A', 'B', 'A+B'], YB.map(([lab, lo, hi]) => [lab,
  A.filter((x) => x.p.year >= lo && x.p.year <= hi).length,
  B.filter((x) => x.p.year >= lo && x.p.year <= hi).length,
  POP.filter((x) => x.p.year >= lo && x.p.year <= hi).length]));
console.log('* 20252026 is provisional: CCS 2026 and IMC 2026 have not been held, and IEEE S&P / WWW 2026 are incompletely selected. See literature:corpus.');
 
const topics = [...new Set(POP.map((x) => x.v[1]))];
T(['Topic (hand-assigned, ranking only)', 'A', 'B', 'A+B'], topics.map((t) => [t,
  A.filter((x) => x.v[1] === t).length, B.filter((x) => x.v[1] === t).length,
  POP.filter((x) => x.v[1] === t).length]).sort((a, b) => b[3] - a[3]));
 
// ---------------------------------------------------------------------------
// 2. Platform fields — the roadmap's web column
// ---------------------------------------------------------------------------
H('2. PLATFORM FIELDS: WHY `web` IS THE WRONG FILTER');
const plats = ['web', 'mobile', 'iot', 'other-online-service', 'offline', 'not-applicable'];
T(['platforms[] value', `A+B (N=${nPOP})`, 'share', `corpus (N=${ALL.length})`, 'share'],
  plats.map((v) => {
    const n = POP.filter((x) => x.p.platforms.includes(v)).length;
    const m = ALL.filter((p) => p.platforms.includes(v)).length;
    return [v, n, pct(n, nPOP), m, pct(m, ALL.length)];
  }));
console.log(`\nPapers in the population carrying platform 'web' : ${POP.filter((x) => x.p.platforms.includes('web')).length} of ${nPOP}`);
console.log('Which paper(s):', POP.filter((x) => x.p.platforms.includes('web')).map((x) => `${x.p.venue} ${x.p.year} ${x.p.title.slice(0, 50)}`).join(' | ') || '(none)');
console.log(`\npopulation[].unit == 'iot-devices' : ${POP.filter((x) => x.p.population.some((q) => q.unit === 'iot-devices')).length} of ${nPOP}`);
console.log(`population[].unit == 'mobile-apps' : ${POP.filter((x) => x.p.population.some((q) => q.unit === 'mobile-apps')).length} of ${nPOP}`);
console.log(`population[].unit == 'websites'    : ${POP.filter((x) => x.p.population.some((q) => q.unit === 'websites')).length} of ${nPOP}`);
const CRAWLED = (p) => p.crawlConfig !== null || p.studyTypes.includes('automated-web-crawl');
console.log(`\nIn the corpus-wide 'crawled' population (crawlConfig or automated-web-crawl): ${POP.filter((x) => CRAWLED(x.p)).length} of ${nPOP}`);
console.log('Which:', POP.filter((x) => CRAWLED(x.p)).map((x) => `${x.p.venue} ${x.p.year}`).join(', ') || '(none)');
 
// ---------------------------------------------------------------------------
// 3. Instrumentation
// ---------------------------------------------------------------------------
H('3. HOW THESE PAPERS GET AT THE TRAFFIC');
const TOOLCATS = ['proxy-interception', 'traffic-capture', 'mobile-instrumentation', 'program-analysis', 'crawler-framework', 'browser-automation', 'network-scanner', 'blocklist'];
T(['tool category', `A (N=${nA})`, `A+B (N=${nPOP})`], TOOLCATS.map((c) => [c,
  A.filter((x) => x.p.tools.some((t) => t.category === c && t.usedOrMentioned === 'used')).length,
  POP.filter((x) => x.p.tools.some((t) => t.category === c && t.usedOrMentioned === 'used')).length]));
 
// Named instruments, folded by lower-cased alphanumeric skeleton.
const skel = (s) => s.toLowerCase().replace(/[^a-z0-9]+/g, '');
const ALIAS = { mitmproxy: 'mitmproxy', mitmdump: 'mitmproxy', mitmweb: 'mitmproxy', wireshark: 'Wireshark', tshark: 'Wireshark', tcpdump: 'tcpdump',
  frida: 'Frida', adb: 'adb', androiddebugbridge: 'adb', charles: 'Charles Proxy', charlesproxy: 'Charles Proxy', pihole: 'Pi-hole',
  virustotal: 'VirusTotal', apktool: 'apktool', jadx: 'jadx', flowdroid: 'FlowDroid', libscout: 'LibScout', easylist: 'EasyList',
  scapy: 'Scapy', selenium: 'Selenium', openwpm: 'OpenWPM', mercury: 'Mercury', pingpong: 'PingPong', appium: 'Appium',
  raspberrypi: 'Raspberry Pi', monkey: 'UI/Application Exerciser Monkey', uiautomator: 'UIAutomator' };
const instr = new Map(); const residue = new Map();
for (const x of POP) for (const t of x.p.tools) {
  if (t.usedOrMentioned !== 'used') continue;
  const k = skel(t.name); const disp = ALIAS[k];
  const target = disp ? instr : residue;
  const key = disp ?? t.name;
  if (!target.has(key)) target.set(key, new Set());
  target.get(key).add(x.p.slug);
}
T(['instrument (alias-folded)', `papers (of ${nPOP})`], [...instr.entries()].map(([k, v]) => [k, v.size]).sort((a, b) => b[1] - a[1]));
console.log(`\nUNMAPPED RESIDUE of the alias fold — ${residue.size} distinct raw tool names used by the ${nPOP} papers, printed in full:`);
console.log([...residue.entries()].sort((a, b) => b[1].size - a[1].size || a[0].localeCompare(b[0])).map(([k, v]) => `${v.size}× ${k}`).join(' | '));
 
console.log('\nInterception evidence, per Tier A paper (full text, whitespace-collapsed):');
const IPROBE = {
  'mitm/proxy': /\bmitmproxy\b|\bmitm\b|\bman-in-the-middle\b|\bCharles\b|\bproxy\b/i,
  'own CA / root cert': /\broot certificate\b|\bcustom (?:CA|certificate)\b|\binstall(?:ed|ing)? (?:our|a) (?:own )?(?:CA|certificate|root)\b|\bCA certificate\b/i,
  'pinning / cert failure': /\bcertificate pinning\b|\bpinn(?:ing|ed)\b|\bcertificate validation\b/i,
  'undecryptable reported': /\bcould not (?:be )?decrypt\b|\bfail(?:ed|s|ure) to decrypt\b|\bdecryption fail\w*\b|\bunable to (?:decrypt|intercept)\b/i,
  'DNS-level capture': /\bDNS (?:queries|traffic|resolution|logs)\b|\bPi-hole\b|\bdnsmasq\b/i,
  'router/AP capture': /\b(?:wireless )?access point\b|\bRaspberry Pi\b|\brouter\b|\btcpdump\b|\bWireshark\b/i,
  'HDMI / screen capture': /\bHDMI\b|\bscreen ?(?:shot|capture|recording)\b|\bframe grabber\b/i,
  'remote-control automation': /\bADB\b|\bremote control\b|\bIR blaster\b|\bkey ?events?\b|\bExternal Control\b/i,
};
T(['paper', ...Object.keys(IPROBE)], A.map((x) => {
  const t = fulltext(x.p);
  return [`${x.p.venue} ${x.p.year}`, ...Object.values(IPROBE).map((re) => ((t.match(new RegExp(re.source, re.flags + 'g')) || []).length || '—'))];
}));
for (const [lab, re] of Object.entries(IPROBE)) {
  const n = A.filter((x) => re.test(fulltext(x.p))).length;
  console.log(`${lab.padEnd(26)} ${n} of ${nA} Tier A   |   ${POP.filter((x) => re.test(fulltext(x.p))).length} of ${nPOP} A+B`);
}
 
H('3b. VANTAGE');
const locs = new Map();
for (const x of POP) for (const v of x.p.vantage) for (const l of v.locations) {
  if (isSentinel(l)) continue;
  const k = l.trim(); if (!locs.has(k)) locs.set(k, new Set()); locs.get(k).add(x.p.slug);
}
T(['vantage location (verbatim, unfolded)', 'papers'], [...locs.entries()].map(([k, v]) => [k, v.size]).sort((a, b) => b[1] - a[1]));
const withV = POP.filter((x) => x.p.vantage.length > 0);
const statedV = withV.filter((x) => x.p.vantage.some((v) => v.locations.some((l) => !isSentinel(l))));
console.log(`\npapers with a vantage tuple: ${withV.length} of ${nPOP}; of those, stating a location: ${statedV.length} (${pct(statedV.length, withV.length)})`);
 
// ---------------------------------------------------------------------------
// 4. Sampling
// ---------------------------------------------------------------------------
H('4. WHAT THEY SAMPLE — THE NO-TRANCO PROBLEM');
const srcs = new Map();
for (const x of POP) for (const q of x.p.population) {
  if (isSentinel(q.sourceList)) continue;
  const k = q.sourceList.trim(); if (!srcs.has(k)) srcs.set(k, new Set()); srcs.get(k).add(x.p.slug);
}
T(['population.sourceList (verbatim, unfolded — ranking only)', 'papers'], [...srcs.entries()].map(([k, v]) => [k.slice(0, 58), v.size]).sort((a, b) => b[1] - a[1]));
T(['population.unit', 'papers'], [...new Set(POP.flatMap((x) => x.p.population.map((q) => q.unit)))]
  .map((u) => [u, POP.filter((x) => x.p.population.some((q) => q.unit === u)).length]).sort((a, b) => b[1] - a[1]));
T(['population.samplingMethod', 'papers'], [...new Set(POP.flatMap((x) => x.p.population.map((q) => q.samplingMethod)))]
  .map((m) => [m, POP.filter((x) => x.p.population.some((q) => q.samplingMethod === m)).length]).sort((a, b) => b[1] - a[1]));
// Labelled by slug, not `venue year`: there are two IMC 2011 papers in this
// population and only one of them has iot-devices tuples, so a venue+year label
// is ambiguous exactly where a reader would want to check it.
const devSets = POP.map((x) => ({
  k: `${x.p.venue} ${x.p.year} ${x.p.slug.slice(0, 38)}`,
  ns: x.p.population.filter((q) => q.unit === 'iot-devices' && q.n !== null).map((q) => q.n),
})).filter((r) => r.ns.length);
T(['paper', 'iot-devices n values stated'], devSets.map((r) => [r.k, r.ns.join(', ')]));
 
// The page states this distribution in prose, so it is derived here rather than
// eyeballed off the table above. A hand-typed version of this list shipped in
// the first draft and omitted 57 — the PETS 2020 testbed, the population's own
// flagship paper.
const allNs = devSets.flatMap((r) => r.ns).sort((a, b) => a - b);
const largestPer = devSets.map((r) => Math.max(...r.ns)).sort((a, b) => a - b);
const median = (arr) => (arr.length % 2
  ? arr[(arr.length - 1) / 2]
  : (arr[arr.length / 2 - 1] + arr[arr.length / 2]) / 2);
console.log(`\npapers stating an iot-devices size : ${devSets.length} of ${nPOP}`);
console.log(`stated size values                : ${allNs.length}`);
console.log(`all values, sorted                : ${allNs.join(', ')}`);
console.log(`median of all stated values       : ${median(allNs)}`);
console.log(`largest set per paper, sorted     : ${largestPer.join(', ')}`);
console.log(`median of largest-per-paper       : ${median(largestPer)}`);
for (const cut of [100, 200, 1000]) {
  console.log(`papers whose LARGEST set <= ${String(cut).padStart(4)}   : ${largestPer.filter((v) => v <= cut).length} of ${devSets.length}`);
}
console.log('papers whose largest set is over 200, i.e. not a lab bench:');
for (const r of devSets.filter((r) => Math.max(...r.ns) > 200)) console.log(`  ${r.k}  -> ${r.ns.join(', ')}`);
const listv = POP.flatMap((x) => x.p.population).filter((q) => q.listVersion !== null).length;
console.log(`\npopulation tuples in A+B stating a listVersion: ${listv} of ${POP.flatMap((x) => x.p.population).length}`);
 
// ---------------------------------------------------------------------------
// 5. Measured results
// ---------------------------------------------------------------------------
H('5. MEASURED RESULTS (detection[].prevalence, Tier A only)');
let dt = 0, dp = 0;
for (const x of A) {
  const withPrev = x.p.detection.filter((d) => d.prevalence !== null);
  dt += x.p.detection.length; dp += withPrev.length;
  console.log(`\n--- ${x.p.venue} ${x.p.year} | ${x.p.title}`);
  for (const d of withPrev) console.log(`  * ${d.phenomenon}\n      technique : ${d.technique}\n      metric    : ${d.metric}\n      prevalence: ${d.prevalence}\n      quote (${d.evidence.section}): ${d.evidence.quote}`);
}
console.log(`\ndetection tuples in Tier A: ${dt}; carrying a prevalence: ${dp} (${pct(dp, dt)})`);
const dpAll = ALL.flatMap((p) => p.detection);
console.log(`corpus-wide: ${dpAll.length} tuples, ${dpAll.filter((d) => d.prevalence !== null).length} carry a prevalence (${pct(dpAll.filter((d) => d.prevalence !== null).length, dpAll.length)})`);
 
// ---------------------------------------------------------------------------
// 6. Reporting gaps
// ---------------------------------------------------------------------------
H('6. WHERE THIS LITERATURE GOES QUIET');
const CRAWL_CORPUS = ALL.filter(CRAWLED);
const crawlPOP = POP.filter((x) => x.p.crawlConfig !== null);
// `ethics` and `crawlConfig` are nullable by schema: a paper with no crawl has
// no crawlConfig, and 894 corpus papers carry no ethics object at all. Each
// block below names the population that HAS the object, so a sentinel is never
// counted as an answer and a missing object is never counted as a sentinel.
T(['crawlConfig field (papers with a crawlConfig object)', `A+B (N=${crawlPOP.length})`, `corpus crawled (N=${CRAWL_CORPUS.filter((p) => p.crawlConfig !== null).length})`],
  [['consentAction', 'interactionDepth', 'statefulness', 'browsers']].flat().map((f) => {
    const cc = CRAWL_CORPUS.filter((p) => p.crawlConfig !== null);
    const st = (p) => (f === 'browsers' ? p.crawlConfig.browsers.length > 0 : !isSentinel(p.crawlConfig[f]));
    const a = crawlPOP.filter((x) => st(x.p)).length, b = cc.filter(st).length;
    return [f, `${a} of ${crawlPOP.length} (${pct(a, crawlPOP.length)})`, `${b} of ${cc.length} (${pct(b, cc.length)})`];
  }));
console.log('\ncrawlConfig values in the A+B population, verbatim:');
for (const x of crawlPOP) console.log(`  ${x.p.venue} ${x.p.year}  consent=${x.p.crawlConfig.consentAction}  depth=${x.p.crawlConfig.interactionDepth}  state=${x.p.crawlConfig.statefulness}  browsers=[${x.p.crawlConfig.browsers.join(',')}]  robotsTxt=${x.p.ethics === null ? '(no ethics object)' : x.p.ethics.robotsTxt}`);
 
const ethPOP = POP.filter((x) => x.p.ethics !== null);
const ethEmpCorpus = ALL.filter((p) => p.isEmpirical && p.ethics !== null);
const emp = POP.filter((x) => x.p.isEmpirical);
const EMP_CORPUS = ALL.filter((p) => p.isEmpirical);
const artPOP = POP.filter((x) => x.p.artifacts !== null);
const artEmpCorpus = ALL.filter((p) => p.isEmpirical && p.artifacts !== null);
console.log(`\nA+B papers carrying an ethics object: ${ethPOP.length} of ${nPOP};  artifacts object: ${artPOP.length} of ${nPOP}`);
T(['field', 'A+B', 'corpus empirical'], [
  ['ethics.reviewOutcome stated',
    `${ethPOP.filter((x) => !isSentinel(x.p.ethics.reviewOutcome)).length} of ${ethPOP.length} (${pct(ethPOP.filter((x) => !isSentinel(x.p.ethics.reviewOutcome)).length, ethPOP.length)})`,
    `${ethEmpCorpus.filter((p) => !isSentinel(p.ethics.reviewOutcome)).length} of ${ethEmpCorpus.length} (${pct(ethEmpCorpus.filter((p) => !isSentinel(p.ethics.reviewOutcome)).length, ethEmpCorpus.length)})`],
  ['artifacts.availability stated',
    `${artPOP.filter((x) => !isSentinel(x.p.artifacts.availability)).length} of ${artPOP.length} (${pct(artPOP.filter((x) => !isSentinel(x.p.artifacts.availability)).length, artPOP.length)})`,
    `${artEmpCorpus.filter((p) => !isSentinel(p.artifacts.availability)).length} of ${artEmpCorpus.length} (${pct(artEmpCorpus.filter((p) => !isSentinel(p.artifacts.availability)).length, artEmpCorpus.length)})`],
  ['temporal.spanStart stated',
    `${emp.filter((x) => x.p.temporal.some((t) => t.spanStart !== null)).length} of ${emp.length} (${pct(emp.filter((x) => x.p.temporal.some((t) => t.spanStart !== null)).length, emp.length)})`,
    `${EMP_CORPUS.filter((p) => p.temporal.some((t) => t.spanStart !== null)).length} of ${EMP_CORPUS.length} (${pct(EMP_CORPUS.filter((p) => p.temporal.some((t) => t.spanStart !== null)).length, EMP_CORPUS.length)})`],
]);
function linkOf(a) {
  if (a.codeUrl) return a.codeUrl;
  if (a.dataUrl) return a.dataUrl;
  if (!a.links || a.links.length === 0) return '—';
  const l = a.links[0];
  if (typeof l === 'string') throw new Error(`artifacts.links[0] is a string, expected an object: ${l}`);
  return `${l.url}  (${l.kind}; ${l.what}; authors=${l.belongsToAuthors})`;
}
 
console.log('\nArtifact links released, Tier A:');
for (const x of A) console.log(x.p.artifacts === null
  ? `  ${x.p.venue} ${x.p.year}  (no artifacts object extracted)`
  // artifacts.links[] entries are OBJECTS ({url, kind, what, belongsToAuthors}).
  // Falling back to links[0] itself string-coerced to "[object Object]" and that
  // literal was published three times on the provenance page \u2014 found by a review
  // pass on 2026-09-13. Take .url, and print what the link is for.
  : `  ${x.p.venue} ${x.p.year}  ${x.p.artifacts.availability.padEnd(22)} ${linkOf(x.p.artifacts)}`);
 
// ---------------------------------------------------------------------------
// ---------------------------------------------------------------------------
// 6b. Two probes the page cites in footnotes. Both were added on 2026-09-13
//     after review found the page asserting negatives no probe supported.
// ---------------------------------------------------------------------------
H('6b. TWO PROBES THE PAGE CITES');
 
// (i) The TLS decryption hole. The page originally ran only NARROW, matched 1
//     of 13 Tier A papers, and published "almost nobody reports the hole" —
//     while quoting two of the sentences NARROW misses. A narrowing probe must
//     return a SUBSET of the loose one; both counts are printed so the claim
//     can be read off the right width.
const HOLE_NARROW = /could not decrypt|failed to decrypt|decryption fail|unable to (decrypt|intercept)/i;
const HOLE_WIDE = new RegExp([
  /(could not|cannot|can ?not|unable to|failed to|no way to)[^.]{0,90}(decrypt|intercept|install (our own |a |custom )?(self-signed )?certificat)/.source,
  /(bypass|remove|disabl\w*)[^.]{0,40}certificate pinning/.source,
  /certificate pinning checks/.source,
].join('|'), 'i');
const holeN = A.filter((x) => HOLE_NARROW.test(fulltext(x.p)));
const holeW = A.filter((x) => HOLE_WIDE.test(fulltext(x.p)));
const holeNs = new Set(holeN.map((x) => x.p.slug));
const notContained = holeN.filter((x) => !holeW.some((y) => y.p.slug === x.p.slug));
if (notContained.length)
  throw new Error(`the narrow decryption probe is not a subset of the wide one: ${notContained.map((x) => x.p.slug).join(', ')}`);
if (holeN.length > holeW.length)
  throw new Error(`narrow probe matched ${holeN.length} > wide ${holeW.length}; a narrowing probe must match fewer`);
console.log(`TLS decryption hole, Tier A (n=${A.length}):`);
console.log(`  narrow probe (the one the page published until 2026-09-13): ${holeN.length}`);
console.log(`  wide probe   (the one the page cites now)                 : ${holeW.length}`);
for (const x of holeW) {
  const t = fulltext(x.p), m = t.match(HOLE_WIDE), i = t.indexOf(m[0]);
  console.log(`   ${holeNs.has(x.p.slug) ? 'both  ' : 'wide  '} ${x.p.venue} ${x.p.year} ${x.p.slug}`);
  console.log(`          ...${t.slice(Math.max(0, i - 100), i + 200)}...`);
}
 
// (ii) ATSC 3.0 / NextGen TV. The page says no paper measures it. That was an
//      unprobed negative until now; ATSC was never in PROBES.
const ATSC = /\bATSC\b|\bNextGen ?TV\b|\bNext ?Gen ?TV\b/i;
const atscHits = ALL.filter((p) => ATSC.test(fulltext(p)));
console.log(`\nATSC 3.0 / NextGen TV, full text, all ${ALL.length} corpus papers: ${atscHits.length} match`);
for (const p of atscHits) {
  const t = fulltext(p), m = t.match(ATSC), i = t.indexOf(m[0]);
  console.log(`   ${p.venue} ${p.year} ${p.slug}   [${MAP[p.slug] ? MAP[p.slug][0] : 'not in audit set'}]`);
  console.log(`          ...${t.slice(Math.max(0, i - 130), i + 210)}...`);
}
 
// (iii) The factory-reset gap. crawlConfig.statefulness is empty for all 6
//       papers that have the object; the page needed the full-text picture too.
const RESET = /factory[- ]reset/i;
const resetHits = POP.filter((x) => RESET.test(fulltext(x.p)));
const withCfg = POP.filter((x) => x.p.crawlConfig !== null);
console.log(`\nStatefulness: ${withCfg.length} of ${nPOP} papers have a crawlConfig object;`);
console.log(`  statefulness values among them: ${JSON.stringify(withCfg.map((x) => x.p.crawlConfig.statefulness))}`);
console.log(`  full-text "factory reset" anywhere in the ${nPOP}: ${resetHits.length}`);
for (const x of resetHits) {
  const t = fulltext(x.p), i = t.search(RESET);
  console.log(`   ${MAP[x.p.slug][0]}  ${x.p.venue} ${x.p.year} ${x.p.slug}`);
  console.log(`          ...${t.slice(Math.max(0, i - 130), i + 210)}...`);
}
 
// (iv) LLM use inside the population. The page said "not one paper uses an LLM
//      for anything"; two do, and both tool names are in the printed residue.
console.log('\nLLM and language-model tools used by the population (tools[].category, used only):');
for (const x of POP) {
  for (const t of x.p.tools) {
    if (t.usedOrMentioned !== 'used') continue;
    if (!/^llm$/i.test(t.category) && !/gpt|openai|chatgpt|bert|language model/i.test(t.name)) continue;
    console.log(`   ${x.p.venue} ${x.p.year}  category=${t.category.padEnd(22)} ${t.name}`);
    console.log(`          quote: ${JSON.stringify((t.evidence?.quote || '').slice(0, 200))}`);
  }
}
 
// 7. Quote verification lives in scripts/ctv_quotecheck.py, which checks each
//    quote against BOTH paper.cols.txt and the PDF text layer. Doing it here
//    against .cols alone produced two false NOTFOUNDs on the first run, one of
//    which was a real error in the needle and one of which was a de-columning
//    artefact — see provenance:design:connected_tv.
// ---------------------------------------------------------------------------
console.log('\n' + '='.repeat(78) + '\n7. QUOTES: see scripts/ctv_quotecheck.py (checks .cols AND the PDF text layer)\n' + '='.repeat(78));
ctv_quotecheck.py
#!/usr/bin/env python3
"""Quote check for design:connected_tv and provenance:design:connected_tv.
 
Every phrase either page puts in quotation marks, or leans on for a figure, is
listed here and matched against BOTH renderings of the source paper:
 
  data/fulltext/<year>/<venue>/<slug>/paper.cols.txt   (de-columned text)
  data/fulltext/<year>/<venue>/<slug>/paper.pdf        (pypdf text layer)
 
Both are needed. `.cols` splices some two-column sentences, and the PDF layer
loses reading order elsewhere, so a NOTFOUND in one rendering is not evidence
that a paper does not contain the sentence. On the first run of this check two
quotes failed against `.cols` alone: one was a de-columning splice (All Things
Considered) and the other was a genuinely wrong needle, where two halves of a
sentence in the PETS 2020 paper had been glued together into a number the paper
does not state. The second is why bare-number needles are avoided here.
 
Usage: python3 scripts/ctv_quotecheck.py
"""
import re
import sys
import unicodedata
from pathlib import Path
 
ROOTS = [Path('/workspace/publications_dataset/data/fulltext'),
         Path('/workspace/publications_dataset/fulltext')]
ROOT = next(r for r in ROOTS if r.is_dir())
 
 
def norm(s: str) -> str:
    s = unicodedata.normalize('NFKD', s)
    s = s.replace('­', '').replace('-\n', '')
    s = re.sub(r'[‘’]', "'", s)
    s = re.sub(r'[“”]', '"', s)
    return re.sub(r'\s+', ' ', s).lower()
 
 
# (year, venue, slug, needle)
QUOTES = [
    ('2019', 'CCS', 'watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices',
     'present on 69% of Roku channels and 89% of Amazon Fire TV channels'),
    ('2019', 'CCS', 'watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices',
     'we were able to install our own cert on the device which allowed us to intercept HTTPS requests on 957 of the 1000 channels'),
    ('2019', 'CCS', 'watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices',
     'that leaked the title of the video to a tracking domain'),
    ('2020', 'PETS', 'the-tv-is-smart-and-full-of-trackers-measuring-smart-tv-advertising-and-tracking',
     '1 out of 5 (or fewer) TLS connections for 80% of all apps'),
    ('2020', 'PETS', 'the-tv-is-smart-and-full-of-trackers-measuring-smart-tv-advertising-and-tracking',
     '314 ATS domains that are unique to the Roku dataset'),
    ('2022', 'PETS', 'fingerprintv-fingerprinting-smart-tv-apps',
     'among 80 apps that are made available on all three smart TV platforms, 76% exhibit a different fingerprint on each platform'),
    ('2022', 'PETS', 'watch-over-your-tv-a-security-and-privacy-analysis-of-the-android-tv-ecosystem',
     'The analysis found at least one sensitive data flow in 78% of the files'),
    ('2023', 'NDSS', 'i-still-know-what-you-watched-last-sunday-privacy-of-the-hbbtv-protocol-in-the-european-smart-tv-landscape',
     '26 communicate with trackers before the user has expressed their consent'),
    ('2023', 'NDSS', 'i-still-know-what-you-watched-last-sunday-privacy-of-the-hbbtv-protocol-in-the-european-smart-tv-landscape',
     'only block at maximum 44% in 2021 and 81% in 2022'),
    ('2024', 'IMC', 'watching-tv-with-the-second-party-a-first-look-at-automatic-content-recognition',
     'there is a complete absence of communication with any previously identified ACR domains'),
    ('2024', 'IMC', 'watching-tv-with-the-second-party-a-first-look-at-automatic-content-recognition',
     'smart TVs in the UK and the US contact distinct ACR domains'),
    ('2024', 'IMC', 'watching-tv-with-the-second-party-a-first-look-at-automatic-content-recognition',
     'ACR network traffic exists when watching linear TV and when using smart TV as an external display using HDMI'),
    ('2019', 'USENIX', 'all-things-considered-an-analysis-of-iot-devices-on-home-networks',
     'the most popular vendor, Roku, only accounts for 17.4% of media devices'),
    ('2018', 'IMC', 'understanding-video-management-planes',
     'streaming set-top boxes1 dominate by view-hours'),
    ('2023', 'IMC', 'in-the-room-where-it-happens-characterizing-local-communication-and-threats-in-s',
     'the analysis of the Smart TV ecosystem is left for future work'),
    ('2024', 'IEEE-SP', 'surveilling-the-masses-with-wi-fi-based-positioning-systems',
     'belong to the streaming television equipment manufacturer Roku'),
    ('2021', 'IMC', 'iotls-understanding-tls-usage-in-consumer-iot-devices',
     'such as voice assistants, smart TVs and video doorbells'),
    ('2025', 'USENIX', 'watch-out-your-tv-box-reversing-and-blocking-a-p2p-based-illegal-streaming-ecosy',
     'they are offered only to those who have purchased specific'),
    ('2014', 'USENIX', 'from-the-aether-to-the-ethernet-attacking-the-internet-using-broadcast-digital-t',
     'which requires a minimal budget and infrastructure'),
    ('2024', 'IMC', 'iot-bricks-over-v6-understanding-ipv6-usage-in-smart-homes',
     'only eight out of 93 devices remain functional'),
    ('2019', 'CCS', 'watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices',
     'On Roku, a total of 43 channels failed to properly verify the server'),
    ('2019', 'CCS', 'watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices',
     '794 of the 1000 Roku channels sent at least one request in cleartext'),
    ('2019', 'CCS', 'watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices',
     'We found 9 channels on Roku and 14 channels on the Fire TV'),
    ('2019', 'CCS', 'watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices',
     'an HTTP GET request to "http://ROKU_ DEVICE_IP_ADDRESS:8060/keydown/left " does the same for all Roku devices'),
    ('2020', 'PETS', 'the-tv-is-smart-and-full-of-trackers-measuring-smart-tv-advertising-and-tracking',
     '697 Fire TV apps that expose advertising ID alongside serial number and device ID'),
    ('2022', 'PETS', 'fingerprintv-fingerprinting-smart-tv-apps',
     '96% (N = 961) of the top'),
    ('2022', 'PETS', 'watch-over-your-tv-a-security-and-privacy-analysis-of-the-android-tv-ecosystem',
     '75% of the apps contain analytics libraries and 77% contain advertising libraries'),
    ('2021', 'USENIX', 'android-smarttvs-vulnerability-discovery-via-log-guided-fuzzing',
     '37 unique vulnerabilities, including 11 high-impact cyber threats, 10 new memory corruptions, and 16 visual and auditory anomalies'),
    ('2023', 'USENIX', 'homespy-the-invisible-sniffer-of-infrared-remote-control-of-smart-tvs',
     'The accuracy increases to 70% for Top3 and 77% for Top5'),
    ('2024', 'NDSS', 'acoustic-keystroke-leakage-on-smart-televisions',
     'up to 60.19% of common passwords'),
    ('2024', 'IMC', 'watching-tv-with-the-second-party-a-first-look-at-automatic-content-recognition',
     'the fact that we observe network traffic every 15 seconds suggests'),
    ('2011', 'IMC', 'understanding-couch-potatoes-measurement-and-modeling-of-interactive-usage-of-ip',
     'The average number of set-top boxes provisioned was approximately 3 million over this period'),
    ('2019', 'USENIX', 'all-things-considered-an-analysis-of-iot-devices-on-home-networks',
     'are the most common type of device in seven of the eleven regions'),
]
 
pdf_cache: dict[Path, str] = {}
 
 
def pdf_text(path: Path) -> str:
    if path not in pdf_cache:
        try:
            import pypdf
            reader = pypdf.PdfReader(str(path))
            pdf_cache[path] = norm('\n'.join(p.extract_text() or '' for p in reader.pages))
        except Exception as exc:            # a missing text layer is a result, not a crash
            pdf_cache[path] = ''
            print(f'   (pypdf failed on {path}: {exc})')
    return pdf_cache[path]
 
 
def main() -> int:
    bad = []
    cols_only = pdf_only = both = 0
    for year, venue, slug, needle in QUOTES:
        d = ROOT / year / venue / slug
        cols = norm((d / 'paper.cols.txt').read_text(encoding='utf8'))
        n = norm(needle)
        in_cols = n in cols
        in_pdf = n in pdf_text(d / 'paper.pdf')
        where = ('cols+pdf' if in_cols and in_pdf else
                 'cols only' if in_cols else 'PDF only' if in_pdf else 'NOTFOUND')
        if in_cols and in_pdf:
            both += 1
        elif in_cols:
            cols_only += 1
        elif in_pdf:
            pdf_only += 1
        else:
            bad.append((slug, needle))
        print(f'{where:<9} {venue} {year} {slug[:44]:<44} "{needle[:66]}"')
    print(f'\n{len(QUOTES)} quotes: {both} in both renderings, {cols_only} in .cols only, '
          f'{pdf_only} in the PDF only, {len(bad)} in neither.')
    for slug, needle in bad:
        print(f'  NOTFOUND {slug}: "{needle}"')
    return 1 if bad else 0
 
 
if __name__ == '__main__':
    sys.exit(main())

13. Full report output, unedited

report_connected_tv-output.txt
==============================================================================
0. CANDIDATE POOL AND AUDIT
==============================================================================
corpus papers with full text scanned : 5855   (no paper.cols.txt: 4)
roadmap title+summary probe alone    : 16
  ...of which carry platform 'web'   : 1
gate 1, any TV vocabulary            : 142
gate 2, the hand-audited set         : 103
  A  television is the study object  : 13
  B  television inside a device set  : 22
  ADJ adjacent, cited not counted    : 16
  OUT off topic                      : 52
POPULATION (A+B)                     : 35
precision of the audit set           : 34.0%
precision of the roadmap's 16        : 62.5%
population papers the roadmap probe MISSES: 25 of 35
 
==============================================================================
0b. THE POPULATION, PAPER BY PAPER
==============================================================================
Tier  Year  Venue    Topic                 Title
----  ----  -------  --------------------  ----------------------------------------------------------------------------------------------------------------
A     2011  IMC      delivery-performance  Understanding couch potatoes: measurement and modeling of interactive usage of IPTV at large scale
A     2011  IMC      delivery-performance  Q-score: proactive service quality assessment in a large IPTV system
A     2014  USENIX   broadcast             From the Aether to the Ethernet—Attacking the Internet using Broadcast Digital Television
A     2019  CCS      tracking              Watching You Watch: The Tracking Ecosystem of Over-the-Top TV Streaming Devices
A     2020  PETS     tracking              The TV is Smart and Full of Trackers: Measuring Smart TV Advertising and Tracking
A     2021  USENIX   vulnerability         Android SmartTVs Vulnerability Discovery via Log-Guided Fuzzing
A     2022  PETS     tracking              FingerprinTV: Fingerprinting Smart TV Apps
A     2022  PETS     app-analysis          Watch Over Your TV: A Security and Privacy Analysis of the Android TV Ecosystem
A     2023  NDSS     broadcast             I Still Know What You Watched Last Sunday: Privacy of the HbbTV Protocol in the European Smart TV Landscape
A     2023  USENIX   side-channel          HOMESPY: The Invisible Sniffer of Infrared Remote Control of Smart TVs
A     2024  IMC      acr                   Watching TV with the Second-Party: A First Look at Automatic Content Recognition Tracking in Smart TVs
A     2024  NDSS     side-channel          Acoustic Keystroke Leakage on Smart Televisions
A     2025  USENIX   piracy                Watch Out Your TV Box: Reversing and Blocking a P2P-based Illegal Streaming Ecosystem
B     2018  IMC      delivery-performance  Understanding Video Management Planes
B     2019  IMC      iot-device-set        Information Exposure From Consumer IoT Devices: A Multidimensional, Network-Informed Measurement Approach
B     2019  USENIX   iot-device-set        All Things Considered: An Analysis of IoT Devices on Home Networks
B     2020  IMC      iot-device-set        A Haystack Full of Needles: Scalable Detection of IoT Devices in the Wild
B     2020  NDSS     iot-device-set        Packet-Level Signatures for Smart Home Devices
B     2020  NDSS     iot-device-set        Et Tu Alexa? When Commodity WiFi Devices Turn into Adversarial Motion Sensors
B     2020  USENIX   iot-device-set        You Are What You Broadcast: Identification of Mobile and IoT Devices from (Public) WiFi
B     2021  IMC      iot-device-set        IoTLS: understanding TLS usage in consumer IoT devices
B     2021  PETS     iot-device-set        Blocking Without Breaking: Identification and Mitigation of Non-Essential IoT Traffic
B     2022  CCS      app-analysis          Understanding IoT Security from a Market-Scale Perspective
B     2022  PETS     iot-device-set        Analyzing the Feasibility and Generalizability of Fingerprinting Internet of Things Devices
B     2022  USENIX   iot-device-set        Lumos: Identifying and Localizing Diverse Hidden IoT Devices in an Unfamiliar Environment
B     2023  IMC      iot-device-set        Behind the Scenes: Uncovering TLS and Server Certificate Practice of IoT Device Vendors in the Wild
B     2023  IMC      iot-device-set        In the Room Where It Happens: Characterizing Local Communication and Threats in Smart Homes
B     2024  IEEE-SP  device-population     Surveilling the Masses with Wi-Fi-Based Positioning Systems
B     2024  IMC      iot-device-set        IoT Bricks Over v6: Understanding IPv6 Usage in Smart Homes
B     2024  IMC      delivery-performance  Characterizing User Platforms for Video Streaming in Broadband Networks
B     2024  PETS     iot-device-set        Connecting the Dots: Tracing Data Endpoints in IoT Devices
B     2025  NDSS     iot-device-set        Evaluating Machine Learning-Based IoT Device Identification Models for Security Applications
B     2025  USENIX   vulnerability         Tracking You from a Thousand Miles Away! Turning a Bluetooth Device into an Apple AirTag Without Root Privileges
B     2026  NDSS     vulnerability         BLERP: BLE Re-Pairing Attacks and Defenses
B     2026  USENIX   policy-compliance     Missing, Present and Conflicting: A Large Scale Analysis of IoT Update Information in the EU Market
 
==============================================================================
0c. ADJACENT AND OUT — the audit trail for what was NOT counted
==============================================================================
V    Year  Venue    Title                                                           Reason
---  ----  -------  --------------------------------------------------------------  -----------------------------------------------------------------------------------------------------------------------------------------------
ADJ  2011  IMC      Measurement and analysis of a large scale commercial mobile in  "TV" delivered to mobile handsets, not to a television.
ADJ  2012  IMC      Watching videos from everywhere: a study of the PPTV mobile Vo  Mobile VoD; no television endpoint.
ADJ  2013  IMC      Analyzing the potential benefits of CDN augmentation strategie  CDN augmentation for video workloads; no TV endpoint.
ADJ  2013  IMC      Peer-assisted content distribution in Akamai netsession         Peer-assisted CDN; set-top-box mention is background.
ADJ  2016  IMC      Performance Characterization of a Commercial Video Streaming S  Streaming service performance from browser/CDN vantage; no TV-specific result.
ADJ  2016  IMC      Anatomy of a Personalized Livestreaming System                  Livestreaming (Periscope) system measurement; no TV endpoint.
ADJ  2017  PETS     On the Privacy and Security of the Ultrasound Ecosystem         Ultrasonic cross-device tracking (uXDT): beacons emitted by TV adverts and picked up by phone SDKs. The TV is the emitter, never measured.
ADJ  2019  WWW      Exploiting Diversity in Android TLS Implementations for Mobile  Android app traffic classification; TLS-fingerprint method later reused on TV apps.
ADJ  2020  USENIX   Void: A fast and light voice liveness detection system          A Samsung Smart TV is used as a replay LOUDSPEAKER; the TV is apparatus, not the measured object.
ADJ  2022  NDSS     A Lightweight IoT Cryptojacking Detection Mechanism in Heterog  Authors implement their own cryptojacking PoC on an LG webOS TV to test a detector; no deployed-TV population.
ADJ  2022  USENIX   OVRseen: Auditing Network Traffic and Privacy Policies in Ocul  VR headsets; smart TVs used as the comparison ecosystem. Same lab, same pipeline shape.
ADJ  2023  IMC      Tracking, Profiling, and Ad Targeting in the Alexa Echo Smart   Smart speaker study; TVs cited as the comparable prior ecosystem, not measured. The closest methodological sibling.
ADJ  2023  PETS     Your DRM Can Watch You Too: Exploring the Privacy Implications  Widevine EME in browsers and Android; TVs named as another Widevine host, not measured.
ADJ  2024  IMC      Cost-Saving Streaming: Unlocking the Potential of Alternative   Edge-node economics for streaming delivery; no TV endpoint measured.
ADJ  2025  PETS     Unmasking the Shadows: A Cross-Country Study of Online Trackin  Illegal movie streaming WEBSITES crawled with a browser; a web-tracking study, not a TV study.
ADJ  2025  USENIX   Endangered Privacy: Large-Scale Monitoring of Video Streaming   Video identification from encrypted MPEG-DASH traffic. The roadmap listed it as CTV spine; it measures the SERVICE and its traffic, never a TV.
OUT  2010  IMC      What happened in my network: mining network events from router  IPTV named as the service carried; router syslogs are the object.
OUT  2011  CCS      On the vulnerability of FPGA bitstream encryption against powe  Set-top box named as an FPGA application domain.
OUT  2011  IMC      Broadcast yourself: understanding YouTube uploaders             IPTV appears once in related work.
OUT  2014  CCS      (Nothing else) MATor(s): Monitoring the Anonymity of Tor's Pat  "ACR" homonym.
OUT  2015  USENIX   A Placement Vulnerability Study in Multi-Tenant Public Clouds   Title probe matched "streaming"; cloud VM placement.
OUT  2015  USENIX   Rocking Drones with Intentional Sound Noise on Gyroscopic Sens  Passing mention.
OUT  2016  CCS      SandScout: Automatic Detection of Flaws in iOS Sandbox Profile  Apple TV named as a device that runs iOS/tvOS; iOS sandbox is the object.
OUT  2016  IMC      Entropy/IP: Uncovering Structure in IPv6 Addresses              "ACR" homonym.
OUT  2016  USENIX   You Are Who You Know and How You Behave: Attribute Inference A  IPTV homonym.
OUT  2017  CCS      POSTER: Watch Out Your Smart Watch When Paired                  Tizen here is the smartwatch platform, not the TV one.
OUT  2017  PETS     Why can’t users choose their identity providers on the web?     "ACR" homonym.
OUT  2017  USENIX   Same-Origin Policy: Evaluation in Modern Browsers               Passing mention of TV browsers.
OUT  2017  WWW      FLOCK: Combating Astroturfing on Livestreaming Platforms        Title probe matched "streaming platform"; astroturfing detection on Twitch-like sites.
OUT  2018  CCS      Medical Devices are at Risk: Information Security on Diagnosti  "ACR" = American College of Radiology.
OUT  2019  IEEE-SP  Drones' Cryptanalysis - Smashing Cryptography with a Flicker    IPTV/ACR homonyms.
OUT  2019  NDSS     cleaning-up-the-internet-of-evil-things-real-world-evidence-on  One infected set-top box in a Mirai remediation table; no TV finding.
OUT  2019  NDSS     latex-gloves-protecting-browser-extensions-from-probing-and-re  The Chromecast browser EXTENSION, not the device.
OUT  2019  USENIX   A Billion Open Interfaces for Eve and Mallory: MitM, DoS, and   tvOS listed among Apple OSes; AWDL is the object.
OUT  2019  WWW      Snapshot-based Loading Acceleration of Web Apps with Nondeterm  Tizen/webOS named as embedded web-app platforms; benchmark is web apps.
OUT  2020  CCS      iDEA: Static Analysis on the Security of Apple Kernel Drivers   tvOS is one of four Apple OSes scanned; no TV-specific result.
OUT  2020  PETS     Smart Devices in Airbnbs: Considering Privacy and Security for  Survey; smart TV is a scenario option.
OUT  2022  IMC      Deep dive into the IoT backend ecosystem                        Backend infrastructure; TV mentions are motivation and a citation to FingerprinTV.
OUT  2022  PETS     A Multi-Region Investigation of the Perceptions and Use of Sma  Survey; smart TV is a related-work citation and an ownership option.
OUT  2022  PETS     Exploring the Privacy Concerns of Bystanders in Smart Homes fr  Survey; smart TV is an example in a prompt.
OUT  2023  CCS      IoTFlow: Inferring IoT Device Behavior at Scale through Static  Companion-app analysis; no TV breakout.
OUT  2023  IEEE-SP  Characterizing Everyday Misuse of Smart Home Devices            Survey of 483 people; smart TV is an ownership option, not a measured device.
OUT  2023  IEEE-SP  WebSpec: Towards Machine-Checked Analysis of Browser Security   "ACR" homonym.
OUT  2023  IEEE-SP  UTopia: Automatic Generation of Fuzz Driver using Unit Tests    Tizen as an open-source project under test; no TV device.
OUT  2023  PETS     No Privacy Among Spies: Assessing the Functionality and Insecu  Android stalkerware; "ACR" homonym.
OUT  2023  USENIX   Examining Consumer Reviews to Understand Security and Privacy   Review-text analysis; set-top box is a Mirai product category, no TV measurement.
OUT  2023  USENIX   Examining Power Dynamics and User Privacy in Smart Technology   Interview study; TVs are participant device inventories.
OUT  2023  USENIX   Internet Service Providers' and Individuals' Attitudes, Barrie  Interview and survey; TV is a device-ownership row.
OUT  2023  USENIX   "It's the Equivalent of Feeling Like You're in Jail”: Lessons   Interview study on IPV; TV is a reported abuse vector, not measured.
OUT  2023  USENIX   Measuring Up to (Reasonable) Consumer Expectations: Providing   Vignette survey; Vizio appears only in a news citation.
OUT  2023  USENIX   Abuse Vectors: A Framework for Conceptualizing IoT-Enabled Int  Qualitative framework; TV is an example abuse vector.
OUT  2023  USENIX   Exploring Tenants' Preferences of Privacy Negotiation in Airbn  Vignette survey; smart TV is a device-type option.
OUT  2023  WWW      SISSI: An Architecture for Semantic Interoperable Self-Soverei  "ACR" homonym (authentication context reference).
OUT  2024  IEEE-SP  SoK: Technical Implementation and Human Impact of Internet Pri  SoK; TV work cited, not measured.
OUT  2024  PETS     A Bilingual Longitudinal Analysis of Privacy Policies Measurin  ACR homonym: "ACR" is not automatic content recognition here.
OUT  2024  PETS     Contextualizing Interpersonal Data Sharing in Smart Homes       Vignette survey; "viewing history from your smart TV" is a question stem.
OUT  2024  PETS     "My Best Friend's Husband Sees and Knows Everything": A Cross-  Survey; smart TV is a free-text mention count.
OUT  2024  USENIX   Co-Designing a Mobile App for Bystander Privacy Protection in   Interview study; TV names are participant-reported device inventories.
OUT  2025  IEEE-SP  Analyzing the iOS Local Network Permission from a Technical an  Chromecast is one of four IoT devices used to trigger the permission; no TV result.
OUT  2025  IEEE-SP  Hey, Your Secrets Leaked! Detecting and Characterizing Secret   IPTV homonym in leaked-credential data.
OUT  2025  NDSS     Non-intrusive and Unconstrained Keystroke Inference in VR Plat  VR; smart TV appears only as a citation to HomeSpy.
OUT  2025  PETS     Help Me Help You: Privacy Considerations for Third Party IoT D  Vignette survey; TVs appear in a device-category prompt.
OUT  2025  PETS     Who Cares? Contextual Privacy Judgments from Owner and Bystand  Survey; smart TV is a device-category option.
OUT  2025  USENIX   Regulating Smart Device Support Periods: User Expectations and  Survey; Smart TV is a self-reported ownership category.
OUT  2026  IEEE-SP  Privacy Perspectives and Practices of Chinese Smart Home Produ  Interview study; smart TV is a company product-line row.
OUT  2026  NDSS     TBTrackerX: Fantastic Trigger Bots and Where to Find Malicious  IPTV spam homonym.
OUT  2026  PETS     Dead Domains, Living Data: A Privacy Risk Analysis of Domain L  Android apps; one expired-domain example happens to also ship on Roku.
OUT  2026  USENIX   PANGOLIN: Fuzzing Multilingual IoT Firmware with LLM-Driven Co  "SmartTVs" is a citation to the 2021 fuzzing paper, used as a baseline name.
 
==============================================================================
1. THE POPULATION BY VENUE, YEAR AND TOPIC
==============================================================================
Venue    A  B  A+B  Venue papers in corpus  Share of venue
-------  -  -  ---  ----------------------  --------------
IMC      3  8  11   638                     1.7%
USENIX   4  5  9    1410                    0.6%
NDSS     2  4  6    701                     0.9%
PETS     3  3  6    510                     1.2%
CCS      1  1  2    990                     0.2%
IEEE-SP  0  1  1    767                     0.1%
WWW      0  0  0    843                     0.0%
Window      A  B  A+B
----------  -  -  ---
2010–2014   3  0  3
2015–2018   0  1  1
2019–2021   3  8  11
2022–2024   6  9  15
2025–2026*  1  4  5
* 2025–2026 is provisional: CCS 2026 and IMC 2026 have not been held, and IEEE S&P / WWW 2026 are incompletely selected. See literature:corpus.
Topic (hand-assigned, ranking only)  A  B   A+B
-----------------------------------  -  --  ---
iot-device-set                       0  15  15
delivery-performance                 2  2   4
tracking                             3  0   3
vulnerability                        1  2   3
broadcast                            2  0   2
app-analysis                         1  1   2
side-channel                         2  0   2
acr                                  1  0   1
piracy                               1  0   1
device-population                    0  1   1
policy-compliance                    0  1   1
 
==============================================================================
2. PLATFORM FIELDS: WHY `web` IS THE WRONG FILTER
==============================================================================
platforms[] value     A+B (N=35)  share  corpus (N=5859)  share
--------------------  ----------  -----  ---------------  -----
web                   1           2.9%   1622             27.7%
mobile                6           17.1%  1075             18.3%
iot                   26          74.3%  436              7.4%
other-online-service  12          34.3%  2429             41.5%
offline               3           8.6%   2139             36.5%
not-applicable        0           0.0%   56               1.0%
 
Papers in the population carrying platform 'web' : 1 of 35
Which paper(s): USENIX 2026 Missing, Present and Conflicting: A Large Scale An
 
population[].unit == 'iot-devices' : 20 of 35
population[].unit == 'mobile-apps' : 5 of 35
population[].unit == 'websites'    : 1 of 35
 
In the corpus-wide 'crawled' population (crawlConfig or automated-web-crawl): 6 of 35
Which: CCS 2019, PETS 2022, USENIX 2023, CCS 2022, PETS 2024, USENIX 2026
 
==============================================================================
3. HOW THESE PAPERS GET AT THE TRAFFIC
==============================================================================
tool category           A (N=13)  A+B (N=35)
----------------------  --------  ----------
proxy-interception      3         4
traffic-capture         5         15
mobile-instrumentation  6         11
program-analysis        4         5
crawler-framework       1         2
browser-automation      1         1
network-scanner         1         4
blocklist               1         1
instrument (alias-folded)  papers (of 35)
-------------------------  --------------
Wireshark                  10
tcpdump                    9
mitmproxy                  4
Frida                      3
adb                        3
VirusTotal                 2
LibScout                   2
OpenWPM                    1
PingPong                   1
Mercury                    1
FlowDroid                  1
Charles Proxy              1
jadx                       1
Raspberry Pi               1
Scapy                      1
 
UNMAPPED RESIDUE of the alias fold — 191 distinct raw tool names used by the 35 papers, printed in full:
3× DBSCAN | 2× Censys | 2× dnsmasq | 2× Google voice synthesizer | 2× IoT Inspector | 2× nmap | 2× OpenSSL | 2× random forest | 2× scikit-learn | 2× t-SNE | 2× TF-IDF | 2× WHOIS | 1× Adam | 1× adb_shell | 1× Afatech AF9015 | 1× agglomerative clustering | 1× Anaconda | 1× Analysis Scripts | 1× Androguard | 1× Android Debug Bridge (adb) | 1× Android Debug Bridge (ADB) | 1× Android Studio APK Analyzer | 1× AntMonitor | 1× apk-mitm | 1× apksigner | 1× AppCensus | 1× Apple trust store | 1× Apple Wi-Fi geolocation API | 1× Application Exerciser Monkey | 1× Apriori | 1× arecord | 1× ARKit | 1× Avalpa OpenCaster | 1× BeautifulSoup | 1× BeEF Toolkit | 1× BERT | 1× BiLSTM | 1× Bing | 1× Bleak | 1× Bumble | 1× Chapoly1305/FindMy | 1× ChatGPT (OpenAI's TextCompletion API) | 1× Chrome | 1× CICFlowmeter | 1× CogniCrypt | 1× Common CA Database | 1× Conviva | 1× cosine distance | 1× Criminal IP | 1× crt.sh | 1× cryptography/fernet | 1× CryptoGuard | 1× curl | 1× DekTec DTU-215 | 1× DekTec StreamXpress | 1× DICE coefficient | 1× Dijkstra's algorithm | 1× DNSDB | 1× DPDK | 1× DroidBot | 1× fastText | 1× FCC database of digital TV towers | 1× Flight Radar 24 | 1× Forward feature selection (FFS) | 1× Fourier transform | 1× generic deep neural network | 1× GNU TLS | 1× Google Play API | 1× Google Public DNS | 1× Google search | 1× Google Search | 1× Google Voice synthesizer | 1× GPS Tracks | 1× Gradient Boosting Decision Tree | 1× GSDMM | 1× HDMI Video Capture Device | 1× HiDes UT-100c | 1× Hurricane Electric IPv6-over-IPv4 tunnel | 1× IDA Pro | 1× IDAPython | 1× IEEE Organizationally Unique Identifier registry | 1× IFTTT | 1× Intel RealSense Camera T265 | 1× InternalBlue | 1× IP2Location | 1× IPFIX | 1× iptables | 1× IRDB | 1× irgen | 1× IrScrutinizer | 1× Java | 1× Keras | 1× Latent Dirichlet Allocation | 1× LightGBM | 1× logistic regression (custom) | 1× MakeHex | 1× MAPS | 1× Maven Repository | 1× MaxMind | 1× MaxMind GeoLite2 | 1× MaxMind geolocation database | 1× Mbed TLS | 1× MbedTLS | 1× McAfee | 1× median absolute deviation (MAD) | 1× Microsoft trust store | 1× Mon(IoT)r | 1× Monkey Application Exerciser | 1× Monkey Application Exerciser for Android Studio | 1× Monte Carlo sampling | 1× Mother of all Ad-Blocking | 1× Mozilla trust store | 1× Naïve Bayes | 1× NASA SEDAC Metropolitan Statistical Areas dataset | 1× nDPI | 1× nearest-neighbor classifier | 1× Nessus | 1× Netdisco | 1× NetFlow | 1× Netify | 1× Nexmon | 1× NFF-Go | 1× NimBLE | 1× Non-Negative Matrix Factorization (NMF) | 1× NoxPlayer | 1× Objection | 1× OpenAI Text Completion API | 1× OpenCaster | 1× OpenDNS | 1× OpenWRT | 1× OpenWrt/LEDE | 1× OPP-115 | 1× Oracle Java | 1× passive network telescope | 1× Passport | 1× Pi-hole Default blocklist | 1× PostgreSQL | 1× PrivBERT | 1× Prodigy | 1× ProVerif | 1× pyshark | 1× Python | 1× Python requests/2.31.0 | 1× Python TLS implementation | 1× Radare2 | 1× Random Forest | 1× Randoop | 1× Raspberry Pi 3 | 1× Raspberry Pi 4 | 1× Redis | 1× RedOrbit HbbTV Emulator | 1× Remote Central Forums | 1× RIPE IPmap | 1× Roku External Control Protocol | 1× SciPy | 1× Secure Transport | 1× SHAP | 1× Similarweb | 1× Snorkel | 1× Softflowd | 1× SoSci Survey | 1× spaCy | 1× spaCy en_core_web_lg | 1× StopAd smart TV blocklist | 1× TensorFlow | 1× The Big Blocklist Collection (Firebog) | 1× TP-Link power plugs | 1× traceroute | 1× TrafficPassthrough | 1× Trigger Scripts | 1× TSDuck | 1× Tuya Smart app | 1× TV Fool | 1× tvbus.exe | 1× Unity | 1× Validation Scripts | 1× VLC Player | 1× VS1838B | 1× WALA | 1× WiFi Inspector | 1× WiGLE | 1× WiGLE API | 1× WireShark/tshark | 1× wolfSSL | 1× WolfSSL | 1× word2vec | 1× XCUITest | 1× XGBoost | 1× YAF | 1× Yersinia | 1× Zeek
 
Interception evidence, per Tier A paper (full text, whitespace-collapsed):
paper        mitm/proxy  own CA / root cert  pinning / cert failure  undecryptable reported  DNS-level capture  router/AP capture  HDMI / screen capture  remote-control automation
-----------  ----------  ------------------  ----------------------  ----------------------  -----------------  -----------------  ---------------------  -------------------------
IMC 2011     —           —                   —                       —                       —                  1                  —                      1
IMC 2011     1           —                   —                       —                       —                  1                  —                      —
USENIX 2014  3           —                   —                       —                       —                  4                  —                      1
CCS 2019     53          1                   12                      —                       27                 5                  5                      35
PETS 2020    1           —                   2                       9                       14                 12                 —                      4
USENIX 2021  8           —                   —                       —                       —                  1                  14                     1
PETS 2022    —           —                   —                       —                       4                  11                 —                      4
PETS 2022    9           —                   2                       —                       —                  —                  2                      1
NDSS 2023    16          —                   3                       —                       11                 18                 3                      2
USENIX 2023  1           —                   —                       —                       —                  4                  1                      41
IMC 2024     2           —                   —                       —                       —                  3                  15                     4
NDSS 2024    —           —                   —                       —                       —                  —                  —                      3
USENIX 2025  2           —                   —                       —                       5                  1                  —                      5
mitm/proxy                 10 of 13 Tier A   |   22 of 35 A+B
own CA / root cert         1 of 13 Tier A   |   4 of 35 A+B
pinning / cert failure     4 of 13 Tier A   |   6 of 35 A+B
undecryptable reported     1 of 13 Tier A   |   1 of 35 A+B
DNS-level capture          5 of 13 Tier A   |   12 of 35 A+B
router/AP capture          11 of 13 Tier A   |   30 of 35 A+B
HDMI / screen capture      6 of 13 Tier A   |   9 of 35 A+B
remote-control automation  12 of 13 Tier A   |   16 of 35 A+B
 
==============================================================================
3b. VANTAGE
==============================================================================
vantage location (verbatim, unfolded)  papers
-------------------------------------  ------
United States                          5
Italy                                  2
Germany                                2
France                                 2
Austria                                2
Finland                                2
US                                     2
UK                                     2
Europe                                 2
Portugal                               2
Sweden                                 2
Norway                                 2
United Kingdom                         1
241 countries and territories          1
Japan                                  1
Korea                                  1
China                                  1
US (North Carolina)                    1
office space                           1
Apartment 1                            1
Apartment 2                            1
lab space                              1
U.S.                                   1
Asia                                   1
New York, U.S.                         1
Frankfurt, Europe                      1
Singapore, Asia                        1
Sydney, NSW, Australia                 1
worldwide                              1
USA                                    1
university building                    1
single-family house                    1
e-bike route                           1
flight                                 1
Belgium                                1
Bulgaria                               1
Croatia                                1
Cyprus                                 1
Czech Republic                         1
Denmark                                1
Estonia                                1
Greece                                 1
Hungary                                1
Iceland                                1
Ireland                                1
Latvia                                 1
Liechtenstein                          1
Lithuania                              1
Luxembourg                             1
Malta                                  1
Netherlands                            1
Poland                                 1
Romania                                1
Slovakia                               1
Slovenia                               1
Spain                                  1
 
papers with a vantage tuple: 34 of 35; of those, stating a location: 18 (52.9%)
 
==============================================================================
4. WHAT THEY SAMPLE — THE NO-TRANCO PROBLEM
==============================================================================
population.sourceList (verbatim, unfolded — ranking only)   papers
----------------------------------------------------------  ------
custom seed list                                            7
Roku Channel Store                                          3
Google Play Store                                           2
IoT Inspector                                               2
DS1: operational IPTV traces                                1
DS2: operational IPTV traces                                1
DS3: operational IPTV traces                                1
large commercial IPTV service provider in the United State  1
NASA SEDAC Metropolitan Statistical Areas dataset           1
FCC database of digital TV towers in the United States      1
station coverage maps supplied by TV Fool                   1
Amazon Fire TV channel store                                1
Roku-Top1K                                                  1
FireTV-Top1K                                                1
residential gateways                                        1
Amazon curated list "Top Featured" apps                     1
buyers' guides in North America and Europe                  1
Apple iTunes Preview / Apple App Store                      1
Fire TV app store                                           1
AndroZoo and APKMirror                                      1
custom channel selection                                    1
social media platforms                                      1
IRDB and Remote Central Forums                              1
password-lists and username-lists                           1
custom testbed                                              1
custom user-study recruitment                               1
2014 PhpBB password leak                                    1
RockYou password leak                                       1
generated synthetic credit card details                     1
news headlines                                              1
Conviva                                                     1
Avast WiFi Inspector                                        1
Censys                                                      1
custom IoT testbeds                                         1
large European ISP                                          1
major European IXP                                          1
custom device selection                                     1
Mon(IoT)r dataset                                           1
UNSW Smart Home Traffic Dataset                             1
YourThings Smart Home Traffic Dataset                       1
UNB Simulated Office-Space Traffic Dataset                  1
11 typical offices and apartments that are accessible to u  1
custom test scenes                                          1
custom-collected WiFi networks                              1
Androzoo                                                    1
Google Play                                                 1
YourThings Dataset                                          1
HomeSnitch Dataset                                          1
PingPong Dataset                                            1
Mon(IoT)r Dataset                                           1
UNSW Dataset                                                1
Our Dataset                                                 1
custom device set                                           1
lab dataset                                                 1
custom lab setup                                            1
MonIoTr Lab                                                 1
AndroZoo                                                    1
IoT Inspector dataset                                       1
IEEE OUI database                                           1
WiGLE                                                       1
popular US stores, including Amazon.com                     1
custom laboratory experimental setup                        1
custom campus network                                       1
custom home network                                         1
UNSW IoT Analytics                                          1
YourThings IoTFinder                                        1
custom SERP corpus                                          1
OPP-115                                                     1
custom testbeds                                             1
custom test devices                                         1
population.unit     papers
------------------  ------
iot-devices         20
other               17
mobile-apps         5
human-participants  3
documents           3
ip-addresses        1
autonomous-systems  1
network-flows       1
domains             1
websites            1
web-pages           1
population.samplingMethod  papers
-------------------------  ------
purposive                  19
pre-existing-dataset       11
convenience                10
exhaustive                 7
top-n                      5
seed-and-crawl             5
random                     5
not-stated                 4
stratified                 3
paper                                               iot-devices n values stated
--------------------------------------------------  ---------------------------
IMC 2011 q-score-proactive-service-quality-asse     7000000, 140000
PETS 2020 the-tv-is-smart-and-full-of-trackers-m    57
USENIX 2021 android-smarttvs-vulnerability-discove  11
IMC 2024 watching-tv-with-the-second-party-a-fi     2
IMC 2019 information-exposure-from-consumer-iot     81
USENIX 2019 all-things-considered-an-analysis-of-i  83000000, 500000, 1000
IMC 2020 a-haystack-full-of-needles-scalable-de     96
NDSS 2020 packet-level-signatures-for-smart-home    19, 55, 26, 45
NDSS 2020 et-tu-alexa-when-commodity-wifi-device    31
USENIX 2020 you-are-what-you-broadcast-identificat  31850, 423, 26478
IMC 2021 iotls-understanding-tls-usage-in-consu     40
PETS 2021 blocking-without-breaking-identificati    31
PETS 2022 analyzing-the-feasibility-and-generali    45, 28, 18, 70, 19, 8
USENIX 2022 lumos-identifying-and-localizing-diver  44
IMC 2023 behind-the-scenes-uncovering-tls-and-s     2014, 113, 7
IMC 2023 in-the-room-where-it-happens-character     93, 13487
IMC 2024 iot-bricks-over-v6-understanding-ipv6-     93
PETS 2024 connecting-the-dots-tracing-data-endpo    25123, 54950, 30, 66
NDSS 2025 evaluating-machine-learning-based-iot-    90
NDSS 2026 blerp-ble-re-pairing-attacks-and-defen    22
 
papers stating an iot-devices size : 20 of 35
stated size values                : 39
all values, sorted                : 2, 7, 8, 11, 18, 19, 19, 22, 26, 28, 30, 31, 31, 40, 44, 45, 45, 55, 57, 66, 70, 81, 90, 93, 93, 96, 113, 423, 1000, 2014, 13487, 25123, 26478, 31850, 54950, 140000, 500000, 7000000, 83000000
median of all stated values       : 66
largest set per paper, sorted     : 2, 11, 22, 31, 31, 40, 44, 55, 57, 70, 81, 90, 93, 96, 2014, 13487, 31850, 54950, 7000000, 83000000
median of largest-per-paper       : 75.5
papers whose LARGEST set <=  100   : 14 of 20
papers whose LARGEST set <=  200   : 14 of 20
papers whose LARGEST set <= 1000   : 14 of 20
papers whose largest set is over 200, i.e. not a lab bench:
  IMC 2011 q-score-proactive-service-quality-asse  -> 7000000, 140000
  USENIX 2019 all-things-considered-an-analysis-of-i  -> 83000000, 500000, 1000
  USENIX 2020 you-are-what-you-broadcast-identificat  -> 31850, 423, 26478
  IMC 2023 behind-the-scenes-uncovering-tls-and-s  -> 2014, 113, 7
  IMC 2023 in-the-room-where-it-happens-character  -> 93, 13487
  PETS 2024 connecting-the-dots-tracing-data-endpo  -> 25123, 54950, 30, 66
 
population tuples in A+B stating a listVersion: 42 of 105
 
==============================================================================
5. MEASURED RESULTS (detection[].prevalence, Tier A only)
==============================================================================
 
--- IMC 2011 | Understanding couch potatoes: measurement and modeling of interactive usage of IPTV at large scale.
  * Stream control operations
      technique : Analyzed logged control events and constructed a finite state machine.
      metric    : relative proportion of operations
      prevalence: FastForward, play, and Replay comprised 45% of total events.
      quote (results): The sum of FastForward (FF), play, and Replay comprise 45% of the total.
  * Video popularity
      technique : Counted requests and tracked rank persistence over time.
      metric    : rank-frequency distribution and top-N retention
      prevalence: Top-100 and top-300 video popularity dropped rapidly over 60 days.
      quote (results): For the next 60 days, we counted how many of those popular videos remained among the top 300 (or 100) most popular.
  * Stream-control server impact
      technique : Compared actual traces with and without control events in a discrete-event simulator.
      metric    : peak server bandwidth
      prevalence: Peak bandwidth increased from about 17 Gbps to about 20 Gbps.
      quote (results): We observe that server bandwidth increases from about 17 Gbps to about 20 Gbps at the peak when stream control operations are accounted for
 
--- IMC 2011 | Q-score: proactive service quality assessment in a large IPTV system.
  * customer service problem prediction
      technique : Ridge regression over network KPIs and aggregated customer trouble tickets
      metric    : false negative rate and false positive rate
      prevalence: predict 60% of service problems reported by customers with only 0.1% false positive rate
      quote (results): Q-score is able to predict 60% of service problems reported by customers with only 0.1% misclassification (i.e., false positive rate).
  * proactive service degradation detection
      technique : Evaluated Q-score accuracy after increasing network-event-to-feedback skip intervals
      metric    : lead time before customer reports
      prevalence: 9 hours of lead time preserved 0.1% FPR, with FNR increasing from 30% to 40%
      quote (results): we find 9 hours of lead time is at the feasible level, as observing 9 hours of skip interval preserves 0.1% of FPR only sacrificing 10% of FNR
 
--- USENIX 2014 | From the Aether to the Ethernet—Attacking the Internet using Broadcast Digital Television
  * RF injection into DVB-T broadcasts
      technique : Intercepted, modified, and retransmitted DVB streams on the original frequency.
      metric    : coverage radius and area
      prevalence: 1 W amplifier: 477 m radius and 1.4 km²; 25 W: 2385 m radius and 35 km²
      quote (results): Using this formula shows that with a 1 W (30 dBm) amplifier ... cover a region with radius of 477 m, or an area of 1.4 km2.
  * Malicious HbbTV application execution
      technique : Injected AIT and HTML payloads into multiplexed DVB streams.
      metric    : successful attack capabilities on one smart TV
      prevalence: Invisible execution, screen takeover, intranet scanning, TV crash, and external-web-server denial of service
      quote (evaluation): Using our test setup, we were able to create HbbTV applications which ran invisibly in the background, as well as applications which completely took over the TV screen.
  * Urban attack scalability
      technique : Cross-correlated population density, tower data, and propagation coverage maps.
      metric    : number of stations and potentially affected devices
      prevalence: More than 20,000 devices in a single attack; up to 10 stations in parts of New York City
      quote (results): In certain locations in the Inwood area, where the population density is 50,000 persons per km2, the attacker can infect 10 different stations
  * HbbTV Internet and intranet attacks
      technique : Ran malicious JavaScript through injected HbbTV applications.
      metric    : verified attack types
      prevalence: Port scanning, fraudulent login display, malformed-image crash, and denial of service
      quote (evaluation): We verified that we were able to access servers both on the Internet at large and on the local intranet.
 
--- CCS 2019 | Watching You Watch: The Tracking Ecosystem of Over-the-Top TV Streaming Devices.
  * known tracker contact
      technique : Matched contacted hosts and domains against five tracking lists.
      metric    : share of channels contacting known trackers
      prevalence: 69% of Roku channels and 89% of Amazon Fire TV channels
      quote (abstract): traffic to known trackers present on 69% of Roku channels and 89% of Amazon Fire TV channels.
  * identifier leakage
      technique : Searched HTTP URLs, headers, cookies, and bodies for encoded, hashed, or cleartext identifiers.
      metric    : requests or identifiers containing leaks
      prevalence: 4,452 of 6,142 Roku requests containing AD ID or serial number were tracker-flagged; 3,427 of 8,433 Amazon identifiers were cleartext
      quote (results): We searched for various encoding and hashing combinations using the method described by Englehardt et al. [22].
  * unencrypted HTTP traffic
      technique : Counted requests sent over port 80 in captured PCAPs.
      metric    : share of channels sending cleartext requests
      prevalence: 794 of 1,000 Roku channels and 762 of 1,000 Fire TV channels
      quote (results): Analyzing the requests sent over port 80 we found that 794 of the 1000 Roku channels sent at least one request in cleartext.
  * video title leakage
      technique : Searched captured traffic for encodings of manually identified video titles.
      metric    : channels leaking titles to tracking domains
      prevalence: 9 of 100 Roku channels and 14 of 100 Fire TV channels
      quote (results): We found 9 channels on Roku and 14 channels on the Fire TV, among the 100 channels we randomly selected on each device, that leaked the title of the video to a tracking domain.
  * TLS certificate validation
      technique : Attempted MITM interception and measured successfully decrypted channels.
      metric    : channels with intercepted HTTPS
      prevalence: 957 Fire TV channels and 43 Roku channels
      quote (results): On Amazon Fire TV, we were able to install our own cert on the device which allowed us to intercept HTTPS requests on 957 of the 1000 channels.
  * remote-control API vulnerability
      technique : Reverse engineered API traffic and tested malicious cross-origin web requests.
      metric    : vulnerability assessment
      prevalence: Roku API exposed identifiers, channel control, installed-channel lists, and SSID
      quote (results): We set up a page to demonstrate the attack and verified that a malicious web page visited by Roku users ... can abuse the External Control API.
 
--- PETS 2020 | The TV is Smart and Full of Trackers: Measuring Smart TV Advertising and Tracking
  * ATS domains
      technique : Service labels and union of DNS blocklists
      metric    : share of domains or apps contacting ATSes
      prevalence: about 10% of Roku and Fire TV apps contact 20+ and 10+ ATS domains, respectively
      quote (results): about 10% of the Roku and Fire TV apps contact 20+ and 10+ ATS domains, respectively.
  * Platform segmentation
      technique : Compared FQDN, eSLD, and parent-organization overlap
      metric    : dataset and eSLD overlap
      prevalence: 314 ATS domains unique to Roku, 285 unique to Fire TV, and 227 overlapping
      quote (results): we identify 314 ATS domains that are unique to the Roku dataset, 285 that are unique to the Fire TV dataset, and an overlap of 227 between the two datasets.
  * DNS blocklist coverage
      technique : Matched contacted FQDNs against four blocklists
      metric    : block rate
      prevalence: Firebog blocked 22% of Roku and 27% of Fire TV testbed FQDNs
      quote (results): TF, closely followed by MoaAB and PD, blocks the highest fraction of domains across all of the platforms in both the in the wild and testbed datasets.
  * Missed ads and app breakage
      technique : Repeated manual app interaction under each blocklist
      metric    : ads missed and functionality breakage
      prevalence: all blocklists produced non-trivial false positives and false negatives
      quote (results): All blocklists suffer from a non-trivial amount of visually observable FPs and FNs.
  * PII exposure
      technique : Searched HTTP headers and URI paths for raw and hashed identifiers
      metric    : apps, eSLDs, and blocked FQDNs
      prevalence: hundreds of apps exfiltrated PII to third parties and platform-specific parties
      quote (conclusion): Hundreds of Roku and Fire TV apps expose PII, mostly to third parties and the platform-specific party.
  * Joint advertising and static identifiers
      technique : Detected co-occurrence of advertising IDs with serial or device IDs
      metric    : number of apps
      prevalence: 697 Fire TV apps sent advertising ID alongside serial number and device ID
      quote (results): Aside from the 697 Fire TV apps that expose advertising ID alongside serial number and device ID discussed earlier.
  * TLS interception failure
      technique : Compared TLS connections with connections containing decrypted HTTP packets
      metric    : decryption failure rate
      prevalence: failure was at most 20% of TLS connections for 80% of Fire TV apps
      quote (appendix): decryption fails for 1 out of 5 (or fewer) TLS connections for 80% of all apps.
 
--- USENIX 2021 | Android SmartTVs Vulnerability Discovery via Log-Guided Fuzzing
  * SmartTV API vulnerabilities
      technique : Log-guided dynamic fuzzing with cyber and physical feedback
      metric    : unique vulnerabilities
      prevalence: 37 unique vulnerabilities across 11 Android TVBoxes
      quote (abstract): Our analysis led to the automatic discovery of 37 unique vulnerabilities, including 11 high-impact cyber threats, 10 new memory corruptions, and 16 visual and auditory anomalies.
  * Memory corruptions
      technique : Execution-log monitoring for crashes and anomalous states
      metric    : number of vulnerabilities
      prevalence: 10 memory corruptions
      quote (evaluation): We discovered 37 security-critical flaws leading to various cyber attacks (11), physical disturbances (16) and memory corruptions (10).
  * Visual and auditory anomalies
      technique : External HDMI capture and before-after signal comparison
      metric    : number of anomalies
      prevalence: 16 visual and auditory anomalies
      quote (introduction): Our analysis led to the automatic discovery of 37 unique vulnerabilities, including 11 high-impact cyber threats, 10 new memory corruptions, and 16 visual and auditory anomalies.
  * Input-validation messages
      technique : CNN classification of log messages trained from Android ROMs
      metric    : classifier accuracy and recall
      prevalence: 46% of APIs triggered at least one input validation
      quote (evaluation): As shown, on average 87% APIs triggered at least 1 log message and 46% triggered at least 1 input validation.
 
--- PETS 2022 | FingerprinTV: Fingerprinting Smart TV Apps
  * domain-based fingerprints
      technique : Extracted domains recurring in all ten launch samples; clustered app fingerprints.
      metric    : prevalence and distinctiveness
      prevalence: 96% Apple TV, 88% Fire TV, and 100% Roku apps exhibited DBFs.
      quote (results): We find that 96% (N = 961) of the top-1000 Apple TV apps exhibit a DBF; 88% (N = 884) ... Fire TV ... and 100% ... Roku
  * packet-pair-based fingerprints
      technique : Extracted identical-size directional packet pairs using PingPong.
      metric    : prevalence and distinctiveness
      prevalence: 68% Apple TV, 95% Fire TV, and 100% Roku apps exhibited PBFs.
      quote (results): We find that 68% (N = 678) of the top-1000 Apple TV apps exhibit a PBF; 95% (N = 952) ... Fire TV ... and 100% ... Roku
  * TLS-based fingerprints
      technique : Extracted recurring TLS ClientHello fingerprints using Mercury.
      metric    : prevalence and distinctiveness
      prevalence: 95% Apple TV, 86% Fire TV, and 100% Roku apps exhibited TBFs; distinctiveness was 3%, 7%, and 1%.
      quote (results): only 3%, 7%, and 1% of the Apple TV, Fire TV, and Roku apps that exhibit TBFs, exhibit distinct TBFs.
  * smart TV app fingerprinting
      technique : Combined DBF and PBF fingerprints and evaluated cluster uniqueness.
      metric    : prevalence and distinctiveness
      prevalence: DBF-and-PBF fingerprints were distinct for 89% Apple TV, 95% Fire TV, and 76% Roku apps exhibiting both.
      quote (results): the fingerprint is distinct for 89% (599) of the Apple TV apps, 95% (802) of the Fire TV apps, and 76% (760) of the Roku apps
  * platform-specific fingerprints
      technique : Compared fingerprints for apps matched across all three platforms.
      metric    : share of matched apps with platform differences
      prevalence: 76% of 80 apps available on all three platforms exhibited different fingerprints on each platform.
      quote (introduction): among 80 apps that are made available on all three smart TV platforms, 76% exhibit a different fingerprint on each platform
  * identical fingerprints
      technique : Examined developer identities and parent organizations within non-singleton clusters.
      metric    : share attributable to same developer
      prevalence: DBF-sharing apps attributable solely to the same developer: 37% Apple TV, 57% Fire TV, and 20% Roku before consolidation.
      quote (results): For Apple TV, 142 of the 397 apps (37%) that share their DBF with other app(s) only share it with other apps from the same developer.
 
--- PETS 2022 | Watch Over Your TV: A Security and Privacy Analysis of the Android TV Ecosystem
  * third-party library prevalence
      technique : Libscout plus package-prefix clustering and manual classification
      metric    : share of apps containing libraries
      prevalence: Social Media libraries in 88%, analytics in 75%, and advertising in 77% of apps
      quote (results): We detected Social Media libraries in 88% of the apps, with multiple Facebook libraries occupying the top 5. Similarly, we found that 75% of the apps contain analytics libraries and 77% contain advertising libraries.
  * sensitive data flows
      technique : Customized Flowdroid taint analysis
      metric    : share of APKs with at least one sensitive flow
      prevalence: 78% of files
      quote (results): The analysis found at least one sensitive data flow in 78% of the files.
  * static identifier use
      technique : Static taint-flow source and sink analysis
      metric    : APKs with identifier flows
      prevalence: 3031 APKs used UUID-generated globally unique identifiers; 285 APKs exposed MAC-address or SSID identifiers
      quote (results): The most used identifier is a globally unique ID (GUID) generated with the java.util.UUID package (3031 APKs).
  * network data leakage
      technique : Charles-captured traces manually searched for sensitive values
      metric    : TV apps sending data to first parties, trackers, or CDNs
      prevalence: 42% of explored TV apps used static identifiers; media metadata appeared in 65%
      quote (results): We captured traffic of 21 TV apps and 22 mobile apps out of 30 Popular-Streaming apps.
  * malware
      technique : VirusTotal multi-engine detection threshold
      metric    : APKs flagged by antivirus engines
      prevalence: 34 APKs were flagged by more than 5 engines; 34 were discussed as flagged by more than 10 engines
      quote (results): We detected 34 APKs in our dataset that are flagged as malware by more than 10 engines in VirusTotal (VT).
  * socket communication
      technique : Cross-reference analysis of socket API methods
      metric    : APKs containing Socket APIs
      prevalence: 2646 APKs, or 56%
      quote (results): Overall, we detected 2646 APKs (56%) including Socket APIs.
  * Nearby API authentication
      technique : Search for authentication-token reads in callbacks
      metric    : APKs implementing authentication
      prevalence: None of the APKs using NearbyConnection implemented authentication
      quote (results): Unfortunately, none of the APKs that use the NearbyConnection API implement authentication.
  * permission rationale display
      technique : Bytecode pattern matching and context heuristics
      metric    : apps showing rationale for dangerous permissions
      prevalence: Only small percentages showed context, including 3% for coarse location and 5% for record audio
      quote (results): Our results indicate that only a small percentage of TV apps display the rationale behind dangerous permission
 
--- NDSS 2023 | I Still Know What You Watched Last Sunday: Privacy of the HbbTV Protocol in the European Smart TV Landscape
  * Tracking before consent
      technique : Inspected contacted domains and consent-phase traffic
      metric    : share of TV channels
      prevalence: 26 of 36 channels communicated with trackers before consent
      quote (discussion): All the 36 TV channels we analyzed contact at least one tracking domain; further, 26 communicate with trackers before the user has expressed their consent.
  * Tracking pixels
      technique : Identified returned 1×1 pixel image objects
      metric    : share of TV channels
      prevalence: 20 of 36 channels (56%) adopted tracking pixels
      quote (discussion): 20 of the 36 TV channels (56%) we analyzed adopt the invisible 'tracking pixel' to profile users.
  * Periodic tracking requests
      technique : Computed mean and standard deviation of inter-request times
      metric    : average time between requests
      prevalence: SportItalia approximately every 70 seconds; RDS approximately every 14 seconds
      quote (results): SportItalia makes requests to Smartclip ... around every 70 seconds, while RDS contacts Google Analytics around every 14 seconds.
  * Plaintext sensitive traffic
      technique : Inspected unencrypted HTTP packets and payloads
      metric    : presence of sensitive data in HTTP
      prevalence: Most traffic captures included HTTP; HSE exposed login information
      quote (discussion): We found HTTP communication in most of our traffic captures; such traffic contained sensitive information such as device IDs, visitor IDs, country codes, and ISP information.
  * Denylist coverage
      technique : Matched manually identified tracking domains against Pi-hole and EasyList
      metric    : fraction of tracking domains blocked
      prevalence: At maximum 44% in 2021 and 81% in 2022
      quote (discussion): Commonly used tracking denylists only block at maximum 44% in 2021 and 81% in 2022 of the domains in our traffic captures and marked as tracking.
  * User risk awareness
      technique : Surveyed awareness and coded open-ended responses
      metric    : share mentioning no security or privacy risk
      prevalence: 68% of 132 Smart TV-related respondents mentioned no risk
      quote (discussion): Out of the 132 participants in the Smart TV and HbbTV awareness survey, 68% could not mention any security or privacy risk
 
--- USENIX 2023 | HOMESPY: The Invisible Sniffer of Infrared Remote Control of Smart TVs
  * Off-path infrared signal sniffing
      technique : Mounted a COTS receiver at positions across four room layouts.
      metric    : IR-key extraction accuracy
      prevalence: 98.9%, 90.8%, 83.1%, and 75.0% for layouts A–D
      quote (results): The result is shown in Table 4. In particular, for layout A, 98.9% of the keys (including the repeat keys) can be correctly sniffed
  * IR command decoding
      technique : Matched recovered timings against a 75,901-code hash database.
      metric    : unique device mappings
      prevalence: 98% of devices had unique D-pad, OK, or BACK mappings
      quote (results): Among the 1303 devices, only 26 devices have overlapped D-pad, OK or BACK keys, which means that for 98% of all devices have their unique mappings.
  * Sensitive input extraction
      technique : Mapped D-pad sequences onto virtual-keyboard coordinates and filtered candidates.
      metric    : Top-1, Top-3, and Top-5 accuracy
      prevalence: 47% Top-1, 70% Top-3, and 77% Top-5
      quote (results): The accuracy increases to 70% for Top3 and 77% for Top5.
  * Virtual-keyboard activity detection
      technique : Applied OK-count and BACK-absence thresholds within time windows.
      metric    : activity coverage
      prevalence: 100% of keyboard-input activities
      quote (results): The experiment result shows H OME S PY could cover 100% of the activities about keyboard input.
 
--- IMC 2024 | Watching TV with the Second-Party: A First Look at Automatic Content Recognition Tracking in Smart TVs.
  * ACR network traffic
      technique : Filtered captured DNS domains containing the string “acr”.
      metric    : traffic presence, bytes, frequency, and packet timing
      prevalence: ACR traffic existed during linear TV and HDMI scenarios.
      quote (introduction): ACR network traffic exists when watching linear TV and when using smart TV as an external display using HDMI
  * ACR traffic after opt-out
      technique : Compared traffic across opted-in and opted-out experimental phases.
      metric    : presence or absence of communication with ACR domains
      prevalence: Opting out produced a complete absence of communication with previously identified ACR domains.
      quote (results): once opt-out is exercised (Table 1), there is a complete absence of communication with any previously identified ACR domains
  * Geographic ACR differences
      technique : Compared contacted domains and geolocated their server IP addresses.
      metric    : domain identity, server location, and traffic levels
      prevalence: UK and US televisions contacted distinct ACR domains; US FAST streaming generated ACR traffic unlike the UK.
      quote (introduction): smart TVs in the UK and the US contact distinct ACR domains
  * Login-status effect
      technique : Compared logged-in and logged-out phases for each television.
      metric    : CDFs of bytes transferred and traffic periodicity
      prevalence: User login status appeared to have no material impact on ACR traffic behavior.
      quote (results): user login status appears to have no material impact on the ACR network traffic behavior.
  * ACR traffic by viewing scenario
      technique : Ran six one-hour scenarios on each television.
      metric    : traffic volume, frequency, and periodicity
      prevalence: Linear and HDMI had the highest ACR traffic for both brands in the UK.
      quote (results): For both LG (a) and Samsung (b) TVs, the scenarios with the highest ACR traffic are Linear and HDMI.
 
--- NDSS 2024 | Acoustic Keystroke Leakage on Smart Televisions
  * acoustic keystroke leakage
      technique : Matched reference Smart TV sounds and extracted movement-count sequences.
      metric    : top-K recovery accuracy
      prevalence: up to 60.19% of common passwords within 100 guesses for realistic Samsung users
      quote (introduction): For ten subjects typing into real applications on a Samsung TV, the attack recovers 53.33% of CCNs, 33.33% of full credit card details, and up to 60.19% of common passwords.
  * keyboard-instance splitting
      technique : Used SystemSelect sounds or timing outliers between adjacent movements.
      metric    : correctly identified instances
      prevalence: 98 of 100 passwords
      quote (results): This method correctly identifies 98 / 100 passwords, where the two failures result from long pauses while typing.
  * credit-card entry detection
      technique : Matched consecutive keyboard-instance lengths to payment-field sizes.
      metric    : detected interactions
      prevalence: 29 of 30 human credit-card interactions
      quote (results): The attack successfully identifies that the user enters credit card details in 29 of the 30 total interactions.
  * password-entry classification
      technique : Random Forest classified dynamic-suggestion behavior from movement histograms.
      metric    : accuracy
      prevalence: 99.02% on the human password set
      quote (results): Further, we infer when a subject types a password (§IV-B1), and this classifier has an accuracy of 99.02% on this set.
  * suboptimal keyboard paths
      technique : Used pauses and movement timing to prioritize tolerated path deviations.
      metric    : optimal-path rate
      prevalence: 89.35% for human credit-card typing on Samsung; 45.35% on AppleTV passwords
      quote (results): When typing CCNs, users take the optimal path between keys only 89.35% of the time.
  * direction inference
      technique : Detected rapid horizontal scrolls using median adjacent-movement timing.
      metric    : recovery impact
      prevalence: Direction inference never harms recovery rates
      quote (methodology): We find that direction inference never harms the recovery (§VI-E).
 
--- USENIX 2025 | Watch Out Your TV Box: Reversing and Blocking a P2P-based Illegal Streaming Ecosystem
  * EVPAD P2P users
      technique : Queried Broker peer lists and deduplicated composite IP identifiers.
      metric    : unique peers
      prevalence: 131,175 unique peers over two months
      quote (results): Consequently, our analysis revealed that a total of 131,175 unique peers were active in EVBOX's P2P network had over the two-month observation period.
  * Operational servers
      technique : Extracted server IPs from reversed app data and peer responses.
      metric    : operational servers
      prevalence: 78 operational servers
      quote (conclusion): With EVPAD devices, we identified 131,175 users across 116 countries and 78 operational servers located in the United States, Japan, Singapore, Hong Kong, and other countries.
  * Concurrent streaming users
      technique : Repeatedly collected users across all channels over two months.
      metric    : users across all channels per measurement
      prevalence: minimum 13,438, maximum 19,889, average 16,997
      quote (results): our observations of concurrent users across all channels over 30 measurements spanning two months—with a minimum of 13,438, maximum of 19,889, and an average of 16,997 users
  * VoD content
      technique : Changed category identifiers and collected encrypted VoD lists.
      metric    : collected content items
      prevalence: 24,934 pieces of content
      quote (results): By changing the ID values for the main and subcategories, we collected all VoD lists, resulting in a total of 24,934 pieces of content.
  * Country distribution
      technique : Geolocated collected peer IP addresses.
      metric    : countries represented
      prevalence: 116 countries
      quote (results): Upon identifying the country associated with each IP address, we found that users were distributed across 116 countries [27].
  * Authentication bypass
      technique : Replayed copied device attributes in NoxPlayer.
      metric    : successful service access
      prevalence: emulator successfully accessed and streamed live channels
      quote (evaluation): We verified this vulnerability by using the NoxPlayer to mimic an authenticated device. As shown in Figure 8, the emulator successfully accessed and streamed live channels
  * P2P denial of service
      technique : Sent a crafted TCP packet to a connected EVPAD peer.
      metric    : service termination
      prevalence: a single crafted TCP packet terminated the target service
      quote (evaluation): From one device, we crafted and sent a TCP socket-based packet to the other device... the EVPAD streaming service on the target device immediately terminated.
 
detection tuples in Tier A: 70; carrying a prevalence: 68 (97.1%)
corpus-wide: 27241 tuples, 26316 carry a prevalence (96.6%)
 
==============================================================================
6. WHERE THIS LITERATURE GOES QUIET
==============================================================================
crawlConfig field (papers with a crawlConfig object)  A+B (N=6)       corpus crawled (N=1080)
----------------------------------------------------  --------------  -----------------------
consentAction                                         0 of 6 (0.0%)   349 of 1080 (32.3%)
interactionDepth                                      5 of 6 (83.3%)  841 of 1080 (77.9%)
statefulness                                          0 of 6 (0.0%)   219 of 1080 (20.3%)
browsers                                              0 of 6 (0.0%)   529 of 1080 (49.0%)
 
crawlConfig values in the A+B population, verbatim:
  CCS 2019  consent=not-applicable  depth=single-target-page  state=not-stated  browsers=[]  robotsTxt=not-stated
  PETS 2022  consent=not-applicable  depth=deep-crawl  state=not-stated  browsers=[]  robotsTxt=not-stated
  USENIX 2023  consent=not-stated  depth=not-stated  state=not-stated  browsers=[]  robotsTxt=not-stated
  CCS 2022  consent=not-applicable  depth=single-target-page  state=not-stated  browsers=[]  robotsTxt=not-stated
  PETS 2024  consent=not-stated  depth=single-target-page  state=not-stated  browsers=[]  robotsTxt=not-stated
  USENIX 2026  consent=not-stated  depth=landing-plus-subpages  state=not-stated  browsers=[]  robotsTxt=not-stated
 
A+B papers carrying an ethics object: 34 of 35;  artifacts object: 34 of 35
field                          A+B               corpus empirical
-----------------------------  ----------------  --------------------
ethics.reviewOutcome stated    14 of 34 (41.2%)  1728 of 4472 (38.6%)
artifacts.availability stated  27 of 34 (79.4%)  2890 of 4854 (59.5%)
temporal.spanStart stated      27 of 34 (79.4%)  2882 of 5118 (56.3%)
 
Artifact links released, Tier A:
  IMC 2011  public                 www.research.att.com/∼kkrama/papers/streamcontrol.pdf
  IMC 2011  (no artifacts object extracted)
  USENIX 2014  none-mentioned         http://www.avalpa.com/the-key-values/15-free-software/33-opencaster  (other; OpenCaster software; authors=false)
  CCS 2019  promised-not-yet-available —
  PETS 2020  promised-not-yet-available http://athinagroup.eng.uci.edu/projects/smarttv/  (project-page; Project page for tools and testbed datasets; authors=true)
  USENIX 2021  none-mentioned         https://sites.google.com/site/smarttvdemos/  (project-page; Demonstration website for discovered attacks; authors=true)
  PETS 2022  promised-not-yet-available https://github.com/UCI-Networking-Group/fingerprintv
  PETS 2022  public                 https://gitlab.com/s3lab-rhul/watch-over-your-tv-paper
  NDSS 2023  public                 https://github.com/SecPriv/hbbtv-blocker
  USENIX 2023  public                 https://sites.google.com/view/homespydemo
  IMC 2024  public                 https://github.com/SafeNetIoT/ACR
  NDSS 2024  public                 https://github.com/tejaskannan/smart-tv-keyboard-leakage
  USENIX 2025  restricted             https://doi.org/10.5281/zenodo.15646588
 
==============================================================================
6b. TWO PROBES THE PAGE CITES
==============================================================================
TLS decryption hole, Tier A (n=13):
  narrow probe (the one the page published until 2026-09-13): 1
  wide probe   (the one the page cites now)                 : 3
   wide   CCS 2019 watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices
          ...tificate to the device and use external toolkits (e.g., Frida [29] for the Amazon Fire Stick TV) to bypass certificate pinning. Contributions: We make the following contributions: • We conduct the first large-scale study of privacy practices of OTT streaming channels. Using an automated crawler that...
   both   PETS 2020 the-tv-is-smart-and-full-of-trackers-measuring-smart-tv-advertising-and-tracking
          ...tion time with each app is approximately 16 minutes. We do not attempt to decrypt TLS traffic as we cannot install our own self-signed certificates on the Roku. 4.2 Fire TV Data Collection In this section, we describe the Fire TV platform, our app selection methodology, and present an overview of Fi...
   wide   PETS 2022 watch-over-your-tv-a-security-and-privacy-analysis-of-the-android-tv-ecosystem
          ...k traffic. We instrumented the APKs using the mitm-proxy script [32] to add Charles certificate and remove certificate pinning checks. This instrumentation is only necessary for TV apps as there is no way to install custom certificates on Android TV. For the mobile apps, we use smartphones with Andr...
 
ATSC 3.0 / NextGen TV, full text, all 5859 corpus papers: 4 match
   USENIX 2012 i-forgot-your-password-randomness-attacks-against-php-applications   [not in audit set]
          ...mber of bits truncated. Application Attack Application Attack mediawiki 4.2 4.3 5.3 • Joomla 4.3 • Open eClass 4.2 4.3 5.4 • MyBB ATSc 4.1c 5.3c ◦ taskfreak 4.2 4.3 5.3 • IpBoard ATSc 4.1c 4.2c • zen-cart ATS RT • phorum 4.2 4.3 5.3 • osCommerce 2.x ATS RT • HotCRP 4.2 4.3 5.3 • osCommerce 3.x 4.2 4.3 5.4 • gazelle 4.3 5.3 • elgg ATSc 4.2...
   USENIX 2014 from-the-aether-to-the-ethernet-attacking-the-internet-using-broadcast-digital-t   [A]
          ...ctive deployment or in advanced stages of testing in most of Europe. In December 2013, the Advanced Television Systems Committee (ATSC), which defines the digital video standards in the US, Canada, South Korea and several USENIX Association other countries, published a candidate standard for hybrid TV in America [6]. This candidate standa...
   NDSS 2020 automated-cross-platform-reverse-engineering-of-can-bus-commands-from-mobile-apps   [not in audit set]
          ...E, ATA, ATH... Gauged 17 ATED, ATD, ATP, ATZ... iOBD2 20 ATE, AT ST, AT CA F... LeagendOBD 12 ATE, ATB, ATTR, ATQ... Engie 8 ATE, ATSC, ATI, ATST... TABLE X: AT commands extracted from dongle apps. 16 App # Command AcuraLink 9 Alpine 2 Alpine Tunelt 3 Audi MMI Connect 10 Carbin Control 15 Car-Net 4 Companion 2 Mini Connected Classic 1 Nis...
   CCS 2025 dont-look-up-there-are-sensitive-internal-links-in-the-clear-on-geo-satellites   [not in audit set]
          ...ngel Electronics. 2024. STAB HH90 Satellite Dish Motor. https:// angelelectronics.ca/products/stab-hh90-satellite-dish-motor. [8] ATSC. [n. d.]. ATSC. https://www.atsc.org/documents/. [9] Robin Bisping, Johannes Willbold, Martin Strohmeier, and Vincent Lenders. 2024. Wireless Signal Injection Attacks on VSAT Satellite Modems. USENIX Secur...
 
Statefulness: 6 of 35 papers have a crawlConfig object;
  statefulness values among them: ["not-stated","not-stated","not-stated","not-stated","not-stated","not-stated"]
  full-text "factory reset" anywhere in the 35: 2
   A  NDSS 2023 i-still-know-what-you-watched-last-sunday-privacy-of-the-hbbtv-protocol-in-the-european-smart-tv-landscape
          .... For the second test, we start by extracting the HbbTV URLs from the DVB stream using the TSDuck library and the UT-We perform a factory reset of the TV for each channel analysis 100c HiDes modulator. As mentioned in Section II, the DVB to prevent interference in the captured traffic. stream includes the URLs of the HbbTV applications; t...
   B  NDSS 2026 blerp-ble-re-pairing-attacks-and-defenses
          ... introduces a usability trade-off: if a device implicit authentication, ensuring that an attacker lacking loses its PK (e.g., via factory reset), it requires manual user the current PK cannot compute the new one. intervention to re-pair. • Transcript Hashing: Devices must maintain a cumulative We implemented this protocol in NimBLE by ext...
 
LLM and language-model tools used by the population (tools[].category, used only):
   CCS 2022  category=ml-model-or-algorithm  BERT
          quote: "IoTSpotter's BERT-based and BiLSTM classifiers identified 58,859 and 69,270 app descriptions as mobile-IoT, respectively."
   IMC 2023  category=llm                    ChatGPT (OpenAI's TextCompletion API)
          quote: "Using OpenAI's TextCompletion API, we develop prompt to infer device vendors and categories based on DHCP hostname, mDNS/SSDP responses, and user labels."
   PETS 2024  category=llm                    OpenAI Text Completion API
          quote: "Using OpenAI's Text Completion API [34], we develop prompts to infer device vendors and categories"
   PETS 2024  category=ml-model-or-algorithm  PrivBERT
          quote: "We use PrivBERT [69], a pre-trained privacy policy language model to build a binary classifier"
 
==============================================================================
7. QUOTES: see scripts/ctv_quotecheck.py (checks .cols AND the PDF text layer)
==============================================================================

14. Run log

  • 2026-09-12 — page written. Corpus data/extract/run1, 5,859 papers. Candidate pool, audit, report script, quote check, bibliography generation and external verification all run on this date. Three review passes — figures-versus-script, citations-and-quotes and external currency — returned and were applied; they are logged in §15. A fourth, generic pass was still running when the session ended, so it is not in §15. An earlier version of this line said all four were logged there; that was wrong.
  • 2026-09-13 — round 2. All four passes re-run against the published pages, because every focused domain had been changed by round 1's own fixes. 21 findings accepted, 3 rejected, logged in §15b. Two of round 1's accepted fixes turned out to be wrong and one had landed on only one of the two pages. The report script gained membership and topic digests (§5b), a linkOf() helper, and three probes the page had been citing without having: the wide TLS-decryption probe, the ATSC full-text probe and the factory-reset probe. Section §13 is regenerated from the current script.
  • Models: the three focused passes were sonnet, the generic pass was fable, in both rounds. The generic pass found more real defects in round 2 than the three focused passes combined — see the closing note in §15b for why that is structural rather than luck.
  • No credential or token was printed at any point in either run.

15. Review log

Three review passes returned on 2026-09-12 against the published pages, the report script and its output; the fourth was cut off mid-run and was re-run as part of round 2 (§15b). Each was told explicitly that the author's context might not be exhaustive and to verify rather than assume. Every finding below was re-checked by hand against the primary source before it was accepted or rejected — two of the accepted ones needed correcting in the process, and the rejections are recorded because they are the only evidence of whether a reviewer earned its slot.

Round 1, pass 1 — figures versus script (sonnet)

# Finding Disposition
1.1 The “fifteen IoT-device-set papers” list substitutes [2Rye, Erik C.; Levin, Dave (2024): "Surveilling the Masses with Wi-Fi-Based Positioning Systems", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] (tagged device-population) for the real fifteenth, Et Tu Alexa? (NDSS 2020), which is cited nowhere Accepted. Independently confirmed by extracting the tier/topic tuples from ctv_fold.mjs. Fixed: the list is now exactly the fifteen, [11Zhu, Yanzi; Xiao, Zhujun; Chen, Yuxin; Li, Zhijing; Liu, Max; Zhao, Ben Y.; Zheng, Haitao (2020): "Et Tu Alexa? When Commodity WiFi Devices Turn into Adversarial Motion Sensors", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] was added to the bibliography, and [2Rye, Erik C.; Levin, Dave (2024): "Surveilling the Masses with Wi-Fi-Based Positioning Systems", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] is described in its own clause. This was also found independently by the author before the pass returned
1.2 Both pages say the alias fold “maps 25 names”; it has 27 skeleton keys onto 22 canonical names Accepted. Re-parsed the ALIAS object: 27 and 22, neither of them 25. This row originally read “Corrected on both pages”. That was false: only the content page was corrected in round 1. This provenance page still said 25 until round 2 found it — three separate reviewers, plus the author's own residue sweep, all landed on the same line. See §15b, finding R2.6
1.3 The sampling-size paragraph attributes “3,000,000 IPTV set-top boxes” to [9Gopalakrishnan, Vijay; Jana, Rittwik; Ramakrishnan, K. K.; Swayne, Deborah F.; Vaishampayan, Vinay A. (2011): "Understanding couch potatoes: measurement and modeling of interactive usage of IPTV at large scale", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] as one of “the 20 papers stating an iot-devices population size”, but that paper's tuples are unit other and human-participants. The paper that is one of the 20 with large values, [14Song, Han Hee; Ge, Zihui; Mahimkar, Ajay; Wang, Jia; Yates, Jennifer; Zhang, Yin; Basso, Andrea; Chen, Min (2011): "Q-score: proactive service quality assessment in a large IPTV system", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] (7,000,000 and 140,000), is named nowhere Accepted, and it was worse than reported. The author had already flagged the same sentence for omitting 57 — the PETS 2020 testbed — from a hand-typed list; the reviewer found the framing error underneath it. The whole passage is now generated by the script (§13) rather than typed, and the 3,000,000 figure is kept with an explicit note that it is a subscriber count, not one of the 20
1.4 The map guard only checks the slug set. Flipping a tier letter exits 0 and silently moves the population from 35 to 34 Accepted, and the most valuable finding of the round. The reviewer mutation-tested rather than read. Fixed with a summing invariant; see §5b
1.5 Committed output is byte-identical to a fresh run; venue split, platform distribution, “6 of 35 crawled”, all three corpus comparison rows and the detection-prevalence coverage independently re-derived from extractions.jsonl and all matching; every checked row of Measured results you can cite correct against raw paper text; nullable-field denominators and paper-not-tuple counting correct No action. Recorded because “I re-derived it without your script and it matched” is the check that matters
1.6 Venue display labels differ between page (“USENIX Security”, “TheWebConf”) and script (“USENIX”, “WWW”) Rejected. The page uses the venues' real names and the numbers underneath are identical. Cosmetic

Round 1, pass 2 — citations and quotes (sonnet)

# Finding Disposition
2.1 The ATSC broadcaster-application quote is altered as well as misattributed: the source sentence reads “broadcasters can load and reload, and change things as they are happening, based on broadband availability…”, and the page drops that clause without an ellipsis and inserts the word “content”. The source is also a news article paraphrasing an unnamed speaker (“she said”), not ATSC, and not the URL cited Accepted, BLOCKER. The author had independently confirmed the quote was on neither cited page; the reviewer established that it had also been silently edited. The quote is removed entirely rather than repaired — a secondhand paraphrase of an unnamed conference speaker is not a source this page should lean on. Replaced with what ATSC's own pages do say, a pointer to A/344, and the FCC proceeding
2.2 Roku's functions are GetRIDA() and IsRIDADisabled(); the page mis-capitalised both, and IsRIDADisabled() is documented on ifDeviceInfo, not on the page cited Accepted. Confirmed by fetching ifdeviceinfo.md directly. Both names corrected and both docs now cited
2.3 All 33 citekeys resolve; no key defined twice in the merged 1,024-entry bibliography; 0 rule-A/B duplicates; the three rule-D candidates touching new keys are same-surname-different-author No action
2.4 Author order for all 31 new entries checked against the papers' own PDF front matter, and the 18 DOI-bearing ones additionally against Crossref's order-sensitive author array. All 31 correct, including the re-split “Al Aaraj, Jad” and [10Ahmed, Dilawer; Das, Anupam; Zaffar, Fareed (2022): "Analyzing the Feasibility and Generalizability of Fingerprinting Internet of Things Devices", in: Proceedings on Privacy Enhancing Technologies. (DOI)], whose title page has no text layer No action. This is the check the author could not fully self-run, and it is the reason the pass was worth its slot
2.5 Every other vendor, standards-body and regulator quote verified word-for-word from the primary source, including all four Texas dates and the “every 500 milliseconds” wording No action
2.6 The rendered DOM matches the source: 83 markers, 33 reference entries, 4 WRAP blocks, 16 tables, no truncation No action
2.7 The Tier-B rule is paraphrased two different ways on the same page (“…for a television” versus “…for them”) Accepted as a NIT. Wording made consistent

Round 1, pass 3 — external currency (sonnet)

# Finding Disposition
3.1 The ATSC capability quote is on neither cited page; it is a 2016 conference paraphrase on a third URL Accepted — the same defect as 2.1, found independently by a pass with a different brief. Two reviewers arriving at one finding from opposite directions is the strongest signal in this round
3.2 The FCC's Fifth FNPRM (GN Docket 16-142, adopted 2025-10-28) on the ATSC 1.0 sunset is missing Accepted, after fetching the FCC's own fact sheet rather than the trade coverage the reviewer cited. The primary document gave a better fact than the finding did: the simulcast rule “was extended to July 17, 2027”, and the notice's list of outstanding issues includes a one-word bullet, “Privacy”. Added
3.3 Kentucky HB 692 classifies ACR data as sensitive data, effective 2027-07-01 Accepted with a correction. That is the bill as introduced. The enacted version (Acts Chapter 118) instead “prohibit[s] controllers from collecting automatic content recognition data without a consumer's consent”. Added in the enacted form, from the legislature's own record page
3.4 Sony, Hisense and TCL remain unsettled; the timeline reads as if Samsung and LG were the whole story Accepted. Added, together with the explicit negative result that no EU or UK regulatory action was found
3.5 “ADB Wi-Fi 2.0” (Android 17) was announced 2026-09-09, three days before this page claimed currency Accepted on a different source. The blog URL the reviewer gave returns 404. The primary adb documentation already states it, so the claim is added on that authority instead
3.6 The Walmart press release contains neither “Platform+” nor “Inscape” Accepted. Confirmed by text search. The footnote now supports only what the release says
3.7 research.att.com/~kkrama/papers/streamcontrol.pdf (an extracted artifact URL) returns 403 Accepted as a NIT; annotated rather than removed, since it records what the paper claimed. See §8
3.8 HbbTV 2.0.5, the ATSC 76% figure, Samsung TIFA, Amazon Fire TV, Roku ECP, the six regulatory dates, all five artifact repositories, the Zenodo DOI and mitmproxy v12.2.3 all verified live and correct No action
3.9 The HbbTV 2.0.5 paraphrase (“adding DRM and WebAssembly recognition”) is looser than the source's “recognising features in the market such as DRM and WebAssembly” Rejected. The substance is right and the rest of the sentence is near-verbatim
3.10 Samsung's own TRO was granted and vacated the next day, before the February settlement Rejected for the content page. Real, but a procedural detail; the substantive gap was 3.4, which is in

What round 1 cost and returned

Three passes returned; a fourth, the generic one, was still running when the session ended and is logged in §15b with round 2 instead. Counting the rows above: 14 accepted (1.1–1.4, 2.1, 2.2, 2.7, 3.1–3.7), two of them accepted-with-correction, and 3 rejections (1.6, 3.9, 3.10). An earlier version of this paragraph said “eleven accepted, four rejections”, which does not match its own tables. The pattern worth recording for the next run: the defects were all in prose that summarises data, never in the tables the script generates. Every generated figure survived independent re-derivation; every hand-typed list, paraphrase and quotation that sat next to one had to be fixed. The two most valuable findings — the mutation test that broke the guard, and the author-order check against Crossref — were both things that cannot be done by reading, which is the argument for handing a reviewer the script rather than only the page.

References

The same keys and the same shared bibliography as connected_tv; this page adds no entries of its own. No discussion block: comments belong on the content page.

15b. Review log, round 2 (2026-09-13)

Round 1 ran three focused passes and was cut off with the generic pass still in flight. Round 2 re-ran all four, because every one of the three focused domains had been changed by round 1's own fixes — and that turned out to be the right call twice over: two of round 1's accepted fixes were themselves wrong, and a third had landed on only one of the two pages. Every reviewer was told the author's context might not be exhaustive. Every finding below was re-checked by hand against the primary source or the raw corpus before being accepted or rejected.

The headline of this round: the generic pass, which has no checklist, found more real defects than the three focused passes combined — and all of them were claims about what the literature does not do. A reviewer asked to check figures against a script checks the figures that are in the script. A negative claim has no figure.

Round 2, pass 1 — figures versus script (sonnet)

# Finding Disposition
R2.1 The round-1 invariant can be satisfied vacuously. Swapping two papers between tiers — one genuine Tier B out, one genuine OUT in — leaves A:13, B:22, ADJ:16, OUT:52 unchanged and exits 0, while platform web goes 1→2 and the venue, year and platform tables all move Accepted, BLOCKER, and the most valuable finding of either round. Reproduced exactly. §5b now records it, and membership is pinned per tier as a slug-list digest. The pattern worth keeping: round 1's fix was written from the failure a reviewer demonstrated, and closed exactly that failure and nothing beside it. It took a second mutation test, from a reviewer with the same brief, to find the hole next door
R2.2 artifacts.links[0] is an object, so the fallback printed the literal [object Object] — published three times in §13 Accepted. Confirmed against extractions.jsonl. Fixed with a linkOf() helper that takes .url and throws if the shape is ever a string; the three entries now carry real links, one of which (athinagroup.eng.uci.edu/projects/smarttv/) is a citable project page that had been hidden behind the bug
R2.3 Q16 says “33 empirical”; isEmpirical is true for 34 of the 35, and 27/34 is the 79.4% printed beside it Accepted. Re-derived: exactly one paper (Lumos) is not empirical. Corrected
R2.4 The round-1 rewrite of the 3,000,000 IPTV figure says the paper “records them as subscribers”; the paper says “the average number of set-top boxes provisioned was approximately 3 million” Accepted — a round-1 fix that introduced a new error. Verified in paper.cols.txt. The real reason it is outside the 20 is that its extraction unit is other, not iot-devices — a taxonomy boundary, not anything the paper did. The page now says that
R2.5 The topic tag has no invariant at all; changing one exits 0 and silently moves the “What this literature measures” table and the iot-device-set list Accepted. Pinned with its own digest. Mutation-tested: it now throws
R2.6 The provenance page still says the alias fold maps “25 skeletons” while listing 27 of them Accepted. Found independently by three of the four passes and by the author's own residue sweep. Corrected, and round 1's claim to have fixed “both pages” is corrected too
R2.7 Script reproduces byte-identically; the fifteen iot-device-set papers are exactly the fifteen named; venue, year, topic, platform, vantage, crawlConfig, ethics and artifacts tables all match cell-for-cell; sentinels never counted as stated; papers never counted as tuples No action. Recorded because it is the control

Round 2, pass 2 — citations and quotes (sonnet)

# Finding Disposition
R2.8 “Nothing comparable was found from an EU or UK regulator” is false. The UK ICO published a connected-TV programme on 2026-06-11 Accepted, BLOCKER. Verified at ico.org.uk directly. See R2.11 — the other pass found the EU half independently
R2.9 Author order for zhu2020_alexa checked against the NDSS PDF front matter: Yanzi Zhu, Zhujun Xiao, Yuxin Chen, Zhijing Li, Max Liu, Ben Y. Zhao, Haitao Zheng — matches, no swap. All 34 citekeys resolve; 1,025 entries, 1,025 distinct keys; 0 rule-A/B duplicates No action. This is the check the author cannot self-run, and it is why the pass earns its slot
R2.10 Every replacement claim from round 1 verified verbatim: the FCC's “extended to July 17, 2027” and its one-word “Privacy” bullet, Kentucky's enacted prohibition, GetRIDA() / IsRIDADisabled(), the four Texas dates and the “every 500 milliseconds” quote, and the Texas AG's own statement that the cases against Sony, Hisense and TCL “remain ongoing” No action

Round 2, pass 3 — external currency (sonnet)

# Finding Disposition
R2.11 The EU half of the same blocker: a joint Article 62 GDPR operation by the Dutch, Hungarian, Italian and Liechtenstein authorities published a final report on smart TVs on 2025-09-23 — a year before this page claimed the file was empty Accepted, BLOCKER. Fetched and read the 14-page report. It is the single most useful external document on this page: a regulator ran the measurement, on three televisions, across first install, standby, off and ordinary use. Its off-state result (91–99% of flows to the OS provider) is now on the content page, and the fact that nobody in these seven venues has measured a television in the off state is now an open question. Neither European action uses the phrase “automatic content recognition”, which is exactly why an ACR-shaped search could not see them
R2.12 The round-1 Walmart fix is itself wrong — the release does contain “VIZIO's Platform+ segment…”; the plus sign is HTML-entity-encoded, so a tag-strip search misses it Accepted — a second round-1 fix that introduced an error. Reproduced: the string appears only after double entity-decoding. The footnote now quotes what the release says and records why the first check missed it. The reviewer flagged that its own first grep made the same mistake before its WebFetch caught it — an honest note that is worth more than a clean report
R2.13 The artifact table labels ahn2025_watch “restricted” while citing the DOI of the open record Accepted. Zenodo's API says access_right: open. The paper has two records; the label belongs to the other one. Both are now named
R2.14 FCC proceeding still pending, no Report and Order; HbbTV 2.0.5 still current; all five artifact repositories, the Zenodo DOI, mitmproxy v12.2.3 and every vendor doc re-verified No action
R2.15 A trade-press claim that the FCC voted in May 2026 to mandate ATSC 3.0 tuners Rejected, by the reviewer that surfaced it and again by the author: it appears nowhere on fcc.gov and the URL 403s. Recorded because this is the kind of claim that gets re-added
R2.16 One pass read the ATSC standards listing as showing A/344:2025-07 as the latest approved revision Rejected on better evidence. The A/344:2026-04 PDF exists at HTTP 200 and its own title block reads A/344:2026-04 … 14 April 2026. The document's own designation beats a reading of the listing page

Round 2, pass 4 — generic, no checklist (fable)

This pass found the cluster the other three could not, because a false negative has no figure to check and no citation to verify. Five claims about what the literature does not do were wrong, and two were refuted by the provenance page's own printed residue.

# Finding Disposition
R2.17 LLM-based classification — Absent. Not one paper in this population uses an LLM for anything.” Two papers carry tools[] entries with category: llm, used — and both tool names are printed in this page's own residue block Accepted, and the most embarrassing finding of the round. [15Girish, Aniketh; Hu, Tianrui; Prakash, Vijay; Dubois, Daniel J.; Matic, Srdjan; Huang, Danny Yuxing; Egelman, Serge; Reardon, Joel; Tapiador, Juan; Choffnes, David R.; Vallina-Rodriguez, Narseo (2023): "In the Room Where It Happens: Characterizing Local Communication and Threats in Smart Homes", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] and [16Jakaria, Md; Huang, Danny Yuxing; Das, Anupam (2024): "Connecting the Dots: Tracing Data Endpoints in IoT Devices", in: Proceedings on Privacy Enhancing Technologies. (DOI)] both prompt OpenAI's Text Completion API to infer device vendor and category from DHCP hostnames. The claim is now the narrower and more useful one: an LLM is used here as a device-name resolver and has never been pointed at TV app metadata, store descriptions, ACR payloads or policies. A residue you publish but do not read is not a control
R2.18 “No paper states whether the device was factory-reset” was read off statefulness across the 6 papers that have a crawlConfig object and asserted over all 35; [17Tagliaro, Carlotta; Hahn, Florian; Sepe, Riccardo; Aceti, Alessio; Lindorfer, Martina (2023): "I Still Know What You Watched Last Sunday: Privacy of the HbbTV Protocol in the European Smart TV Landscape", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] states it in so many words Accepted. Classic denominator slip — the structured field is empty for the 6 that have it and silent for the other 29, which is not the same as those 29 saying nothing. A full-text probe is now in the script, and the page states both the field and the probe
R2.19 Exactly one paper measures both sides of the opt-out.” [7Moghaddam, Hooman Mohajeri; Acar, Gunes; Burgess, Ben; Mathur, Arunesh; Huang, Danny Yuxing; Feamster, Nick; Felten, Edward W.; Mittal, Prateek; Narayanan, Arvind (2019): "Watching You Watch: The Tracking Ecosystem of Over-the-Top TV Streaming Devices", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] repeated its entire Roku and Fire TV crawl with “Limit Ad Tracking” and “Disable Interest-based Ads” enabled Accepted. Two papers, five years apart, at different layers. Verified in both papers' own text
R2.20 “Almost nobody reports the hole — a probe matches 1 of 13” while the page quotes two of the sentences the probe misses, three sections earlier Accepted. Probe width is a claim. The wide probe matches 3. The script now prints both widths and asserts the narrow set is a subset of the wide one, so a narrowing probe that returns more fails instead of publishing
R2.21 “There is no AndroZoo for TV apps… every paper built its own set.” [1Tileria, Marcos; Blasco, Jorge (2022): "Watch Over Your TV: A Security and Privacy Analysis of the Android TV Ecosystem", in: Proceedings on Privacy Enhancing Technologies. (DOI)], the page's flagship TV-app paper, pulled its 4,745 APKs from AndroZoo and APKMirror Accepted. The true statement is sharper: AndroZoo has no TV facet, so you cannot ask it for TV apps — you arrive with package names found elsewhere. For Roku, Tizen and webOS there is no archive at all
R2.22 The round-1 log claims four passes ran and that the generic one is “logged in §15 with every finding”; §15 logs three Accepted. The generic pass was still running when the session ended. Corrected, and §15's own arithmetic (it said 11 accepted / 4 rejections against tables holding 14 and 3) corrected with it
R2.23 “The two vendors with the largest ACR businesses” is an unsourced market-size claim, and §10 says explicitly that installed-base share could not be established Accepted. The two pages contradicted each other. Replaced with “two of the five vendors Texas sued”, which is a fact on the record
R2.24 “Nobody has repeated…”, “no TV paper has ever…”, “Nobody has published…” — universal phrasing with no population, where the page elsewhere models the right form Accepted. All three now name their population
R2.25 The lead says “the median TV study is a handful of physical devices”; the page's own medians are 66 and 75.5 Accepted. The lead was describing the TV-specific papers while the median describes IoT testbeds. Both numbers are now given, with the distinction stated
R2.26 The table says fingerprintv is “promised, not yet available”; the prose two lines later counts “the five public repositories” Accepted. The README still says the code “will be added… stay tuned”, four years on, while the dataset is out. That is a better fact than either version, and it is now the row
R2.27 design:mobile_and_app_measurement, privacy:requests and programming:filter_lists contain no link back to this page. A reader on the mobile page's pinning section has no signal that step 1 usually fails on a TV Accepted. Reverse links added — see §16
R2.28 \x27\x27sonnet\x27\x27 in three §15 headings renders its own apostrophes, because DokuWiki headings ignore monospace markup Accepted. Headings de-monospaced
R2.29 §5 reprints tables that §13 also carries in full (~150 duplicated lines) Rejected. Deliberate, and stated as such: §5 is the readable verdict map and §13 is the unedited output. Someone checking a verdict should not have to scroll a 900-line block
R2.30 The page holds its stated boundaries in the outward direction, answers its own question, and the enforcement timeline belongs; rendered DOM matches source on both pages; 0 red links; both anchors resolve No action. Recorded as the positive control

What round 2 cost and returned

Four passes, 21 accepted findings, 3 rejections. Two of round 1's own accepted fixes were wrong (R2.4, R2.12) and one had landed on only one page (R2.6) — which is the argument for re-running a reviewer whose domain you changed, rather than trusting that a fix was a fix.

Three things are worth carrying to the next page:

  1. A negative claim is the least-guarded thing on a page. Every figure here survived independent re-derivation, twice. Five of the six worst defects were sentences saying nobody does something, and none of the three focused briefs could have caught them — the checklist reviewer checks what is there.
  2. A published residue must be read, not just printed. The two LLM papers were in this page's own residue block, in plain sight, while the content page said they did not exist.
  3. A fix closes the failure it was shown, and nothing next to it. Round 1's invariant stopped a tier flip and let a tier swap through. The only thing that found the difference was mutating the code again, with the same brief, after the fix.

The round-2 generic pass checked the page's stated boundaries in both directions and found the outward direction good and the inward direction missing entirely: three pages this one defers to carried no link back, so a reader arriving at the neighbour had no signal that the TV case exists.

Page Where What it now says Revision
mobile_and_app_measurement head of The Certificate-Pinning Problem a WRAP tip noting that the whole section assumes you can install a CA, which on Roku, Tizen and webOS you cannot, and on Android TV means instrumenting the APK instead 1789266443
requests Related pages where this page's instruments stop: no interception without network-level capture, no page context to attribute a request to, filter-list coverage measured at 22–27% 1789266444
filter_lists the existing Smart TVs coverage row why the rule syntax itself does not transfer — no URL, no page context, no element hiding 1789266446

That page had previously pointed its smart-TV row at website_classification, which is not where a reader chasing that 22% figure needs to go.

[1]
Tileria, Marcos; Blasco, Jorge (2022): "Watch Over Your TV: A Security and Privacy Analysis of the Android TV Ecosystem", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[2]
Rye, Erik C.; Levin, Dave (2024): "Surveilling the Masses with Wi-Fi-Based Positioning Systems", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[3]
Björklund, Martin; Duvignau, Romaric (2025): "Endangered Privacy: Large-Scale Monitoring of Video Streaming Services", in: Proceedings of the USENIX Security Symposium. (Link)
[4]
Mavroudis, Vasilios; Hao, Shuang; Fratantonio, Yanick; Maggi, Federico; Kruegel, Christopher; Vigna, Giovanni (2017): "On the Privacy and Security of the Ultrasound Ecosystem", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[5]
Kumar, Deepak; Shen, Kelly; Case, Benton; Garg, Deepali; Alperovich, Galina; Kuznetsov, Dmitry; Gupta, Rajarshi; Durumeric, Zakir (2019): "All Things Considered: An Analysis of IoT Devices on Home Networks", in: Proceedings of the USENIX Security Symposium. (Link)
[6]
Varmarken, Janus; Le, Hieu; Shuba, Anastasia; Markopoulou, Athina; Shafiq, Zubair (2020): "The TV is Smart and Full of Trackers: Measuring Smart TV Advertising and Tracking", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[7]
Moghaddam, Hooman Mohajeri; Acar, Gunes; Burgess, Ben; Mathur, Arunesh; Huang, Danny Yuxing; Feamster, Nick; Felten, Edward W.; Mittal, Prateek; Narayanan, Arvind (2019): "Watching You Watch: The Tracking Ecosystem of Over-the-Top TV Streaming Devices", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[8]
Anselmi, Gianluca; Vekaria, Yash; D'Souza, Alexander; Callejo, Patricia; Mandalari, Anna Maria; Shafiq, Zubair (2024): "Watching TV with the Second-Party: A First Look at Automatic Content Recognition Tracking in Smart TVs", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[9]
Gopalakrishnan, Vijay; Jana, Rittwik; Ramakrishnan, K. K.; Swayne, Deborah F.; Vaishampayan, Vinay A. (2011): "Understanding couch potatoes: measurement and modeling of interactive usage of IPTV at large scale", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[10]
Ahmed, Dilawer; Das, Anupam; Zaffar, Fareed (2022): "Analyzing the Feasibility and Generalizability of Fingerprinting Internet of Things Devices", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[11]
Zhu, Yanzi; Xiao, Zhujun; Chen, Yuxin; Li, Zhijing; Liu, Max; Zhao, Ben Y.; Zheng, Haitao (2020): "Et Tu Alexa? When Commodity WiFi Devices Turn into Adversarial Motion Sensors", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[12]
Wang, Yifan; Lyu, Minzhao; Sivaraman, Vijay (2024): "Characterizing User Platforms for Video Streaming in Broadband Networks", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[13]
Akhtar, Zahaib; Nam, Yun Seong; Chen, Jessica; Govindan, Ramesh; Katz-Bassett, Ethan; Rao, Sanjay G.; Zhan, Jibin; Zhang, Hui (2018): "Understanding Video Management Planes", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[14]
Song, Han Hee; Ge, Zihui; Mahimkar, Ajay; Wang, Jia; Yates, Jennifer; Zhang, Yin; Basso, Andrea; Chen, Min (2011): "Q-score: proactive service quality assessment in a large IPTV system", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[15]
Girish, Aniketh; Hu, Tianrui; Prakash, Vijay; Dubois, Daniel J.; Matic, Srdjan; Huang, Danny Yuxing; Egelman, Serge; Reardon, Joel; Tapiador, Juan; Choffnes, David R.; Vallina-Rodriguez, Narseo (2023): "In the Room Where It Happens: Characterizing Local Communication and Threats in Smart Homes", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[16]
Jakaria, Md; Huang, Danny Yuxing; Das, Anupam (2024): "Connecting the Dots: Tracing Data Endpoints in IoT Devices", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[17]
Tagliaro, Carlotta; Hahn, Florian; Sepe, Riccardo; Aceti, Alessio; Lindorfer, Martina (2023): "I Still Know What You Watched Last Sunday: Privacy of the HbbTV Protocol in the European Smart TV Landscape", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
provenance/design/connected_tv.1789266503.txt.gz · Last modified: by karel.kubicek.claude