Table of Contents
Provenance: Connected TV
Working log for connected_tv. Every figure on that page is produced by one script, printed here with its denominator, and every quote it uses is machine-checked against both renderings of its source paper. Corpus-level caveats — the seven venues, the funnel, the provisional 2025–2026 years — are on corpus and are not restated.
Run date 2026-09-12. Corpus at the time: 5,859 papers with extracted full text, data/extract/run1/extractions.jsonl, seven venues, 2010–2026.
1. Why this page exists, and what it is not
roadmap queued design:connected_tv on 2026-09-07 with a 16-paper candidate set from scripts/gap_probe_roadmap.mjs (family ctv_streaming), and roadmap §5 recorded a condition on it: “Connected TV will need its population derived from platform fields rather than the probe, because the probe's web-platform column is exactly the wrong filter for it.” That condition was honoured, and it turned out to understate the problem in one direction and overstate it in another.
- The title probe's precision is 62.5% — 10 of its 16 are in the final population.
- Its recall is much worse: it misses 25 of the 35. The largest TV app analysis in the corpus, [1Tileria, Marcos; Blasco, Jorge (2022): "Watch Over Your TV: A Security and Privacy Analysis of the Android TV Ecosystem", in: Proceedings on Privacy Enhancing Technologies. (DOI)] (4,745 Android TV APKs), is not in the 16 at all; neither is [2Rye, Erik C.; Levin, Dave (2024): "Surveilling the Masses with Wi-Fi-Based Positioning Systems", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)], in which Roku devices are two of the five most common vendor prefixes in a 490-million-row dataset.
- “Derive from platform fields” cannot mean “filter on a platform field”. There is no
tvvalue in theplatformsenum, and the closest one,iot, holds 436 corpus papers of which the overwhelming majority are smart speakers, cameras, plugs and firmware. Platform fields are used here as evidence about the population, not as the selector: 26 of the 35 carryiot, 1 carriesweb, and only 6 are inside the corpus-widecrawledpopulation. That is the page's thesis, measured.
The roadmap row also named one paper that is not in the population. It listed Endangered Privacy (USENIX 2025) [3Björklund, Martin; Duvignau, Romaric (2025): "Endangered Privacy: Large-Scale Monitoring of Video Streaming Services", in: Proceedings of the USENIX Security Symposium. (Link)] among the six-paper “spine”. Reading it, it identifies videos from encrypted MPEG-DASH traffic against Amazon Prime Video, Max and SVT Play; its populations are 242,364 video manifests, 900 sampled titles and the VNAT flow dataset. No television is measured anywhere in it. It is recorded as ADJ with that reason. The roadmap row is left as written, because it is a record of what was believed on 2026-09-07; this is the correction.
2. The inclusion rule, fixed before any figure
Written into scripts/ctv_fold.mjs before the first count was taken:
- A television-class endpoint is a smart TV set, a TV operating system (Android TV / Google TV, tvOS, Tizen, webOS, Roku OS, Fire OS), a streaming stick, box or set-top box, an app running on one, or the broadcast path (HbbTV / DVB) into one.
- A — the paper's central object of measurement is a television-class endpoint.
- B — television-class devices are part of a broader measured population and the paper reports at least one result broken out for them.
- ADJ — adjacent. Video streaming measured off a TV (browser DRM, piracy websites, encrypted-traffic video fingerprinting), or a TV used as apparatus rather than measured.
- OUT — the TV name is a passing reference, a survey answer option, a related-work sentence, or a homonym.
Two boundary decisions are worth naming because a reasonable person would draw them differently.
- A TV used as apparatus is OUT, even when it has its own results row. [4Mavroudis, Vasilios; Hao, Shuang; Fratantonio, Yanick; Maggi, Federico; Kruegel, Christopher; Vigna, Giovanni (2017): "On the Privacy and Security of the Ultrasound Ecosystem", in: Proceedings on Privacy Enhancing Technologies. (DOI)]'s ultrasound work and Void (USENIX 2020, a Samsung Smart TV used as a replay loudspeaker with its own 24,282-sample row) both fail the rule for the same reason: the television generates a stimulus, it is not measured. Under a purely mechanical “named device with a results row” rule, Void would be Tier B.
- Survey and interview studies are OUT even when every participant owns a TV. This removes about twenty smart-home qualitative papers. They are real research about televisions; they are not measurements of one, and including them would have made the “15 of 35 are IoT device sets” finding meaningless.
3. Every query, with its population
| # | Question | Population | Answer |
|---|---|---|---|
| Q1 | How many corpus papers have full text to probe? | all 5,859 | 5,855 scanned; 4 have no paper.cols.txt |
| Q2 | What does the roadmap's title+summary probe return? | all 5,859 | 16; 1 carries web |
| Q3 | Wide TV-vocabulary gate (gate 1) | all 5,855 scanned | 142 |
| Q4 | Audit set (gate 2) | gate 1 | 103 |
| Q5 | Population after hand audit | the 103 | 35 (A 13, B 22); ADJ 16, OUT 52; precision 34.0% |
| Q6 | Population papers the roadmap probe misses | the 35 | 25 |
| Q7 | Population by venue | the 35 | IMC 11, USENIX 9, NDSS 6, PETS 6, CCS 2, IEEE S&P 1, WWW 0 |
| Q8 | platforms[] distribution | the 35 | iot 26, other-online-service 12, mobile 6, offline 3, web 1 |
| Q9 | Inside the corpus crawled population | the 35 | 6 |
| Q10 | population[].unit | the 35 | iot-devices 20, other 17, mobile-apps 5 |
| Q11 | Named instruments, alias-folded, used only | the 35 | Wireshark 10, tcpdump 9, mitmproxy 4, Frida 3, adb 3 |
| Q12 | Interception-evidence probes | 13 Tier A / 35 A+B | router/AP 11/30, mitm 10/22, DNS 5/12, HDMI 6/9, remote-control 12/16, undecryptable reported, narrow probe 1, wide probe 3 (see §6b) |
| Q13 | crawlConfig fields stated | 6 papers with a crawlConfig | interactionDepth 5, consentAction 0, statefulness 0, browsers 0 |
| Q14 | ethics.reviewOutcome stated | 34 with an ethics object | 14 (41.2%) vs corpus-empirical 1,728 of 4,472 (38.6%) |
| Q15 | artifacts.availability stated | 34 with an artifacts object | 27 (79.4%) vs 2,890 of 4,854 (59.5%) |
| Q16 | temporal.spanStart stated | 34 empirical | 27 (79.4%) vs 2,882 of 5,118 (56.3%) |
| Q17 | detection[].prevalence coverage | 70 Tier A tuples | 68 carry a prevalence (97.1%); corpus-wide 26,316 of 27,241 (96.6%) |
| Q18 | population[].listVersion stated | 105 population tuples in the 35 | 42 |
| Q19 | Vantage location stated | 34 with a vantage tuple | 18 (52.9%) |
Denominators that are easy to get wrong here, spelled out. ethics and artifacts are nullable in the schema — one of the 35 papers has neither object — so Q14–Q16 divide by the papers that carry the object, never by 35 and never by the corpus. crawlConfig is null for 29 of the 35, so Q13 divides by 6, and its zeros are read on the page rather than reported bare. No TV figure on either page divides by 5,859. The one place that denominator appears is the corpus-share column of the platform table on the content page, which exists precisely to say what share of the whole corpus each platform value has — that column divides by 5,859 by design, and is labelled as doing so. An earlier version of this sentence said “nothing on either page divides by 5,859”, which was wrong.
4. The candidate pool, and why it has two gates
Full text is whitespace-collapsed (soft hyphens stripped, hyphen-newline joined, runs of whitespace reduced to one space) before any regex runs. Without that a phrase broken across a line silently fails to match.
Nine probes over every paper.cols.txt:
smarttv /\bsmart[-\s]?TVs?\b/i ctv /\bconnected[-\s]TVs?\b|\bCTV\b/i ott /\bover[-\s]the[-\s]top\b|\bOTT\b/i hbbtv /\bHbbTV\b|\bhybrid broadcast broadband\b/i acr /\bautomatic content recognition\b|\bACR\b/i platformdev /\bRoku\b|\bFire ?TV\b|\bApple ?TV\b|\bChromecast\b|\bAndroid ?TV\b|\bGoogle ?TV\b|\btvOS\b|\bWebOS\b|\bTizen\b|\bset[-\s]?top box(es)?\b/i streamsvc /\bNetflix\b|\bHulu\b|\bDisney\+|\bAmazon Prime Video\b|\bYouTube ?TV\b|\bTwitch\b/i tvapp /\bTV app(s|lication)?\b|\btelevision app(s)?\b/i iptv /\bIPTV\b|\binternet protocol television\b/i
Plus a device-name probe used to find televisions inside broader IoT device sets, which is where 15 of the 35 came from:
DEV /\b(Roku|Fire ?TV|Apple ?TV|Chromecast|Android ?TV|Google ?TV|tvOS|WebOS|Tizen|Vizio|Hisense|Bravia|Nvidia Shield|Samsung(?: Smart)? TV|LG(?: Smart)? TV|TCL|smart[- ]?TVs?|set[- ]?top box(?:es)?)\b/gi
Gate 1 (142 papers) — core >= 2 || acr >= 2 || iptv >= 2 || hbbtv >= 1 || tvapp >= 1 || titleHit, where core sums smarttv + ctv + hbbtv + platformdev + tvapp.
Gate 2, the audit set (103 papers) — gate 1 narrowed by devN >= 4 || brands >= 3 || titleHit || hbbtv > 0 || acr >= 2 || iptv >= 2.
The 39 papers dropped between the gates all have a single-brand, low-count mention — a Tizen in a list of embedded platforms, one Apple TV in an enumeration of Apple hardware. That is a judgement, not a measurement: it was not hand-audited, and if a television study exists that names exactly one TV-class device three times or fewer, this page does not contain it. A wider audit would be the cheapest improvement to make here.
Homonyms found, and what they cost. ACR is the American College of Radiology, an authentication context reference, and an arbitrary abbreviation in a privacy-policy paper; IPTV appears in leaked-credential corpora and in X spam campaigns; Tizen is a smartwatch platform and an open-source project under fuzz testing; WebOS matches both the LG TV OS and unrelated prose. 52 of 103 audit-set papers are OUT, and the biggest single class is the smart-home survey, where “smart TV” is an answer option. The precision of the audit set is 34.0% — for comparison, the ad-archives row on roadmap recorded 14.5% and the authentication row 24.5%.
5. Verdict map — the 68 papers NOT in the population
Published in full, because the rejections are the only record of where the line was drawn.
0c. ADJACENT AND OUT — the audit trail for what was NOT counted ============================================================================== V Year Venue Title Reason --- ---- ------- -------------------------------------------------------------- ----------------------------------------------------------------------------------------------------------------------------------------------- ADJ 2011 IMC Measurement and analysis of a large scale commercial mobile in "TV" delivered to mobile handsets, not to a television. ADJ 2012 IMC Watching videos from everywhere: a study of the PPTV mobile Vo Mobile VoD; no television endpoint. ADJ 2013 IMC Analyzing the potential benefits of CDN augmentation strategie CDN augmentation for video workloads; no TV endpoint. ADJ 2013 IMC Peer-assisted content distribution in Akamai netsession Peer-assisted CDN; set-top-box mention is background. ADJ 2016 IMC Performance Characterization of a Commercial Video Streaming S Streaming service performance from browser/CDN vantage; no TV-specific result. ADJ 2016 IMC Anatomy of a Personalized Livestreaming System Livestreaming (Periscope) system measurement; no TV endpoint. ADJ 2017 PETS On the Privacy and Security of the Ultrasound Ecosystem Ultrasonic cross-device tracking (uXDT): beacons emitted by TV adverts and picked up by phone SDKs. The TV is the emitter, never measured. ADJ 2019 WWW Exploiting Diversity in Android TLS Implementations for Mobile Android app traffic classification; TLS-fingerprint method later reused on TV apps. ADJ 2020 USENIX Void: A fast and light voice liveness detection system A Samsung Smart TV is used as a replay LOUDSPEAKER; the TV is apparatus, not the measured object. ADJ 2022 NDSS A Lightweight IoT Cryptojacking Detection Mechanism in Heterog Authors implement their own cryptojacking PoC on an LG webOS TV to test a detector; no deployed-TV population. ADJ 2022 USENIX OVRseen: Auditing Network Traffic and Privacy Policies in Ocul VR headsets; smart TVs used as the comparison ecosystem. Same lab, same pipeline shape. ADJ 2023 IMC Tracking, Profiling, and Ad Targeting in the Alexa Echo Smart Smart speaker study; TVs cited as the comparable prior ecosystem, not measured. The closest methodological sibling. ADJ 2023 PETS Your DRM Can Watch You Too: Exploring the Privacy Implications Widevine EME in browsers and Android; TVs named as another Widevine host, not measured. ADJ 2024 IMC Cost-Saving Streaming: Unlocking the Potential of Alternative Edge-node economics for streaming delivery; no TV endpoint measured. ADJ 2025 PETS Unmasking the Shadows: A Cross-Country Study of Online Trackin Illegal movie streaming WEBSITES crawled with a browser; a web-tracking study, not a TV study. ADJ 2025 USENIX Endangered Privacy: Large-Scale Monitoring of Video Streaming Video identification from encrypted MPEG-DASH traffic. The roadmap listed it as CTV spine; it measures the SERVICE and its traffic, never a TV. OUT 2010 IMC What happened in my network: mining network events from router IPTV named as the service carried; router syslogs are the object. OUT 2011 CCS On the vulnerability of FPGA bitstream encryption against powe Set-top box named as an FPGA application domain. OUT 2011 IMC Broadcast yourself: understanding YouTube uploaders IPTV appears once in related work. OUT 2014 CCS (Nothing else) MATor(s): Monitoring the Anonymity of Tor's Pat "ACR" homonym. OUT 2015 USENIX A Placement Vulnerability Study in Multi-Tenant Public Clouds Title probe matched "streaming"; cloud VM placement. OUT 2015 USENIX Rocking Drones with Intentional Sound Noise on Gyroscopic Sens Passing mention. OUT 2016 CCS SandScout: Automatic Detection of Flaws in iOS Sandbox Profile Apple TV named as a device that runs iOS/tvOS; iOS sandbox is the object. OUT 2016 IMC Entropy/IP: Uncovering Structure in IPv6 Addresses "ACR" homonym. OUT 2016 USENIX You Are Who You Know and How You Behave: Attribute Inference A IPTV homonym. OUT 2017 CCS POSTER: Watch Out Your Smart Watch When Paired Tizen here is the smartwatch platform, not the TV one. OUT 2017 PETS Why can’t users choose their identity providers on the web? "ACR" homonym. OUT 2017 USENIX Same-Origin Policy: Evaluation in Modern Browsers Passing mention of TV browsers. OUT 2017 WWW FLOCK: Combating Astroturfing on Livestreaming Platforms Title probe matched "streaming platform"; astroturfing detection on Twitch-like sites. OUT 2018 CCS Medical Devices are at Risk: Information Security on Diagnosti "ACR" = American College of Radiology. OUT 2019 IEEE-SP Drones' Cryptanalysis - Smashing Cryptography with a Flicker IPTV/ACR homonyms. OUT 2019 NDSS cleaning-up-the-internet-of-evil-things-real-world-evidence-on One infected set-top box in a Mirai remediation table; no TV finding. OUT 2019 NDSS latex-gloves-protecting-browser-extensions-from-probing-and-re The Chromecast browser EXTENSION, not the device. OUT 2019 USENIX A Billion Open Interfaces for Eve and Mallory: MitM, DoS, and tvOS listed among Apple OSes; AWDL is the object. OUT 2019 WWW Snapshot-based Loading Acceleration of Web Apps with Nondeterm Tizen/webOS named as embedded web-app platforms; benchmark is web apps. OUT 2020 CCS iDEA: Static Analysis on the Security of Apple Kernel Drivers tvOS is one of four Apple OSes scanned; no TV-specific result. OUT 2020 PETS Smart Devices in Airbnbs: Considering Privacy and Security for Survey; smart TV is a scenario option. OUT 2022 IMC Deep dive into the IoT backend ecosystem Backend infrastructure; TV mentions are motivation and a citation to FingerprinTV. OUT 2022 PETS A Multi-Region Investigation of the Perceptions and Use of Sma Survey; smart TV is a related-work citation and an ownership option. OUT 2022 PETS Exploring the Privacy Concerns of Bystanders in Smart Homes fr Survey; smart TV is an example in a prompt. OUT 2023 CCS IoTFlow: Inferring IoT Device Behavior at Scale through Static Companion-app analysis; no TV breakout. OUT 2023 IEEE-SP Characterizing Everyday Misuse of Smart Home Devices Survey of 483 people; smart TV is an ownership option, not a measured device. OUT 2023 IEEE-SP WebSpec: Towards Machine-Checked Analysis of Browser Security "ACR" homonym. OUT 2023 IEEE-SP UTopia: Automatic Generation of Fuzz Driver using Unit Tests Tizen as an open-source project under test; no TV device. OUT 2023 PETS No Privacy Among Spies: Assessing the Functionality and Insecu Android stalkerware; "ACR" homonym. OUT 2023 USENIX Examining Consumer Reviews to Understand Security and Privacy Review-text analysis; set-top box is a Mirai product category, no TV measurement. OUT 2023 USENIX Examining Power Dynamics and User Privacy in Smart Technology Interview study; TVs are participant device inventories. OUT 2023 USENIX Internet Service Providers' and Individuals' Attitudes, Barrie Interview and survey; TV is a device-ownership row. OUT 2023 USENIX "It's the Equivalent of Feeling Like You're in Jail”: Lessons Interview study on IPV; TV is a reported abuse vector, not measured. OUT 2023 USENIX Measuring Up to (Reasonable) Consumer Expectations: Providing Vignette survey; Vizio appears only in a news citation. OUT 2023 USENIX Abuse Vectors: A Framework for Conceptualizing IoT-Enabled Int Qualitative framework; TV is an example abuse vector. OUT 2023 USENIX Exploring Tenants' Preferences of Privacy Negotiation in Airbn Vignette survey; smart TV is a device-type option. OUT 2023 WWW SISSI: An Architecture for Semantic Interoperable Self-Soverei "ACR" homonym (authentication context reference). OUT 2024 IEEE-SP SoK: Technical Implementation and Human Impact of Internet Pri SoK; TV work cited, not measured. OUT 2024 PETS A Bilingual Longitudinal Analysis of Privacy Policies Measurin ACR homonym: "ACR" is not automatic content recognition here. OUT 2024 PETS Contextualizing Interpersonal Data Sharing in Smart Homes Vignette survey; "viewing history from your smart TV" is a question stem. OUT 2024 PETS "My Best Friend's Husband Sees and Knows Everything": A Cross- Survey; smart TV is a free-text mention count. OUT 2024 USENIX Co-Designing a Mobile App for Bystander Privacy Protection in Interview study; TV names are participant-reported device inventories. OUT 2025 IEEE-SP Analyzing the iOS Local Network Permission from a Technical an Chromecast is one of four IoT devices used to trigger the permission; no TV result. OUT 2025 IEEE-SP Hey, Your Secrets Leaked! Detecting and Characterizing Secret IPTV homonym in leaked-credential data. OUT 2025 NDSS Non-intrusive and Unconstrained Keystroke Inference in VR Plat VR; smart TV appears only as a citation to HomeSpy. OUT 2025 PETS Help Me Help You: Privacy Considerations for Third Party IoT D Vignette survey; TVs appear in a device-category prompt. OUT 2025 PETS Who Cares? Contextual Privacy Judgments from Owner and Bystand Survey; smart TV is a device-category option. OUT 2025 USENIX Regulating Smart Device Support Periods: User Expectations and Survey; Smart TV is a self-reported ownership category. OUT 2026 IEEE-SP Privacy Perspectives and Practices of Chinese Smart Home Produ Interview study; smart TV is a company product-line row. OUT 2026 NDSS TBTrackerX: Fantastic Trigger Bots and Where to Find Malicious IPTV spam homonym. OUT 2026 PETS Dead Domains, Living Data: A Privacy Risk Analysis of Domain L Android apps; one expired-domain example happens to also ship on Roku. OUT 2026 USENIX PANGOLIN: Fuzzing Multilingual IoT Firmware with LLM-Driven Co "SmartTVs" is a citation to the 2021 fuzzing paper, used as a baseline name. ==============================================================================
And the 35 that are in:
0b. THE POPULATION, PAPER BY PAPER ============================================================================== Tier Year Venue Topic Title ---- ---- ------- -------------------- ---------------------------------------------------------------------------------------------------------------- A 2011 IMC delivery-performance Understanding couch potatoes: measurement and modeling of interactive usage of IPTV at large scale A 2011 IMC delivery-performance Q-score: proactive service quality assessment in a large IPTV system A 2014 USENIX broadcast From the Aether to the Ethernet—Attacking the Internet using Broadcast Digital Television A 2019 CCS tracking Watching You Watch: The Tracking Ecosystem of Over-the-Top TV Streaming Devices A 2020 PETS tracking The TV is Smart and Full of Trackers: Measuring Smart TV Advertising and Tracking A 2021 USENIX vulnerability Android SmartTVs Vulnerability Discovery via Log-Guided Fuzzing A 2022 PETS tracking FingerprinTV: Fingerprinting Smart TV Apps A 2022 PETS app-analysis Watch Over Your TV: A Security and Privacy Analysis of the Android TV Ecosystem A 2023 NDSS broadcast I Still Know What You Watched Last Sunday: Privacy of the HbbTV Protocol in the European Smart TV Landscape A 2023 USENIX side-channel HOMESPY: The Invisible Sniffer of Infrared Remote Control of Smart TVs A 2024 IMC acr Watching TV with the Second-Party: A First Look at Automatic Content Recognition Tracking in Smart TVs A 2024 NDSS side-channel Acoustic Keystroke Leakage on Smart Televisions A 2025 USENIX piracy Watch Out Your TV Box: Reversing and Blocking a P2P-based Illegal Streaming Ecosystem B 2018 IMC delivery-performance Understanding Video Management Planes B 2019 IMC iot-device-set Information Exposure From Consumer IoT Devices: A Multidimensional, Network-Informed Measurement Approach B 2019 USENIX iot-device-set All Things Considered: An Analysis of IoT Devices on Home Networks B 2020 IMC iot-device-set A Haystack Full of Needles: Scalable Detection of IoT Devices in the Wild B 2020 NDSS iot-device-set Packet-Level Signatures for Smart Home Devices B 2020 NDSS iot-device-set Et Tu Alexa? When Commodity WiFi Devices Turn into Adversarial Motion Sensors B 2020 USENIX iot-device-set You Are What You Broadcast: Identification of Mobile and IoT Devices from (Public) WiFi B 2021 IMC iot-device-set IoTLS: understanding TLS usage in consumer IoT devices B 2021 PETS iot-device-set Blocking Without Breaking: Identification and Mitigation of Non-Essential IoT Traffic B 2022 CCS app-analysis Understanding IoT Security from a Market-Scale Perspective B 2022 PETS iot-device-set Analyzing the Feasibility and Generalizability of Fingerprinting Internet of Things Devices B 2022 USENIX iot-device-set Lumos: Identifying and Localizing Diverse Hidden IoT Devices in an Unfamiliar Environment B 2023 IMC iot-device-set Behind the Scenes: Uncovering TLS and Server Certificate Practice of IoT Device Vendors in the Wild B 2023 IMC iot-device-set In the Room Where It Happens: Characterizing Local Communication and Threats in Smart Homes B 2024 IEEE-SP device-population Surveilling the Masses with Wi-Fi-Based Positioning Systems B 2024 IMC iot-device-set IoT Bricks Over v6: Understanding IPv6 Usage in Smart Homes B 2024 IMC delivery-performance Characterizing User Platforms for Video Streaming in Broadband Networks B 2024 PETS iot-device-set Connecting the Dots: Tracing Data Endpoints in IoT Devices B 2025 NDSS iot-device-set Evaluating Machine Learning-Based IoT Device Identification Models for Security Applications B 2025 USENIX vulnerability Tracking You from a Thousand Miles Away! Turning a Bluetooth Device into an Apple AirTag Without Root Privileges B 2026 NDSS vulnerability BLERP: BLE Re-Pairing Attacks and Defenses B 2026 USENIX policy-compliance Missing, Present and Conflicting: A Large Scale Analysis of IoT Update Information in the EU Market ==============================================================================
5b. The guard on this map, mutation-tested
A review pass on 2026-09-12 mutation-tested the map rather than reading it, and found the guard was weaker than it looked.
- Deleting an entry — the script throws:
audit set has 1 slug(s) with no verdict in ctv_fold.mjs. Working as documented. - Flipping a tier letter on a paper that stays in the candidate set —
BtoOUTon Tracking You from a Thousand Miles Away — the script exited 0, silently recomputed the population as 34 instead of 35 andOUTas 53 instead of 52, and every percentage on the content page would have moved with it. No assertion fired, because the only check was a slug-set difference.
Fixed the same day, by pinning the split to what connected_tv publishes:
const PUBLISHED_SPLIT = { A: 13, B: 22, ADJ: 16, OUT: 52 };
And that fix was itself insufficient — found by the next review pass, on 2026-09-13. A count is not a membership. Mutating two entries at once, swapping a genuine Tier B paper out and a genuine OUT paper in, leaves all four counts identical and exits 0, while the population silently acquires a paper that is not about television at all. The reviewer demonstrated it with BLERP: BLE Re-Pairing Attacks and Defenses (B to OUT) against a bilingual privacy-policy paper whose only TV content is an ACR homonym (OUT to B): platform web went from 1 to 2, PETS 6 to 7, NDSS 6 to 5, and the iot share from 74.3% to 71.4% — every one of them a published figure, none of them guarded.
So membership itself is now pinned, per tier, as a digest of the sorted slug list — and the topic tag, which had no guard at all and drives the iot-device-set list on the content page, is pinned the same way:
const PUBLISHED_MEMBERS = {
A: 'b7410e5f7a33a91e', B: '4a71c82990115cdf', ADJ: '1eb63ba64295cda8', OUT: '22530ed51f0ceb35',
};
const PUBLISHED_TOPICS = 'fdb92e2a4c3ab263';
Re-run on 2026-09-13, all four mutations now fail and the unmutated script exits 0:
| Mutation | Before 2026-09-13 | Now |
|---|---|---|
| Delete a map entry | throws | throws (audit set has 1 slug(s) with no verdict) |
| Flip one tier letter | throws (since 2026-09-12) | throws (verdict split moved) |
| Swap two papers between tiers | exit 0, population silently wrong | throws (tier B membership changed) |
| Change a topic tag only | exit 0, the iot-device-set list silently wrong | throws (topic assignments … changed) |
The lesson worth carrying: the round-1 fix was written by reading the failure the reviewer demonstrated, and it closed exactly that failure and nothing adjacent to it. Only a second mutation test, by a second reviewer with the same brief, found the hole next to it.
6. Folding, and the residue in full
Two folds are used, and only one aggregates anything.
Tool-name alias fold. report_connected_tv.mjs maps a lower-cased alphanumeric skeleton of each tools[].name onto a canonical display name, for 27 skeletons mapping onto 22 canonical names: mitmproxy / mitmdump / mitmweb to mitmproxy, wireshark / tshark to Wireshark, adb / androiddebugbridge to adb, charles / charlesproxy to Charles Proxy, plus one-to-one entries for tcpdump, Frida, Pi-hole, VirusTotal, apktool, jadx, FlowDroid, LibScout, EasyList, Scapy, Selenium, OpenWPM, Mercury, PingPong, Appium, Raspberry Pi, Monkey and UIAutomator. Only usedOrMentioned == “used” tuples are counted, and the unit is the paper.
The unmapped residue is 191 distinct raw tool names across the 35 papers. Printed in full, because a residue that lives only in a local file is a residue nobody reads — and because in this case reading it is how the broadcast-side instruments were found:
3× DBSCAN | 2× Censys | 2× dnsmasq | 2× Google voice synthesizer | 2× IoT Inspector | 2× nmap | 2× OpenSSL 2× random forest | 2× scikit-learn | 2× t-SNE | 2× TF-IDF | 2× WHOIS | 1× Adam | 1× adb_shell 1× Afatech AF9015 | 1× agglomerative clustering | 1× Anaconda | 1× Analysis Scripts | 1× Androguard 1× Android Debug Bridge (adb) | 1× Android Debug Bridge (ADB) | 1× Android Studio APK Analyzer 1× AntMonitor | 1× apk-mitm | 1× apksigner | 1× AppCensus | 1× Apple trust store 1× Apple Wi-Fi geolocation API | 1× Application Exerciser Monkey | 1× Apriori | 1× arecord | 1× ARKit 1× Avalpa OpenCaster | 1× BeautifulSoup | 1× BeEF Toolkit | 1× BERT | 1× BiLSTM | 1× Bing | 1× Bleak 1× Bumble | 1× Chapoly1305/FindMy | 1× ChatGPT (OpenAI's TextCompletion API) | 1× Chrome | 1× CICFlowmeter 1× CogniCrypt | 1× Common CA Database | 1× Conviva | 1× cosine distance | 1× Criminal IP | 1× crt.sh 1× cryptography/fernet | 1× CryptoGuard | 1× curl | 1× DekTec DTU-215 | 1× DekTec StreamXpress 1× DICE coefficient | 1× Dijkstra's algorithm | 1× DNSDB | 1× DPDK | 1× DroidBot | 1× fastText 1× FCC database of digital TV towers | 1× Flight Radar 24 | 1× Forward feature selection (FFS) 1× Fourier transform | 1× generic deep neural network | 1× GNU TLS | 1× Google Play API 1× Google Public DNS | 1× Google search | 1× Google Search | 1× Google Voice synthesizer | 1× GPS Tracks 1× Gradient Boosting Decision Tree | 1× GSDMM | 1× HDMI Video Capture Device | 1× HiDes UT-100c 1× Hurricane Electric IPv6-over-IPv4 tunnel | 1× IDA Pro | 1× IDAPython 1× IEEE Organizationally Unique Identifier registry | 1× IFTTT | 1× Intel RealSense Camera T265 1× InternalBlue | 1× IP2Location | 1× IPFIX | 1× iptables | 1× IRDB | 1× irgen | 1× IrScrutinizer | 1× Java 1× Keras | 1× Latent Dirichlet Allocation | 1× LightGBM | 1× logistic regression (custom) | 1× MakeHex 1× MAPS | 1× Maven Repository | 1× MaxMind | 1× MaxMind GeoLite2 | 1× MaxMind geolocation database 1× Mbed TLS | 1× MbedTLS | 1× McAfee | 1× median absolute deviation (MAD) | 1× Microsoft trust store 1× Mon(IoT)r | 1× Monkey Application Exerciser | 1× Monkey Application Exerciser for Android Studio 1× Monte Carlo sampling | 1× Mother of all Ad-Blocking | 1× Mozilla trust store | 1× Naïve Bayes 1× NASA SEDAC Metropolitan Statistical Areas dataset | 1× nDPI | 1× nearest-neighbor classifier | 1× Nessus 1× Netdisco | 1× NetFlow | 1× Netify | 1× Nexmon | 1× NFF-Go | 1× NimBLE 1× Non-Negative Matrix Factorization (NMF) | 1× NoxPlayer | 1× Objection | 1× OpenAI Text Completion API 1× OpenCaster | 1× OpenDNS | 1× OpenWRT | 1× OpenWrt/LEDE | 1× OPP-115 | 1× Oracle Java 1× passive network telescope | 1× Passport | 1× Pi-hole Default blocklist | 1× PostgreSQL | 1× PrivBERT 1× Prodigy | 1× ProVerif | 1× pyshark | 1× Python | 1× Python requests/2.31.0 | 1× Python TLS implementation 1× Radare2 | 1× Random Forest | 1× Randoop | 1× Raspberry Pi 3 | 1× Raspberry Pi 4 | 1× Redis 1× RedOrbit HbbTV Emulator | 1× Remote Central Forums | 1× RIPE IPmap | 1× Roku External Control Protocol 1× SciPy | 1× Secure Transport | 1× SHAP | 1× Similarweb | 1× Snorkel | 1× Softflowd | 1× SoSci Survey 1× spaCy | 1× spaCy en_core_web_lg | 1× StopAd smart TV blocklist | 1× TensorFlow 1× The Big Blocklist Collection (Firebog) | 1× TP-Link power plugs | 1× traceroute | 1× TrafficPassthrough 1× Trigger Scripts | 1× TSDuck | 1× Tuya Smart app | 1× TV Fool | 1× tvbus.exe | 1× Unity 1× Validation Scripts | 1× VLC Player | 1× VS1838B | 1× WALA | 1× WiFi Inspector | 1× WiGLE | 1× WiGLE API 1× WireShark/tshark | 1× wolfSSL | 1× WolfSSL | 1× word2vec | 1× XCUITest | 1× XGBoost | 1× YAF 1× Yersinia | 1× Zeek
Nothing in the residue was silently merged and nothing was dropped: the page's instrument table is exactly the 27-alias slice, and the residue is everything else. The residue is the interesting half here. Avalpa OpenCaster, TSDuck, DekTec DTU-215, HiDes UT-100c, Afatech AF9015 and RedOrbit HbbTV Emulator are DVB modulation and stream-authoring tools with no counterpart anywhere else on this wiki; IRDB, irgen, IrScrutinizer, MakeHex and VS1838B are infrared remote tooling; HDMI Video Capture Device, Roku External Control Protocol, tvbus.exe and NoxPlayer are TV-specific automation. A fold that had merged these into “other” would have hidden the page's most useful finding.
No fold is applied to vantage.locations or population.sourceList, and both are published unfolded on the content page as rankings only, never as percentages. This is deliberate: at n=35 the folding error that corpus measures on the vantage field (280 versus 498 for the United States, corpus-wide) is not worth introducing, and the raw strings — “Apartment 1”, “lab space”, “241 countries and territories”, “e-bike route” — are themselves informative about what a TV vantage point is.
7. Quotes: checked against both renderings
scripts/ctv_quotecheck.py matches every phrase either page quotes, or leans on for a figure, against both paper.cols.txt and the PDF text layer via pypdf. Both are needed, and this run proves why:
- 36 of 41 located in both renderings. (It was 28 of 33 until 2026-09-13; round 2 added eight needles, all of which land in both renderings — see below.)
- 2 in
.colsonly — the FingerprinTV DBF sentence and the Roku ECP URL, which the PDF text layer scrambles. - 3 in the PDF only — de-columning splices. The clearest is [5Kumar, Deepak; Shen, Kelly; Case, Benton; Garg, Deepali; Alperovich, Galina; Kuznetsov, Dmitry; Gupta, Rajarshi; Durumeric, Zakir (2019): "All Things Considered: An Analysis of IoT Devices on Home Networks", in: Proceedings of the USENIX Security Symposium. (Link)]: the sentence “the most popular vendor, Roku, only accounts for 17.4% of media devices” is spliced in
.colsinto “nd the most poputions of IoT device types, except when a device type accounts lar vendor, Roku, only accounts for 17.4% of media devices for fewer than 1% of devices”. A.cols-only check would have reported a correct quote as NOTFOUND. - 0 in neither.
Eight needles were added on 2026-09-13, for the claims the round-2 review changed. They are worth listing because each one existed on the page before it existed in this checker — which is the drift this file is supposed to prevent:
| Needle | Why it was added |
|---|---|
to bypass certificate pinning ([6Moghaddam, Hooman Mohajeri; Acar, Gunes; Burgess, Ben; Mathur, Arunesh; Huang, Danny Yuxing; Feamster, Nick; Felten, Edward W.; Mittal, Prateek; Narayanan, Arvind (2019): "Watching You Watch: The Tracking Ecosystem of Over-the-Top TV Streaming Devices", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)]) | the widened TLS-decryption probe; the narrow probe missed this sentence while the page quoted it |
we cannot install our own self-signed certificates on the Roku ([7Varmarken, Janus; Le, Hieu; Shuba, Anastasia; Markopoulou, Athina; Shafiq, Zubair (2020): "The TV is Smart and Full of Trackers: Measuring Smart TV Advertising and Tracking", in: Proceedings on Privacy Enhancing Technologies. (DOI)]) | same |
there is no way to install custom certificates on Android TV ([1Tileria, Marcos; Blasco, Jorge (2022): "Watch Over Your TV: A Security and Privacy Analysis of the Android TV Ecosystem", in: Proceedings on Privacy Enhancing Technologies. (DOI)]) | same |
We perform a factory reset of the TV for each channel analysis ([8Tagliaro, Carlotta; Hahn, Florian; Sepe, Riccardo; Aceti, Alessio; Lindorfer, Martina (2023): "I Still Know What You Watched Last Sunday: Privacy of the HbbTV Protocol in the European Smart TV Landscape", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]) | refutes the page's old “no paper states whether the device was factory-reset”. Note the needle stops where it does: the full sentence is spliced in .cols by the de-columner, which interleaves it with “and the UT-100c HiDes modulator” |
download the last version of each app (as of August 2020) from AndroZoo ([1Tileria, Marcos; Blasco, Jorge (2022): "Watch Over Your TV: A Security and Privacy Analysis of the Android TV Ecosystem", in: Proceedings on Privacy Enhancing Technologies. (DOI)]) | refutes “every paper built its own set” |
we randomly selected 1.5K unique package names from Androzoo ([9Girish, Aniketh; Hu, Tianrui; Prakash, Vijay; Dubois, Daniel J.; Matic, Srdjan; Huang, Danny Yuxing; Egelman, Serge; Reardon, Joel; Tapiador, Juan; Choffnes, David R.; Vallina-Rodriguez, Narseo (2023): "In the Room Where It Happens: Characterizing Local Communication and Threats in Smart Homes", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]) | same |
published a candidate standard for hybrid TV in America ([10Oren, Yossef; Keromytis, Angelos D. (2014): "From the Aether to the Ethernet—Attacking the Internet using Broadcast Digital Television", in: Proceedings of the USENIX Security Symposium. (Link)]) | the single real ATSC mention in all 5,859 papers |
this time enabling the “Limit Ad Tracking” … “Disable Interest-based Ads” … settings ([6Moghaddam, Hooman Mohajeri; Acar, Gunes; Burgess, Ben; Mathur, Arunesh; Huang, Danny Yuxing; Feamster, Nick; Felten, Edward W.; Mittal, Prateek; Narayanan, Arvind (2019): "Watching You Watch: The Tracking Ecosystem of Over-the-Top TV Streaming Devices", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)]) | refutes “exactly one paper measures both sides of the opt-out” |
One needle was genuinely wrong, and the check caught it. The first draft asserted “decryption fails for 1 out of 5 (or fewer) TLS connections for 80% of all apps” against [7Varmarken, Janus; Le, Hieu; Shuba, Anastasia; Markopoulou, Athina; Shafiq, Zubair (2020): "The TV is Smart and Full of Trackers: Measuring Smart TV Advertising and Tracking", in: Proceedings on Privacy Enhancing Technologies. (DOI)]. The paper says “decryption fails for 1 out of 10 (or fewer) TLS connections for 55% of all apps; 1 out of 5 (or fewer) TLS connections for 80% of all apps” — two clauses, and the draft had glued the opening of the first to the end of the second, producing a sentence the paper does not contain. The needle was narrowed to the clause that is actually there. This is the reason bare-number needles are avoided in that script.
Unedited output:
cols+pdf CCS 2019 watching-you-watch-the-tracking-ecosystem-of "present on 69% of Roku channels and 89% of Amazon Fire TV channels" cols+pdf CCS 2019 watching-you-watch-the-tracking-ecosystem-of "we were able to install our own cert on the device which allowed u" cols+pdf CCS 2019 watching-you-watch-the-tracking-ecosystem-of "that leaked the title of the video to a tracking domain" PDF only PETS 2020 the-tv-is-smart-and-full-of-trackers-measuri "1 out of 5 (or fewer) TLS connections for 80% of all apps" cols+pdf PETS 2020 the-tv-is-smart-and-full-of-trackers-measuri "314 ATS domains that are unique to the Roku dataset" cols only PETS 2022 fingerprintv-fingerprinting-smart-tv-apps "among 80 apps that are made available on all three smart TV platfo" cols+pdf PETS 2022 watch-over-your-tv-a-security-and-privacy-an "The analysis found at least one sensitive data flow in 78% of the " cols+pdf NDSS 2023 i-still-know-what-you-watched-last-sunday-pr "26 communicate with trackers before the user has expressed their c" cols+pdf NDSS 2023 i-still-know-what-you-watched-last-sunday-pr "only block at maximum 44% in 2021 and 81% in 2022" cols+pdf IMC 2024 watching-tv-with-the-second-party-a-first-lo "there is a complete absence of communication with any previously i" cols+pdf IMC 2024 watching-tv-with-the-second-party-a-first-lo "smart TVs in the UK and the US contact distinct ACR domains" cols+pdf IMC 2024 watching-tv-with-the-second-party-a-first-lo "ACR network traffic exists when watching linear TV and when using " PDF only USENIX 2019 all-things-considered-an-analysis-of-iot-dev "the most popular vendor, Roku, only accounts for 17.4% of media de" cols+pdf IMC 2018 understanding-video-management-planes "streaming set-top boxes1 dominate by view-hours" cols+pdf IMC 2023 in-the-room-where-it-happens-characterizing- "the analysis of the Smart TV ecosystem is left for future work" cols+pdf IEEE-SP 2024 surveilling-the-masses-with-wi-fi-based-posi "belong to the streaming television equipment manufacturer Roku" cols+pdf IMC 2021 iotls-understanding-tls-usage-in-consumer-io "such as voice assistants, smart TVs and video doorbells" cols+pdf USENIX 2025 watch-out-your-tv-box-reversing-and-blocking "they are offered only to those who have purchased specific" cols+pdf USENIX 2014 from-the-aether-to-the-ethernet-attacking-th "which requires a minimal budget and infrastructure" cols+pdf IMC 2024 iot-bricks-over-v6-understanding-ipv6-usage- "only eight out of 93 devices remain functional" cols+pdf CCS 2019 watching-you-watch-the-tracking-ecosystem-of "On Roku, a total of 43 channels failed to properly verify the serv" cols+pdf CCS 2019 watching-you-watch-the-tracking-ecosystem-of "794 of the 1000 Roku channels sent at least one request in clearte" cols+pdf CCS 2019 watching-you-watch-the-tracking-ecosystem-of "We found 9 channels on Roku and 14 channels on the Fire TV" cols only CCS 2019 watching-you-watch-the-tracking-ecosystem-of "an HTTP GET request to "http://ROKU_ DEVICE_IP_ADDRESS:8060/keydow" cols+pdf PETS 2020 the-tv-is-smart-and-full-of-trackers-measuri "697 Fire TV apps that expose advertising ID alongside serial numbe" cols+pdf PETS 2022 fingerprintv-fingerprinting-smart-tv-apps "96% (N = 961) of the top" PDF only PETS 2022 watch-over-your-tv-a-security-and-privacy-an "75% of the apps contain analytics libraries and 77% contain advert" cols+pdf USENIX 2021 android-smarttvs-vulnerability-discovery-via "37 unique vulnerabilities, including 11 high-impact cyber threats," cols+pdf USENIX 2023 homespy-the-invisible-sniffer-of-infrared-re "The accuracy increases to 70% for Top3 and 77% for Top5" cols+pdf NDSS 2024 acoustic-keystroke-leakage-on-smart-televisi "up to 60.19% of common passwords" cols+pdf IMC 2024 watching-tv-with-the-second-party-a-first-lo "the fact that we observe network traffic every 15 seconds suggests" cols+pdf IMC 2011 understanding-couch-potatoes-measurement-and "The average number of set-top boxes provisioned was approximately " cols+pdf USENIX 2019 all-things-considered-an-analysis-of-iot-dev "are the most common type of device in seven of the eleven regions" cols+pdf CCS 2019 watching-you-watch-the-tracking-ecosystem-of "to bypass certificate pinning" cols+pdf PETS 2020 the-tv-is-smart-and-full-of-trackers-measuri "we cannot install our own self-signed certificates on the Roku" cols+pdf PETS 2022 watch-over-your-tv-a-security-and-privacy-an "there is no way to install custom certificates on Android TV" cols+pdf NDSS 2023 i-still-know-what-you-watched-last-sunday-pr "We perform a factory reset of the TV for each channel analysis" cols+pdf PETS 2022 watch-over-your-tv-a-security-and-privacy-an "download the last version of each app (as of August 2020) from And" cols+pdf IMC 2023 in-the-room-where-it-happens-characterizing- "we randomly selected 1.5K unique package names from Androzoo" cols+pdf USENIX 2014 from-the-aether-to-the-ethernet-attacking-th "published a candidate standard for hybrid TV in America" cols+pdf CCS 2019 watching-you-watch-the-tracking-ecosystem-of "this time enabling the "Limit Ad Tracking" (Roku) and the "Disable" 41 quotes: 36 in both renderings, 2 in .cols only, 3 in the PDF only, 0 in neither.
8. External and industry sources
Every one fetched on 2026-09-12, and every one a primary source: a vendor's own developer documentation, a standards body, or a regulator's own press release. None of the figures on the content page comes from a news article, a vendor blog post or a comparison site.
| Claim on the page | Source | How verified |
|---|---|---|
| HbbTV 2.0.5, published 2026-02-25, incremental over 2.0.4 (March 2023) | HbbTV Association specifications page | fetched; the version table lists 1.0 (2010) through 2.0.5 (2026-02-25) |
| Roku ECP is “a simple RESTful API accessed using HTTP on port 8060”, no authentication documented | Roku developer docs, External Control API | fetched; the port and the query/device-info endpoint quoted verbatim. Cross-checked against [6Moghaddam, Hooman Mohajeri; Acar, Gunes; Burgess, Ben; Mathur, Arunesh; Huang, Danny Yuxing; Feamster, Nick; Felten, Edward W.; Mittal, Prateek; Narayanan, Arvind (2019): "Watching You Watch: The Tracking Ecosystem of Over-the-Top TV Streaming Devices", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], which uses the same URL form |
RIDA, GetRIDA() / IsRIDADisabled(), 30-day temporary ID under limit-ad-tracking | Roku developer docs: integrating-roku-advertising-framework for GetRIDA() and the 30-day ID, ifDeviceInfo for IsRIDADisabled() (which is not on the RAF page) | fetched; re-verified 2026-09-13 |
TIFA, getTIFA() / isLATEnabled(), resettable, “no connection to any PII … or DUID” | Samsung Smart TV developer docs | fetched from the unique-identifiers-for-smarttv guide |
Fire TV Advertising ID, advertising_id / limit_ad_tracking, Fire OS 5.2.1.1+ on TV | Amazon Developer Policy Center, Advertising ID Policy | fetched |
Wireless adb needs Android 13 (API 33) for TV, against Android 11 for phones | Android developer docs, adb page | fetched; the TV/WearOS threshold is stated separately from the phone one |
| FTC/NJ–VIZIO, $2.2m, 11 million televisions, second-by-second, delete pre-2016-03-01 data | US FTC press release, 2017-02-06 | fetched |
| Walmart completed the VIZIO acquisition 2024-12-03 | Walmart corporate newsroom | fetched |
| Texas sues Sony, Samsung, LG, Hisense, TCL, 2025-12-15; “every 500 milliseconds” | Texas Attorney General press release | fetched with curl and a browser User-Agent (WebFetch returns HTTP 402 on this host); date and quote read from the rendered page |
| Hisense TRO, 2025-12-17 | Texas Attorney General press release | fetched the same way |
| Samsung agreement, 2026-02-26; LG agreement, 2026-05-11 | Texas Attorney General press releases | fetched the same way; the Samsung URL is not the one a search result suggested and 404s under the guessed slug |
| ATSC 3.0 reaches “more than 76% of U.S. households”; broadcaster applications support profile-based personalisation | ATSC deployments and NextGen TV pages | fetched; the page's own deployment map is dated July 2026 |
| Artifact repository currency (5 repositories, none touched in 2025–2026); mitmproxy v12.2.3, 2026-05-12 | GitHub REST API | repos/<owner>/<repo> for archived and pushed_at, plus the default branch's newest commit date, because pushed_at counts any branch |
Rejected, and why. A search for recent ACR measurement returned several consumer-facing articles (a “how to disable ACR in 2026” listicle, a cybersecurity blog summarising the IMC paper, a compliance vendor's education page) which between them asserted the LG-15-seconds and Samsung-per-minute cadences, the Texas lawsuit and the Samsung settlement. None was used. The cadences were taken from [11Anselmi, Gianluca; Vekaria, Yash; D'Souza, Alexander; Callejo, Patricia; Mandalari, Anna Maria; Shafiq, Zubair (2024): "Watching TV with the Second-Party: A First Look at Automatic Content Recognition Tracking in Smart TVs", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]'s own text, and the enforcement dates from the Attorney General's releases. One of those articles also attributed the “every 500 milliseconds” figure to the research paper; it is the regulator's pleading, and the paper's own 500 ms figure is an estimate of Samsung's capture rate. The content page keeps those two apart deliberately.
One dead link, and it is in the corpus's own data. The artifact URL the extraction recovered for [12Gopalakrishnan, Vijay; Jana, Rittwik; Ramakrishnan, K. K.; Swayne, Deborah F.; Vaishampayan, Vinay A. (2011): "Understanding couch potatoes: measurement and modeling of interactive usage of IPTV at large scale", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] — www.research.att.com/~kkrama/papers/streamcontrol.pdf — returns HTTP 403 over both http and https and with the tilde encoded. It appears in this page's artifact listing because it is what the paper stated, not because anything on either page rests on it; it is left in place as a record of the paper's own claim. Every URL actually cited as evidence on connected_tv returns 200.
Two further checks worth recording. The Mon(IoT)r testbed software is live at github.com/djdubois/moniotr-core (last pushed 2024-08-09) but the lab's tools page does not publish a smart-TV dataset for download, so the page describes testbed captures as a route without promising a TV dataset exists to fetch. And the PETS landing pages were used to recover author lists for five entries the corpus index lacks; [13Ahmed, Dilawer; Das, Anupam; Zaffar, Fareed (2022): "Analyzing the Feasibility and Generalizability of Fingerprinting Internet of Things Devices", in: Proceedings on Privacy Enhancing Technologies. (DOI)]'s authors were additionally cross-checked against Crossref because its stored PDF has no usable text layer on the title page.
8b. Sources added or corrected in round 2 (2026-09-13)
| Claim | Primary source, and how it was checked | Result |
|---|---|---|
ATSC 3.0 broadcaster applications: advertisingId, filterCode, receiver cookies | ATSC A/344:2026-04, ATSC 3.0 Interactive Content, 14 April 2026. PDF downloaded (HTTP 200, 2,363,014 bytes, 202 pages), text extracted with pypdf, each needle matched in whitespace-collapsed text | All five quotes FOUND verbatim. data collection returns 0 hits in the whole standard — the phrase the withdrawn 2016 quote used |
| A/344 current revision | the A/344 document page lists 2026-04 (14 April 2026) above 2026-02; the 2026-04 PDF's own title block reads A/344:2026-04 | 2026-04, not the 2026-02 the author first downloaded. One round-2 reviewer read the standards listing as showing 2025-07 as the latest approved; the PDF's own designation settles it |
| EU: joint Article 62 GDPR operation on smart TVs | Autoriteit Persoonsgegevens landing page (HTTP 200) and the report PDF (HTTP 200, 1,005,269 bytes, 14 pages), both needing a browser User-Agent and a Referer header | Real. NL/HU/IT/LI, published 23 September 2025. The off-state percentages (97.52 / 98.84 / 91.10) read out of the PDF's own table |
| UK: ICO connected-TV programme | ICO news release, 11 June 2026 (HTTP 403 to a bare fetcher, HTTP 200 with a browser User-Agent) | Real. Quote and attribution to William Malcolm verified in the fetched HTML |
Walmart press release names Platform+ | the release (HTTP 200, 180,202 bytes), searched after double HTML-entity decoding | It does. The round-1 footnote said it did not. The plus sign is entity-encoded, so a tag-strip-only search misses it — this is why the check has to decode entities, not just strip tags. Inscape is genuinely absent |
ahn2025_watch Zenodo access | Zenodo REST API, records/15646588 (HTTP 200): access_right: open, CC-BY-4.0, title ends [Public Artifact] | The paper has two records; the extraction's restricted is true of zenodo.15602938 only. Round 1 attached the label to the open DOI |
| FingerprinTV code release | README fetched raw (HTTP 200) | “The FingerprinTV dataset has already been released.” / “Once it is ready for release to the public, the code will be added to this repository. Please stay tuned.” Four years on |
| FCC Fifth FNPRM is still pending | Checked for a Report and Order; none found as of 2026-09-13. The adopted item is FCC 25-72, adopted 2025-10-28, released 2025-10-29 | The page's fact-sheet citation stands. One reviewer surfaced a trade-press claim of a May 2026 FCC vote mandating ATSC 3.0 tuners; it appears nowhere on fcc.gov and is rejected |
Rejected in round 2. A blog URL for “ADB Wi-Fi 2.0” offered in round 1 returns 404 and stays rejected; the claim rests on developer.android.com/tools/adb instead. The trade-press “FCC voted 3–2 in May 2026” story is rejected as above. Samsung's briefly-vacated TRO is still judged a procedural detail and stays off the content page.
9. Bibliography
32 entries were added to bibliography in this sitting — 31 in the first pass and [14Zhu, Yanzi; Xiao, Zhujun; Chen, Yuxin; Li, Zhijing; Liu, Max; Zhao, Ben Y.; Zheng, Haitao (2020): "Et Tu Alexa? When Commodity WiFi Devices Turn into Adversarial Motion Sensors", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] added after review (see §15) — generated by scripts/bibgen.mjs from data/corpus2/.meta so authors, titles and DOIs are publisher metadata rather than recall. Checks run before appending:
- Citekey collisions: none of the 31 keys exists in the live bibliography.
- Duplicate-paper scan:
scripts/bib_dedup_scan.pyon the merged file reports 0 definite duplicates (rule A, same DOI; rule B, same squashed title). Three of the 89 pre-existing rule-D candidates involve a new key —hu2024_bricksagainsthu2024_unmasking, andwang2024_characterizingagainst two otherwang2024keys — and all three are different first authors, so none is a duplicate. - Literal
@in a field: none. A raw ASCII@inside a BibTeX field makes the bibtex4dw plugin drop the entry and every marker to it, silently. - Authors: PETS and USENIX records carry no authors in the index. Five PETS entries were filled from the publisher's landing pages (Varmarken et al., Tileria and Blasco, Mandalari et al., Ahmed et al., Mavroudis et al.); eight USENIX entries were resolved by
scripts/fetch_authors.py. First and last author of all 31 were then checked against the paper's own PDF text layer, which passed for 30; the one that failed, [13Ahmed, Dilawer; Das, Anupam; Zaffar, Fareed (2022): "Analyzing the Feasibility and Generalizability of Fingerprinting Internet of Things Devices", in: Proceedings on Privacy Enhancing Technologies. (DOI)], has no text layer on its front matter and was confirmed against Crossref instead. - Two hand corrections to bibgen's output: the citekey
bjrklund2025_endangeredwas corrected tobjorklund2025_endangered(the generator drops the ö rather than transliterating it), and Jad Al Aaraj was re-split fromAaraj, Jad AltoAl Aaraj, Jad. - bibgen's stdout carries QA notes (“no DOI available”, “metadata source: venue-page”, citation counts). Only the
@entries were appended; the notes were stripped by a regex that extracts complete entries, because those notes have previously gone live on the public bibliography page.
This provenance page adds no bibliography entries of its own and uses only keys the content page already uses, plus [3Björklund, Martin; Duvignau, Romaric (2025): "Endangered Privacy: Large-Scale Monitoring of Video Streaming Services", in: Proceedings of the USENIX Security Symposium. (Link)] for the roadmap correction in §1.
10. What could not be established
- How big the true population is. The 35 is a floor. The 39 papers dropped between gate 1 and gate 2 were not read, and no probe can reach a paper that measures a television without naming a TV-class term in its full text.
- Whether the Tier B line is where somebody else would put it. “Reports a result broken out for a television” is a reading, not a field. Moving Void and the cryptojacking proof-of-concept in would make it 37; requiring a TV-specific privacy or security result rather than any result would make it roughly 28.
- What ACR does now. The only measurement is 2024, on two sets, and both of those vendors have since agreed consent changes with a US regulator. The page says the 2024 figures are a pre-order baseline; it does not claim to know the current behaviour, and nothing in the corpus does.
- Anything about ATSC 3.0. Not one corpus paper measures it. The deployment share and the personalisation capabilities on the page are ATSC's own statements about its own standard, labelled as such.
- TV market share by installed base. Repeatedly useful and repeatedly unavailable from a primary source that is not a paid analyst report. The page therefore never says which platform is biggest, only what the literature measured.
- Whether TLS interception has improved on closed platforms since 2019. No paper in the corpus revisits it. The 4.3% figure is quoted with its date attached rather than as a current state.
- A figure for the TV share of household or video traffic that a privacy paper could use as a denominator. [15Wang, Yifan; Lyu, Minzhao; Sivaraman, Vijay (2024): "Characterizing User Platforms for Video Streaming in Broadband Networks", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] and [16Akhtar, Zahaib; Nam, Yun Seong; Chen, Jessica; Govindan, Ramesh; Katz-Bassett, Ethan; Rao, Sanjay G.; Zhan, Jibin; Zhang, Hui (2018): "Understanding Video Management Planes", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] come closest and neither gives one; this is written up as an open question rather than back-calculated.
11. Judgement calls
- A new page rather than a section of mobile_and_app_measurement. Android TV apps genuinely are Android apps, and that page's store, static-analysis and pinning material transfers. But ACR, HbbTV and “you cannot install a certificate at all” have no mobile analogue, and 26 of the 35 papers carry
iotrather thanmobile. The content page points at the mobile page rather than restating it, in four places. - Not a section of platforms. A television is a device, not a platform whose API you negotiate access to. The overlap is the store chart, and that is one row.
- The blocklist-coverage numbers are shared with filter_lists rather than moved. That page already carries a smart-TV row citing [7Varmarken, Janus; Le, Hieu; Shuba, Anastasia; Markopoulou, Athina; Shafiq, Zubair (2020): "The TV is Smart and Full of Trackers: Measuring Smart TV Advertising and Tracking", in: Proceedings on Privacy Enhancing Technologies. (DOI)] with the 22%/27% figures. The content page repeats them once, in the section explaining why filter lists do not transfer, and links there rather than re-deriving.
- IPTV performance work is Tier A, not adjacent. [12Gopalakrishnan, Vijay; Jana, Rittwik; Ramakrishnan, K. K.; Swayne, Deborah F.; Vaishampayan, Vinay A. (2011): "Understanding couch potatoes: measurement and modeling of interactive usage of IPTV at large scale", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] and [17Song, Han Hee; Ge, Zihui; Mahimkar, Ajay; Wang, Jia; Yates, Jennifer; Zhang, Yin; Basso, Andrea; Chen, Min (2011): "Q-score: proactive service quality assessment in a large IPTV system", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] measure set-top boxes, which the rule calls television-class. They are tagged
delivery-performanceso the topic table shows that two of the four oldest papers in the population are not privacy work at all. A privacy-only page would have dropped them and reported a smaller, tidier, less honest literature. - The enforcement timeline is on the content page, not only here. A student planning an ACR measurement in 2026 who does not know about the Texas agreements will measure a consent flow and report it as a default. That is a methodological fact, so it belongs on the content page.
- Per-year counts stop being a series. With 35 papers over 16 years, the page reports four-to-five-year windows and labels 2025–2026 provisional in the table itself, rather than drawing a trend.
12. The scripts
Three files, committed with their real output.
- ctv_fold.mjs
// Population map for `design:connected_tv`. // // The roadmap queued this page against a 16-paper TITLE+SUMMARY candidate set // (scripts/gap_probe_roadmap.mjs, family `ctv_streaming`). That probe is a // floor, and it is also the wrong instrument twice over: it misses papers whose // title never says "TV" (Watch Over Your TV is in it only by accident of the // word "TV"; the Android TV ecosystem paper was NOT in the 16), and its `web` // column — 1 of 16 — is not a filter anyone should apply here, because a // television is not a web-platform measurement in the extraction's sense. // // So the population is derived instead from a full-text recall probe over all // 5,859 `paper.cols.txt` files (scripts/_ctv_probe1.mjs / _ctv_probe2.mjs), // gated into an audit set, and then HAND-AUDITED against a written rule. // // THE RULE, fixed before any figure was computed: // // A "television-class endpoint" is a smart TV set, a TV operating system // (Android TV / Google TV, tvOS, Tizen, webOS, Roku OS, Fire OS), a streaming // stick / box / set-top box, an app running on one, or the broadcast path // (HbbTV / DVB) delivered into one. // // A the paper's central object of measurement is a television-class endpoint // (its traffic, apps, firmware, broadcast channel or user interaction). // B television-class devices are part of a broader measured population AND // the paper reports at least one result broken out for them. // ADJ adjacent: cite where relevant, do not count. Video streaming measured // off a TV (browser DRM, piracy websites, encrypted-traffic video // fingerprinting), or a TV used as apparatus rather than measured. // OUT the TV name is a passing reference, a survey answer option, a related- // work sentence, or a homonym. // // `topic` is hand-assigned and only used for a ranking, never a percentage. // // Every slug in the audit set appears here. report_connected_tv.mjs throws if // the audit set and this map disagree, so widening a probe breaks the report // instead of silently moving the page's denominator. export const MAP = { // ---------------------------------------------------------------- Tier A 'from-the-aether-to-the-ethernet-attacking-the-internet-using-broadcast-digital-t': ['A', 'broadcast', 'Injects HbbTV/DVB payloads into smart TVs over the broadcast band; TVs and STBs are the target population.'], 'watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices': ['A', 'tracking', '1,000 Roku channels and 1,000 Fire TV channels crawled on real devices with TLS interception.'], 'the-tv-is-smart-and-full-of-trackers-measuring-smart-tv-advertising-and-tracking': ['A', 'tracking', 'Roku and Fire TV app traffic, testbed plus in-the-wild; the reference smart-TV tracking measurement.'], 'android-smarttvs-vulnerability-discovery-via-log-guided-fuzzing': ['A', 'vulnerability', 'Log-guided fuzzing of 11 Android TV devices; the TV firmware is the object.'], 'fingerprintv-fingerprinting-smart-tv-apps': ['A', 'tracking', 'Top-1000 apps on each of Apple TV, Fire TV and Roku; network fingerprints of TV apps.'], 'watch-over-your-tv-a-security-and-privacy-analysis-of-the-android-tv-ecosystem': ['A', 'app-analysis', '4,745 Android TV APKs statically analysed plus 21 apps intercepted. NOT in the roadmap candidate set.'], 'i-still-know-what-you-watched-last-sunday-privacy-of-the-hbbtv-protocol-in-the-european-smart-tv-landscape': ['A', 'broadcast', '36 European HbbTV channels on real TVs, plus a 174-respondent awareness survey.'], 'homespy-the-invisible-sniffer-of-infrared-remote-control-of-smart-tvs': ['A', 'side-channel', 'IR remote-control signals of smart TVs sniffed by neighbouring IoT devices.'], 'acoustic-keystroke-leakage-on-smart-televisions': ['A', 'side-channel', 'On-screen-keyboard keystrokes recovered from TV audio on Apple and Samsung TVs.'], 'watching-tv-with-the-second-party-a-first-look-at-automatic-content-recognition': ['A', 'acr', 'Two smart TVs (LG, Samsung) in UK and US; ACR traffic across six viewing scenarios.'], 'watch-out-your-tv-box-reversing-and-blocking-a-p2p-based-illegal-streaming-ecosy': ['A', 'piracy', 'Reverses the EVPAD illegal-streaming set-top box and its P2P ecosystem.'], 'understanding-couch-potatoes-measurement-and-modeling-of-interactive-usage-of-ip': ['A', 'delivery-performance', 'Two years of interaction traces from ~3M IPTV set-top boxes.'], 'q-score-proactive-service-quality-assessment-in-a-large-iptv-system': ['A', 'delivery-performance', 'IPTV service quality inferred for millions of set-top boxes from network measurements and STB logs.'], // ---------------------------------------------------------------- Tier B 'information-exposure-from-consumer-iot-devices-a-multidimensional-network-inform': ['B', 'iot-device-set', '81-device lab; Apple TV, Fire TV, LG TV, Roku TV and Samsung TV are named devices with per-device destinations.'], 'a-haystack-full-of-needles-scalable-detection-of-iot-devices-in-the-wild': ['B', 'iot-device-set', 'Video category = Apple TV, Fire TV, LG TV, Roku TV, Samsung TV; detected in IXP flow data.'], 'iotls-understanding-tls-usage-in-consumer-iot-devices': ['B', 'iot-device-set', 'TV category n=5 (Fire TV, Samsung TV, LG TV, Roku TV, Apple TV) with per-device TLS results.'], 'blocking-without-breaking-identification-and-mitigation-of-non-essential-iot-tra': ['B', 'iot-device-set', 'Fire TV and Roku TV in the 31-device set with per-device destination and breakage results.'], 'packet-level-signatures-for-smart-home-devices': ['B', 'iot-device-set', 'Smart TVs among the devices from which packet-level signatures were extracted.'], 'analyzing-the-feasibility-and-generalizability-of-fingerprinting-internet-of-thi': ['B', 'iot-device-set', 'Roku TV reported as its own row with per-device fingerprinting accuracy and a confusion analysis.'], 'behind-the-scenes-uncovering-tls-and-server-certificate-practice-of-iot-device-v': ['B', 'iot-device-set', 'A dedicated "smart TV and local device" capture is analysed as a case study.'], 'in-the-room-where-it-happens-characterizing-local-communication-and-threats-in-s': ['B', 'iot-device-set', 'Smart TVs are one of eight testbed device categories; the paper also flags TV apps as a local-network threat.'], 'iot-bricks-over-v6-understanding-ipv6-usage-in-smart-homes': ['B', 'iot-device-set', '93 devices; three smart TVs among the eight that still work on IPv6-only, with per-device domains.'], 'connecting-the-dots-tracing-data-endpoints-in-iot-devices': ['B', 'iot-device-set', 'Roku and Samsung Smart TV among the fingerprinted device set; User-Agent and OUI evidence quoted for TVs.'], 'evaluating-machine-learning-based-iot-device-identification-models-for-security-applications': ['B', 'iot-device-set', 'Smart TV is a labelled class (Fire TV, Apple TV, LG webOS TV, Roku TV, Samsung SmartTV) with per-device idleness results.'], 'lumos-identifying-and-localizing-diverse-hidden-iot-devices-in-an-unfamiliar-env': ['B', 'iot-device-set', 'TV class = Vizio, Panasonic, TCL in the 44-device set; unseen smart TVs discussed in the field test.'], 'et-tu-alexa-when-commodity-wifi-devices-turn-into-adversarial-motion-sensors': ['B', 'iot-device-set', 'Chromecast, Apple TV and Roku form the "Smart TV (& Sticks)" row with its own packet-rate measurement.'], 'all-things-considered-an-analysis-of-iot-devices-on-home-networks': ['B', 'iot-device-set', '15.5M homes; Media/TV is a device category and Roku is named with a 17.4% within-category share.'], 'you-are-what-you-broadcast-identification-of-mobile-and-iot-devices-from-public': ['B', 'iot-device-set', 'Apple TV identified from mDNS/DHCP/SSDP views, with a false-positive case study on Apple TV.'], 'characterizing-user-platforms-for-video-streaming-in-broadband-networks': ['B', 'delivery-performance', 'Smart TV is one of the classified device types for 100M+ video flows; gives the TV share of streaming.'], 'understanding-video-management-planes': ['B', 'delivery-performance', 'Streaming set-top boxes (Roku, Fire TV, Apple TV) are a platform category and dominate by view-hours.'], 'missing-present-and-conflicting-a-large-scale-analysis-of-iot-update-information': ['B', 'policy-compliance', 'Smart TVs are one of five device types crawled across 58 EU stores, with a TV-specific disclosure result.'], 'understanding-iot-security-from-a-market-scale-perspective': ['B', 'app-analysis', 'Market-scale mobile-IoT app analysis; Fire TV and HiSense TV appear as identified products with findings.'], 'blerp-ble-re-pairing-attacks-and-defenses': ['B', 'vulnerability', 'TCL 43P638 Android TV is row 12 of the tested-device table with its own attack outcome.'], 'tracking-you-from-a-thousand-miles-away-turning-a-bluetooth-device-into-an-apple': ['B', 'vulnerability', 'Sony Bravia A80J (Android TV 10) is in the tested-device table with its own result.'], 'surveilling-the-masses-with-wi-fi-based-positioning-systems': ['B', 'device-population', 'Roku streaming devices are two of the five most-observed BSSID OUIs in a 490M-entry Wi-Fi positioning dataset; Roku-specific shares reported.'], // ---------------------------------------------------------------- Adjacent 'endangered-privacy-large-scale-monitoring-of-video-streaming-services': ['ADJ', 'streaming-service', 'Video identification from encrypted MPEG-DASH traffic. The roadmap listed it as CTV spine; it measures the SERVICE and its traffic, never a TV.'], 'unmasking-the-shadows-a-cross-country-study-of-online-tracking-in-illegal-movie': ['ADJ', 'streaming-service', 'Illegal movie streaming WEBSITES crawled with a browser; a web-tracking study, not a TV study.'], 'your-drm-can-watch-you-too-exploring-the-privacy-implications-of-browsers-mis-im': ['ADJ', 'streaming-service', 'Widevine EME in browsers and Android; TVs named as another Widevine host, not measured.'], 'cost-saving-streaming-unlocking-the-potential-of-alternative-edge-node-resources': ['ADJ', 'streaming-service', 'Edge-node economics for streaming delivery; no TV endpoint measured.'], 'measurement-and-analysis-of-a-large-scale-commercial-mobile-internet-tv-system': ['ADJ', 'streaming-service', '"TV" delivered to mobile handsets, not to a television.'], 'watching-videos-from-everywhere-a-study-of-the-pptv-mobile-vod-system': ['ADJ', 'streaming-service', 'Mobile VoD; no television endpoint.'], 'performance-characterization-of-a-commercial-video-streaming-service': ['ADJ', 'streaming-service', 'Streaming service performance from browser/CDN vantage; no TV-specific result.'], 'analyzing-the-potential-benefits-of-cdn-augmentation-strategies-for-internet-vid': ['ADJ', 'streaming-service', 'CDN augmentation for video workloads; no TV endpoint.'], 'anatomy-of-a-personalized-livestreaming-system': ['ADJ', 'streaming-service', 'Livestreaming (Periscope) system measurement; no TV endpoint.'], 'peer-assisted-content-distribution-in-akamai-netsession': ['ADJ', 'streaming-service', 'Peer-assisted CDN; set-top-box mention is background.'], 'auto-draft-196': // slug is a placeholder in the index; title = A Lightweight IoT Cryptojacking Detection Mechanism in Heterogeneous Smart Home Networks (NDSS 2022) ['ADJ', 'vulnerability', 'Authors implement their own cryptojacking PoC on an LG webOS TV to test a detector; no deployed-TV population.'], 'void-a-fast-and-light-voice-liveness-detection-system': ['ADJ', 'apparatus', 'A Samsung Smart TV is used as a replay LOUDSPEAKER; the TV is apparatus, not the measured object.'], 'tracking-profiling-and-ad-targeting-in-the-alexa-echo-smart-speaker-ecosystem': ['ADJ', 'iot-device-set', 'Smart speaker study; TVs cited as the comparable prior ecosystem, not measured. The closest methodological sibling.'], 'ovrseen-auditing-network-traffic-and-privacy-policies-in-oculus-vr': ['ADJ', 'iot-device-set', 'VR headsets; smart TVs used as the comparison ecosystem. Same lab, same pipeline shape.'], 'exploiting-diversity-in-android-tls-implementations-for-mobile-app-traffic-class': ['ADJ', 'app-analysis', 'Android app traffic classification; TLS-fingerprint method later reused on TV apps.'], 'on-the-privacy-and-security-of-the-ultrasound-ecosystem': ['ADJ', 'acr', 'Ultrasonic cross-device tracking (uXDT): beacons emitted by TV adverts and picked up by phone SDKs. The TV is the emitter, never measured.'], // ---------------------------------------------------------------- Out 'a-bilingual-longitudinal-analysis-of-privacy-policies-measuring-the-impacts-of-t': ['OUT','','ACR homonym: "ACR" is not automatic content recognition here.'], 'a-billion-open-interfaces-for-eve-and-mallory-mitm-dos-and-tracking-attacks-on-i': ['OUT','','tvOS listed among Apple OSes; AWDL is the object.'], 'a-multi-region-investigation-of-the-perceptions-and-use-of-smart-home-devices': ['OUT','','Survey; smart TV is a related-work citation and an ownership option.'], 'a-placement-vulnerability-study-in-multi-tenant-public-clouds': ['OUT','','Title probe matched "streaming"; cloud VM placement.'], 'analyzing-the-ios-local-network-permission-from-a-technical-and-user-perspective': ['OUT','','Chromecast is one of four IoT devices used to trigger the permission; no TV result.'], 'broadcast-yourself-understanding-youtube-uploaders': ['OUT','','IPTV appears once in related work.'], 'characterizing-everyday-misuse-of-smart-home-devices': ['OUT','','Survey of 483 people; smart TV is an ownership option, not a measured device.'], 'co-designing-a-mobile-app-for-bystander-privacy-protection-in-jordanian-smart-ho': ['OUT','','Interview study; TV names are participant-reported device inventories.'], 'contextualizing-interpersonal-data-sharing-in-smart-homes': ['OUT','','Vignette survey; "viewing history from your smart TV" is a question stem.'], 'dead-domains-living-data-a-privacy-risk-analysis-of-domain-lifecycle-in-android': ['OUT','','Android apps; one expired-domain example happens to also ship on Roku.'], 'deep-dive-into-the-iot-backend-ecosystem': ['OUT','','Backend infrastructure; TV mentions are motivation and a citation to FingerprinTV.'], 'drones-cryptanalysis-smashing-cryptography-with-a-flicker': ['OUT','','IPTV/ACR homonyms.'], 'entropy-ip-uncovering-structure-in-ipv6-addresses': ['OUT','','"ACR" homonym.'], 'examining-consumer-reviews-to-understand-security-and-privacy-issues-in-the-mark': ['OUT','','Review-text analysis; set-top box is a Mirai product category, no TV measurement.'], 'examining-power-dynamics-and-user-privacy-in-smart-technology-use-among-jordania': ['OUT','','Interview study; TVs are participant device inventories.'], 'exploring-the-privacy-concerns-of-bystanders-in-smart-homes-from-the-perspective': ['OUT','','Survey; smart TV is an example in a prompt.'], 'flock-combating-astroturfing-on-livestreaming-platforms': ['OUT','','Title probe matched "streaming platform"; astroturfing detection on Twitch-like sites.'], 'help-me-help-you-privacy-considerations-for-third-party-iot-device-repair': ['OUT','','Vignette survey; TVs appear in a device-category prompt.'], 'hey-your-secrets-leaked-detecting-and-characterizing-secret-leakage-in-the-wild': ['OUT','','IPTV homonym in leaked-credential data.'], 'idea-static-analysis-on-the-security-of-apple-kernel-drivers': ['OUT','','tvOS is one of four Apple OSes scanned; no TV-specific result.'], 'internet-service-providers-and-individuals-attitudes-barriers-and-incentives-to': ['OUT','','Interview and survey; TV is a device-ownership row.'], 'iotflow-inferring-iot-device-behavior-at-scale-through-static-mobile-companion-a': ['OUT','','Companion-app analysis; no TV breakout.'], 'its-the-equivalent-of-feeling-like-youre-in-jail-lessons-from-firsthand-and-seco': ['OUT','','Interview study on IPV; TV is a reported abuse vector, not measured.'], 'measuring-up-to-reasonable-consumer-expectations-providing-an-empirical-basis-fo': ['OUT','','Vignette survey; Vizio appears only in a news citation.'], 'medical-devices-are-at-risk-information-security-on-diagnostic-imaging-system': ['OUT','','"ACR" = American College of Radiology.'], 'my-best-friends-husband-sees-and-knows-everything-a-cross-contextual-and-cross-c': ['OUT','','Survey; smart TV is a free-text mention count.'], 'no-privacy-among-spies-assessing-the-functionality-and-insecurity-of-consumer-an': ['OUT','','Android stalkerware; "ACR" homonym.'], 'non-intrusive-and-unconstrained-keystroke-inference-in-vr-platforms-via-infrared-side-channel': ['OUT','','VR; smart TV appears only as a citation to HomeSpy.'], 'nothing-else-mator-s-monitoring-the-anonymity-of-tors-path-selection': ['OUT','','"ACR" homonym.'], 'on-the-vulnerability-of-fpga-bitstream-encryption-against-power-analysis-attacks': ['OUT','','Set-top box named as an FPGA application domain.'], 'privacy-perspectives-and-practices-of-chinese-smart-home-product-teams': ['OUT','','Interview study; smart TV is a company product-line row.'], 'regulating-smart-device-support-periods-user-expectations-and-the-european-cyber': ['OUT','','Survey; Smart TV is a self-reported ownership category.'], 'rocking-drones-with-intentional-sound-noise-on-gyroscopic-sensors': ['OUT','','Passing mention.'], 'same-origin-policy-evaluation-in-modern-browsers': ['OUT','','Passing mention of TV browsers.'], 'sandscout-automatic-detection-of-flaws-in-ios-sandbox-profiles': ['OUT','','Apple TV named as a device that runs iOS/tvOS; iOS sandbox is the object.'], 'sissi-an-architecture-for-semantic-interoperable-self-sovereign-identity-based-a': ['OUT','','"ACR" homonym (authentication context reference).'], 'smart-devices-in-airbnbs-considering-privacy-and-security-for-both-guests-and-ho': ['OUT','','Survey; smart TV is a scenario option.'], 'snapshot-based-loading-acceleration-of-web-apps-with-nondeterministic-javascript': ['OUT','','Tizen/webOS named as embedded web-app platforms; benchmark is web apps.'], 'sok-technical-implementation-and-human-impact-of-internet-privacy-regulations': ['OUT','','SoK; TV work cited, not measured.'], 'tbtrackerx-fantastic-trigger-bots-and-where-to-find-malicious-campaigns-on-x': ['OUT','','IPTV spam homonym.'], 'webspec-towards-machine-checked-analysis-of-browser-security-mechanisms': ['OUT','','"ACR" homonym.'], 'what-happened-in-my-network-mining-network-events-from-router-syslogs': ['OUT','','IPTV named as the service carried; router syslogs are the object.'], 'who-cares-contextual-privacy-judgments-from-owner-and-bystander-perspectives-in': ['OUT','','Survey; smart TV is a device-category option.'], 'why-can-t-users-choose-their-identity-providers-on-the-web': ['OUT','','"ACR" homonym.'], 'you-are-who-you-know-and-how-you-behave-attribute-inference-attacks-via-users-so': ['OUT','','IPTV homonym.'], 'poster-watch-out-your-smart-watch-when-paired': ['OUT','','Tizen here is the smartwatch platform, not the TV one.'], 'cleaning-up-the-internet-of-evil-things-real-world-evidence-on-isp-and-consumer-efforts-to-remove-mirai': ['OUT','','One infected set-top box in a Mirai remediation table; no TV finding.'], 'latex-gloves-protecting-browser-extensions-from-probing-and-revelation-attacks': ['OUT','','The Chromecast browser EXTENSION, not the device.'], 'abuse-vectors-a-framework-for-conceptualizing-iot-enabled-interpersonal-abuse': ['OUT','','Qualitative framework; TV is an example abuse vector.'], 'exploring-tenants-preferences-of-privacy-negotiation-in-airbnb': ['OUT','','Vignette survey; smart TV is a device-type option.'], 'pangolin-fuzzing-multilingual-iot-firmware-with-llm-driven-code-analysis': ['OUT','','"SmartTVs" is a citation to the 2021 fuzzing paper, used as a baseline name.'], 'utopia-automatic-generation-of-fuzz-driver-using-unit-tests': ['OUT','','Tizen as an open-source project under test; no TV device.'], }; export const tier = (t) => Object.entries(MAP).filter(([, v]) => v[0] === t).map(([k]) => k);
- report_connected_tv.mjs
// Report script for `design:connected_tv`. // // Every figure on that page is printed here with its own denominator. The page // carries no number this script cannot produce. // // Structure: // 0. Rebuild the candidate pool from the raw corpus and assert it still // matches scripts/ctv_fold.mjs exactly, in both directions. // 1. The population, by tier, venue, year and topic. // 2. Why the roadmap's `web` column was the wrong filter: platform fields. // 3. How these papers get at the traffic (tools, interception, vantage). // 4. What they sample (population sources, n, units) — the no-Tranco problem. // 5. Measured results, quoted from detection[].prevalence. // 6. Where the field goes quiet, on the CTV population vs the corpus. // 7. Quote verification against paper.cols.txt. // // Usage: node scripts/report_connected_tv.mjs // node scripts/report_connected_tv.mjs --format wiki import fs from 'node:fs'; import path from 'node:path'; import { createHash } from 'node:crypto'; import { loadExtractions, dataRoot, isSentinel, pct, table, wikiTable } from './lib.mjs'; import { MAP } from './ctv_fold.mjs'; const ROOT = dataRoot(); const WIKI = process.argv.includes('--format') && process.argv[process.argv.indexOf('--format') + 1] === 'wiki'; const H = (s) => console.log('\n' + '='.repeat(78) + '\n' + s + '\n' + '='.repeat(78)); const T = (h, r) => console.log(WIKI ? wikiTable(h, r) : table(h, r)); const ALL = loadExtractions(); const ftPath = (p) => path.join(ROOT, 'fulltext', String(p.year), p.venue, p.slug, 'paper.cols.txt'); const collapse = (s) => s.replace(//g, '').replace(/-\n/g, '').replace(/\s+/g, ' '); const ftCache = new Map(); const fulltext = (p) => { if (!ftCache.has(p.slug)) { const f = ftPath(p); ftCache.set(p.slug, fs.existsSync(f) ? collapse(fs.readFileSync(f, 'utf8')) : ''); } return ftCache.get(p.slug); }; // --------------------------------------------------------------------------- // 0. Rebuild the candidate pool and check the hand map against it. // --------------------------------------------------------------------------- const PROBES = { smarttv: /\bsmart[-\s]?TVs?\b/i, ctv: /\bconnected[-\s]TVs?\b|\bCTV\b/i, ott: /\bover[-\s]the[-\s]top\b|\bOTT\b/i, hbbtv: /\bHbbTV\b|\bhybrid broadcast broadband\b/i, acr: /\bautomatic content recognition\b|\bACR\b/i, platformdev: /\bRoku\b|\bFire ?TV\b|\bApple ?TV\b|\bChromecast\b|\bAndroid ?TV\b|\bGoogle ?TV\b|\btvOS\b|\bWebOS\b|\bTizen\b|\bset[-\s]?top box(es)?\b/i, streamsvc: /\bNetflix\b|\bHulu\b|\bDisney\+|\bAmazon Prime Video\b|\bYouTube ?TV\b|\bTwitch\b/i, tvapp: /\bTV app(s|lication)?\b|\btelevision app(s)?\b/i, iptv: /\bIPTV\b|\binternet protocol television\b/i, }; // The roadmap's own title+summary probe, reproduced verbatim from // scripts/gap_probe_roadmap.mjs so the two can be compared. const ROADMAP_TITLE = /smart ?TV|connected TV|\bCTV\b|roku|streaming (device|platform|service)|set-?top box/i; // Device-name probe used to find TVs inside broader IoT device sets. const DEV = /\b(Roku|Fire ?TV|Apple ?TV|Chromecast|Android ?TV|Google ?TV|tvOS|WebOS|Tizen|Vizio|Hisense|Bravia|Nvidia Shield|Samsung(?: Smart)? TV|LG(?: Smart)? TV|TCL|smart[- ]?TVs?|set[- ]?top box(?:es)?)\b/gi; let noFulltext = 0; const scored = []; for (const p of ALL) { const txt = fulltext(p); if (!txt) { noFulltext += 1; continue; } const c = {}; for (const [k, re] of Object.entries(PROBES)) c[k] = (txt.match(new RegExp(re.source, re.flags + 'g')) || []).length; const core = c.smarttv + c.ctv + c.hbbtv + c.platformdev + c.tvapp; const brands = new Set((txt.match(DEV) || []).map((x) => x.toLowerCase().replace(/\s+/g, ''))); const devN = (txt.match(DEV) || []).length; const titleHit = ROADMAP_TITLE.test(`${p.title ?? ''} • ${p.summary ?? ''}`); scored.push({ p, c, core, brands, devN, titleHit }); } // GATE 1 (wide): anything with a real amount of TV vocabulary anywhere. const gate1 = scored.filter((s) => s.core >= 2 || s.c.acr >= 2 || s.c.iptv >= 2 || s.c.hbbtv >= 1 || s.c.tvapp >= 1 || s.titleHit); // GATE 2 (audit set): gate 1 narrowed to what is worth reading in full. const AUDIT = gate1.filter((s) => s.devN >= 4 || s.brands.size >= 3 || s.titleHit || s.c.hbbtv > 0 || s.c.acr >= 2 || s.c.iptv >= 2); const auditSlugs = new Set(AUDIT.map((s) => s.p.slug)); const mapSlugs = new Set(Object.keys(MAP)); const missingFromMap = [...auditSlugs].filter((s) => !mapSlugs.has(s)); const missingFromAudit = [...mapSlugs].filter((s) => !auditSlugs.has(s)); if (missingFromMap.length) throw new Error(`audit set has ${missingFromMap.length} slug(s) with no verdict in ctv_fold.mjs:\n ${missingFromMap.join('\n ')}`); if (missingFromAudit.length) throw new Error(`ctv_fold.mjs has ${missingFromAudit.length} slug(s) the audit set no longer selects:\n ${missingFromAudit.join('\n ')}`); const bySlug = new Map(ALL.map((p) => [p.slug, p])); const verdict = (t) => Object.entries(MAP).filter(([, v]) => v[0] === t).map(([s]) => ({ p: bySlug.get(s), v: MAP[s] })) .sort((a, b) => a.p.year - b.p.year || a.p.venue.localeCompare(b.p.venue)); const A = verdict('A'), B = verdict('B'), ADJ = verdict('ADJ'), OUT = verdict('OUT'); const POP = [...A, ...B]; // the page's population const nA = A.length, nB = B.length, nPOP = POP.length; // The slug-set check above only catches a slug appearing or disappearing. A // mutation test on 2026-09-12 showed that FLIPPING a tier letter — 'B' to 'OUT' // on a paper that stays in the candidate set — passed every assertion and // silently moved the population from 35 to 34, changing every percentage on the // page. So the split itself is pinned to what design:connected_tv publishes. // Changing the map deliberately means changing these four numbers too. const PUBLISHED_SPLIT = { A: 13, B: 22, ADJ: 16, OUT: 52 }; const actualSplit = { A: nA, B: nB, ADJ: ADJ.length, OUT: OUT.length }; for (const k of Object.keys(PUBLISHED_SPLIT)) { if (actualSplit[k] !== PUBLISHED_SPLIT[k]) throw new Error( `verdict split moved: ${k} is ${actualSplit[k]}, design:connected_tv publishes ${PUBLISHED_SPLIT[k]}. ` + `Full split now ${JSON.stringify(actualSplit)} against published ${JSON.stringify(PUBLISHED_SPLIT)}. ` + `If the map change is intended, update PUBLISHED_SPLIT and every figure on the page.` ); } const splitTotal = nA + nB + ADJ.length + OUT.length; if (splitTotal !== AUDIT.length) throw new Error(`verdicts sum to ${splitTotal} but the audit set is ${AUDIT.length}`); // A second mutation test, on 2026-09-13, broke the four counts above: SWAPPING // two papers between tiers — one genuine Tier B out, one genuine OUT in — // leaves A/B/ADJ/OUT unchanged and exits 0, while the population silently gains // a paper that is not about television (platform 'web' went 1 -> 2, the venue // and year tables both moved). Counts are not membership. So the membership // itself is pinned, per tier, as a digest of the sorted slug list. const digest = (rows) => createHash('sha256') .update(rows.map((x) => x.p.slug).sort().join('\n')).digest('hex').slice(0, 16); const PUBLISHED_MEMBERS = { A: 'b7410e5f7a33a91e', B: '4a71c82990115cdf', ADJ: '1eb63ba64295cda8', OUT: '22530ed51f0ceb35', }; const actualMembers = { A: digest(A), B: digest(B), ADJ: digest(ADJ), OUT: digest(OUT) }; for (const k of Object.keys(PUBLISHED_MEMBERS)) { if (actualMembers[k] !== PUBLISHED_MEMBERS[k]) throw new Error( `tier ${k} membership changed (count is still ${actualSplit[k]}, so the count check above could not see it). ` + `digest is ${actualMembers[k]}, design:connected_tv was published against ${PUBLISHED_MEMBERS[k]}. ` + `Full digests now ${JSON.stringify(actualMembers)}. If the map change is intended, update ` + `PUBLISHED_MEMBERS and re-check every per-venue, per-year, per-platform and per-topic figure on the page.` ); } // The topic tag has no count to pin — it is hand-assigned and published as a // ranking. It still drives the "Fifteen of the 35 papers are IoT studies" list, // so its membership is pinned the same way. const topicDigest = createHash('sha256').update( Object.entries(MAP).filter(([, v]) => v[0] === 'A' || v[0] === 'B') .map(([s, v]) => `${s}\t${v[1]}`).sort().join('\n')).digest('hex').slice(0, 16); const PUBLISHED_TOPICS = 'fdb92e2a4c3ab263'; if (topicDigest !== PUBLISHED_TOPICS) throw new Error( `topic assignments for the 35-paper population changed: digest ${topicDigest}, published against ${PUBLISHED_TOPICS}. ` + `Update PUBLISHED_TOPICS and re-check the "What this literature measures" table and the iot-device-set list on the page.` ); H('0. CANDIDATE POOL AND AUDIT'); console.log(`corpus papers with full text scanned : ${scored.length} (no paper.cols.txt: ${noFulltext})`); console.log(`roadmap title+summary probe alone : ${scored.filter((s) => s.titleHit).length}`); console.log(` ...of which carry platform 'web' : ${scored.filter((s) => s.titleHit && s.p.platforms.includes('web')).length}`); console.log(`gate 1, any TV vocabulary : ${gate1.length}`); console.log(`gate 2, the hand-audited set : ${AUDIT.length}`); console.log(` A television is the study object : ${nA}`); console.log(` B television inside a device set : ${nB}`); console.log(` ADJ adjacent, cited not counted : ${ADJ.length}`); console.log(` OUT off topic : ${OUT.length}`); console.log(`POPULATION (A+B) : ${nPOP}`); console.log(`precision of the audit set : ${pct(nPOP, AUDIT.length)}`); console.log(`precision of the roadmap's 16 : ${pct(AUDIT.filter((s) => s.titleHit && MAP[s.p.slug] && ['A','B'].includes(MAP[s.p.slug][0])).length, scored.filter((s) => s.titleHit).length)}`); const titleSet = new Set(scored.filter((s) => s.titleHit).map((s) => s.p.slug)); console.log(`population papers the roadmap probe MISSES: ${POP.filter((x) => !titleSet.has(x.p.slug)).length} of ${nPOP}`); H('0b. THE POPULATION, PAPER BY PAPER'); T(['Tier', 'Year', 'Venue', 'Topic', 'Title'], POP.map((x) => [x.v[0], x.p.year, x.p.venue, x.v[1], x.p.title.replace(/\.$/, '')])); H('0c. ADJACENT AND OUT — the audit trail for what was NOT counted'); T(['V', 'Year', 'Venue', 'Title', 'Reason'], [...ADJ, ...OUT].map((x) => [x.v[0], x.p.year, x.p.venue, x.p.title.slice(0, 62).replace(/\.$/, ''), x.v[2]])); // --------------------------------------------------------------------------- // 1. Shape of the population // --------------------------------------------------------------------------- H('1. THE POPULATION BY VENUE, YEAR AND TOPIC'); const venues = [...new Set(ALL.map((p) => p.venue))].sort(); T(['Venue', 'A', 'B', 'A+B', 'Venue papers in corpus', 'Share of venue'], venues.map((v) => { const a = A.filter((x) => x.p.venue === v).length, b = B.filter((x) => x.p.venue === v).length; const tot = ALL.filter((p) => p.venue === v).length; return [v, a, b, a + b, tot, pct(a + b, tot)]; }).sort((x, y) => y[3] - x[3])); const YB = [['2010–2014', 2010, 2014], ['2015–2018', 2015, 2018], ['2019–2021', 2019, 2021], ['2022–2024', 2022, 2024], ['2025–2026*', 2025, 2026]]; T(['Window', 'A', 'B', 'A+B'], YB.map(([lab, lo, hi]) => [lab, A.filter((x) => x.p.year >= lo && x.p.year <= hi).length, B.filter((x) => x.p.year >= lo && x.p.year <= hi).length, POP.filter((x) => x.p.year >= lo && x.p.year <= hi).length])); console.log('* 2025–2026 is provisional: CCS 2026 and IMC 2026 have not been held, and IEEE S&P / WWW 2026 are incompletely selected. See literature:corpus.'); const topics = [...new Set(POP.map((x) => x.v[1]))]; T(['Topic (hand-assigned, ranking only)', 'A', 'B', 'A+B'], topics.map((t) => [t, A.filter((x) => x.v[1] === t).length, B.filter((x) => x.v[1] === t).length, POP.filter((x) => x.v[1] === t).length]).sort((a, b) => b[3] - a[3])); // --------------------------------------------------------------------------- // 2. Platform fields — the roadmap's web column // --------------------------------------------------------------------------- H('2. PLATFORM FIELDS: WHY `web` IS THE WRONG FILTER'); const plats = ['web', 'mobile', 'iot', 'other-online-service', 'offline', 'not-applicable']; T(['platforms[] value', `A+B (N=${nPOP})`, 'share', `corpus (N=${ALL.length})`, 'share'], plats.map((v) => { const n = POP.filter((x) => x.p.platforms.includes(v)).length; const m = ALL.filter((p) => p.platforms.includes(v)).length; return [v, n, pct(n, nPOP), m, pct(m, ALL.length)]; })); console.log(`\nPapers in the population carrying platform 'web' : ${POP.filter((x) => x.p.platforms.includes('web')).length} of ${nPOP}`); console.log('Which paper(s):', POP.filter((x) => x.p.platforms.includes('web')).map((x) => `${x.p.venue} ${x.p.year} ${x.p.title.slice(0, 50)}`).join(' | ') || '(none)'); console.log(`\npopulation[].unit == 'iot-devices' : ${POP.filter((x) => x.p.population.some((q) => q.unit === 'iot-devices')).length} of ${nPOP}`); console.log(`population[].unit == 'mobile-apps' : ${POP.filter((x) => x.p.population.some((q) => q.unit === 'mobile-apps')).length} of ${nPOP}`); console.log(`population[].unit == 'websites' : ${POP.filter((x) => x.p.population.some((q) => q.unit === 'websites')).length} of ${nPOP}`); const CRAWLED = (p) => p.crawlConfig !== null || p.studyTypes.includes('automated-web-crawl'); console.log(`\nIn the corpus-wide 'crawled' population (crawlConfig or automated-web-crawl): ${POP.filter((x) => CRAWLED(x.p)).length} of ${nPOP}`); console.log('Which:', POP.filter((x) => CRAWLED(x.p)).map((x) => `${x.p.venue} ${x.p.year}`).join(', ') || '(none)'); // --------------------------------------------------------------------------- // 3. Instrumentation // --------------------------------------------------------------------------- H('3. HOW THESE PAPERS GET AT THE TRAFFIC'); const TOOLCATS = ['proxy-interception', 'traffic-capture', 'mobile-instrumentation', 'program-analysis', 'crawler-framework', 'browser-automation', 'network-scanner', 'blocklist']; T(['tool category', `A (N=${nA})`, `A+B (N=${nPOP})`], TOOLCATS.map((c) => [c, A.filter((x) => x.p.tools.some((t) => t.category === c && t.usedOrMentioned === 'used')).length, POP.filter((x) => x.p.tools.some((t) => t.category === c && t.usedOrMentioned === 'used')).length])); // Named instruments, folded by lower-cased alphanumeric skeleton. const skel = (s) => s.toLowerCase().replace(/[^a-z0-9]+/g, ''); const ALIAS = { mitmproxy: 'mitmproxy', mitmdump: 'mitmproxy', mitmweb: 'mitmproxy', wireshark: 'Wireshark', tshark: 'Wireshark', tcpdump: 'tcpdump', frida: 'Frida', adb: 'adb', androiddebugbridge: 'adb', charles: 'Charles Proxy', charlesproxy: 'Charles Proxy', pihole: 'Pi-hole', virustotal: 'VirusTotal', apktool: 'apktool', jadx: 'jadx', flowdroid: 'FlowDroid', libscout: 'LibScout', easylist: 'EasyList', scapy: 'Scapy', selenium: 'Selenium', openwpm: 'OpenWPM', mercury: 'Mercury', pingpong: 'PingPong', appium: 'Appium', raspberrypi: 'Raspberry Pi', monkey: 'UI/Application Exerciser Monkey', uiautomator: 'UIAutomator' }; const instr = new Map(); const residue = new Map(); for (const x of POP) for (const t of x.p.tools) { if (t.usedOrMentioned !== 'used') continue; const k = skel(t.name); const disp = ALIAS[k]; const target = disp ? instr : residue; const key = disp ?? t.name; if (!target.has(key)) target.set(key, new Set()); target.get(key).add(x.p.slug); } T(['instrument (alias-folded)', `papers (of ${nPOP})`], [...instr.entries()].map(([k, v]) => [k, v.size]).sort((a, b) => b[1] - a[1])); console.log(`\nUNMAPPED RESIDUE of the alias fold — ${residue.size} distinct raw tool names used by the ${nPOP} papers, printed in full:`); console.log([...residue.entries()].sort((a, b) => b[1].size - a[1].size || a[0].localeCompare(b[0])).map(([k, v]) => `${v.size}× ${k}`).join(' | ')); console.log('\nInterception evidence, per Tier A paper (full text, whitespace-collapsed):'); const IPROBE = { 'mitm/proxy': /\bmitmproxy\b|\bmitm\b|\bman-in-the-middle\b|\bCharles\b|\bproxy\b/i, 'own CA / root cert': /\broot certificate\b|\bcustom (?:CA|certificate)\b|\binstall(?:ed|ing)? (?:our|a) (?:own )?(?:CA|certificate|root)\b|\bCA certificate\b/i, 'pinning / cert failure': /\bcertificate pinning\b|\bpinn(?:ing|ed)\b|\bcertificate validation\b/i, 'undecryptable reported': /\bcould not (?:be )?decrypt\b|\bfail(?:ed|s|ure) to decrypt\b|\bdecryption fail\w*\b|\bunable to (?:decrypt|intercept)\b/i, 'DNS-level capture': /\bDNS (?:queries|traffic|resolution|logs)\b|\bPi-hole\b|\bdnsmasq\b/i, 'router/AP capture': /\b(?:wireless )?access point\b|\bRaspberry Pi\b|\brouter\b|\btcpdump\b|\bWireshark\b/i, 'HDMI / screen capture': /\bHDMI\b|\bscreen ?(?:shot|capture|recording)\b|\bframe grabber\b/i, 'remote-control automation': /\bADB\b|\bremote control\b|\bIR blaster\b|\bkey ?events?\b|\bExternal Control\b/i, }; T(['paper', ...Object.keys(IPROBE)], A.map((x) => { const t = fulltext(x.p); return [`${x.p.venue} ${x.p.year}`, ...Object.values(IPROBE).map((re) => ((t.match(new RegExp(re.source, re.flags + 'g')) || []).length || '—'))]; })); for (const [lab, re] of Object.entries(IPROBE)) { const n = A.filter((x) => re.test(fulltext(x.p))).length; console.log(`${lab.padEnd(26)} ${n} of ${nA} Tier A | ${POP.filter((x) => re.test(fulltext(x.p))).length} of ${nPOP} A+B`); } H('3b. VANTAGE'); const locs = new Map(); for (const x of POP) for (const v of x.p.vantage) for (const l of v.locations) { if (isSentinel(l)) continue; const k = l.trim(); if (!locs.has(k)) locs.set(k, new Set()); locs.get(k).add(x.p.slug); } T(['vantage location (verbatim, unfolded)', 'papers'], [...locs.entries()].map(([k, v]) => [k, v.size]).sort((a, b) => b[1] - a[1])); const withV = POP.filter((x) => x.p.vantage.length > 0); const statedV = withV.filter((x) => x.p.vantage.some((v) => v.locations.some((l) => !isSentinel(l)))); console.log(`\npapers with a vantage tuple: ${withV.length} of ${nPOP}; of those, stating a location: ${statedV.length} (${pct(statedV.length, withV.length)})`); // --------------------------------------------------------------------------- // 4. Sampling // --------------------------------------------------------------------------- H('4. WHAT THEY SAMPLE — THE NO-TRANCO PROBLEM'); const srcs = new Map(); for (const x of POP) for (const q of x.p.population) { if (isSentinel(q.sourceList)) continue; const k = q.sourceList.trim(); if (!srcs.has(k)) srcs.set(k, new Set()); srcs.get(k).add(x.p.slug); } T(['population.sourceList (verbatim, unfolded — ranking only)', 'papers'], [...srcs.entries()].map(([k, v]) => [k.slice(0, 58), v.size]).sort((a, b) => b[1] - a[1])); T(['population.unit', 'papers'], [...new Set(POP.flatMap((x) => x.p.population.map((q) => q.unit)))] .map((u) => [u, POP.filter((x) => x.p.population.some((q) => q.unit === u)).length]).sort((a, b) => b[1] - a[1])); T(['population.samplingMethod', 'papers'], [...new Set(POP.flatMap((x) => x.p.population.map((q) => q.samplingMethod)))] .map((m) => [m, POP.filter((x) => x.p.population.some((q) => q.samplingMethod === m)).length]).sort((a, b) => b[1] - a[1])); // Labelled by slug, not `venue year`: there are two IMC 2011 papers in this // population and only one of them has iot-devices tuples, so a venue+year label // is ambiguous exactly where a reader would want to check it. const devSets = POP.map((x) => ({ k: `${x.p.venue} ${x.p.year} ${x.p.slug.slice(0, 38)}`, ns: x.p.population.filter((q) => q.unit === 'iot-devices' && q.n !== null).map((q) => q.n), })).filter((r) => r.ns.length); T(['paper', 'iot-devices n values stated'], devSets.map((r) => [r.k, r.ns.join(', ')])); // The page states this distribution in prose, so it is derived here rather than // eyeballed off the table above. A hand-typed version of this list shipped in // the first draft and omitted 57 — the PETS 2020 testbed, the population's own // flagship paper. const allNs = devSets.flatMap((r) => r.ns).sort((a, b) => a - b); const largestPer = devSets.map((r) => Math.max(...r.ns)).sort((a, b) => a - b); const median = (arr) => (arr.length % 2 ? arr[(arr.length - 1) / 2] : (arr[arr.length / 2 - 1] + arr[arr.length / 2]) / 2); console.log(`\npapers stating an iot-devices size : ${devSets.length} of ${nPOP}`); console.log(`stated size values : ${allNs.length}`); console.log(`all values, sorted : ${allNs.join(', ')}`); console.log(`median of all stated values : ${median(allNs)}`); console.log(`largest set per paper, sorted : ${largestPer.join(', ')}`); console.log(`median of largest-per-paper : ${median(largestPer)}`); for (const cut of [100, 200, 1000]) { console.log(`papers whose LARGEST set <= ${String(cut).padStart(4)} : ${largestPer.filter((v) => v <= cut).length} of ${devSets.length}`); } console.log('papers whose largest set is over 200, i.e. not a lab bench:'); for (const r of devSets.filter((r) => Math.max(...r.ns) > 200)) console.log(` ${r.k} -> ${r.ns.join(', ')}`); const listv = POP.flatMap((x) => x.p.population).filter((q) => q.listVersion !== null).length; console.log(`\npopulation tuples in A+B stating a listVersion: ${listv} of ${POP.flatMap((x) => x.p.population).length}`); // --------------------------------------------------------------------------- // 5. Measured results // --------------------------------------------------------------------------- H('5. MEASURED RESULTS (detection[].prevalence, Tier A only)'); let dt = 0, dp = 0; for (const x of A) { const withPrev = x.p.detection.filter((d) => d.prevalence !== null); dt += x.p.detection.length; dp += withPrev.length; console.log(`\n--- ${x.p.venue} ${x.p.year} | ${x.p.title}`); for (const d of withPrev) console.log(` * ${d.phenomenon}\n technique : ${d.technique}\n metric : ${d.metric}\n prevalence: ${d.prevalence}\n quote (${d.evidence.section}): ${d.evidence.quote}`); } console.log(`\ndetection tuples in Tier A: ${dt}; carrying a prevalence: ${dp} (${pct(dp, dt)})`); const dpAll = ALL.flatMap((p) => p.detection); console.log(`corpus-wide: ${dpAll.length} tuples, ${dpAll.filter((d) => d.prevalence !== null).length} carry a prevalence (${pct(dpAll.filter((d) => d.prevalence !== null).length, dpAll.length)})`); // --------------------------------------------------------------------------- // 6. Reporting gaps // --------------------------------------------------------------------------- H('6. WHERE THIS LITERATURE GOES QUIET'); const CRAWL_CORPUS = ALL.filter(CRAWLED); const crawlPOP = POP.filter((x) => x.p.crawlConfig !== null); // `ethics` and `crawlConfig` are nullable by schema: a paper with no crawl has // no crawlConfig, and 894 corpus papers carry no ethics object at all. Each // block below names the population that HAS the object, so a sentinel is never // counted as an answer and a missing object is never counted as a sentinel. T(['crawlConfig field (papers with a crawlConfig object)', `A+B (N=${crawlPOP.length})`, `corpus crawled (N=${CRAWL_CORPUS.filter((p) => p.crawlConfig !== null).length})`], [['consentAction', 'interactionDepth', 'statefulness', 'browsers']].flat().map((f) => { const cc = CRAWL_CORPUS.filter((p) => p.crawlConfig !== null); const st = (p) => (f === 'browsers' ? p.crawlConfig.browsers.length > 0 : !isSentinel(p.crawlConfig[f])); const a = crawlPOP.filter((x) => st(x.p)).length, b = cc.filter(st).length; return [f, `${a} of ${crawlPOP.length} (${pct(a, crawlPOP.length)})`, `${b} of ${cc.length} (${pct(b, cc.length)})`]; })); console.log('\ncrawlConfig values in the A+B population, verbatim:'); for (const x of crawlPOP) console.log(` ${x.p.venue} ${x.p.year} consent=${x.p.crawlConfig.consentAction} depth=${x.p.crawlConfig.interactionDepth} state=${x.p.crawlConfig.statefulness} browsers=[${x.p.crawlConfig.browsers.join(',')}] robotsTxt=${x.p.ethics === null ? '(no ethics object)' : x.p.ethics.robotsTxt}`); const ethPOP = POP.filter((x) => x.p.ethics !== null); const ethEmpCorpus = ALL.filter((p) => p.isEmpirical && p.ethics !== null); const emp = POP.filter((x) => x.p.isEmpirical); const EMP_CORPUS = ALL.filter((p) => p.isEmpirical); const artPOP = POP.filter((x) => x.p.artifacts !== null); const artEmpCorpus = ALL.filter((p) => p.isEmpirical && p.artifacts !== null); console.log(`\nA+B papers carrying an ethics object: ${ethPOP.length} of ${nPOP}; artifacts object: ${artPOP.length} of ${nPOP}`); T(['field', 'A+B', 'corpus empirical'], [ ['ethics.reviewOutcome stated', `${ethPOP.filter((x) => !isSentinel(x.p.ethics.reviewOutcome)).length} of ${ethPOP.length} (${pct(ethPOP.filter((x) => !isSentinel(x.p.ethics.reviewOutcome)).length, ethPOP.length)})`, `${ethEmpCorpus.filter((p) => !isSentinel(p.ethics.reviewOutcome)).length} of ${ethEmpCorpus.length} (${pct(ethEmpCorpus.filter((p) => !isSentinel(p.ethics.reviewOutcome)).length, ethEmpCorpus.length)})`], ['artifacts.availability stated', `${artPOP.filter((x) => !isSentinel(x.p.artifacts.availability)).length} of ${artPOP.length} (${pct(artPOP.filter((x) => !isSentinel(x.p.artifacts.availability)).length, artPOP.length)})`, `${artEmpCorpus.filter((p) => !isSentinel(p.artifacts.availability)).length} of ${artEmpCorpus.length} (${pct(artEmpCorpus.filter((p) => !isSentinel(p.artifacts.availability)).length, artEmpCorpus.length)})`], ['temporal.spanStart stated', `${emp.filter((x) => x.p.temporal.some((t) => t.spanStart !== null)).length} of ${emp.length} (${pct(emp.filter((x) => x.p.temporal.some((t) => t.spanStart !== null)).length, emp.length)})`, `${EMP_CORPUS.filter((p) => p.temporal.some((t) => t.spanStart !== null)).length} of ${EMP_CORPUS.length} (${pct(EMP_CORPUS.filter((p) => p.temporal.some((t) => t.spanStart !== null)).length, EMP_CORPUS.length)})`], ]); function linkOf(a) { if (a.codeUrl) return a.codeUrl; if (a.dataUrl) return a.dataUrl; if (!a.links || a.links.length === 0) return '—'; const l = a.links[0]; if (typeof l === 'string') throw new Error(`artifacts.links[0] is a string, expected an object: ${l}`); return `${l.url} (${l.kind}; ${l.what}; authors=${l.belongsToAuthors})`; } console.log('\nArtifact links released, Tier A:'); for (const x of A) console.log(x.p.artifacts === null ? ` ${x.p.venue} ${x.p.year} (no artifacts object extracted)` // artifacts.links[] entries are OBJECTS ({url, kind, what, belongsToAuthors}). // Falling back to links[0] itself string-coerced to "[object Object]" and that // literal was published three times on the provenance page \u2014 found by a review // pass on 2026-09-13. Take .url, and print what the link is for. : ` ${x.p.venue} ${x.p.year} ${x.p.artifacts.availability.padEnd(22)} ${linkOf(x.p.artifacts)}`); // --------------------------------------------------------------------------- // --------------------------------------------------------------------------- // 6b. Two probes the page cites in footnotes. Both were added on 2026-09-13 // after review found the page asserting negatives no probe supported. // --------------------------------------------------------------------------- H('6b. TWO PROBES THE PAGE CITES'); // (i) The TLS decryption hole. The page originally ran only NARROW, matched 1 // of 13 Tier A papers, and published "almost nobody reports the hole" — // while quoting two of the sentences NARROW misses. A narrowing probe must // return a SUBSET of the loose one; both counts are printed so the claim // can be read off the right width. const HOLE_NARROW = /could not decrypt|failed to decrypt|decryption fail|unable to (decrypt|intercept)/i; const HOLE_WIDE = new RegExp([ /(could not|cannot|can ?not|unable to|failed to|no way to)[^.]{0,90}(decrypt|intercept|install (our own |a |custom )?(self-signed )?certificat)/.source, /(bypass|remove|disabl\w*)[^.]{0,40}certificate pinning/.source, /certificate pinning checks/.source, ].join('|'), 'i'); const holeN = A.filter((x) => HOLE_NARROW.test(fulltext(x.p))); const holeW = A.filter((x) => HOLE_WIDE.test(fulltext(x.p))); const holeNs = new Set(holeN.map((x) => x.p.slug)); const notContained = holeN.filter((x) => !holeW.some((y) => y.p.slug === x.p.slug)); if (notContained.length) throw new Error(`the narrow decryption probe is not a subset of the wide one: ${notContained.map((x) => x.p.slug).join(', ')}`); if (holeN.length > holeW.length) throw new Error(`narrow probe matched ${holeN.length} > wide ${holeW.length}; a narrowing probe must match fewer`); console.log(`TLS decryption hole, Tier A (n=${A.length}):`); console.log(` narrow probe (the one the page published until 2026-09-13): ${holeN.length}`); console.log(` wide probe (the one the page cites now) : ${holeW.length}`); for (const x of holeW) { const t = fulltext(x.p), m = t.match(HOLE_WIDE), i = t.indexOf(m[0]); console.log(` ${holeNs.has(x.p.slug) ? 'both ' : 'wide '} ${x.p.venue} ${x.p.year} ${x.p.slug}`); console.log(` ...${t.slice(Math.max(0, i - 100), i + 200)}...`); } // (ii) ATSC 3.0 / NextGen TV. The page says no paper measures it. That was an // unprobed negative until now; ATSC was never in PROBES. const ATSC = /\bATSC\b|\bNextGen ?TV\b|\bNext ?Gen ?TV\b/i; const atscHits = ALL.filter((p) => ATSC.test(fulltext(p))); console.log(`\nATSC 3.0 / NextGen TV, full text, all ${ALL.length} corpus papers: ${atscHits.length} match`); for (const p of atscHits) { const t = fulltext(p), m = t.match(ATSC), i = t.indexOf(m[0]); console.log(` ${p.venue} ${p.year} ${p.slug} [${MAP[p.slug] ? MAP[p.slug][0] : 'not in audit set'}]`); console.log(` ...${t.slice(Math.max(0, i - 130), i + 210)}...`); } // (iii) The factory-reset gap. crawlConfig.statefulness is empty for all 6 // papers that have the object; the page needed the full-text picture too. const RESET = /factory[- ]reset/i; const resetHits = POP.filter((x) => RESET.test(fulltext(x.p))); const withCfg = POP.filter((x) => x.p.crawlConfig !== null); console.log(`\nStatefulness: ${withCfg.length} of ${nPOP} papers have a crawlConfig object;`); console.log(` statefulness values among them: ${JSON.stringify(withCfg.map((x) => x.p.crawlConfig.statefulness))}`); console.log(` full-text "factory reset" anywhere in the ${nPOP}: ${resetHits.length}`); for (const x of resetHits) { const t = fulltext(x.p), i = t.search(RESET); console.log(` ${MAP[x.p.slug][0]} ${x.p.venue} ${x.p.year} ${x.p.slug}`); console.log(` ...${t.slice(Math.max(0, i - 130), i + 210)}...`); } // (iv) LLM use inside the population. The page said "not one paper uses an LLM // for anything"; two do, and both tool names are in the printed residue. console.log('\nLLM and language-model tools used by the population (tools[].category, used only):'); for (const x of POP) { for (const t of x.p.tools) { if (t.usedOrMentioned !== 'used') continue; if (!/^llm$/i.test(t.category) && !/gpt|openai|chatgpt|bert|language model/i.test(t.name)) continue; console.log(` ${x.p.venue} ${x.p.year} category=${t.category.padEnd(22)} ${t.name}`); console.log(` quote: ${JSON.stringify((t.evidence?.quote || '').slice(0, 200))}`); } } // 7. Quote verification lives in scripts/ctv_quotecheck.py, which checks each // quote against BOTH paper.cols.txt and the PDF text layer. Doing it here // against .cols alone produced two false NOTFOUNDs on the first run, one of // which was a real error in the needle and one of which was a de-columning // artefact — see provenance:design:connected_tv. // --------------------------------------------------------------------------- console.log('\n' + '='.repeat(78) + '\n7. QUOTES: see scripts/ctv_quotecheck.py (checks .cols AND the PDF text layer)\n' + '='.repeat(78));
- ctv_quotecheck.py
#!/usr/bin/env python3 """Quote check for design:connected_tv and provenance:design:connected_tv. Every phrase either page puts in quotation marks, or leans on for a figure, is listed here and matched against BOTH renderings of the source paper: data/fulltext/<year>/<venue>/<slug>/paper.cols.txt (de-columned text) data/fulltext/<year>/<venue>/<slug>/paper.pdf (pypdf text layer) Both are needed. `.cols` splices some two-column sentences, and the PDF layer loses reading order elsewhere, so a NOTFOUND in one rendering is not evidence that a paper does not contain the sentence. On the first run of this check two quotes failed against `.cols` alone: one was a de-columning splice (All Things Considered) and the other was a genuinely wrong needle, where two halves of a sentence in the PETS 2020 paper had been glued together into a number the paper does not state. The second is why bare-number needles are avoided here. Usage: python3 scripts/ctv_quotecheck.py """ import re import sys import unicodedata from pathlib import Path ROOTS = [Path('/workspace/publications_dataset/data/fulltext'), Path('/workspace/publications_dataset/fulltext')] ROOT = next(r for r in ROOTS if r.is_dir()) def norm(s: str) -> str: s = unicodedata.normalize('NFKD', s) s = s.replace('', '').replace('-\n', '') s = re.sub(r'[‘’]', "'", s) s = re.sub(r'[“”]', '"', s) return re.sub(r'\s+', ' ', s).lower() # (year, venue, slug, needle) QUOTES = [ ('2019', 'CCS', 'watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices', 'present on 69% of Roku channels and 89% of Amazon Fire TV channels'), ('2019', 'CCS', 'watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices', 'we were able to install our own cert on the device which allowed us to intercept HTTPS requests on 957 of the 1000 channels'), ('2019', 'CCS', 'watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices', 'that leaked the title of the video to a tracking domain'), ('2020', 'PETS', 'the-tv-is-smart-and-full-of-trackers-measuring-smart-tv-advertising-and-tracking', '1 out of 5 (or fewer) TLS connections for 80% of all apps'), ('2020', 'PETS', 'the-tv-is-smart-and-full-of-trackers-measuring-smart-tv-advertising-and-tracking', '314 ATS domains that are unique to the Roku dataset'), ('2022', 'PETS', 'fingerprintv-fingerprinting-smart-tv-apps', 'among 80 apps that are made available on all three smart TV platforms, 76% exhibit a different fingerprint on each platform'), ('2022', 'PETS', 'watch-over-your-tv-a-security-and-privacy-analysis-of-the-android-tv-ecosystem', 'The analysis found at least one sensitive data flow in 78% of the files'), ('2023', 'NDSS', 'i-still-know-what-you-watched-last-sunday-privacy-of-the-hbbtv-protocol-in-the-european-smart-tv-landscape', '26 communicate with trackers before the user has expressed their consent'), ('2023', 'NDSS', 'i-still-know-what-you-watched-last-sunday-privacy-of-the-hbbtv-protocol-in-the-european-smart-tv-landscape', 'only block at maximum 44% in 2021 and 81% in 2022'), ('2024', 'IMC', 'watching-tv-with-the-second-party-a-first-look-at-automatic-content-recognition', 'there is a complete absence of communication with any previously identified ACR domains'), ('2024', 'IMC', 'watching-tv-with-the-second-party-a-first-look-at-automatic-content-recognition', 'smart TVs in the UK and the US contact distinct ACR domains'), ('2024', 'IMC', 'watching-tv-with-the-second-party-a-first-look-at-automatic-content-recognition', 'ACR network traffic exists when watching linear TV and when using smart TV as an external display using HDMI'), ('2019', 'USENIX', 'all-things-considered-an-analysis-of-iot-devices-on-home-networks', 'the most popular vendor, Roku, only accounts for 17.4% of media devices'), ('2018', 'IMC', 'understanding-video-management-planes', 'streaming set-top boxes1 dominate by view-hours'), ('2023', 'IMC', 'in-the-room-where-it-happens-characterizing-local-communication-and-threats-in-s', 'the analysis of the Smart TV ecosystem is left for future work'), ('2024', 'IEEE-SP', 'surveilling-the-masses-with-wi-fi-based-positioning-systems', 'belong to the streaming television equipment manufacturer Roku'), ('2021', 'IMC', 'iotls-understanding-tls-usage-in-consumer-iot-devices', 'such as voice assistants, smart TVs and video doorbells'), ('2025', 'USENIX', 'watch-out-your-tv-box-reversing-and-blocking-a-p2p-based-illegal-streaming-ecosy', 'they are offered only to those who have purchased specific'), ('2014', 'USENIX', 'from-the-aether-to-the-ethernet-attacking-the-internet-using-broadcast-digital-t', 'which requires a minimal budget and infrastructure'), ('2024', 'IMC', 'iot-bricks-over-v6-understanding-ipv6-usage-in-smart-homes', 'only eight out of 93 devices remain functional'), ('2019', 'CCS', 'watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices', 'On Roku, a total of 43 channels failed to properly verify the server'), ('2019', 'CCS', 'watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices', '794 of the 1000 Roku channels sent at least one request in cleartext'), ('2019', 'CCS', 'watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices', 'We found 9 channels on Roku and 14 channels on the Fire TV'), ('2019', 'CCS', 'watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices', 'an HTTP GET request to "http://ROKU_ DEVICE_IP_ADDRESS:8060/keydown/left " does the same for all Roku devices'), ('2020', 'PETS', 'the-tv-is-smart-and-full-of-trackers-measuring-smart-tv-advertising-and-tracking', '697 Fire TV apps that expose advertising ID alongside serial number and device ID'), ('2022', 'PETS', 'fingerprintv-fingerprinting-smart-tv-apps', '96% (N = 961) of the top'), ('2022', 'PETS', 'watch-over-your-tv-a-security-and-privacy-analysis-of-the-android-tv-ecosystem', '75% of the apps contain analytics libraries and 77% contain advertising libraries'), ('2021', 'USENIX', 'android-smarttvs-vulnerability-discovery-via-log-guided-fuzzing', '37 unique vulnerabilities, including 11 high-impact cyber threats, 10 new memory corruptions, and 16 visual and auditory anomalies'), ('2023', 'USENIX', 'homespy-the-invisible-sniffer-of-infrared-remote-control-of-smart-tvs', 'The accuracy increases to 70% for Top3 and 77% for Top5'), ('2024', 'NDSS', 'acoustic-keystroke-leakage-on-smart-televisions', 'up to 60.19% of common passwords'), ('2024', 'IMC', 'watching-tv-with-the-second-party-a-first-look-at-automatic-content-recognition', 'the fact that we observe network traffic every 15 seconds suggests'), ('2011', 'IMC', 'understanding-couch-potatoes-measurement-and-modeling-of-interactive-usage-of-ip', 'The average number of set-top boxes provisioned was approximately 3 million over this period'), ('2019', 'USENIX', 'all-things-considered-an-analysis-of-iot-devices-on-home-networks', 'are the most common type of device in seven of the eleven regions'), # --- added 2026-09-13, for claims the round-2 review changed ------------- # The widened TLS-decryption probe: the three Tier A sentences the page now # quotes. The narrow probe matched only the second of these, which is how a # "1 of 13" reached the page while two of the three sat in its own prose. ('2019', 'CCS', 'watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices', 'to bypass certificate pinning'), ('2020', 'PETS', 'the-tv-is-smart-and-full-of-trackers-measuring-smart-tv-advertising-and-tracking', 'we cannot install our own self-signed certificates on the Roku'), ('2022', 'PETS', 'watch-over-your-tv-a-security-and-privacy-analysis-of-the-android-tv-ecosystem', 'there is no way to install custom certificates on Android TV'), # The factory-reset sentence. In .cols this one is SPLICED by the # de-columner — "and the UT-We perform a factory reset of the TV for each # channel analysis 100c HiDes modulator" — so only the first clause is # contiguous there. The PDF layer carries the whole sentence. ('2023', 'NDSS', 'i-still-know-what-you-watched-last-sunday-privacy-of-the-hbbtv-protocol-in-the-european-smart-tv-landscape', 'We perform a factory reset of the TV for each channel analysis'), # AndroZoo: the page said every TV-app paper built its own set. ('2022', 'PETS', 'watch-over-your-tv-a-security-and-privacy-analysis-of-the-android-tv-ecosystem', 'download the last version of each app (as of August 2020) from AndroZoo'), ('2023', 'IMC', 'in-the-room-where-it-happens-characterizing-local-communication-and-threats-in-s', 'we randomly selected 1.5K unique package names from Androzoo'), # The one real ATSC mention in the whole 5,859-paper corpus. ('2014', 'USENIX', 'from-the-aether-to-the-ethernet-attacking-the-internet-using-broadcast-digital-t', 'published a candidate standard for hybrid TV in America'), # Both sides of the opt-out: the paper the page had omitted. ('2019', 'CCS', 'watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices', 'this time enabling the "Limit Ad Tracking" (Roku) and the "Disable Interest-based Ads" (Amazon Fire TV) settings'), ] pdf_cache: dict[Path, str] = {} def pdf_text(path: Path) -> str: if path not in pdf_cache: try: import pypdf reader = pypdf.PdfReader(str(path)) pdf_cache[path] = norm('\n'.join(p.extract_text() or '' for p in reader.pages)) except Exception as exc: # a missing text layer is a result, not a crash pdf_cache[path] = '' print(f' (pypdf failed on {path}: {exc})') return pdf_cache[path] def main() -> int: bad = [] cols_only = pdf_only = both = 0 for year, venue, slug, needle in QUOTES: d = ROOT / year / venue / slug cols = norm((d / 'paper.cols.txt').read_text(encoding='utf8')) n = norm(needle) in_cols = n in cols in_pdf = n in pdf_text(d / 'paper.pdf') where = ('cols+pdf' if in_cols and in_pdf else 'cols only' if in_cols else 'PDF only' if in_pdf else 'NOTFOUND') if in_cols and in_pdf: both += 1 elif in_cols: cols_only += 1 elif in_pdf: pdf_only += 1 else: bad.append((slug, needle)) print(f'{where:<9} {venue} {year} {slug[:44]:<44} "{needle[:66]}"') print(f'\n{len(QUOTES)} quotes: {both} in both renderings, {cols_only} in .cols only, ' f'{pdf_only} in the PDF only, {len(bad)} in neither.') for slug, needle in bad: print(f' NOTFOUND {slug}: "{needle}"') return 1 if bad else 0 if __name__ == '__main__': sys.exit(main())
13. Full report output, unedited
- report_connected_tv-output.txt
============================================================================== 0. CANDIDATE POOL AND AUDIT ============================================================================== corpus papers with full text scanned : 5855 (no paper.cols.txt: 4) roadmap title+summary probe alone : 16 ...of which carry platform 'web' : 1 gate 1, any TV vocabulary : 142 gate 2, the hand-audited set : 103 A television is the study object : 13 B television inside a device set : 22 ADJ adjacent, cited not counted : 16 OUT off topic : 52 POPULATION (A+B) : 35 precision of the audit set : 34.0% precision of the roadmap's 16 : 62.5% population papers the roadmap probe MISSES: 25 of 35 ============================================================================== 0b. THE POPULATION, PAPER BY PAPER ============================================================================== Tier Year Venue Topic Title ---- ---- ------- -------------------- ---------------------------------------------------------------------------------------------------------------- A 2011 IMC delivery-performance Understanding couch potatoes: measurement and modeling of interactive usage of IPTV at large scale A 2011 IMC delivery-performance Q-score: proactive service quality assessment in a large IPTV system A 2014 USENIX broadcast From the Aether to the Ethernet—Attacking the Internet using Broadcast Digital Television A 2019 CCS tracking Watching You Watch: The Tracking Ecosystem of Over-the-Top TV Streaming Devices A 2020 PETS tracking The TV is Smart and Full of Trackers: Measuring Smart TV Advertising and Tracking A 2021 USENIX vulnerability Android SmartTVs Vulnerability Discovery via Log-Guided Fuzzing A 2022 PETS tracking FingerprinTV: Fingerprinting Smart TV Apps A 2022 PETS app-analysis Watch Over Your TV: A Security and Privacy Analysis of the Android TV Ecosystem A 2023 NDSS broadcast I Still Know What You Watched Last Sunday: Privacy of the HbbTV Protocol in the European Smart TV Landscape A 2023 USENIX side-channel HOMESPY: The Invisible Sniffer of Infrared Remote Control of Smart TVs A 2024 IMC acr Watching TV with the Second-Party: A First Look at Automatic Content Recognition Tracking in Smart TVs A 2024 NDSS side-channel Acoustic Keystroke Leakage on Smart Televisions A 2025 USENIX piracy Watch Out Your TV Box: Reversing and Blocking a P2P-based Illegal Streaming Ecosystem B 2018 IMC delivery-performance Understanding Video Management Planes B 2019 IMC iot-device-set Information Exposure From Consumer IoT Devices: A Multidimensional, Network-Informed Measurement Approach B 2019 USENIX iot-device-set All Things Considered: An Analysis of IoT Devices on Home Networks B 2020 IMC iot-device-set A Haystack Full of Needles: Scalable Detection of IoT Devices in the Wild B 2020 NDSS iot-device-set Packet-Level Signatures for Smart Home Devices B 2020 NDSS iot-device-set Et Tu Alexa? When Commodity WiFi Devices Turn into Adversarial Motion Sensors B 2020 USENIX iot-device-set You Are What You Broadcast: Identification of Mobile and IoT Devices from (Public) WiFi B 2021 IMC iot-device-set IoTLS: understanding TLS usage in consumer IoT devices B 2021 PETS iot-device-set Blocking Without Breaking: Identification and Mitigation of Non-Essential IoT Traffic B 2022 CCS app-analysis Understanding IoT Security from a Market-Scale Perspective B 2022 PETS iot-device-set Analyzing the Feasibility and Generalizability of Fingerprinting Internet of Things Devices B 2022 USENIX iot-device-set Lumos: Identifying and Localizing Diverse Hidden IoT Devices in an Unfamiliar Environment B 2023 IMC iot-device-set Behind the Scenes: Uncovering TLS and Server Certificate Practice of IoT Device Vendors in the Wild B 2023 IMC iot-device-set In the Room Where It Happens: Characterizing Local Communication and Threats in Smart Homes B 2024 IEEE-SP device-population Surveilling the Masses with Wi-Fi-Based Positioning Systems B 2024 IMC iot-device-set IoT Bricks Over v6: Understanding IPv6 Usage in Smart Homes B 2024 IMC delivery-performance Characterizing User Platforms for Video Streaming in Broadband Networks B 2024 PETS iot-device-set Connecting the Dots: Tracing Data Endpoints in IoT Devices B 2025 NDSS iot-device-set Evaluating Machine Learning-Based IoT Device Identification Models for Security Applications B 2025 USENIX vulnerability Tracking You from a Thousand Miles Away! Turning a Bluetooth Device into an Apple AirTag Without Root Privileges B 2026 NDSS vulnerability BLERP: BLE Re-Pairing Attacks and Defenses B 2026 USENIX policy-compliance Missing, Present and Conflicting: A Large Scale Analysis of IoT Update Information in the EU Market ============================================================================== 0c. ADJACENT AND OUT — the audit trail for what was NOT counted ============================================================================== V Year Venue Title Reason --- ---- ------- -------------------------------------------------------------- ----------------------------------------------------------------------------------------------------------------------------------------------- ADJ 2011 IMC Measurement and analysis of a large scale commercial mobile in "TV" delivered to mobile handsets, not to a television. ADJ 2012 IMC Watching videos from everywhere: a study of the PPTV mobile Vo Mobile VoD; no television endpoint. ADJ 2013 IMC Analyzing the potential benefits of CDN augmentation strategie CDN augmentation for video workloads; no TV endpoint. ADJ 2013 IMC Peer-assisted content distribution in Akamai netsession Peer-assisted CDN; set-top-box mention is background. ADJ 2016 IMC Performance Characterization of a Commercial Video Streaming S Streaming service performance from browser/CDN vantage; no TV-specific result. ADJ 2016 IMC Anatomy of a Personalized Livestreaming System Livestreaming (Periscope) system measurement; no TV endpoint. ADJ 2017 PETS On the Privacy and Security of the Ultrasound Ecosystem Ultrasonic cross-device tracking (uXDT): beacons emitted by TV adverts and picked up by phone SDKs. The TV is the emitter, never measured. ADJ 2019 WWW Exploiting Diversity in Android TLS Implementations for Mobile Android app traffic classification; TLS-fingerprint method later reused on TV apps. ADJ 2020 USENIX Void: A fast and light voice liveness detection system A Samsung Smart TV is used as a replay LOUDSPEAKER; the TV is apparatus, not the measured object. ADJ 2022 NDSS A Lightweight IoT Cryptojacking Detection Mechanism in Heterog Authors implement their own cryptojacking PoC on an LG webOS TV to test a detector; no deployed-TV population. ADJ 2022 USENIX OVRseen: Auditing Network Traffic and Privacy Policies in Ocul VR headsets; smart TVs used as the comparison ecosystem. Same lab, same pipeline shape. ADJ 2023 IMC Tracking, Profiling, and Ad Targeting in the Alexa Echo Smart Smart speaker study; TVs cited as the comparable prior ecosystem, not measured. The closest methodological sibling. ADJ 2023 PETS Your DRM Can Watch You Too: Exploring the Privacy Implications Widevine EME in browsers and Android; TVs named as another Widevine host, not measured. ADJ 2024 IMC Cost-Saving Streaming: Unlocking the Potential of Alternative Edge-node economics for streaming delivery; no TV endpoint measured. ADJ 2025 PETS Unmasking the Shadows: A Cross-Country Study of Online Trackin Illegal movie streaming WEBSITES crawled with a browser; a web-tracking study, not a TV study. ADJ 2025 USENIX Endangered Privacy: Large-Scale Monitoring of Video Streaming Video identification from encrypted MPEG-DASH traffic. The roadmap listed it as CTV spine; it measures the SERVICE and its traffic, never a TV. OUT 2010 IMC What happened in my network: mining network events from router IPTV named as the service carried; router syslogs are the object. OUT 2011 CCS On the vulnerability of FPGA bitstream encryption against powe Set-top box named as an FPGA application domain. OUT 2011 IMC Broadcast yourself: understanding YouTube uploaders IPTV appears once in related work. OUT 2014 CCS (Nothing else) MATor(s): Monitoring the Anonymity of Tor's Pat "ACR" homonym. OUT 2015 USENIX A Placement Vulnerability Study in Multi-Tenant Public Clouds Title probe matched "streaming"; cloud VM placement. OUT 2015 USENIX Rocking Drones with Intentional Sound Noise on Gyroscopic Sens Passing mention. OUT 2016 CCS SandScout: Automatic Detection of Flaws in iOS Sandbox Profile Apple TV named as a device that runs iOS/tvOS; iOS sandbox is the object. OUT 2016 IMC Entropy/IP: Uncovering Structure in IPv6 Addresses "ACR" homonym. OUT 2016 USENIX You Are Who You Know and How You Behave: Attribute Inference A IPTV homonym. OUT 2017 CCS POSTER: Watch Out Your Smart Watch When Paired Tizen here is the smartwatch platform, not the TV one. OUT 2017 PETS Why can’t users choose their identity providers on the web? "ACR" homonym. OUT 2017 USENIX Same-Origin Policy: Evaluation in Modern Browsers Passing mention of TV browsers. OUT 2017 WWW FLOCK: Combating Astroturfing on Livestreaming Platforms Title probe matched "streaming platform"; astroturfing detection on Twitch-like sites. OUT 2018 CCS Medical Devices are at Risk: Information Security on Diagnosti "ACR" = American College of Radiology. OUT 2019 IEEE-SP Drones' Cryptanalysis - Smashing Cryptography with a Flicker IPTV/ACR homonyms. OUT 2019 NDSS cleaning-up-the-internet-of-evil-things-real-world-evidence-on One infected set-top box in a Mirai remediation table; no TV finding. OUT 2019 NDSS latex-gloves-protecting-browser-extensions-from-probing-and-re The Chromecast browser EXTENSION, not the device. OUT 2019 USENIX A Billion Open Interfaces for Eve and Mallory: MitM, DoS, and tvOS listed among Apple OSes; AWDL is the object. OUT 2019 WWW Snapshot-based Loading Acceleration of Web Apps with Nondeterm Tizen/webOS named as embedded web-app platforms; benchmark is web apps. OUT 2020 CCS iDEA: Static Analysis on the Security of Apple Kernel Drivers tvOS is one of four Apple OSes scanned; no TV-specific result. OUT 2020 PETS Smart Devices in Airbnbs: Considering Privacy and Security for Survey; smart TV is a scenario option. OUT 2022 IMC Deep dive into the IoT backend ecosystem Backend infrastructure; TV mentions are motivation and a citation to FingerprinTV. OUT 2022 PETS A Multi-Region Investigation of the Perceptions and Use of Sma Survey; smart TV is a related-work citation and an ownership option. OUT 2022 PETS Exploring the Privacy Concerns of Bystanders in Smart Homes fr Survey; smart TV is an example in a prompt. OUT 2023 CCS IoTFlow: Inferring IoT Device Behavior at Scale through Static Companion-app analysis; no TV breakout. OUT 2023 IEEE-SP Characterizing Everyday Misuse of Smart Home Devices Survey of 483 people; smart TV is an ownership option, not a measured device. OUT 2023 IEEE-SP WebSpec: Towards Machine-Checked Analysis of Browser Security "ACR" homonym. OUT 2023 IEEE-SP UTopia: Automatic Generation of Fuzz Driver using Unit Tests Tizen as an open-source project under test; no TV device. OUT 2023 PETS No Privacy Among Spies: Assessing the Functionality and Insecu Android stalkerware; "ACR" homonym. OUT 2023 USENIX Examining Consumer Reviews to Understand Security and Privacy Review-text analysis; set-top box is a Mirai product category, no TV measurement. OUT 2023 USENIX Examining Power Dynamics and User Privacy in Smart Technology Interview study; TVs are participant device inventories. OUT 2023 USENIX Internet Service Providers' and Individuals' Attitudes, Barrie Interview and survey; TV is a device-ownership row. OUT 2023 USENIX "It's the Equivalent of Feeling Like You're in Jail”: Lessons Interview study on IPV; TV is a reported abuse vector, not measured. OUT 2023 USENIX Measuring Up to (Reasonable) Consumer Expectations: Providing Vignette survey; Vizio appears only in a news citation. OUT 2023 USENIX Abuse Vectors: A Framework for Conceptualizing IoT-Enabled Int Qualitative framework; TV is an example abuse vector. OUT 2023 USENIX Exploring Tenants' Preferences of Privacy Negotiation in Airbn Vignette survey; smart TV is a device-type option. OUT 2023 WWW SISSI: An Architecture for Semantic Interoperable Self-Soverei "ACR" homonym (authentication context reference). OUT 2024 IEEE-SP SoK: Technical Implementation and Human Impact of Internet Pri SoK; TV work cited, not measured. OUT 2024 PETS A Bilingual Longitudinal Analysis of Privacy Policies Measurin ACR homonym: "ACR" is not automatic content recognition here. OUT 2024 PETS Contextualizing Interpersonal Data Sharing in Smart Homes Vignette survey; "viewing history from your smart TV" is a question stem. OUT 2024 PETS "My Best Friend's Husband Sees and Knows Everything": A Cross- Survey; smart TV is a free-text mention count. OUT 2024 USENIX Co-Designing a Mobile App for Bystander Privacy Protection in Interview study; TV names are participant-reported device inventories. OUT 2025 IEEE-SP Analyzing the iOS Local Network Permission from a Technical an Chromecast is one of four IoT devices used to trigger the permission; no TV result. OUT 2025 IEEE-SP Hey, Your Secrets Leaked! Detecting and Characterizing Secret IPTV homonym in leaked-credential data. OUT 2025 NDSS Non-intrusive and Unconstrained Keystroke Inference in VR Plat VR; smart TV appears only as a citation to HomeSpy. OUT 2025 PETS Help Me Help You: Privacy Considerations for Third Party IoT D Vignette survey; TVs appear in a device-category prompt. OUT 2025 PETS Who Cares? Contextual Privacy Judgments from Owner and Bystand Survey; smart TV is a device-category option. OUT 2025 USENIX Regulating Smart Device Support Periods: User Expectations and Survey; Smart TV is a self-reported ownership category. OUT 2026 IEEE-SP Privacy Perspectives and Practices of Chinese Smart Home Produ Interview study; smart TV is a company product-line row. OUT 2026 NDSS TBTrackerX: Fantastic Trigger Bots and Where to Find Malicious IPTV spam homonym. OUT 2026 PETS Dead Domains, Living Data: A Privacy Risk Analysis of Domain L Android apps; one expired-domain example happens to also ship on Roku. OUT 2026 USENIX PANGOLIN: Fuzzing Multilingual IoT Firmware with LLM-Driven Co "SmartTVs" is a citation to the 2021 fuzzing paper, used as a baseline name. ============================================================================== 1. THE POPULATION BY VENUE, YEAR AND TOPIC ============================================================================== Venue A B A+B Venue papers in corpus Share of venue ------- - - --- ---------------------- -------------- IMC 3 8 11 638 1.7% USENIX 4 5 9 1410 0.6% NDSS 2 4 6 701 0.9% PETS 3 3 6 510 1.2% CCS 1 1 2 990 0.2% IEEE-SP 0 1 1 767 0.1% WWW 0 0 0 843 0.0% Window A B A+B ---------- - - --- 2010–2014 3 0 3 2015–2018 0 1 1 2019–2021 3 8 11 2022–2024 6 9 15 2025–2026* 1 4 5 * 2025–2026 is provisional: CCS 2026 and IMC 2026 have not been held, and IEEE S&P / WWW 2026 are incompletely selected. See literature:corpus. Topic (hand-assigned, ranking only) A B A+B ----------------------------------- - -- --- iot-device-set 0 15 15 delivery-performance 2 2 4 tracking 3 0 3 vulnerability 1 2 3 broadcast 2 0 2 app-analysis 1 1 2 side-channel 2 0 2 acr 1 0 1 piracy 1 0 1 device-population 0 1 1 policy-compliance 0 1 1 ============================================================================== 2. PLATFORM FIELDS: WHY `web` IS THE WRONG FILTER ============================================================================== platforms[] value A+B (N=35) share corpus (N=5859) share -------------------- ---------- ----- --------------- ----- web 1 2.9% 1622 27.7% mobile 6 17.1% 1075 18.3% iot 26 74.3% 436 7.4% other-online-service 12 34.3% 2429 41.5% offline 3 8.6% 2139 36.5% not-applicable 0 0.0% 56 1.0% Papers in the population carrying platform 'web' : 1 of 35 Which paper(s): USENIX 2026 Missing, Present and Conflicting: A Large Scale An population[].unit == 'iot-devices' : 20 of 35 population[].unit == 'mobile-apps' : 5 of 35 population[].unit == 'websites' : 1 of 35 In the corpus-wide 'crawled' population (crawlConfig or automated-web-crawl): 6 of 35 Which: CCS 2019, PETS 2022, USENIX 2023, CCS 2022, PETS 2024, USENIX 2026 ============================================================================== 3. HOW THESE PAPERS GET AT THE TRAFFIC ============================================================================== tool category A (N=13) A+B (N=35) ---------------------- -------- ---------- proxy-interception 3 4 traffic-capture 5 15 mobile-instrumentation 6 11 program-analysis 4 5 crawler-framework 1 2 browser-automation 1 1 network-scanner 1 4 blocklist 1 1 instrument (alias-folded) papers (of 35) ------------------------- -------------- Wireshark 10 tcpdump 9 mitmproxy 4 Frida 3 adb 3 VirusTotal 2 LibScout 2 OpenWPM 1 PingPong 1 Mercury 1 FlowDroid 1 Charles Proxy 1 jadx 1 Raspberry Pi 1 Scapy 1 UNMAPPED RESIDUE of the alias fold — 191 distinct raw tool names used by the 35 papers, printed in full: 3× DBSCAN | 2× Censys | 2× dnsmasq | 2× Google voice synthesizer | 2× IoT Inspector | 2× nmap | 2× OpenSSL | 2× random forest | 2× scikit-learn | 2× t-SNE | 2× TF-IDF | 2× WHOIS | 1× Adam | 1× adb_shell | 1× Afatech AF9015 | 1× agglomerative clustering | 1× Anaconda | 1× Analysis Scripts | 1× Androguard | 1× Android Debug Bridge (adb) | 1× Android Debug Bridge (ADB) | 1× Android Studio APK Analyzer | 1× AntMonitor | 1× apk-mitm | 1× apksigner | 1× AppCensus | 1× Apple trust store | 1× Apple Wi-Fi geolocation API | 1× Application Exerciser Monkey | 1× Apriori | 1× arecord | 1× ARKit | 1× Avalpa OpenCaster | 1× BeautifulSoup | 1× BeEF Toolkit | 1× BERT | 1× BiLSTM | 1× Bing | 1× Bleak | 1× Bumble | 1× Chapoly1305/FindMy | 1× ChatGPT (OpenAI's TextCompletion API) | 1× Chrome | 1× CICFlowmeter | 1× CogniCrypt | 1× Common CA Database | 1× Conviva | 1× cosine distance | 1× Criminal IP | 1× crt.sh | 1× cryptography/fernet | 1× CryptoGuard | 1× curl | 1× DekTec DTU-215 | 1× DekTec StreamXpress | 1× DICE coefficient | 1× Dijkstra's algorithm | 1× DNSDB | 1× DPDK | 1× DroidBot | 1× fastText | 1× FCC database of digital TV towers | 1× Flight Radar 24 | 1× Forward feature selection (FFS) | 1× Fourier transform | 1× generic deep neural network | 1× GNU TLS | 1× Google Play API | 1× Google Public DNS | 1× Google search | 1× Google Search | 1× Google Voice synthesizer | 1× GPS Tracks | 1× Gradient Boosting Decision Tree | 1× GSDMM | 1× HDMI Video Capture Device | 1× HiDes UT-100c | 1× Hurricane Electric IPv6-over-IPv4 tunnel | 1× IDA Pro | 1× IDAPython | 1× IEEE Organizationally Unique Identifier registry | 1× IFTTT | 1× Intel RealSense Camera T265 | 1× InternalBlue | 1× IP2Location | 1× IPFIX | 1× iptables | 1× IRDB | 1× irgen | 1× IrScrutinizer | 1× Java | 1× Keras | 1× Latent Dirichlet Allocation | 1× LightGBM | 1× logistic regression (custom) | 1× MakeHex | 1× MAPS | 1× Maven Repository | 1× MaxMind | 1× MaxMind GeoLite2 | 1× MaxMind geolocation database | 1× Mbed TLS | 1× MbedTLS | 1× McAfee | 1× median absolute deviation (MAD) | 1× Microsoft trust store | 1× Mon(IoT)r | 1× Monkey Application Exerciser | 1× Monkey Application Exerciser for Android Studio | 1× Monte Carlo sampling | 1× Mother of all Ad-Blocking | 1× Mozilla trust store | 1× Naïve Bayes | 1× NASA SEDAC Metropolitan Statistical Areas dataset | 1× nDPI | 1× nearest-neighbor classifier | 1× Nessus | 1× Netdisco | 1× NetFlow | 1× Netify | 1× Nexmon | 1× NFF-Go | 1× NimBLE | 1× Non-Negative Matrix Factorization (NMF) | 1× NoxPlayer | 1× Objection | 1× OpenAI Text Completion API | 1× OpenCaster | 1× OpenDNS | 1× OpenWRT | 1× OpenWrt/LEDE | 1× OPP-115 | 1× Oracle Java | 1× passive network telescope | 1× Passport | 1× Pi-hole Default blocklist | 1× PostgreSQL | 1× PrivBERT | 1× Prodigy | 1× ProVerif | 1× pyshark | 1× Python | 1× Python requests/2.31.0 | 1× Python TLS implementation | 1× Radare2 | 1× Random Forest | 1× Randoop | 1× Raspberry Pi 3 | 1× Raspberry Pi 4 | 1× Redis | 1× RedOrbit HbbTV Emulator | 1× Remote Central Forums | 1× RIPE IPmap | 1× Roku External Control Protocol | 1× SciPy | 1× Secure Transport | 1× SHAP | 1× Similarweb | 1× Snorkel | 1× Softflowd | 1× SoSci Survey | 1× spaCy | 1× spaCy en_core_web_lg | 1× StopAd smart TV blocklist | 1× TensorFlow | 1× The Big Blocklist Collection (Firebog) | 1× TP-Link power plugs | 1× traceroute | 1× TrafficPassthrough | 1× Trigger Scripts | 1× TSDuck | 1× Tuya Smart app | 1× TV Fool | 1× tvbus.exe | 1× Unity | 1× Validation Scripts | 1× VLC Player | 1× VS1838B | 1× WALA | 1× WiFi Inspector | 1× WiGLE | 1× WiGLE API | 1× WireShark/tshark | 1× wolfSSL | 1× WolfSSL | 1× word2vec | 1× XCUITest | 1× XGBoost | 1× YAF | 1× Yersinia | 1× Zeek Interception evidence, per Tier A paper (full text, whitespace-collapsed): paper mitm/proxy own CA / root cert pinning / cert failure undecryptable reported DNS-level capture router/AP capture HDMI / screen capture remote-control automation ----------- ---------- ------------------ ---------------------- ---------------------- ----------------- ----------------- --------------------- ------------------------- IMC 2011 — — — — — 1 — 1 IMC 2011 1 — — — — 1 — — USENIX 2014 3 — — — — 4 — 1 CCS 2019 53 1 12 — 27 5 5 35 PETS 2020 1 — 2 9 14 12 — 4 USENIX 2021 8 — — — — 1 14 1 PETS 2022 — — — — 4 11 — 4 PETS 2022 9 — 2 — — — 2 1 NDSS 2023 16 — 3 — 11 18 3 2 USENIX 2023 1 — — — — 4 1 41 IMC 2024 2 — — — — 3 15 4 NDSS 2024 — — — — — — — 3 USENIX 2025 2 — — — 5 1 — 5 mitm/proxy 10 of 13 Tier A | 22 of 35 A+B own CA / root cert 1 of 13 Tier A | 4 of 35 A+B pinning / cert failure 4 of 13 Tier A | 6 of 35 A+B undecryptable reported 1 of 13 Tier A | 1 of 35 A+B DNS-level capture 5 of 13 Tier A | 12 of 35 A+B router/AP capture 11 of 13 Tier A | 30 of 35 A+B HDMI / screen capture 6 of 13 Tier A | 9 of 35 A+B remote-control automation 12 of 13 Tier A | 16 of 35 A+B ============================================================================== 3b. VANTAGE ============================================================================== vantage location (verbatim, unfolded) papers ------------------------------------- ------ United States 5 Italy 2 Germany 2 France 2 Austria 2 Finland 2 US 2 UK 2 Europe 2 Portugal 2 Sweden 2 Norway 2 United Kingdom 1 241 countries and territories 1 Japan 1 Korea 1 China 1 US (North Carolina) 1 office space 1 Apartment 1 1 Apartment 2 1 lab space 1 U.S. 1 Asia 1 New York, U.S. 1 Frankfurt, Europe 1 Singapore, Asia 1 Sydney, NSW, Australia 1 worldwide 1 USA 1 university building 1 single-family house 1 e-bike route 1 flight 1 Belgium 1 Bulgaria 1 Croatia 1 Cyprus 1 Czech Republic 1 Denmark 1 Estonia 1 Greece 1 Hungary 1 Iceland 1 Ireland 1 Latvia 1 Liechtenstein 1 Lithuania 1 Luxembourg 1 Malta 1 Netherlands 1 Poland 1 Romania 1 Slovakia 1 Slovenia 1 Spain 1 papers with a vantage tuple: 34 of 35; of those, stating a location: 18 (52.9%) ============================================================================== 4. WHAT THEY SAMPLE — THE NO-TRANCO PROBLEM ============================================================================== population.sourceList (verbatim, unfolded — ranking only) papers ---------------------------------------------------------- ------ custom seed list 7 Roku Channel Store 3 Google Play Store 2 IoT Inspector 2 DS1: operational IPTV traces 1 DS2: operational IPTV traces 1 DS3: operational IPTV traces 1 large commercial IPTV service provider in the United State 1 NASA SEDAC Metropolitan Statistical Areas dataset 1 FCC database of digital TV towers in the United States 1 station coverage maps supplied by TV Fool 1 Amazon Fire TV channel store 1 Roku-Top1K 1 FireTV-Top1K 1 residential gateways 1 Amazon curated list "Top Featured" apps 1 buyers' guides in North America and Europe 1 Apple iTunes Preview / Apple App Store 1 Fire TV app store 1 AndroZoo and APKMirror 1 custom channel selection 1 social media platforms 1 IRDB and Remote Central Forums 1 password-lists and username-lists 1 custom testbed 1 custom user-study recruitment 1 2014 PhpBB password leak 1 RockYou password leak 1 generated synthetic credit card details 1 news headlines 1 Conviva 1 Avast WiFi Inspector 1 Censys 1 custom IoT testbeds 1 large European ISP 1 major European IXP 1 custom device selection 1 Mon(IoT)r dataset 1 UNSW Smart Home Traffic Dataset 1 YourThings Smart Home Traffic Dataset 1 UNB Simulated Office-Space Traffic Dataset 1 11 typical offices and apartments that are accessible to u 1 custom test scenes 1 custom-collected WiFi networks 1 Androzoo 1 Google Play 1 YourThings Dataset 1 HomeSnitch Dataset 1 PingPong Dataset 1 Mon(IoT)r Dataset 1 UNSW Dataset 1 Our Dataset 1 custom device set 1 lab dataset 1 custom lab setup 1 MonIoTr Lab 1 AndroZoo 1 IoT Inspector dataset 1 IEEE OUI database 1 WiGLE 1 popular US stores, including Amazon.com 1 custom laboratory experimental setup 1 custom campus network 1 custom home network 1 UNSW IoT Analytics 1 YourThings IoTFinder 1 custom SERP corpus 1 OPP-115 1 custom testbeds 1 custom test devices 1 population.unit papers ------------------ ------ iot-devices 20 other 17 mobile-apps 5 human-participants 3 documents 3 ip-addresses 1 autonomous-systems 1 network-flows 1 domains 1 websites 1 web-pages 1 population.samplingMethod papers ------------------------- ------ purposive 19 pre-existing-dataset 11 convenience 10 exhaustive 7 top-n 5 seed-and-crawl 5 random 5 not-stated 4 stratified 3 paper iot-devices n values stated -------------------------------------------------- --------------------------- IMC 2011 q-score-proactive-service-quality-asse 7000000, 140000 PETS 2020 the-tv-is-smart-and-full-of-trackers-m 57 USENIX 2021 android-smarttvs-vulnerability-discove 11 IMC 2024 watching-tv-with-the-second-party-a-fi 2 IMC 2019 information-exposure-from-consumer-iot 81 USENIX 2019 all-things-considered-an-analysis-of-i 83000000, 500000, 1000 IMC 2020 a-haystack-full-of-needles-scalable-de 96 NDSS 2020 packet-level-signatures-for-smart-home 19, 55, 26, 45 NDSS 2020 et-tu-alexa-when-commodity-wifi-device 31 USENIX 2020 you-are-what-you-broadcast-identificat 31850, 423, 26478 IMC 2021 iotls-understanding-tls-usage-in-consu 40 PETS 2021 blocking-without-breaking-identificati 31 PETS 2022 analyzing-the-feasibility-and-generali 45, 28, 18, 70, 19, 8 USENIX 2022 lumos-identifying-and-localizing-diver 44 IMC 2023 behind-the-scenes-uncovering-tls-and-s 2014, 113, 7 IMC 2023 in-the-room-where-it-happens-character 93, 13487 IMC 2024 iot-bricks-over-v6-understanding-ipv6- 93 PETS 2024 connecting-the-dots-tracing-data-endpo 25123, 54950, 30, 66 NDSS 2025 evaluating-machine-learning-based-iot- 90 NDSS 2026 blerp-ble-re-pairing-attacks-and-defen 22 papers stating an iot-devices size : 20 of 35 stated size values : 39 all values, sorted : 2, 7, 8, 11, 18, 19, 19, 22, 26, 28, 30, 31, 31, 40, 44, 45, 45, 55, 57, 66, 70, 81, 90, 93, 93, 96, 113, 423, 1000, 2014, 13487, 25123, 26478, 31850, 54950, 140000, 500000, 7000000, 83000000 median of all stated values : 66 largest set per paper, sorted : 2, 11, 22, 31, 31, 40, 44, 55, 57, 70, 81, 90, 93, 96, 2014, 13487, 31850, 54950, 7000000, 83000000 median of largest-per-paper : 75.5 papers whose LARGEST set <= 100 : 14 of 20 papers whose LARGEST set <= 200 : 14 of 20 papers whose LARGEST set <= 1000 : 14 of 20 papers whose largest set is over 200, i.e. not a lab bench: IMC 2011 q-score-proactive-service-quality-asse -> 7000000, 140000 USENIX 2019 all-things-considered-an-analysis-of-i -> 83000000, 500000, 1000 USENIX 2020 you-are-what-you-broadcast-identificat -> 31850, 423, 26478 IMC 2023 behind-the-scenes-uncovering-tls-and-s -> 2014, 113, 7 IMC 2023 in-the-room-where-it-happens-character -> 93, 13487 PETS 2024 connecting-the-dots-tracing-data-endpo -> 25123, 54950, 30, 66 population tuples in A+B stating a listVersion: 42 of 105 ============================================================================== 5. MEASURED RESULTS (detection[].prevalence, Tier A only) ============================================================================== --- IMC 2011 | Understanding couch potatoes: measurement and modeling of interactive usage of IPTV at large scale. * Stream control operations technique : Analyzed logged control events and constructed a finite state machine. metric : relative proportion of operations prevalence: FastForward, play, and Replay comprised 45% of total events. quote (results): The sum of FastForward (FF), play, and Replay comprise 45% of the total. * Video popularity technique : Counted requests and tracked rank persistence over time. metric : rank-frequency distribution and top-N retention prevalence: Top-100 and top-300 video popularity dropped rapidly over 60 days. quote (results): For the next 60 days, we counted how many of those popular videos remained among the top 300 (or 100) most popular. * Stream-control server impact technique : Compared actual traces with and without control events in a discrete-event simulator. metric : peak server bandwidth prevalence: Peak bandwidth increased from about 17 Gbps to about 20 Gbps. quote (results): We observe that server bandwidth increases from about 17 Gbps to about 20 Gbps at the peak when stream control operations are accounted for --- IMC 2011 | Q-score: proactive service quality assessment in a large IPTV system. * customer service problem prediction technique : Ridge regression over network KPIs and aggregated customer trouble tickets metric : false negative rate and false positive rate prevalence: predict 60% of service problems reported by customers with only 0.1% false positive rate quote (results): Q-score is able to predict 60% of service problems reported by customers with only 0.1% misclassification (i.e., false positive rate). * proactive service degradation detection technique : Evaluated Q-score accuracy after increasing network-event-to-feedback skip intervals metric : lead time before customer reports prevalence: 9 hours of lead time preserved 0.1% FPR, with FNR increasing from 30% to 40% quote (results): we find 9 hours of lead time is at the feasible level, as observing 9 hours of skip interval preserves 0.1% of FPR only sacrificing 10% of FNR --- USENIX 2014 | From the Aether to the Ethernet—Attacking the Internet using Broadcast Digital Television * RF injection into DVB-T broadcasts technique : Intercepted, modified, and retransmitted DVB streams on the original frequency. metric : coverage radius and area prevalence: 1 W amplifier: 477 m radius and 1.4 km²; 25 W: 2385 m radius and 35 km² quote (results): Using this formula shows that with a 1 W (30 dBm) amplifier ... cover a region with radius of 477 m, or an area of 1.4 km2. * Malicious HbbTV application execution technique : Injected AIT and HTML payloads into multiplexed DVB streams. metric : successful attack capabilities on one smart TV prevalence: Invisible execution, screen takeover, intranet scanning, TV crash, and external-web-server denial of service quote (evaluation): Using our test setup, we were able to create HbbTV applications which ran invisibly in the background, as well as applications which completely took over the TV screen. * Urban attack scalability technique : Cross-correlated population density, tower data, and propagation coverage maps. metric : number of stations and potentially affected devices prevalence: More than 20,000 devices in a single attack; up to 10 stations in parts of New York City quote (results): In certain locations in the Inwood area, where the population density is 50,000 persons per km2, the attacker can infect 10 different stations * HbbTV Internet and intranet attacks technique : Ran malicious JavaScript through injected HbbTV applications. metric : verified attack types prevalence: Port scanning, fraudulent login display, malformed-image crash, and denial of service quote (evaluation): We verified that we were able to access servers both on the Internet at large and on the local intranet. --- CCS 2019 | Watching You Watch: The Tracking Ecosystem of Over-the-Top TV Streaming Devices. * known tracker contact technique : Matched contacted hosts and domains against five tracking lists. metric : share of channels contacting known trackers prevalence: 69% of Roku channels and 89% of Amazon Fire TV channels quote (abstract): traffic to known trackers present on 69% of Roku channels and 89% of Amazon Fire TV channels. * identifier leakage technique : Searched HTTP URLs, headers, cookies, and bodies for encoded, hashed, or cleartext identifiers. metric : requests or identifiers containing leaks prevalence: 4,452 of 6,142 Roku requests containing AD ID or serial number were tracker-flagged; 3,427 of 8,433 Amazon identifiers were cleartext quote (results): We searched for various encoding and hashing combinations using the method described by Englehardt et al. [22]. * unencrypted HTTP traffic technique : Counted requests sent over port 80 in captured PCAPs. metric : share of channels sending cleartext requests prevalence: 794 of 1,000 Roku channels and 762 of 1,000 Fire TV channels quote (results): Analyzing the requests sent over port 80 we found that 794 of the 1000 Roku channels sent at least one request in cleartext. * video title leakage technique : Searched captured traffic for encodings of manually identified video titles. metric : channels leaking titles to tracking domains prevalence: 9 of 100 Roku channels and 14 of 100 Fire TV channels quote (results): We found 9 channels on Roku and 14 channels on the Fire TV, among the 100 channels we randomly selected on each device, that leaked the title of the video to a tracking domain. * TLS certificate validation technique : Attempted MITM interception and measured successfully decrypted channels. metric : channels with intercepted HTTPS prevalence: 957 Fire TV channels and 43 Roku channels quote (results): On Amazon Fire TV, we were able to install our own cert on the device which allowed us to intercept HTTPS requests on 957 of the 1000 channels. * remote-control API vulnerability technique : Reverse engineered API traffic and tested malicious cross-origin web requests. metric : vulnerability assessment prevalence: Roku API exposed identifiers, channel control, installed-channel lists, and SSID quote (results): We set up a page to demonstrate the attack and verified that a malicious web page visited by Roku users ... can abuse the External Control API. --- PETS 2020 | The TV is Smart and Full of Trackers: Measuring Smart TV Advertising and Tracking * ATS domains technique : Service labels and union of DNS blocklists metric : share of domains or apps contacting ATSes prevalence: about 10% of Roku and Fire TV apps contact 20+ and 10+ ATS domains, respectively quote (results): about 10% of the Roku and Fire TV apps contact 20+ and 10+ ATS domains, respectively. * Platform segmentation technique : Compared FQDN, eSLD, and parent-organization overlap metric : dataset and eSLD overlap prevalence: 314 ATS domains unique to Roku, 285 unique to Fire TV, and 227 overlapping quote (results): we identify 314 ATS domains that are unique to the Roku dataset, 285 that are unique to the Fire TV dataset, and an overlap of 227 between the two datasets. * DNS blocklist coverage technique : Matched contacted FQDNs against four blocklists metric : block rate prevalence: Firebog blocked 22% of Roku and 27% of Fire TV testbed FQDNs quote (results): TF, closely followed by MoaAB and PD, blocks the highest fraction of domains across all of the platforms in both the in the wild and testbed datasets. * Missed ads and app breakage technique : Repeated manual app interaction under each blocklist metric : ads missed and functionality breakage prevalence: all blocklists produced non-trivial false positives and false negatives quote (results): All blocklists suffer from a non-trivial amount of visually observable FPs and FNs. * PII exposure technique : Searched HTTP headers and URI paths for raw and hashed identifiers metric : apps, eSLDs, and blocked FQDNs prevalence: hundreds of apps exfiltrated PII to third parties and platform-specific parties quote (conclusion): Hundreds of Roku and Fire TV apps expose PII, mostly to third parties and the platform-specific party. * Joint advertising and static identifiers technique : Detected co-occurrence of advertising IDs with serial or device IDs metric : number of apps prevalence: 697 Fire TV apps sent advertising ID alongside serial number and device ID quote (results): Aside from the 697 Fire TV apps that expose advertising ID alongside serial number and device ID discussed earlier. * TLS interception failure technique : Compared TLS connections with connections containing decrypted HTTP packets metric : decryption failure rate prevalence: failure was at most 20% of TLS connections for 80% of Fire TV apps quote (appendix): decryption fails for 1 out of 5 (or fewer) TLS connections for 80% of all apps. --- USENIX 2021 | Android SmartTVs Vulnerability Discovery via Log-Guided Fuzzing * SmartTV API vulnerabilities technique : Log-guided dynamic fuzzing with cyber and physical feedback metric : unique vulnerabilities prevalence: 37 unique vulnerabilities across 11 Android TVBoxes quote (abstract): Our analysis led to the automatic discovery of 37 unique vulnerabilities, including 11 high-impact cyber threats, 10 new memory corruptions, and 16 visual and auditory anomalies. * Memory corruptions technique : Execution-log monitoring for crashes and anomalous states metric : number of vulnerabilities prevalence: 10 memory corruptions quote (evaluation): We discovered 37 security-critical flaws leading to various cyber attacks (11), physical disturbances (16) and memory corruptions (10). * Visual and auditory anomalies technique : External HDMI capture and before-after signal comparison metric : number of anomalies prevalence: 16 visual and auditory anomalies quote (introduction): Our analysis led to the automatic discovery of 37 unique vulnerabilities, including 11 high-impact cyber threats, 10 new memory corruptions, and 16 visual and auditory anomalies. * Input-validation messages technique : CNN classification of log messages trained from Android ROMs metric : classifier accuracy and recall prevalence: 46% of APIs triggered at least one input validation quote (evaluation): As shown, on average 87% APIs triggered at least 1 log message and 46% triggered at least 1 input validation. --- PETS 2022 | FingerprinTV: Fingerprinting Smart TV Apps * domain-based fingerprints technique : Extracted domains recurring in all ten launch samples; clustered app fingerprints. metric : prevalence and distinctiveness prevalence: 96% Apple TV, 88% Fire TV, and 100% Roku apps exhibited DBFs. quote (results): We find that 96% (N = 961) of the top-1000 Apple TV apps exhibit a DBF; 88% (N = 884) ... Fire TV ... and 100% ... Roku * packet-pair-based fingerprints technique : Extracted identical-size directional packet pairs using PingPong. metric : prevalence and distinctiveness prevalence: 68% Apple TV, 95% Fire TV, and 100% Roku apps exhibited PBFs. quote (results): We find that 68% (N = 678) of the top-1000 Apple TV apps exhibit a PBF; 95% (N = 952) ... Fire TV ... and 100% ... Roku * TLS-based fingerprints technique : Extracted recurring TLS ClientHello fingerprints using Mercury. metric : prevalence and distinctiveness prevalence: 95% Apple TV, 86% Fire TV, and 100% Roku apps exhibited TBFs; distinctiveness was 3%, 7%, and 1%. quote (results): only 3%, 7%, and 1% of the Apple TV, Fire TV, and Roku apps that exhibit TBFs, exhibit distinct TBFs. * smart TV app fingerprinting technique : Combined DBF and PBF fingerprints and evaluated cluster uniqueness. metric : prevalence and distinctiveness prevalence: DBF-and-PBF fingerprints were distinct for 89% Apple TV, 95% Fire TV, and 76% Roku apps exhibiting both. quote (results): the fingerprint is distinct for 89% (599) of the Apple TV apps, 95% (802) of the Fire TV apps, and 76% (760) of the Roku apps * platform-specific fingerprints technique : Compared fingerprints for apps matched across all three platforms. metric : share of matched apps with platform differences prevalence: 76% of 80 apps available on all three platforms exhibited different fingerprints on each platform. quote (introduction): among 80 apps that are made available on all three smart TV platforms, 76% exhibit a different fingerprint on each platform * identical fingerprints technique : Examined developer identities and parent organizations within non-singleton clusters. metric : share attributable to same developer prevalence: DBF-sharing apps attributable solely to the same developer: 37% Apple TV, 57% Fire TV, and 20% Roku before consolidation. quote (results): For Apple TV, 142 of the 397 apps (37%) that share their DBF with other app(s) only share it with other apps from the same developer. --- PETS 2022 | Watch Over Your TV: A Security and Privacy Analysis of the Android TV Ecosystem * third-party library prevalence technique : Libscout plus package-prefix clustering and manual classification metric : share of apps containing libraries prevalence: Social Media libraries in 88%, analytics in 75%, and advertising in 77% of apps quote (results): We detected Social Media libraries in 88% of the apps, with multiple Facebook libraries occupying the top 5. Similarly, we found that 75% of the apps contain analytics libraries and 77% contain advertising libraries. * sensitive data flows technique : Customized Flowdroid taint analysis metric : share of APKs with at least one sensitive flow prevalence: 78% of files quote (results): The analysis found at least one sensitive data flow in 78% of the files. * static identifier use technique : Static taint-flow source and sink analysis metric : APKs with identifier flows prevalence: 3031 APKs used UUID-generated globally unique identifiers; 285 APKs exposed MAC-address or SSID identifiers quote (results): The most used identifier is a globally unique ID (GUID) generated with the java.util.UUID package (3031 APKs). * network data leakage technique : Charles-captured traces manually searched for sensitive values metric : TV apps sending data to first parties, trackers, or CDNs prevalence: 42% of explored TV apps used static identifiers; media metadata appeared in 65% quote (results): We captured traffic of 21 TV apps and 22 mobile apps out of 30 Popular-Streaming apps. * malware technique : VirusTotal multi-engine detection threshold metric : APKs flagged by antivirus engines prevalence: 34 APKs were flagged by more than 5 engines; 34 were discussed as flagged by more than 10 engines quote (results): We detected 34 APKs in our dataset that are flagged as malware by more than 10 engines in VirusTotal (VT). * socket communication technique : Cross-reference analysis of socket API methods metric : APKs containing Socket APIs prevalence: 2646 APKs, or 56% quote (results): Overall, we detected 2646 APKs (56%) including Socket APIs. * Nearby API authentication technique : Search for authentication-token reads in callbacks metric : APKs implementing authentication prevalence: None of the APKs using NearbyConnection implemented authentication quote (results): Unfortunately, none of the APKs that use the NearbyConnection API implement authentication. * permission rationale display technique : Bytecode pattern matching and context heuristics metric : apps showing rationale for dangerous permissions prevalence: Only small percentages showed context, including 3% for coarse location and 5% for record audio quote (results): Our results indicate that only a small percentage of TV apps display the rationale behind dangerous permission --- NDSS 2023 | I Still Know What You Watched Last Sunday: Privacy of the HbbTV Protocol in the European Smart TV Landscape * Tracking before consent technique : Inspected contacted domains and consent-phase traffic metric : share of TV channels prevalence: 26 of 36 channels communicated with trackers before consent quote (discussion): All the 36 TV channels we analyzed contact at least one tracking domain; further, 26 communicate with trackers before the user has expressed their consent. * Tracking pixels technique : Identified returned 1×1 pixel image objects metric : share of TV channels prevalence: 20 of 36 channels (56%) adopted tracking pixels quote (discussion): 20 of the 36 TV channels (56%) we analyzed adopt the invisible 'tracking pixel' to profile users. * Periodic tracking requests technique : Computed mean and standard deviation of inter-request times metric : average time between requests prevalence: SportItalia approximately every 70 seconds; RDS approximately every 14 seconds quote (results): SportItalia makes requests to Smartclip ... around every 70 seconds, while RDS contacts Google Analytics around every 14 seconds. * Plaintext sensitive traffic technique : Inspected unencrypted HTTP packets and payloads metric : presence of sensitive data in HTTP prevalence: Most traffic captures included HTTP; HSE exposed login information quote (discussion): We found HTTP communication in most of our traffic captures; such traffic contained sensitive information such as device IDs, visitor IDs, country codes, and ISP information. * Denylist coverage technique : Matched manually identified tracking domains against Pi-hole and EasyList metric : fraction of tracking domains blocked prevalence: At maximum 44% in 2021 and 81% in 2022 quote (discussion): Commonly used tracking denylists only block at maximum 44% in 2021 and 81% in 2022 of the domains in our traffic captures and marked as tracking. * User risk awareness technique : Surveyed awareness and coded open-ended responses metric : share mentioning no security or privacy risk prevalence: 68% of 132 Smart TV-related respondents mentioned no risk quote (discussion): Out of the 132 participants in the Smart TV and HbbTV awareness survey, 68% could not mention any security or privacy risk --- USENIX 2023 | HOMESPY: The Invisible Sniffer of Infrared Remote Control of Smart TVs * Off-path infrared signal sniffing technique : Mounted a COTS receiver at positions across four room layouts. metric : IR-key extraction accuracy prevalence: 98.9%, 90.8%, 83.1%, and 75.0% for layouts A–D quote (results): The result is shown in Table 4. In particular, for layout A, 98.9% of the keys (including the repeat keys) can be correctly sniffed * IR command decoding technique : Matched recovered timings against a 75,901-code hash database. metric : unique device mappings prevalence: 98% of devices had unique D-pad, OK, or BACK mappings quote (results): Among the 1303 devices, only 26 devices have overlapped D-pad, OK or BACK keys, which means that for 98% of all devices have their unique mappings. * Sensitive input extraction technique : Mapped D-pad sequences onto virtual-keyboard coordinates and filtered candidates. metric : Top-1, Top-3, and Top-5 accuracy prevalence: 47% Top-1, 70% Top-3, and 77% Top-5 quote (results): The accuracy increases to 70% for Top3 and 77% for Top5. * Virtual-keyboard activity detection technique : Applied OK-count and BACK-absence thresholds within time windows. metric : activity coverage prevalence: 100% of keyboard-input activities quote (results): The experiment result shows H OME S PY could cover 100% of the activities about keyboard input. --- IMC 2024 | Watching TV with the Second-Party: A First Look at Automatic Content Recognition Tracking in Smart TVs. * ACR network traffic technique : Filtered captured DNS domains containing the string “acr”. metric : traffic presence, bytes, frequency, and packet timing prevalence: ACR traffic existed during linear TV and HDMI scenarios. quote (introduction): ACR network traffic exists when watching linear TV and when using smart TV as an external display using HDMI * ACR traffic after opt-out technique : Compared traffic across opted-in and opted-out experimental phases. metric : presence or absence of communication with ACR domains prevalence: Opting out produced a complete absence of communication with previously identified ACR domains. quote (results): once opt-out is exercised (Table 1), there is a complete absence of communication with any previously identified ACR domains * Geographic ACR differences technique : Compared contacted domains and geolocated their server IP addresses. metric : domain identity, server location, and traffic levels prevalence: UK and US televisions contacted distinct ACR domains; US FAST streaming generated ACR traffic unlike the UK. quote (introduction): smart TVs in the UK and the US contact distinct ACR domains * Login-status effect technique : Compared logged-in and logged-out phases for each television. metric : CDFs of bytes transferred and traffic periodicity prevalence: User login status appeared to have no material impact on ACR traffic behavior. quote (results): user login status appears to have no material impact on the ACR network traffic behavior. * ACR traffic by viewing scenario technique : Ran six one-hour scenarios on each television. metric : traffic volume, frequency, and periodicity prevalence: Linear and HDMI had the highest ACR traffic for both brands in the UK. quote (results): For both LG (a) and Samsung (b) TVs, the scenarios with the highest ACR traffic are Linear and HDMI. --- NDSS 2024 | Acoustic Keystroke Leakage on Smart Televisions * acoustic keystroke leakage technique : Matched reference Smart TV sounds and extracted movement-count sequences. metric : top-K recovery accuracy prevalence: up to 60.19% of common passwords within 100 guesses for realistic Samsung users quote (introduction): For ten subjects typing into real applications on a Samsung TV, the attack recovers 53.33% of CCNs, 33.33% of full credit card details, and up to 60.19% of common passwords. * keyboard-instance splitting technique : Used SystemSelect sounds or timing outliers between adjacent movements. metric : correctly identified instances prevalence: 98 of 100 passwords quote (results): This method correctly identifies 98 / 100 passwords, where the two failures result from long pauses while typing. * credit-card entry detection technique : Matched consecutive keyboard-instance lengths to payment-field sizes. metric : detected interactions prevalence: 29 of 30 human credit-card interactions quote (results): The attack successfully identifies that the user enters credit card details in 29 of the 30 total interactions. * password-entry classification technique : Random Forest classified dynamic-suggestion behavior from movement histograms. metric : accuracy prevalence: 99.02% on the human password set quote (results): Further, we infer when a subject types a password (§IV-B1), and this classifier has an accuracy of 99.02% on this set. * suboptimal keyboard paths technique : Used pauses and movement timing to prioritize tolerated path deviations. metric : optimal-path rate prevalence: 89.35% for human credit-card typing on Samsung; 45.35% on AppleTV passwords quote (results): When typing CCNs, users take the optimal path between keys only 89.35% of the time. * direction inference technique : Detected rapid horizontal scrolls using median adjacent-movement timing. metric : recovery impact prevalence: Direction inference never harms recovery rates quote (methodology): We find that direction inference never harms the recovery (§VI-E). --- USENIX 2025 | Watch Out Your TV Box: Reversing and Blocking a P2P-based Illegal Streaming Ecosystem * EVPAD P2P users technique : Queried Broker peer lists and deduplicated composite IP identifiers. metric : unique peers prevalence: 131,175 unique peers over two months quote (results): Consequently, our analysis revealed that a total of 131,175 unique peers were active in EVBOX's P2P network had over the two-month observation period. * Operational servers technique : Extracted server IPs from reversed app data and peer responses. metric : operational servers prevalence: 78 operational servers quote (conclusion): With EVPAD devices, we identified 131,175 users across 116 countries and 78 operational servers located in the United States, Japan, Singapore, Hong Kong, and other countries. * Concurrent streaming users technique : Repeatedly collected users across all channels over two months. metric : users across all channels per measurement prevalence: minimum 13,438, maximum 19,889, average 16,997 quote (results): our observations of concurrent users across all channels over 30 measurements spanning two months—with a minimum of 13,438, maximum of 19,889, and an average of 16,997 users * VoD content technique : Changed category identifiers and collected encrypted VoD lists. metric : collected content items prevalence: 24,934 pieces of content quote (results): By changing the ID values for the main and subcategories, we collected all VoD lists, resulting in a total of 24,934 pieces of content. * Country distribution technique : Geolocated collected peer IP addresses. metric : countries represented prevalence: 116 countries quote (results): Upon identifying the country associated with each IP address, we found that users were distributed across 116 countries [27]. * Authentication bypass technique : Replayed copied device attributes in NoxPlayer. metric : successful service access prevalence: emulator successfully accessed and streamed live channels quote (evaluation): We verified this vulnerability by using the NoxPlayer to mimic an authenticated device. As shown in Figure 8, the emulator successfully accessed and streamed live channels * P2P denial of service technique : Sent a crafted TCP packet to a connected EVPAD peer. metric : service termination prevalence: a single crafted TCP packet terminated the target service quote (evaluation): From one device, we crafted and sent a TCP socket-based packet to the other device... the EVPAD streaming service on the target device immediately terminated. detection tuples in Tier A: 70; carrying a prevalence: 68 (97.1%) corpus-wide: 27241 tuples, 26316 carry a prevalence (96.6%) ============================================================================== 6. WHERE THIS LITERATURE GOES QUIET ============================================================================== crawlConfig field (papers with a crawlConfig object) A+B (N=6) corpus crawled (N=1080) ---------------------------------------------------- -------------- ----------------------- consentAction 0 of 6 (0.0%) 349 of 1080 (32.3%) interactionDepth 5 of 6 (83.3%) 841 of 1080 (77.9%) statefulness 0 of 6 (0.0%) 219 of 1080 (20.3%) browsers 0 of 6 (0.0%) 529 of 1080 (49.0%) crawlConfig values in the A+B population, verbatim: CCS 2019 consent=not-applicable depth=single-target-page state=not-stated browsers=[] robotsTxt=not-stated PETS 2022 consent=not-applicable depth=deep-crawl state=not-stated browsers=[] robotsTxt=not-stated USENIX 2023 consent=not-stated depth=not-stated state=not-stated browsers=[] robotsTxt=not-stated CCS 2022 consent=not-applicable depth=single-target-page state=not-stated browsers=[] robotsTxt=not-stated PETS 2024 consent=not-stated depth=single-target-page state=not-stated browsers=[] robotsTxt=not-stated USENIX 2026 consent=not-stated depth=landing-plus-subpages state=not-stated browsers=[] robotsTxt=not-stated A+B papers carrying an ethics object: 34 of 35; artifacts object: 34 of 35 field A+B corpus empirical ----------------------------- ---------------- -------------------- ethics.reviewOutcome stated 14 of 34 (41.2%) 1728 of 4472 (38.6%) artifacts.availability stated 27 of 34 (79.4%) 2890 of 4854 (59.5%) temporal.spanStart stated 27 of 34 (79.4%) 2882 of 5118 (56.3%) Artifact links released, Tier A: IMC 2011 public www.research.att.com/∼kkrama/papers/streamcontrol.pdf IMC 2011 (no artifacts object extracted) USENIX 2014 none-mentioned http://www.avalpa.com/the-key-values/15-free-software/33-opencaster (other; OpenCaster software; authors=false) CCS 2019 promised-not-yet-available — PETS 2020 promised-not-yet-available http://athinagroup.eng.uci.edu/projects/smarttv/ (project-page; Project page for tools and testbed datasets; authors=true) USENIX 2021 none-mentioned https://sites.google.com/site/smarttvdemos/ (project-page; Demonstration website for discovered attacks; authors=true) PETS 2022 promised-not-yet-available https://github.com/UCI-Networking-Group/fingerprintv PETS 2022 public https://gitlab.com/s3lab-rhul/watch-over-your-tv-paper NDSS 2023 public https://github.com/SecPriv/hbbtv-blocker USENIX 2023 public https://sites.google.com/view/homespydemo IMC 2024 public https://github.com/SafeNetIoT/ACR NDSS 2024 public https://github.com/tejaskannan/smart-tv-keyboard-leakage USENIX 2025 restricted https://doi.org/10.5281/zenodo.15646588 ============================================================================== 6b. TWO PROBES THE PAGE CITES ============================================================================== TLS decryption hole, Tier A (n=13): narrow probe (the one the page published until 2026-09-13): 1 wide probe (the one the page cites now) : 3 wide CCS 2019 watching-you-watch-the-tracking-ecosystem-of-over-the-top-tv-streaming-devices ...tificate to the device and use external toolkits (e.g., Frida [29] for the Amazon Fire Stick TV) to bypass certificate pinning. Contributions: We make the following contributions: • We conduct the first large-scale study of privacy practices of OTT streaming channels. Using an automated crawler that... both PETS 2020 the-tv-is-smart-and-full-of-trackers-measuring-smart-tv-advertising-and-tracking ...tion time with each app is approximately 16 minutes. We do not attempt to decrypt TLS traffic as we cannot install our own self-signed certificates on the Roku. 4.2 Fire TV Data Collection In this section, we describe the Fire TV platform, our app selection methodology, and present an overview of Fi... wide PETS 2022 watch-over-your-tv-a-security-and-privacy-analysis-of-the-android-tv-ecosystem ...k traffic. We instrumented the APKs using the mitm-proxy script [32] to add Charles certificate and remove certificate pinning checks. This instrumentation is only necessary for TV apps as there is no way to install custom certificates on Android TV. For the mobile apps, we use smartphones with Andr... ATSC 3.0 / NextGen TV, full text, all 5859 corpus papers: 4 match USENIX 2012 i-forgot-your-password-randomness-attacks-against-php-applications [not in audit set] ...mber of bits truncated. Application Attack Application Attack mediawiki 4.2 4.3 5.3 • Joomla 4.3 • Open eClass 4.2 4.3 5.4 • MyBB ATSc 4.1c 5.3c ◦ taskfreak 4.2 4.3 5.3 • IpBoard ATSc 4.1c 4.2c • zen-cart ATS RT • phorum 4.2 4.3 5.3 • osCommerce 2.x ATS RT • HotCRP 4.2 4.3 5.3 • osCommerce 3.x 4.2 4.3 5.4 • gazelle 4.3 5.3 • elgg ATSc 4.2... USENIX 2014 from-the-aether-to-the-ethernet-attacking-the-internet-using-broadcast-digital-t [A] ...ctive deployment or in advanced stages of testing in most of Europe. In December 2013, the Advanced Television Systems Committee (ATSC), which defines the digital video standards in the US, Canada, South Korea and several USENIX Association other countries, published a candidate standard for hybrid TV in America [6]. This candidate standa... NDSS 2020 automated-cross-platform-reverse-engineering-of-can-bus-commands-from-mobile-apps [not in audit set] ...E, ATA, ATH... Gauged 17 ATED, ATD, ATP, ATZ... iOBD2 20 ATE, AT ST, AT CA F... LeagendOBD 12 ATE, ATB, ATTR, ATQ... Engie 8 ATE, ATSC, ATI, ATST... TABLE X: AT commands extracted from dongle apps. 16 App # Command AcuraLink 9 Alpine 2 Alpine Tunelt 3 Audi MMI Connect 10 Carbin Control 15 Car-Net 4 Companion 2 Mini Connected Classic 1 Nis... CCS 2025 dont-look-up-there-are-sensitive-internal-links-in-the-clear-on-geo-satellites [not in audit set] ...ngel Electronics. 2024. STAB HH90 Satellite Dish Motor. https:// angelelectronics.ca/products/stab-hh90-satellite-dish-motor. [8] ATSC. [n. d.]. ATSC. https://www.atsc.org/documents/. [9] Robin Bisping, Johannes Willbold, Martin Strohmeier, and Vincent Lenders. 2024. Wireless Signal Injection Attacks on VSAT Satellite Modems. USENIX Secur... Statefulness: 6 of 35 papers have a crawlConfig object; statefulness values among them: ["not-stated","not-stated","not-stated","not-stated","not-stated","not-stated"] full-text "factory reset" anywhere in the 35: 2 A NDSS 2023 i-still-know-what-you-watched-last-sunday-privacy-of-the-hbbtv-protocol-in-the-european-smart-tv-landscape .... For the second test, we start by extracting the HbbTV URLs from the DVB stream using the TSDuck library and the UT-We perform a factory reset of the TV for each channel analysis 100c HiDes modulator. As mentioned in Section II, the DVB to prevent interference in the captured traffic. stream includes the URLs of the HbbTV applications; t... B NDSS 2026 blerp-ble-re-pairing-attacks-and-defenses ... introduces a usability trade-off: if a device implicit authentication, ensuring that an attacker lacking loses its PK (e.g., via factory reset), it requires manual user the current PK cannot compute the new one. intervention to re-pair. • Transcript Hashing: Devices must maintain a cumulative We implemented this protocol in NimBLE by ext... LLM and language-model tools used by the population (tools[].category, used only): CCS 2022 category=ml-model-or-algorithm BERT quote: "IoTSpotter's BERT-based and BiLSTM classifiers identified 58,859 and 69,270 app descriptions as mobile-IoT, respectively." IMC 2023 category=llm ChatGPT (OpenAI's TextCompletion API) quote: "Using OpenAI's TextCompletion API, we develop prompt to infer device vendors and categories based on DHCP hostname, mDNS/SSDP responses, and user labels." PETS 2024 category=llm OpenAI Text Completion API quote: "Using OpenAI's Text Completion API [34], we develop prompts to infer device vendors and categories" PETS 2024 category=ml-model-or-algorithm PrivBERT quote: "We use PrivBERT [69], a pre-trained privacy policy language model to build a binary classifier" ============================================================================== 7. QUOTES: see scripts/ctv_quotecheck.py (checks .cols AND the PDF text layer) ==============================================================================
14. Run log
- 2026-09-12 — page written. Corpus
data/extract/run1, 5,859 papers. Candidate pool, audit, report script, quote check, bibliography generation and external verification all run on this date. Three review passes — figures-versus-script, citations-and-quotes and external currency — returned and were applied; they are logged in §15. A fourth, generic pass was still running when the session ended, so it is not in §15. An earlier version of this line said all four were logged there; that was wrong. - 2026-09-13 — round 2. All four passes re-run against the published pages, because every focused domain had been changed by round 1's own fixes. 21 findings accepted, 3 rejected, logged in §15b. Two of round 1's accepted fixes turned out to be wrong and one had landed on only one of the two pages. The report script gained membership and topic digests (§5b), a
linkOf()helper, and three probes the page had been citing without having: the wide TLS-decryption probe, the ATSC full-text probe and the factory-reset probe. Section §13 is regenerated from the current script. - Models: the three focused passes were
sonnet, the generic pass wasfable, in both rounds. The generic pass found more real defects in round 2 than the three focused passes combined — see the closing note in §15b for why that is structural rather than luck. - No credential or token was printed at any point in either run.
15. Review log
Three review passes returned on 2026-09-12 against the published pages, the report script and its output; the fourth was cut off mid-run and was re-run as part of round 2 (§15b). Each was told explicitly that the author's context might not be exhaustive and to verify rather than assume. Every finding below was re-checked by hand against the primary source before it was accepted or rejected — two of the accepted ones needed correcting in the process, and the rejections are recorded because they are the only evidence of whether a reviewer earned its slot.
Round 1, pass 1 — figures versus script (sonnet)
| # | Finding | Disposition |
|---|---|---|
| 1.1 | The “fifteen IoT-device-set papers” list substitutes [2Rye, Erik C.; Levin, Dave (2024): "Surveilling the Masses with Wi-Fi-Based Positioning Systems", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] (tagged device-population) for the real fifteenth, Et Tu Alexa? (NDSS 2020), which is cited nowhere | Accepted. Independently confirmed by extracting the tier/topic tuples from ctv_fold.mjs. Fixed: the list is now exactly the fifteen, [14Zhu, Yanzi; Xiao, Zhujun; Chen, Yuxin; Li, Zhijing; Liu, Max; Zhao, Ben Y.; Zheng, Haitao (2020): "Et Tu Alexa? When Commodity WiFi Devices Turn into Adversarial Motion Sensors", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] was added to the bibliography, and [2Rye, Erik C.; Levin, Dave (2024): "Surveilling the Masses with Wi-Fi-Based Positioning Systems", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] is described in its own clause. This was also found independently by the author before the pass returned |
| 1.2 | Both pages say the alias fold “maps 25 names”; it has 27 skeleton keys onto 22 canonical names | Accepted. Re-parsed the ALIAS object: 27 and 22, neither of them 25. This row originally read “Corrected on both pages”. That was false: only the content page was corrected in round 1. This provenance page still said 25 until round 2 found it — three separate reviewers, plus the author's own residue sweep, all landed on the same line. See §15b, finding R2.6 |
| 1.3 | The sampling-size paragraph attributes “3,000,000 IPTV set-top boxes” to [12Gopalakrishnan, Vijay; Jana, Rittwik; Ramakrishnan, K. K.; Swayne, Deborah F.; Vaishampayan, Vinay A. (2011): "Understanding couch potatoes: measurement and modeling of interactive usage of IPTV at large scale", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] as one of “the 20 papers stating an iot-devices population size”, but that paper's tuples are unit other and human-participants. The paper that is one of the 20 with large values, [17Song, Han Hee; Ge, Zihui; Mahimkar, Ajay; Wang, Jia; Yates, Jennifer; Zhang, Yin; Basso, Andrea; Chen, Min (2011): "Q-score: proactive service quality assessment in a large IPTV system", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] (7,000,000 and 140,000), is named nowhere | Accepted, and it was worse than reported. The author had already flagged the same sentence for omitting 57 — the PETS 2020 testbed — from a hand-typed list; the reviewer found the framing error underneath it. The whole passage is now generated by the script (§13) rather than typed, and the 3,000,000 figure is kept with an explicit note that it is a subscriber count, not one of the 20 |
| 1.4 | The map guard only checks the slug set. Flipping a tier letter exits 0 and silently moves the population from 35 to 34 | Accepted, and the most valuable finding of the round. The reviewer mutation-tested rather than read. Fixed with a summing invariant; see §5b |
| 1.5 | Committed output is byte-identical to a fresh run; venue split, platform distribution, “6 of 35 crawled”, all three corpus comparison rows and the detection-prevalence coverage independently re-derived from extractions.jsonl and all matching; every checked row of Measured results you can cite correct against raw paper text; nullable-field denominators and paper-not-tuple counting correct | No action. Recorded because “I re-derived it without your script and it matched” is the check that matters |
| 1.6 | Venue display labels differ between page (“USENIX Security”, “TheWebConf”) and script (“USENIX”, “WWW”) | Rejected. The page uses the venues' real names and the numbers underneath are identical. Cosmetic |
Round 1, pass 2 — citations and quotes (sonnet)
| # | Finding | Disposition |
|---|---|---|
| 2.1 | The ATSC broadcaster-application quote is altered as well as misattributed: the source sentence reads “broadcasters can load and reload, and change things as they are happening, based on broadband availability…”, and the page drops that clause without an ellipsis and inserts the word “content”. The source is also a news article paraphrasing an unnamed speaker (“she said”), not ATSC, and not the URL cited | Accepted, BLOCKER. The author had independently confirmed the quote was on neither cited page; the reviewer established that it had also been silently edited. The quote is removed entirely rather than repaired — a secondhand paraphrase of an unnamed conference speaker is not a source this page should lean on. Replaced with what ATSC's own pages do say, a pointer to A/344, and the FCC proceeding |
| 2.2 | Roku's functions are GetRIDA() and IsRIDADisabled(); the page mis-capitalised both, and IsRIDADisabled() is documented on ifDeviceInfo, not on the page cited | Accepted. Confirmed by fetching ifdeviceinfo.md directly. Both names corrected and both docs now cited |
| 2.3 | All 33 citekeys resolve; no key defined twice in the merged 1,024-entry bibliography; 0 rule-A/B duplicates; the three rule-D candidates touching new keys are same-surname-different-author | No action |
| 2.4 | Author order for all 31 new entries checked against the papers' own PDF front matter, and the 18 DOI-bearing ones additionally against Crossref's order-sensitive author array. All 31 correct, including the re-split “Al Aaraj, Jad” and [13Ahmed, Dilawer; Das, Anupam; Zaffar, Fareed (2022): "Analyzing the Feasibility and Generalizability of Fingerprinting Internet of Things Devices", in: Proceedings on Privacy Enhancing Technologies. (DOI)], whose title page has no text layer | No action. This is the check the author could not fully self-run, and it is the reason the pass was worth its slot |
| 2.5 | Every other vendor, standards-body and regulator quote verified word-for-word from the primary source, including all four Texas dates and the “every 500 milliseconds” wording | No action |
| 2.6 | The rendered DOM matches the source: 83 markers, 33 reference entries, 4 WRAP blocks, 16 tables, no truncation | No action |
| 2.7 | The Tier-B rule is paraphrased two different ways on the same page (“…for a television” versus “…for them”) | Accepted as a NIT. Wording made consistent |
Round 1, pass 3 — external currency (sonnet)
| # | Finding | Disposition |
|---|---|---|
| 3.1 | The ATSC capability quote is on neither cited page; it is a 2016 conference paraphrase on a third URL | Accepted — the same defect as 2.1, found independently by a pass with a different brief. Two reviewers arriving at one finding from opposite directions is the strongest signal in this round |
| 3.2 | The FCC's Fifth FNPRM (GN Docket 16-142, adopted 2025-10-28) on the ATSC 1.0 sunset is missing | Accepted, after fetching the FCC's own fact sheet rather than the trade coverage the reviewer cited. The primary document gave a better fact than the finding did: the simulcast rule “was extended to July 17, 2027”, and the notice's list of outstanding issues includes a one-word bullet, “Privacy”. Added |
| 3.3 | Kentucky HB 692 classifies ACR data as sensitive data, effective 2027-07-01 | Accepted with a correction. That is the bill as introduced. The enacted version (Acts Chapter 118) instead “prohibit[s] controllers from collecting automatic content recognition data without a consumer's consent”. Added in the enacted form, from the legislature's own record page |
| 3.4 | Sony, Hisense and TCL remain unsettled; the timeline reads as if Samsung and LG were the whole story | Accepted. Added, together with the explicit negative result that no EU or UK regulatory action was found |
| 3.5 | “ADB Wi-Fi 2.0” (Android 17) was announced 2026-09-09, three days before this page claimed currency | Accepted on a different source. The blog URL the reviewer gave returns 404. The primary adb documentation already states it, so the claim is added on that authority instead |
| 3.6 | The Walmart press release contains neither “Platform+” nor “Inscape” | Accepted. Confirmed by text search. The footnote now supports only what the release says |
| 3.7 | research.att.com/~kkrama/papers/streamcontrol.pdf (an extracted artifact URL) returns 403 | Accepted as a NIT; annotated rather than removed, since it records what the paper claimed. See §8 |
| 3.8 | HbbTV 2.0.5, the ATSC 76% figure, Samsung TIFA, Amazon Fire TV, Roku ECP, the six regulatory dates, all five artifact repositories, the Zenodo DOI and mitmproxy v12.2.3 all verified live and correct | No action |
| 3.9 | The HbbTV 2.0.5 paraphrase (“adding DRM and WebAssembly recognition”) is looser than the source's “recognising features in the market such as DRM and WebAssembly” | Rejected. The substance is right and the rest of the sentence is near-verbatim |
| 3.10 | Samsung's own TRO was granted and vacated the next day, before the February settlement | Rejected for the content page. Real, but a procedural detail; the substantive gap was 3.4, which is in |
What round 1 cost and returned
Three passes returned; a fourth, the generic one, was still running when the session ended and is logged in §15b with round 2 instead. Counting the rows above: 14 accepted (1.1–1.4, 2.1, 2.2, 2.7, 3.1–3.7), two of them accepted-with-correction, and 3 rejections (1.6, 3.9, 3.10). An earlier version of this paragraph said “eleven accepted, four rejections”, which does not match its own tables. The pattern worth recording for the next run: the defects were all in prose that summarises data, never in the tables the script generates. Every generated figure survived independent re-derivation; every hand-typed list, paraphrase and quotation that sat next to one had to be fixed. The two most valuable findings — the mutation test that broke the guard, and the author-order check against Crossref — were both things that cannot be done by reading, which is the argument for handing a reviewer the script rather than only the page.
References
The same keys and the same shared bibliography as connected_tv; this page adds no entries of its own. No discussion block: comments belong on the content page.
15b. Review log, round 2 (2026-09-13)
Round 1 ran three focused passes and was cut off with the generic pass still in flight. Round 2 re-ran all four, because every one of the three focused domains had been changed by round 1's own fixes — and that turned out to be the right call twice over: two of round 1's accepted fixes were themselves wrong, and a third had landed on only one of the two pages. Every reviewer was told the author's context might not be exhaustive. Every finding below was re-checked by hand against the primary source or the raw corpus before being accepted or rejected.
The headline of this round: the generic pass, which has no checklist, found more real defects than the three focused passes combined — and all of them were claims about what the literature does not do. A reviewer asked to check figures against a script checks the figures that are in the script. A negative claim has no figure.
Round 2, pass 1 — figures versus script (sonnet)
| # | Finding | Disposition |
|---|---|---|
| R2.1 | The round-1 invariant can be satisfied vacuously. Swapping two papers between tiers — one genuine Tier B out, one genuine OUT in — leaves A:13, B:22, ADJ:16, OUT:52 unchanged and exits 0, while platform web goes 1→2 and the venue, year and platform tables all move | Accepted, BLOCKER, and the most valuable finding of either round. Reproduced exactly. §5b now records it, and membership is pinned per tier as a slug-list digest. The pattern worth keeping: round 1's fix was written from the failure a reviewer demonstrated, and closed exactly that failure and nothing beside it. It took a second mutation test, from a reviewer with the same brief, to find the hole next door |
| R2.2 | artifacts.links[0] is an object, so the fallback printed the literal [object Object] — published three times in §13 | Accepted. Confirmed against extractions.jsonl. Fixed with a linkOf() helper that takes .url and throws if the shape is ever a string; the three entries now carry real links, one of which (athinagroup.eng.uci.edu/projects/smarttv/) is a citable project page that had been hidden behind the bug |
| R2.3 | Q16 says “33 empirical”; isEmpirical is true for 34 of the 35, and 27/34 is the 79.4% printed beside it | Accepted. Re-derived: exactly one paper (Lumos) is not empirical. Corrected |
| R2.4 | The round-1 rewrite of the 3,000,000 IPTV figure says the paper “records them as subscribers”; the paper says “the average number of set-top boxes provisioned was approximately 3 million” | Accepted — a round-1 fix that introduced a new error. Verified in paper.cols.txt. The real reason it is outside the 20 is that its extraction unit is other, not iot-devices — a taxonomy boundary, not anything the paper did. The page now says that |
| R2.5 | The topic tag has no invariant at all; changing one exits 0 and silently moves the “What this literature measures” table and the iot-device-set list | Accepted. Pinned with its own digest. Mutation-tested: it now throws |
| R2.6 | The provenance page still says the alias fold maps “25 skeletons” while listing 27 of them | Accepted. Found independently by three of the four passes and by the author's own residue sweep. Corrected, and round 1's claim to have fixed “both pages” is corrected too |
| R2.7 | Script reproduces byte-identically; the fifteen iot-device-set papers are exactly the fifteen named; venue, year, topic, platform, vantage, crawlConfig, ethics and artifacts tables all match cell-for-cell; sentinels never counted as stated; papers never counted as tuples | No action. Recorded because it is the control |
Round 2, pass 2 — citations and quotes (sonnet)
| # | Finding | Disposition |
|---|---|---|
| R2.8 | “Nothing comparable was found from an EU or UK regulator” is false. The UK ICO published a connected-TV programme on 2026-06-11 | Accepted, BLOCKER. Verified at ico.org.uk directly. See R2.11 — the other pass found the EU half independently |
| R2.9 | Author order for zhu2020_alexa checked against the NDSS PDF front matter: Yanzi Zhu, Zhujun Xiao, Yuxin Chen, Zhijing Li, Max Liu, Ben Y. Zhao, Haitao Zheng — matches, no swap. All 34 citekeys resolve; 1,025 entries, 1,025 distinct keys; 0 rule-A/B duplicates | No action. This is the check the author cannot self-run, and it is why the pass earns its slot |
| R2.10 | Every replacement claim from round 1 verified verbatim: the FCC's “extended to July 17, 2027” and its one-word “Privacy” bullet, Kentucky's enacted prohibition, GetRIDA() / IsRIDADisabled(), the four Texas dates and the “every 500 milliseconds” quote, and the Texas AG's own statement that the cases against Sony, Hisense and TCL “remain ongoing” | No action |
Round 2, pass 3 — external currency (sonnet)
| # | Finding | Disposition |
|---|---|---|
| R2.11 | The EU half of the same blocker: a joint Article 62 GDPR operation by the Dutch, Hungarian, Italian and Liechtenstein authorities published a final report on smart TVs on 2025-09-23 — a year before this page claimed the file was empty | Accepted, BLOCKER. Fetched and read the 14-page report. It is the single most useful external document on this page: a regulator ran the measurement, on three televisions, across first install, standby, off and ordinary use. Its off-state result (91–99% of flows to the OS provider) is now on the content page, and the fact that nobody in these seven venues has measured a television in the off state is now an open question. Neither European action uses the phrase “automatic content recognition”, which is exactly why an ACR-shaped search could not see them |
| R2.12 | The round-1 Walmart fix is itself wrong — the release does contain “VIZIO's Platform+ segment…”; the plus sign is HTML-entity-encoded, so a tag-strip search misses it | Accepted — a second round-1 fix that introduced an error. Reproduced: the string appears only after double entity-decoding. The footnote now quotes what the release says and records why the first check missed it. The reviewer flagged that its own first grep made the same mistake before its WebFetch caught it — an honest note that is worth more than a clean report |
| R2.13 | The artifact table labels ahn2025_watch “restricted” while citing the DOI of the open record | Accepted. Zenodo's API says access_right: open. The paper has two records; the label belongs to the other one. Both are now named |
| R2.14 | FCC proceeding still pending, no Report and Order; HbbTV 2.0.5 still current; all five artifact repositories, the Zenodo DOI, mitmproxy v12.2.3 and every vendor doc re-verified | No action |
| R2.15 | A trade-press claim that the FCC voted in May 2026 to mandate ATSC 3.0 tuners | Rejected, by the reviewer that surfaced it and again by the author: it appears nowhere on fcc.gov and the URL 403s. Recorded because this is the kind of claim that gets re-added |
| R2.16 | One pass read the ATSC standards listing as showing A/344:2025-07 as the latest approved revision | Rejected on better evidence. The A/344:2026-04 PDF exists at HTTP 200 and its own title block reads A/344:2026-04 … 14 April 2026. The document's own designation beats a reading of the listing page |
Round 2, pass 4 — generic, no checklist (fable)
This pass found the cluster the other three could not, because a false negative has no figure to check and no citation to verify. Five claims about what the literature does not do were wrong, and two were refuted by the provenance page's own printed residue.
| # | Finding | Disposition |
|---|---|---|
| R2.17 | “LLM-based classification — Absent. Not one paper in this population uses an LLM for anything.” Two papers carry tools[] entries with category: llm, used — and both tool names are printed in this page's own residue block | Accepted, and the most embarrassing finding of the round. [9Girish, Aniketh; Hu, Tianrui; Prakash, Vijay; Dubois, Daniel J.; Matic, Srdjan; Huang, Danny Yuxing; Egelman, Serge; Reardon, Joel; Tapiador, Juan; Choffnes, David R.; Vallina-Rodriguez, Narseo (2023): "In the Room Where It Happens: Characterizing Local Communication and Threats in Smart Homes", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] and [18Jakaria, Md; Huang, Danny Yuxing; Das, Anupam (2024): "Connecting the Dots: Tracing Data Endpoints in IoT Devices", in: Proceedings on Privacy Enhancing Technologies. (DOI)] both prompt OpenAI's Text Completion API to infer device vendor and category from DHCP hostnames. The claim is now the narrower and more useful one: an LLM is used here as a device-name resolver and has never been pointed at TV app metadata, store descriptions, ACR payloads or policies. A residue you publish but do not read is not a control |
| R2.18 | “No paper states whether the device was factory-reset” was read off statefulness across the 6 papers that have a crawlConfig object and asserted over all 35; [8Tagliaro, Carlotta; Hahn, Florian; Sepe, Riccardo; Aceti, Alessio; Lindorfer, Martina (2023): "I Still Know What You Watched Last Sunday: Privacy of the HbbTV Protocol in the European Smart TV Landscape", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] states it in so many words | Accepted. Classic denominator slip — the structured field is empty for the 6 that have it and silent for the other 29, which is not the same as those 29 saying nothing. A full-text probe is now in the script, and the page states both the field and the probe |
| R2.19 | “Exactly one paper measures both sides of the opt-out.” [6Moghaddam, Hooman Mohajeri; Acar, Gunes; Burgess, Ben; Mathur, Arunesh; Huang, Danny Yuxing; Feamster, Nick; Felten, Edward W.; Mittal, Prateek; Narayanan, Arvind (2019): "Watching You Watch: The Tracking Ecosystem of Over-the-Top TV Streaming Devices", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] repeated its entire Roku and Fire TV crawl with “Limit Ad Tracking” and “Disable Interest-based Ads” enabled | Accepted. Two papers, five years apart, at different layers. Verified in both papers' own text |
| R2.20 | “Almost nobody reports the hole — a probe matches 1 of 13” while the page quotes two of the sentences the probe misses, three sections earlier | Accepted. Probe width is a claim. The wide probe matches 3. The script now prints both widths and asserts the narrow set is a subset of the wide one, so a narrowing probe that returns more fails instead of publishing |
| R2.21 | “There is no AndroZoo for TV apps… every paper built its own set.” [1Tileria, Marcos; Blasco, Jorge (2022): "Watch Over Your TV: A Security and Privacy Analysis of the Android TV Ecosystem", in: Proceedings on Privacy Enhancing Technologies. (DOI)], the page's flagship TV-app paper, pulled its 4,745 APKs from AndroZoo and APKMirror | Accepted. The true statement is sharper: AndroZoo has no TV facet, so you cannot ask it for TV apps — you arrive with package names found elsewhere. For Roku, Tizen and webOS there is no archive at all |
| R2.22 | The round-1 log claims four passes ran and that the generic one is “logged in §15 with every finding”; §15 logs three | Accepted. The generic pass was still running when the session ended. Corrected, and §15's own arithmetic (it said 11 accepted / 4 rejections against tables holding 14 and 3) corrected with it |
| R2.23 | “The two vendors with the largest ACR businesses” is an unsourced market-size claim, and §10 says explicitly that installed-base share could not be established | Accepted. The two pages contradicted each other. Replaced with “two of the five vendors Texas sued”, which is a fact on the record |
| R2.24 | “Nobody has repeated…”, “no TV paper has ever…”, “Nobody has published…” — universal phrasing with no population, where the page elsewhere models the right form | Accepted. All three now name their population |
| R2.25 | The lead says “the median TV study is a handful of physical devices”; the page's own medians are 66 and 75.5 | Accepted. The lead was describing the TV-specific papers while the median describes IoT testbeds. Both numbers are now given, with the distinction stated |
| R2.26 | The table says fingerprintv is “promised, not yet available”; the prose two lines later counts “the five public repositories” | Accepted. The README still says the code “will be added… stay tuned”, four years on, while the dataset is out. That is a better fact than either version, and it is now the row |
| R2.27 | design:mobile_and_app_measurement, privacy:requests and programming:filter_lists contain no link back to this page. A reader on the mobile page's pinning section has no signal that step 1 usually fails on a TV | Accepted. Reverse links added — see §16 |
| R2.28 | \x27\x27sonnet\x27\x27 in three §15 headings renders its own apostrophes, because DokuWiki headings ignore monospace markup | Accepted. Headings de-monospaced |
| R2.29 | §5 reprints tables that §13 also carries in full (~150 duplicated lines) | Rejected. Deliberate, and stated as such: §5 is the readable verdict map and §13 is the unedited output. Someone checking a verdict should not have to scroll a 900-line block |
| R2.30 | The page holds its stated boundaries in the outward direction, answers its own question, and the enforcement timeline belongs; rendered DOM matches source on both pages; 0 red links; both anchors resolve | No action. Recorded as the positive control |
What round 2 cost and returned
Four passes, 21 accepted findings, 3 rejections. Two of round 1's own accepted fixes were wrong (R2.4, R2.12) and one had landed on only one page (R2.6) — which is the argument for re-running a reviewer whose domain you changed, rather than trusting that a fix was a fix.
Three things are worth carrying to the next page:
- A negative claim is the least-guarded thing on a page. Every figure here survived independent re-derivation, twice. Five of the six worst defects were sentences saying nobody does something, and none of the three focused briefs could have caught them — the checklist reviewer checks what is there.
- A published residue must be read, not just printed. The two LLM papers were in this page's own residue block, in plain sight, while the content page said they did not exist.
- A fix closes the failure it was shown, and nothing next to it. Round 1's invariant stopped a tier flip and let a tier swap through. The only thing that found the difference was mutating the code again, with the same brief, after the fix.
16. Reverse links added on 2026-09-13
The round-2 generic pass checked the page's stated boundaries in both directions and found the outward direction good and the inward direction missing entirely: three pages this one defers to carried no link back, so a reader arriving at the neighbour had no signal that the TV case exists.
| Page | Where | What it now says | Revision |
|---|---|---|---|
| mobile_and_app_measurement | head of The Certificate-Pinning Problem | a WRAP tip noting that the whole section assumes you can install a CA, which on Roku, Tizen and webOS you cannot, and on Android TV means instrumenting the APK instead | 1789266443 |
| requests | Related pages | where this page's instruments stop: no interception without network-level capture, no page context to attribute a request to, filter-list coverage measured at 22–27% | 1789266444 |
| filter_lists | the existing Smart TVs coverage row | why the rule syntax itself does not transfer — no URL, no page context, no element hiding | 1789266446 |
That page had previously pointed its smart-TV row at website_classification, which is not where a reader chasing that 22% figure needs to go.
- [1]
- Tileria, Marcos; Blasco, Jorge (2022): "Watch Over Your TV: A Security and Privacy Analysis of the Android TV Ecosystem", in: Proceedings on Privacy Enhancing Technologies. (DOI)
- [2]
- Rye, Erik C.; Levin, Dave (2024): "Surveilling the Masses with Wi-Fi-Based Positioning Systems", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
- [3]
- Björklund, Martin; Duvignau, Romaric (2025): "Endangered Privacy: Large-Scale Monitoring of Video Streaming Services", in: Proceedings of the USENIX Security Symposium. (Link)
- [4]
- Mavroudis, Vasilios; Hao, Shuang; Fratantonio, Yanick; Maggi, Federico; Kruegel, Christopher; Vigna, Giovanni (2017): "On the Privacy and Security of the Ultrasound Ecosystem", in: Proceedings on Privacy Enhancing Technologies. (DOI)
- [5]
- Kumar, Deepak; Shen, Kelly; Case, Benton; Garg, Deepali; Alperovich, Galina; Kuznetsov, Dmitry; Gupta, Rajarshi; Durumeric, Zakir (2019): "All Things Considered: An Analysis of IoT Devices on Home Networks", in: Proceedings of the USENIX Security Symposium. (Link)
- [6]
- Moghaddam, Hooman Mohajeri; Acar, Gunes; Burgess, Ben; Mathur, Arunesh; Huang, Danny Yuxing; Feamster, Nick; Felten, Edward W.; Mittal, Prateek; Narayanan, Arvind (2019): "Watching You Watch: The Tracking Ecosystem of Over-the-Top TV Streaming Devices", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
- [7]
- Varmarken, Janus; Le, Hieu; Shuba, Anastasia; Markopoulou, Athina; Shafiq, Zubair (2020): "The TV is Smart and Full of Trackers: Measuring Smart TV Advertising and Tracking", in: Proceedings on Privacy Enhancing Technologies. (DOI)
- [8]
- Tagliaro, Carlotta; Hahn, Florian; Sepe, Riccardo; Aceti, Alessio; Lindorfer, Martina (2023): "I Still Know What You Watched Last Sunday: Privacy of the HbbTV Protocol in the European Smart TV Landscape", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
- [9]
- Girish, Aniketh; Hu, Tianrui; Prakash, Vijay; Dubois, Daniel J.; Matic, Srdjan; Huang, Danny Yuxing; Egelman, Serge; Reardon, Joel; Tapiador, Juan; Choffnes, David R.; Vallina-Rodriguez, Narseo (2023): "In the Room Where It Happens: Characterizing Local Communication and Threats in Smart Homes", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [10]
- Oren, Yossef; Keromytis, Angelos D. (2014): "From the Aether to the Ethernet—Attacking the Internet using Broadcast Digital Television", in: Proceedings of the USENIX Security Symposium. (Link)
- [11]
- Anselmi, Gianluca; Vekaria, Yash; D'Souza, Alexander; Callejo, Patricia; Mandalari, Anna Maria; Shafiq, Zubair (2024): "Watching TV with the Second-Party: A First Look at Automatic Content Recognition Tracking in Smart TVs", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [12]
- Gopalakrishnan, Vijay; Jana, Rittwik; Ramakrishnan, K. K.; Swayne, Deborah F.; Vaishampayan, Vinay A. (2011): "Understanding couch potatoes: measurement and modeling of interactive usage of IPTV at large scale", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [13]
- Ahmed, Dilawer; Das, Anupam; Zaffar, Fareed (2022): "Analyzing the Feasibility and Generalizability of Fingerprinting Internet of Things Devices", in: Proceedings on Privacy Enhancing Technologies. (DOI)
- [14]
- Zhu, Yanzi; Xiao, Zhujun; Chen, Yuxin; Li, Zhijing; Liu, Max; Zhao, Ben Y.; Zheng, Haitao (2020): "Et Tu Alexa? When Commodity WiFi Devices Turn into Adversarial Motion Sensors", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
- [15]
- Wang, Yifan; Lyu, Minzhao; Sivaraman, Vijay (2024): "Characterizing User Platforms for Video Streaming in Broadband Networks", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [16]
- Akhtar, Zahaib; Nam, Yun Seong; Chen, Jessica; Govindan, Ramesh; Katz-Bassett, Ethan; Rao, Sanjay G.; Zhan, Jibin; Zhang, Hui (2018): "Understanding Video Management Planes", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [17]
- Song, Han Hee; Ge, Zihui; Mahimkar, Ajay; Wang, Jia; Yates, Jennifer; Zhang, Yin; Basso, Andrea; Chen, Min (2011): "Q-score: proactive service quality assessment in a large IPTV system", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [18]
- Jakaria, Md; Huang, Danny Yuxing; Das, Anupam (2024): "Connecting the Dots: Tracing Data Endpoints in IoT Devices", in: Proceedings on Privacy Enhancing Technologies. (DOI)
