From 3f197d871bf2b7b243a0b29c5d35f714981207ab Mon Sep 17 00:00:00 2001 From: Max Albrecht <1@178.is> Date: Mon, 17 Aug 2026 19:49:09 +0200 Subject: [PATCH 01/22] =?UTF-8?q?E:=20drop=20the=20empty=20dexd=20submodul?= =?UTF-8?q?e=20=E2=80=94=20the=20crate=20lives=20in-tree?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit packages/dexd was a gitlink to github.com/eins78/dexd, a repository that no longer exists. The player crate is placed directly in the monorepo instead of as a submodule (story decision 2026-08-17): submodule friction — the crate's text, packaging, CI and the OS image change in lockstep, and a submodule boundary got in the way of that. --- .gitmodules | 3 --- packages/dexd | 1 - 2 files changed, 4 deletions(-) delete mode 160000 packages/dexd diff --git a/.gitmodules b/.gitmodules index 508ef42..65b70c7 100644 --- a/.gitmodules +++ b/.gitmodules @@ -1,9 +1,6 @@ [submodule "packages/pi-gen"] path = packages/pi-gen url = https://github.com/eins78/pi-gen -[submodule "packages/dexd"] - path = packages/dexd - url = https://github.com/eins78/dexd [submodule "packages/branding"] path = packages/branding url = https://github.com/KTE/dex-branding diff --git a/packages/dexd b/packages/dexd deleted file mode 160000 index d8fa55f..0000000 --- a/packages/dexd +++ /dev/null @@ -1 +0,0 @@ -Subproject commit d8fa55f268725acc1972d16ab1b992efcff1cf92 From c825da488d62e8d1a56d3bab64f82b9dbd69205b Mon Sep 17 00:00:00 2001 From: Max Albrecht <1@178.is> Date: Mon, 17 Aug 2026 20:50:57 +0200 Subject: [PATCH 02/22] E: extract the dexd crate from the 4K loop experiment (mechanical) Source: home-workspace experiments/2026-08-12-4k-hevc-perfect-loop/dex-loop at commit 71711ec1 (dex-loop's own history, log -1 on that path). Copied with deploy/, renamed dex-loop -> dexd throughout: crate name, binary, Debian package, systemd unit, man pages, lintian overrides, maintainer scripts, changelog head, test binary paths, and every in-source reference (usage text, log/error prefixes, doc comments). The ingest sidecar checker sidecar-check.rs -> dex-sidecar.rs, per the story's naming decision (helpers keep the dex- prefix; the main binary/package/unit becomes dexd). build.rs's git-HEAD paths adjusted for the crate now sitting two levels below the repo root (packages/dexd) instead of one. Not copied: IMPLEMENTATION-PLAN.md, PLAN.md, soak-24h/ -- private experiment records that stay in home-workspace. README.md travels for now (it still says dex-loop throughout) and is replaced with rewritten, outsider-readable docs in Phase 2 of the extraction plan; same for the historically-accurate old-name mentions in deploy/changelog's older entry text and a doc comment in dex-sidecar.rs describing the tool's pre-cargo-workspace history. No prose was rewritten and no behaviour changed: cargo check/clippy/ test --lib, cargo deny check, reuse lint and shellcheck all pass unmodified from source, same as before the move. --- packages/dexd/.gitignore | 3 + packages/dexd/Cargo.lock | 230 ++ packages/dexd/Cargo.toml | 169 ++ packages/dexd/LICENSE | 75 + packages/dexd/LICENSES/MIT-0.txt | 16 + packages/dexd/README.md | 662 ++++++ packages/dexd/REUSE.toml | 23 + packages/dexd/build.rs | 91 + packages/dexd/deny.toml | 86 + packages/dexd/deploy/changelog | 28 + packages/dexd/deploy/dex-wait-hdmi | 40 + packages/dexd/deploy/dexd.service | 162 ++ packages/dexd/deploy/exhibit.json.default | 6 + packages/dexd/deploy/lintian-overrides | 53 + .../dexd/deploy/maintainer-scripts/postinst | 102 + .../dexd/deploy/maintainer-scripts/postrm | 13 + packages/dexd/deploy/man/dex-exhibit-apply.1 | 215 ++ packages/dexd/deploy/man/dex-wait-hdmi.1 | 30 + packages/dexd/deploy/man/dexd.1 | 119 + packages/dexd/src/bin/dex-exhibit-apply.rs | 320 +++ packages/dexd/src/bin/dex-sidecar.rs | 81 + packages/dexd/src/chunk.rs | 185 ++ packages/dexd/src/exhibit.rs | 2108 +++++++++++++++++ packages/dexd/src/ffi_consts.rs | 99 + packages/dexd/src/health.rs | 604 +++++ packages/dexd/src/heartbeat.rs | 540 +++++ packages/dexd/src/lib.rs | 20 + packages/dexd/src/main.rs | 1938 +++++++++++++++ packages/dexd/src/nal.rs | 201 ++ packages/dexd/src/sha256.rs | 110 + packages/dexd/src/sidecar.rs | 549 +++++ packages/dexd/src/watchdog.rs | 641 +++++ packages/dexd/tests/cli.rs | 1501 ++++++++++++ packages/dexd/tests/ffi_constants.rs | 160 ++ 34 files changed, 11180 insertions(+) create mode 100644 packages/dexd/.gitignore create mode 100644 packages/dexd/Cargo.lock create mode 100644 packages/dexd/Cargo.toml create mode 100644 packages/dexd/LICENSE create mode 100644 packages/dexd/LICENSES/MIT-0.txt create mode 100644 packages/dexd/README.md create mode 100644 packages/dexd/REUSE.toml create mode 100644 packages/dexd/build.rs create mode 100644 packages/dexd/deny.toml create mode 100644 packages/dexd/deploy/changelog create mode 100755 packages/dexd/deploy/dex-wait-hdmi create mode 100644 packages/dexd/deploy/dexd.service create mode 100644 packages/dexd/deploy/exhibit.json.default create mode 100644 packages/dexd/deploy/lintian-overrides create mode 100755 packages/dexd/deploy/maintainer-scripts/postinst create mode 100755 packages/dexd/deploy/maintainer-scripts/postrm create mode 100644 packages/dexd/deploy/man/dex-exhibit-apply.1 create mode 100644 packages/dexd/deploy/man/dex-wait-hdmi.1 create mode 100644 packages/dexd/deploy/man/dexd.1 create mode 100644 packages/dexd/src/bin/dex-exhibit-apply.rs create mode 100644 packages/dexd/src/bin/dex-sidecar.rs create mode 100644 packages/dexd/src/chunk.rs create mode 100644 packages/dexd/src/exhibit.rs create mode 100644 packages/dexd/src/ffi_consts.rs create mode 100644 packages/dexd/src/health.rs create mode 100644 packages/dexd/src/heartbeat.rs create mode 100644 packages/dexd/src/lib.rs create mode 100644 packages/dexd/src/main.rs create mode 100644 packages/dexd/src/nal.rs create mode 100644 packages/dexd/src/sha256.rs create mode 100644 packages/dexd/src/sidecar.rs create mode 100644 packages/dexd/src/watchdog.rs create mode 100644 packages/dexd/tests/cli.rs create mode 100644 packages/dexd/tests/ffi_constants.rs diff --git a/packages/dexd/.gitignore b/packages/dexd/.gitignore new file mode 100644 index 0000000..47d9ec2 --- /dev/null +++ b/packages/dexd/.gitignore @@ -0,0 +1,3 @@ +/target +# Written by an external sync step (see build.rs); never checked in. +/.dex-build-id diff --git a/packages/dexd/Cargo.lock b/packages/dexd/Cargo.lock new file mode 100644 index 0000000..f42b745 --- /dev/null +++ b/packages/dexd/Cargo.lock @@ -0,0 +1,230 @@ +# This file is automatically @generated by Cargo. +# It is not intended for manual editing. +version = 4 + +[[package]] +name = "arraydeque" +version = "0.5.1" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "7d902e3d592a523def97af8f317b08ce16b7ab854c1985a0c671e6f15cebc236" + +[[package]] +name = "block-buffer" +version = "0.12.1" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "d2f6c7dbe95a6ed67ad9f18e57daf93a2f034c524b99fd2b76d18fdfeb6660aa" +dependencies = [ + "hybrid-array", +] + +[[package]] +name = "cfg-if" +version = "1.0.4" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "9330f8b2ff13f34540b44e946ef35111825727b38d33286ef986142615121801" + +[[package]] +name = "const-oid" +version = "0.10.2" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "a6ef517f0926dd24a1582492c791b6a4818a4d94e789a334894aa15b0d12f55c" + +[[package]] +name = "cpufeatures" +version = "0.3.0" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "8b2a41393f66f16b0823bb79094d54ac5fbd34ab292ddafb9a0456ac9f87d201" +dependencies = [ + "libc", +] + +[[package]] +name = "crypto-common" +version = "0.2.2" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "ce6e4c961d6cd6c9a86db418387425e8bdeaf05b3c8bc1411e6dca4c252f1453" +dependencies = [ + "hybrid-array", +] + +[[package]] +name = "dexd" +version = "0.1.0" +dependencies = [ + "serde", + "serde_json", + "sha2", + "yaml-rust2", +] + +[[package]] +name = "digest" +version = "0.11.3" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "f1dd6dbb5841937940781866fa1281a1ff7bd3bf827091440879f9994983d5c2" +dependencies = [ + "block-buffer", + "const-oid", + "crypto-common", +] + +[[package]] +name = "foldhash" +version = "0.2.0" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "77ce24cb58228fbb8aa041425bb1050850ac19177686ea6e0f41a70416f56fdb" + +[[package]] +name = "hashbrown" +version = "0.16.1" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "841d1cc9bed7f9236f321df977030373f4a4163ae1a7dbfe1a51a2c1a51d9100" +dependencies = [ + "foldhash", +] + +[[package]] +name = "hashlink" +version = "0.11.1" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "824e001ac4f3012dd16a264bec811403a67ca9deb6c102fc5049b32c4574b35f" +dependencies = [ + "hashbrown", +] + +[[package]] +name = "hybrid-array" +version = "0.4.14" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "707114b52a152fa7bdb290cd7cd5912d9467273b6d74e21b8d81aca1f8533f6b" +dependencies = [ + "typenum", +] + +[[package]] +name = "itoa" +version = "1.0.18" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "8f42a60cbdf9a97f5d2305f08a87dc4e09308d1276d28c869c684d7777685682" + +[[package]] +name = "libc" +version = "0.2.189" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "3eaf3ede3fee6db1a4c2ee091bf8a8b4dccdc6d17f656fb07896ee72867612f2" + +[[package]] +name = "memchr" +version = "2.8.3" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "cf8baf1c55e62ffcace7a9f06f4bd9cd3f0c4beb022d3b367256b91b87513d98" + +[[package]] +name = "proc-macro2" +version = "1.0.107" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "985e7ec9bb745e6ce6535b544d84d6cd6f7ad8bd711c398938ae983b91a766d9" +dependencies = [ + "unicode-ident", +] + +[[package]] +name = "quote" +version = "1.0.47" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "1fbf4db142a473a8d80c26bbf18454ed458bf8d26c8219c331daecfdbd079001" +dependencies = [ + "proc-macro2", +] + +[[package]] +name = "serde" +version = "1.0.229" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "4148590afebada386688f18773da617792bf2ef03ffc1e4cbd2b1d45b023e0ba" +dependencies = [ + "serde_core", +] + +[[package]] +name = "serde_core" +version = "1.0.229" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "67dca2c9c51e58a4791a4b1ed58308b39c64224d349a935ab5039aa360942a48" +dependencies = [ + "serde_derive", +] + +[[package]] +name = "serde_derive" +version = "1.0.229" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "e7a5d71263a5a7d47b41f6b3f06ba276f10cc18b0931f1799f710578e2309348" +dependencies = [ + "proc-macro2", + "quote", + "syn", +] + +[[package]] +name = "serde_json" +version = "1.0.151" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "c841b55ecdae098c80dcae9cf767f6f8a0c2cdb3416bbef72181df4d0fe73f14" +dependencies = [ + "itoa", + "memchr", + "serde", + "serde_core", + "zmij", +] + +[[package]] +name = "sha2" +version = "0.11.0" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "446ba717509524cb3f22f17ecc096f10f4822d76ab5c0b9822c5f9c284e825f4" +dependencies = [ + "cfg-if", + "cpufeatures", + "digest", +] + +[[package]] +name = "syn" +version = "3.0.3" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "53e9bae58849f64dfa4f5d5ae372c8341f7305f82a3868709269343628b659a3" +dependencies = [ + "proc-macro2", + "quote", + "unicode-ident", +] + +[[package]] +name = "typenum" +version = "1.20.1" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "b6f5e870be6c3b371b77fe0ee0bafb859fa4964b4404c27de1d380043c4dda20" + +[[package]] +name = "unicode-ident" +version = "1.0.24" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "e6e4313cd5fcd3dad5cafa179702e2b244f760991f45397d14d4ebf38247da75" + +[[package]] +name = "yaml-rust2" +version = "0.11.0" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "631a50d867fafb7093e709d75aaee9e0e0d5deb934021fcea25ac2fe09edc51e" +dependencies = [ + "arraydeque", + "hashlink", +] + +[[package]] +name = "zmij" +version = "1.0.23" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "29666d0abbfad1e3dc4dcf6144730dd3a3ab225bbbdac83319345b1b44ccfc1b" diff --git a/packages/dexd/Cargo.toml b/packages/dexd/Cargo.toml new file mode 100644 index 0000000..6d3f050 --- /dev/null +++ b/packages/dexd/Cargo.toml @@ -0,0 +1,169 @@ +[package] +name = "dexd" +version = "0.1.0" +edition = "2021" +description = "Gapless HEVC looper: libmpv fed by an endless in-process byte stream" +# The dex project releases its works to the public domain as far as legally and +# practically possible: CC0-1.0 for content, MIT-0 for software. MIT-0 is MIT +# minus the attribution clause — maximally permissive, OSI-approved, and unlike +# CC0 it is accepted by every distributor for CODE (Fedora has disallowed CC0 +# for code since 2022, because CC0 expressly disclaims patent grants). +# +# This describes THE SOURCE ONLY. The .deb links libmpv, which Debian ships as +# GPL-3+ (it links GPL-3+ libsmbclient), so the shipped binary is a combined +# work conveyed under GPL-3+. No conflict — MIT-0 is GPL-compatible, and the two +# statements are about different artifacts. See LICENSE, which the package +# embeds verbatim as its copyright file so the distinction ships with it. +# Machine-readable per-file licensing: REUSE.toml + LICENSES/ (reuse.software). +license = "MIT-0" +# The compiler Debian trixie ships, which is what the Pi has. The device is a +# development host, so the crate must stay buildable with the fleet's own +# toolchain — CI compiles with apt's rustc to enforce it, not rustup's stable. +# Raising this floor is a decision about whether devices can still build their +# own software, not a routine bump. +rust-version = "1.85" + +# Dependencies are DECLARED, not avoided — SPEC §5c. This binary links libmpv +# and, transitively, 228 shared objects; a rule that forbade four small cargo +# crates while linking that was bookkeeping, not restraint. The requirement is +# that the set stays small enough to read: 23 crates, no proc macros. +# +# serde_json Replaces a 585-line hand-rolled JSON subset parser. Its \uXXXX +# and UTF-16 surrogate-pair handling is load-bearing, not +# decoration: Python's json.dumps escapes ALL non-ASCII by +# default, and a sidecar's informational `source` key can hold +# any filename, so refusing escapes would refuse byte-perfect +# assets over an ingest tool's serializer settings. +# serde For the map visitor ONLY — no `derive` feature, so no +# syn/quote/proc-macro2 on the build's critical path. Needed +# because the sidecar grammar REJECTS duplicate keys, which +# serde_json's last-wins Map cannot express. +# sha2 Replaces a hand-rolled SHA-256. Confirmed byte-identical on the +# real bench assets before the swap, so sidecars already in the +# field stay valid — a hash change is a data-format change. +# yaml-rust2 F6's `.yaml` exhibit config, so the file a human edits in a +# venue can carry comments. Measured 2026-08-17 against the three +# live options — yaml-rust2 adds 8 crates and 0 proc macros; +# serde_yaml_ng 10; saphyr 20 and SIX proc macros, which +# deny.toml bans outright. (serde_yaml itself was archived in +# 2024 and is not an option.) `default-features = false` drops +# the `encoding` feature and encoding_rs with it, bringing the +# real cost to +5: an exhibit config is read with +# fs::read_to_string, which already requires UTF-8, so charset +# transcoding would be dead weight. +# +# It also earns its place by REMOVING code rather than adding a +# capability: its loader errors on a duplicate mapping key +# instead of last-wins, which is the single rule the JSON side +# had to hand-roll a serde visitor for. +[dependencies] +serde = "1.0.229" +serde_json = "1.0.151" +sha2 = "0.11.0" +yaml-rust2 = { version = "0.11.0", default-features = false } + +# --- Debian package (SPEC §5c) ---------------------------------------------- +# `cargo deb` builds dexd__arm64.deb. The package exists to make +# the dependency set CHECKED rather than merely written down: libmpv's version +# is verified by apt at install time, on a bench, instead of surfacing as a +# black screen in a gallery. +# +# Build: cargo deb (on arm64; CI does it in a debian:trixie container) +# Install: sudo apt install ./dexd_0.1.0_arm64.deb +[package.metadata.deb] +maintainer = "Max Albrecht <1@178.is>" +copyright = "Max Albrecht" +# Ship LICENSE verbatim as the package's copyright file. Without this, cargo-deb +# emits `License: CC0-1.0` alone — true of the source and misleading about the +# binary, which is a GPL-3+ combined work. A package that understates its own +# obligations is the one licensing bug worth going out of the way to avoid. +license-file = "LICENSE" +extended-description = """ +Gapless 4K HEVC loop player for Raspberry Pi gallery installations. Feeds +libmpv an endless in-process byte stream so the decoder never sees EOF and +never seeks, which is what removes the seam at the loop point.""" +section = "video" +priority = "optional" + +# THE POINT OF THE PACKAGE. "$auto" runs dpkg-shlibdeps over the built binary, +# so Depends is DERIVED from the sonames it actually links -- libmpv2 and the +# rest -- and cannot drift from reality the way a hand-written list would. +# Never replace this with a literal list. +# +# The explicit libmpv2 floor is ADDED to it, not a substitute, because the two +# state different things and only one of them is derivable: +# +# $auto alone produced `libmpv2 (>= 0.19.0)`. That is dpkg-shlibdeps working +# correctly -- it reports the oldest version providing the SYMBOLS we call. +# But nothing we depend on is symbol-shaped: --gpu-hwdec-interop=drmprime- +# overlay is a runtime option, and the C1 recovery fix rests on mpv emitting +# END_FILE(reason=stop) for a `loadfile replace`. Both are BEHAVIOUR, both +# were verified against 0.40 (the version trixie ships and all three reviews +# probed), and a device with 0.19 would install cleanly and then misbehave -- +# the silent-wrongness class this package was built to close. +# +# So: keep $auto for what can be derived, and state what cannot. Raise the +# floor whenever a fix is verified against a newer mpv. +# `adduser` is here because postinst calls it -- lintian catches maintainer +# scripts that use a tool the package does not depend on, which on a minimal +# image is a failing install rather than a missing convenience. +depends = "$auto, libmpv2 (>= 0.40.0), adduser" + +assets = [ + # /usr/bin, not /usr/local/bin: Debian policy reserves /usr/local for the + # local administrator, and a package that writes there is a policy + # violation. The systemd unit was updated to match. + ["target/release/dexd", "usr/bin/", "755"], + # F6's privileged sibling. Shipped (unlike dex-sidecar below) because it + # runs ON THE DEVICE, as root, whenever an operator changes the exhibit -- + # see src/bin/dex-exhibit-apply.rs and man dex-exhibit-apply. + ["target/release/dex-exhibit-apply", "usr/bin/", "755"], + ["deploy/dex-wait-hdmi", "usr/bin/", "755"], + # F6's conffile default. `auto`/`none` is deliberately inert -- it asks + # mpv for the connector-preferred mode and forces nothing at the KMS + # layer -- so a fresh install with the stock config never black-screens a + # panel it has not been told about; see deploy/exhibit.json.default. + ["deploy/exhibit.json.default", "etc/dex/exhibit.json", "644"], + ["README.md", "usr/share/doc/dexd/", "644"], + ["deploy/man/dexd.1", "usr/share/man/man1/", "644"], + ["deploy/man/dex-wait-hdmi.1", "usr/share/man/man1/", "644"], + ["deploy/man/dex-exhibit-apply.1", "usr/share/man/man1/", "644"], + ["deploy/lintian-overrides", "usr/share/lintian/overrides/dexd", "644"], +] + +# dex-sidecar is deliberately NOT shipped: it is an INGEST tool (it validates +# a sidecar against an asset while building one), and ingest happens on a +# workstation, never on the gallery device. + +# F6 -- /etc/dex/exhibit.json is a CONFFILE, not merely a shipped file: dpkg's +# conffile semantics guarantee a hand-edited exhibit config survives a package +# upgrade (and prompts on a genuine conflict) BY MACHINERY, not by convention +# -- the exact property F6 exists to give the display mode, the same F3 +# already gives the sidecar's file layout implicitly by never touching /etc. +conf-files = ["/etc/dex/exhibit.json"] + +maintainer-scripts = "deploy/maintainer-scripts" +# A non-native package without a Debian changelog is a lintian error, and +# the changelog is where a device operator can see what shipped. +changelog = "deploy/changelog" + +[package.metadata.deb.systemd-units] +unit-scripts = "deploy" +# Enable at install so a device that is only ever power-cycled comes back +# playing, but do not start during install: the unit takes DRM master, and an +# operator installing over SSH from a console session should choose the moment. +enable = true +start = false + +# Match release: an unwind crossing the FFI boundary, or main unwinding while +# mpv threads are live, is UB territory. rustc >= 1.81 aborts at the extern "C" +# boundary, but main() unwinding out from under a registered callback is a +# separate hazard, and a dev build should not have a weaker safety story than +# the shipping one. +[profile.dev] +panic = "abort" + +[profile.release] +opt-level = 2 +lto = true +panic = "abort" diff --git a/packages/dexd/LICENSE b/packages/dexd/LICENSE new file mode 100644 index 0000000..1acfde9 --- /dev/null +++ b/packages/dexd/LICENSE @@ -0,0 +1,75 @@ +dexd — licensing +================ + +INTENT +------ +The dex project aims to release all its works to the public domain, as far as is +legally and practically possible. Public domain dedication is not uniformly +recognised across jurisdictions, and the instruments that come closest carry +different trade-offs for software and for content — so "as far as possible" +resolves to two licences rather than one: + + content CC0-1.0 (test cards, video masters, branding) + software MIT-0 (this crate, and dex code generally) + +Both are maximally permissive: no attribution required, no conditions on use, +modification, or redistribution. The split is not a change of intent, it is the +same intent surviving contact with two practical facts: + + * CC0 expressly disclaims any patent grant, and Fedora has therefore + disallowed it for CODE since 2022. Using it for software would forfeit a + major distribution channel for no benefit. + * MIT-0 is MIT with the attribution clause removed. It is OSI-approved, + universally accepted by distributors, and asks nothing of a reuser — the + closest a software licence gets to "do whatever you like" while remaining a + licence rather than a dedication. + +GPL appears in the project only where it was inherited (pi-gen → dex-os), never +by choice. + + Copyright: Max Albrecht + SPDX-License-Identifier: MIT-0 + +Machine-readable licensing follows the REUSE specification (reuse.software): +see REUSE.toml and LICENSES/. `reuse lint` runs in CI. + +WHAT THIS DOES NOT COVER: THE DISTRIBUTED BINARY +------------------------------------------------ +dexd links libmpv, and Debian ships that as GPL-3+: + + "While the mpv source code is distributed mostly under the GPL-2+ or + LGPL-2.1+ licenses, the binaries are distributed under the GPL-3+ license + because they are linked to the GPL-3+ libsmbclient library." + -- /usr/share/doc/libmpv2/copyright, Debian trixie + +So the .deb built from this crate is a combined work and must be conveyed under +GPL-3+, including its offer of source. That is not a conflict, and not a +retreat from the intent above: MIT-0 is GPL-compatible, so this source imposes +no condition on anyone. The two statements are about different things -- + + this source MIT-0 (what we wrote, freely reusable, no conditions) + the shipped binary GPL-3+ (what it links, inherited) + +The binary's terms are not ours to set: they follow from how each distributor +builds mpv. An mpv built without libsmbclient is LGPL-2.1+, and the resulting +combination is correspondingly less restrictive. Anyone taking this crate's +source and building it against such a libmpv carries none of the above. + +-------------------------------------------------------------------------- + +MIT No Attribution + +Copyright 2026 Max Albrecht + +Permission is hereby granted, free of charge, to any person obtaining a copy of this +software and associated documentation files (the "Software"), to deal in the Software +without restriction, including without limitation the rights to use, copy, modify, +merge, publish, distribute, sublicense, and/or sell copies of the Software, and to +permit persons to whom the Software is furnished to do so. + +THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, +INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A +PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT +HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION +OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE +SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. diff --git a/packages/dexd/LICENSES/MIT-0.txt b/packages/dexd/LICENSES/MIT-0.txt new file mode 100644 index 0000000..a4e9dc9 --- /dev/null +++ b/packages/dexd/LICENSES/MIT-0.txt @@ -0,0 +1,16 @@ +MIT No Attribution + +Copyright + +Permission is hereby granted, free of charge, to any person obtaining a copy of this +software and associated documentation files (the "Software"), to deal in the Software +without restriction, including without limitation the rights to use, copy, modify, +merge, publish, distribute, sublicense, and/or sell copies of the Software, and to +permit persons to whom the Software is furnished to do so. + +THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, +INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A +PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT +HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION +OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE +SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. diff --git a/packages/dexd/README.md b/packages/dexd/README.md new file mode 100644 index 0000000..dbe4c3e --- /dev/null +++ b/packages/dexd/README.md @@ -0,0 +1,662 @@ +# dex-loop + +Gapless HEVC looper for the Raspberry Pi. One process, no shell, and a dependency +set small enough to read — all of it declared (SPEC §5c). It links libmpv, so the +honest count is 228 shared objects, not zero. + +```bash +# once, at ingest -- mpv needs a raw elementary stream, not MP4: +ffmpeg -i card.mp4 -c:v copy -bsf:v hevc_mp4toannexb -f hevc loop.265 +# and the binding sidecar (fps + hash travel WITH the asset): +printf '{"fps":"30","sha256":"%s","width":3840,"height":2160}\n' \ + "$(shasum -a 256 loop.265 | cut -d' ' -f1)" > loop.265.json # Linux: sha256sum + +# once, on the device -- WHICH asset and WHICH display mode are both exhibit +# config, not asset metadata (F6). /etc/dex/exhibit.json, or .yaml if you want +# comments; the .deb ships a stock .json. +printf '{"asset":"/opt/dex/loop.265","display_mode":"3840x2160@30"}\n' \ + | sudo tee /etc/dex/exhibit.json + +# then -- no arguments: the exhibit says what to play and how +dex-loop +``` + +## What it does + +Feeds libmpv a byte stream that never ends, by registering a `loop://` protocol +whose read callback wraps back to byte 0 instead of ever returning EOF. + +That single property is the whole point. Every mpv looping mechanism stalls at +the wrap because each one makes the decoder **re-enter** the file; measured on a +Pi 4 against an HDMI capture card: + +| mechanism | hold at the wrap | +|---|---| +| `--loop-file=inf` | 83 ms, every loop | +| `--ab-loop-a/b` | 83 ms, identical | +| `--playlist` + `--prefetch-playlist` | 117-133 ms, worse | +| `ffmpeg -stream_loop` + `vout_drm` | 67-217 ms, 3 per loop | +| `--loop-file=inf` on a raw `.265` | freezes on the last frame | +| **endless stream (this)** | **none** | + +It works because a correctly-built bench asset has a **closed GOP with an IDR at +frame 0**, so presenting byte 0 straight after the last byte is an ordinary +mid-stream keyframe rather than a seek. + +## The stack, and where each layer stops + +``` + ┌────────────────────────────────────────────────────────────────────────┐ + │ dex-loop OURS — ~370 lines of Rust │ + │ loop:// stream callback; read_fn never returns 0, so there is no EOF │ + │ option set · END_FILE treated as fatal · mpv log capture │ + └────────────────────────────────────────────────────────────────────────┘ + │ mpv client API (mpv_stream_cb_add_ro, loadfile) + ┌────────────────────────────────────────────────────────────────────────┐ + │ libmpv 0.40 PRESENTATION + TIMING │ + │ vsync-locked scheduling (--video-sync=display-resample) │ + │ vo=gpu / gpu-context=drm — modeset, atomic commits, page flips │ + │ --gpu-hwdec-interop=drmprime-overlay — hands the frame to a KMS plane│ + └────────────────────────────────────────────────────────────────────────┘ + │ libav* API + ┌────────────────────────────────────────────────────────────────────────┐ + │ FFmpeg — libavformat / libavcodec DEMUX + DECODE │ + │ raw Annex-B HEVC demux (carries no timestamps, hence --fps) │ + │ hwaccel: V4L2 Request API (stateless) │ + └────────────────────────────────────────────────────────────────────────┘ + │ ioctl + ┌────────────────────────────────────────────────────────────────────────┐ + │ Linux kernel │ + │ │ + │ V4L2 stateless decoder ──── dma-buf ────▶ DRM / KMS (vc4) │ + │ rpi-hevc-dec (DRM_PRIME) planes · CRTC pixelvalve-2│ + │ /dev/video19 connector HDMI-A-1 │ + └────────────────────────────────────────────────────────────────────────┘ + │ + ┌────────────────────────────────────────────────────────────────────────┐ + │ Broadcom BCM2711 │ + │ HEVC decode block ──▶ SAND-tiled NV12, in CMA │ + │ │ no conversion: the HVS reads SAND │ + │ HVS (compositor) ──▶ PixelValve ──▶ HDMI PHY @ 297 MHz TMDS │ + └────────────────────────────────────────────────────────────────────────┘ +``` + +**Read the arrows carefully: the decoded frame never travels back up the stack.** +It is written once into CMA by the HEVC block and stays there. What moves upward +is a *dma-buf file descriptor*; libmpv hands that to KMS as a framebuffer, and +the HVS scans the same memory out. Control flows down; pixels move sideways at +the bottom. + +That is the whole performance story, and every configuration that lost was one +that dragged the frame upward: + +| Path | What touches the frame | Result | +|---|---|---| +| `--hwdec=drm-copy` | CPU detiles SAND into linear NV12 | 14.3 fps | +| `--gpu-hwdec-interop=drmprime` (GL) | V3D samples SAND as a texture | ~5 fps | +| **`drmprime-overlay`** | **nothing — fd straight to a KMS plane** | **29.1 fps** | + +It also explains the two upstream gaps we hit: GStreamer's `kmssink` cannot bind +a SAND dma-buf at all (so it falls back to dumb-buffer copies and OOMs at 4K), +and mpv's `--hwdec=drm` with `--vo=drm` silently selects the software decoder +rather than the plane path. + + +## How a frame actually travels + +The layer stack above answers "who owns what". It cannot show the *path*, because +the path is not a stack — it is a loop. Compressed bytes go **down**, a file +descriptor comes **up**, and the pixels never move at all: + +```mermaid +flowchart TB + subgraph APP["dex-loop — ours, ~370 lines"] + LOOP["loop:// read_fn
never returns 0 → no EOF"] + end + + subgraph MPV["libmpv 0.40 — presentation + timing"] + SCHED["vsync scheduling
display-resample"] + VO["vo=gpu · gpu-context=drm
atomic commit / page flip"] + end + + subgraph FF["FFmpeg — libavformat / libavcodec"] + DEMUX["demux raw Annex-B HEVC"] + HWA["hwaccel: V4L2 Request API"] + end + + subgraph KRN["Linux kernel"] + V4L2["rpi-hevc-dec
/dev/video19"] + KMS["DRM / KMS — vc4
plane · CRTC · connector"] + end + + subgraph HW["Broadcom BCM2711"] + HEVC["HEVC decode block"] + CMA[("CMA buffer
SAND-tiled NV12")] + HVS["HVS compositor"] + PHY["PixelValve → HDMI PHY
297 MHz TMDS"] + end + + LOOP -->|compressed bytes| DEMUX + DEMUX --> HWA + HWA -->|slice params + bitstream| V4L2 + V4L2 --> HEVC + HEVC ==>|writes pixels ONCE| CMA + + CMA -.->|dma-buf fd| V4L2 + V4L2 -.-> HWA + HWA -.->|AVFrame w/ DRM_PRIME| VO + SCHED --> VO + VO -.->|"same fd, as a framebuffer
(drmprime-overlay)"| KMS + + CMA ==>|scanned out in place| HVS + KMS -->|plane config| HVS + HVS ==> PHY + + classDef pixels fill:#1f6feb22,stroke:#1f6feb,stroke-width:2px + class CMA,HEVC,HVS,PHY pixels +``` + +**Legend.** Thin arrows are compressed data and control. Dotted arrows carry a +*file descriptor*, not pixels. Thick arrows are the only places actual pixels +move — and note they all live in the bottom box. + +The frame is written into CMA once by the HEVC block and is scanned out from +that same memory by the HVS. Everything in between is passing a handle around. +That is what "zero-copy" means here, concretely. + +## Measured + +4K30 HEVC on a Pi 4 (trixie), captured over HDMI and decoded frame-by-frame: + +``` +dwell={1: 1500} wraps=19 held_at=[] # every capture a distinct frame index +``` + +Identical to `while true; do cat loop.265; done | mpv -`, which is the reference +this replaces. Memory is flat (221 MB RSS, unchanged over 20 s; mpv's demuxer +cache is bounded by `--demuxer-max-bytes`), and a 3.5 h soak of the equivalent +shell pipeline showed no held frames and no growth. + +## Why not the shell pipeline, then + +`while true; do cat loop.265; done | mpv -` is correct but not shippable: + +* a process per loop -- roughly 29,000/day for a 3 s card +* if mpv ever exits, `cat` dies on SIGPIPE and the `while` loop **spins hot**, + burning a core indefinitely. A dex player would go from "video stopped" to + "CPU pegged forever" + +## Why not Python + +The identical design in Python (`../scripts/loop-player.py`, kept for reference) +displays correctly but runs at **0.6x realtime**, with frames held at random +points across the loop -- the signature of a starved feed rather than a seam. +The stream has to deliver ~5 MB/s against hard per-frame deadlines, and a ctypes +callback under the GIL cannot promise that. libmpv's `stream_cb` is a C API, so +in Rust the same callback is a `copy_nonoverlapping`. + +## Options and the sidecar + +`dex-loop ` requires `.json` next to the asset — written at ingest, +binding the two facts that are undetectable when wrong: the frame rate (a raw stream +has no timestamps; a wrong rate plays slow forever with every metric nominal) and the +exact bytes (sha256 — a truncated copy glitches at every wrap). + +```json +{"fps":"30","sha256":"<64 hex>","width":3840,"height":2160} +``` + +`fps` is a string (`"30"`, `"29.97"`, `"30000/1001"`), passed verbatim to mpv. +Required: `fps`, `sha256`. Optional: `width`, `height`, `source`, `encoder_cmd`. +Unknown keys are ignored — but their *values* still must be strings or unsigned +integers, the same subset `fps`/`sha256` use (arrays/booleans/null/nested objects +are refused everywhere, not just on the required keys). String escapes include +`\uXXXX` (standard JSON, surrogate pairs included), so a `source`/`encoder_cmd` +value produced by a default-safe serializer (Python's `json.dumps`, Go's +`encoding/json`) does not break startup on a non-ASCII byte. Anything outside the +subset refuses startup — fail closed. + +| flag | meaning | +|---|---| +| `--fps F` | optional cross-check; must equal the sidecar fps exactly, or startup is refused | +| `--mode WxH@R` | optional cross-check against the exhibit config's `display_mode` (F6); refused if it disagrees, naming both. Under `--bench-no-sidecar` it is the only source (default there: `auto`) | +| `--exhibit-config PATH` | F6: path to the exhibit config, default `/etc/dex/exhibit.json` — see below | +| `--bench-no-sidecar` | BENCH ONLY: skip the sidecar AND the exhibit config, and take `--fps`/`--mode` as given (both flags required alongside `--fps` — the escape hatch is a deliberate two-flag act) | +| `--force-recovery-after-secs N` | T7, BENCH ONLY: force a tier-0 in-place recovery N seconds into playback, whether or not anything has stalled — a live-fire probe for F1's recovery command. Requires `--bench-no-sidecar` (refused otherwise), so it can never end up armed against a real, sidecar-bound deployment asset | +| `--bench-wedge-after-secs N` | F10, BENCH ONLY: deliberately hang the event thread FOREVER N seconds after startup, simulating the one hazard class F1/F9 cannot see (an event-thread hang outside any mpv call) — proves whether a systemd watchdog (`WatchdogSec=`) actually fires. The process never recovers on its own once armed and fired. Requires `--bench-no-sidecar` (refused otherwise) | +| `--opt K=V` | pass any extra mpv option (repeatable) | +| `--no-defaults` | omit the built-in Pi 4 zero-copy option set | + +Startup gates, in order: exhibit config parses → resolved against `--mode` → +kernel cmdline agrees with its `kms_force` → the resolved resolution exists on the +connector's own mode list → asset readable and non-empty → sidecar parses → fps +resolved → sha256 matches → leading NALs are VPS/SPS/PPS + IDR (open-GOP/CRA assets +are refused — the wrap premise is "IDR at frame 0") → every `--opt`/default/`drm-mode`/ +`drm-connector` mpv option is accepted. Display gates run FIRST and before the asset is +even read: they are the cheapest checks and a wrong-panel install is worth catching even +when the asset path is also wrong. Exit codes: **2** = refused before playback (fix the +asset/invocation/exhibit config — including a rejected mpv option; restarting cannot +help), **1** = playback/runtime failure (the supervisor restarts). Every start logs +`dex-loop ()`, the resolved display binding (`display ... (...), +connector ..., kms-force ...`), and a heartbeat line (`wraps=`, `temp=`, `frame-drops=`, +`pos-age=`, `watchdog=`) at boot and every 10 minutes. The drop counters +(`frame-drops=`, `vo-delayed=`) are cumulative since process start and read `n/a` +until the first value arrives from mpv, or `off` if the subscription failed at +startup — never a direct property read, so the heartbeat itself can never block on a +wedged mpv core. `watchdog=` reads `inert` off systemd (Mac/bench/CI — the +overwhelmingly common case) or `armed pings-dropped=N` under a unit with +`WatchdogSec=` set (F10, see **Deployment** below) — see `src/watchdog.rs`'s module +doc for the full design. + +### Exhibit config (F6) + +The display mode is a property of the **installation** (the venue's panel), not the +asset — the same 2160p30 asset plays correctly on a 4K projector and a 1440p desktop +monitor, and one measured sink (an Elgato Cam Link 4K) advertises 4K30 as its own +*preferred* EDID timing while the DRM driver still declines to build the mode +unforced. So the mode lives in an exhibit config — a dpkg **conffile** (a hand +edit survives a package upgrade), parsed with the same hardened flat subset +grammar as the sidecar, but with **unknown keys refused** rather than ignored: this +file has no independent producer to stay compatible with, so a typo (`kms_forse` for +`kms_force`) must be a startup refusal, not a silently dropped force. + +**Two formats, and the file extension decides which.** `.json` is strict JSON — the +format the `.deb` ships, and what a machine should write. `.yaml`/`.yml` is YAML, for +the case this file actually exists to serve: a human editing it in a venue, possibly +on a phone over SSH, who wants a comment next to the value explaining why this panel +needs a force. + +```json +{"display_mode":"3840x2160@30","kms_force":"3840x2160@30","connector":"HDMI-A-1", + "display":"Elgato Cam Link 4K","venue":"gallery east wall","note":"..."} +``` + +```yaml +display_mode: 3840x2160@30 # what mpv is asked for +kms_force: 3840x2160@30 # what the kernel cmdline must carry +connector: HDMI-A-1 +display: Elgato Cam Link 4K +venue: gallery east wall +note: vc4 builds no 4K mode from this sink's EDID unforced +``` + +Both files above mean exactly the same thing, and produce the same struct: only the +~40 lines that turn text into a key/value list differ, and one shared function does +every mapping, default and grammar check for both. That is what keeps the formats from +drifting into two dialects. + +**The exhibit names its own asset.** `asset` is what makes this an *exhibit* file +rather than a display file: an exhibit is a pairing of a venue with an artwork, and +until 2026-08-17 it could only express the venue half — `ExecStart` hardcoded +`/opt/dex/loop.265`, so changing the artwork meant overwriting that one path or +editing a unit the package owns. Now several assets can sit in `/opt/dex` and the +exhibit picks one: + +```yaml +asset: /opt/dex/spring-2026.265 +display_mode: 3840x2160@30 +``` + +`ExecStart` is therefore just `/usr/bin/dex-loop`, with no arguments at all: the unit +says *how* to run the player, and one editable file says *what* it plays and *where*. +A path given on the command line anyway must agree with the config or startup refuses +naming both — the same cross-check `--mode` gets. And if neither names an asset, +startup refuses rather than falling back to `/opt/dex/loop.265`: that fallback would +silently play last season's artwork for someone who mistyped the key, which is the +worst guess this program could make. + +**The extension is honoured, not sniffed.** YAML is a superset of JSON, so parsing +everything with the YAML parser would work — and would be wrong: it would accept +comments and unquoted keys inside a file named `.json`, and that file would then break +`jq`, `python -m json.tool`, and every other consumer that trusts the name. An +extension is a promise about what the bytes are. A `.json` file containing YAML is +therefore refused, while strict JSON inside a `.yaml` file is fine (it *is* YAML) — +which is what lets a machine emit one format under either name. + +By default the player looks for `/etc/dex/exhibit.yaml`, then `/etc/dex/exhibit.json`. +**Exactly one may exist.** Both present is refused naming both, rather than resolved by +precedence — "the other file wins silently" is how someone edits a config all afternoon +while the player reads a different one, the exact drift F6 exists to end. Point +`--exhibit-config` at a specific file to override. + +So **switching to YAML is two commands**, because the package installs the `.json`: + +```bash +sudoedit /etc/dex/exhibit.yaml # write it +sudo rm /etc/dex/exhibit.json # remove the shipped one, or startup refuses +``` + +The refusal names that second command, so getting it wrong costs one restart, not a +debugging session. + +| key | required | meaning | +|---|---|---| +| `asset` | see below | absolute path to the file to play, e.g. `"/opt/dex/loop.265"` — **which artwork this exhibit shows**. Optional in the file, but something must supply it: a path on the command line agrees or contradicts it, and if neither names an asset, startup refuses rather than guessing one | +| `display_mode` | yes | `"auto"` or `"WxH@R"` (R a positive INTEGER, same rule as `kms_force` — bench-verified 2026-08-17: mpv's `--drm-mode` rejects a rational refresh at option parse (-7, a guaranteed restart loop) and silently rounds a decimal to the integer vrefresh, so non-integer forms are refused at config parse instead) — what `dex-loop` asks mpv for via `--drm-mode` | +| `kms_force` | no (default `"none"`) | `"none"` or `"WxH@R"`/`"WxH@RD"` (integer R only — the kernel `video=` grammar has no fractional refresh) — what the kernel cmdline is expected to carry for this connector | +| `connector` | no (default `"HDMI-A-1"`) | which DRM connector, e.g. `"HDMI-A-2"` | +| `display`, `venue`, `note` | no | informational, logged verbatim at every start — this is where the *why* that used to live in a `config.txt` comment block belongs now (in a `.yaml` file, a real comment works too) | + +Values are strings or non-negative integers, one level deep, in either format. YAML +extras that JSON could not express — nested mappings, lists, anchors, `---` multi-doc +streams — are refused, so a config cannot mean something different depending on which +extension it was saved under. One YAML-specific note, measured rather than assumed: +`yaml-rust2` resolves scalars close to the **YAML 1.2 core schema**, so only `true`/`false` +are booleans and the "Norway problem" (`no` → `false`) does not arise here — `kms_force: +no` arrives as the string `"no"` and is refused by the grammar, naming the valid values. +(*Close to*: its null resolution is `""`/`~`/`null` only, so the core schema's `Null` and +`NULL` arrive as ordinary strings. Driven, not read off the spec.) + +**Not** YAML's own JSON schema, which sounds like it should be the tool for the `.json` +path and is not. YAML's three schemas (failsafe, JSON, core) govern only how an untagged +*scalar* resolves to a type; none of them restricts syntax. A YAML parser in JSON-schema +mode still accepts comments, block style and anchors — so it could not deliver the promise +a `.json` name makes, which is the whole job here. Hence a real JSON parser for `.json`. +It would be wrong for `.yaml` too, in the other direction: under the JSON schema a plain +scalar matching none of null/bool/int/float is an *error*, so `display_mode: 3840x2160@30` +would not resolve at all. (Moot anyway — `yaml-rust2` exposes no schema selection.) + +Outside `--bench-no-sidecar`, a missing or invalid exhibit config refuses startup — +the same fail-closed contract F3 has for frame rate: a wrong guess plays wrong forever +with every metric green. Two more gates keep the config honest against reality: + +* **The cmdline gate.** If `kms_force` disagrees with the `video=:...` + token actually present in the *running* kernel's `/proc/cmdline`, startup refuses, + naming both values and `sudo dex-exhibit-apply`. An edited `exhibit.json` with no + matching apply-and-reboot is exactly this disagreement, caught at the next start + instead of black-screening the venue for the run of the show. +* **The sysfs mode pre-flight.** If `display_mode` names a resolution absent from + `/sys/class/drm/card*-/modes`, startup refuses, naming the requested + resolution and what the connector actually offers — the wrong-panel case, caught + before the asset is even opened. **Resolution only:** the kernel's `modes` file + has no refresh column, so the `@R` half of `display_mode` is validated by grammar + alone and is settled by mpv's `--drm-mode` matching at VO init — a refresh the + connector does not offer fails there (bench-verified: `Could not find mode + matching 3840x2160@60`, a 2s-cadence restart loop), with mpv's error in the + journal rather than this gate's. See `man dex-exhibit-apply`. + +Changing the mode is two steps: edit the file, then run `sudo dex-exhibit-apply` +(a separate, privileged binary — `dex-loop` itself runs unprivileged under +`ProtectSystem=strict` and must never write boot config) to reconcile +`cmdline.txt`'s `video=` token, idempotently, preserving every other token and every +other connector's token untouched. It prints `REBOOT REQUIRED` iff `cmdline.txt` +actually changed. Full schema and the swap-the-panel worked example: +`man dex-exhibit-apply`. + +**Migration from the pre-F6 arrangement.** Before this, the mode lived in three +places that drifted out of sync within days: a hand-edited `cmdline.txt`, a systemd +drop-in overriding `--mode` in `dex-loop.service`, and a `config.txt` comment block +carrying the reasoning. `dex-loop.service` no longer passes `--mode`; postinst warns +(non-fatal) if a leftover drop-in still mentions it. Remove it: +`sudo rm /etc/systemd/system/dex-loop.service.d/*.conf && sudo systemctl daemon-reload`. + +The defaults encode the measured zero-copy path: the Pi's decoder emits +Broadcom SAND-tiled NV12, and the display scans SAND out natively **only** +straight onto a KMS plane. `--gpu-hwdec-interop=drmprime-overlay` is what puts +it there; plain `drmprime` imports into GL and is 2x slower (5 fps vs 29). + +## Where this is developed + +The split is imposed by hardware, not chosen: + +| | Mac | Pi 4 | +|---|---|---| +| DRM/KMS — the player's whole output path | **does not exist on macOS** | yes | +| Broadcom HEVC decoder | no | yes | +| Cam Link — the measurement instrument | **yes** | no | +| What runs | pure-logic tests (36), the capture/analysis harness | build, the player, the FULL test matrix (58) | + +So the player can only ever run on the Pi, and it can only ever be *measured* from +the Mac. Neither machine alone is enough, and no amount of tooling changes that. + +**Both are git checkouts of this repo; they reconcile through the remote, never on +disk.** An earlier arrangement rsync'd the crate to the Pi and cost real time: the +Pi copy had no `.git`, so the build-identity work invented a `.dex-build-id` stamp +file to recover what `git rev-parse` already knew. With a real checkout the startup +line reads `dex-loop 0.1.0 (8b8c00ef5eef)` and the stamp file is unnecessary. + +```bash +# on the Pi, once +git clone --no-recurse-submodules https://github.com/KTE/dex.git ~/dex +cd ~/dex && git checkout experiment/4k-hevc-perfect-loop +# submodules skipped deliberately: packages/example-content carries video masters +# the bench does not need. Clone is ~9 MB. +``` + +Generated bench assets (`loop4k.265`, encoded cards) live outside the checkout in +`~/bench/` — they are build products, not repo content. + +## Building + +```bash +sudo apt install rustc cargo libmpv-dev # trixie: rustc 1.85 +cargo build --release # ~30 s on a Pi 4, no dependencies +``` + +## Testing + +```bash +# Mac (no libmpv): type-check everything, run the pure-logic tests +cargo check --all-targets && cargo test --lib + +# Pi (dexpi@dexpi4.local, crate mirrored at ~/bench/dex-loop): full suite +cargo test +``` + +The integration tests (`tests/cli.rs`, `tests/ffi_constants.rs`) link libmpv and spawn +the real binary — Pi only. They never touch the display: every playback-reaching +invocation uses `--no-defaults --opt vo=null --opt vid=no --opt aid=no`, so they are +safe to run while a soak owns the screen (`nice -n 19 cargo test` to keep builds off +the soak's CPU). See IMPLEMENTATION-PLAN.md for the full hardening rationale. + +One deliberate exception: `force_recovery_survives_against_real_mpv` (T7's live-fire +probe for F1's recovery command, PLAN.md) needs REAL decode — the whole point is +observing a health-check tick against playback that is actually advancing — so it +skips `vid=no` and is `#[ignore]`d rather than part of `cargo test`'s default run. +Run it deliberately, by name, on a Pi with nothing else on the display and no soak +using the CPU: + +```bash +cargo test --test cli force_recovery_survives_against_real_mpv -- --ignored --nocapture +``` + +**CI runs this test too**, as its own step in `.github/workflows/dex-loop-deb.yml` +(after `Test`, before `Build package`), by exact name — `debian:trixie`'s software +HEVC decoder makes real decode available with no device, DRM or display needed. It +stays `#[ignore]`d (so a Pi mid-soak's plain `cargo test` still never picks it up); +the CI step is what runs it automatically instead. What this proves: the process +survives its own recovery under software decode with no display. What it does not +prove: the picture actually comes back with `hwdec=drm` / `drmprime-overlay` / the +DRM plane swap on real hardware. + +**2026-08-17, that manual on-Pi run happened.** `--force-recovery-after-secs` +against `~/bench/loop4k.265` with defaults ON (the real `hwdec=drm` / +`drmprime-overlay` / DRM-plane path, not CI's `vo=null` stand-in): the forced +recovery fired, was absorbed cleanly, and the process kept running — sampled +CPU stayed at realtime-decode levels (~25–27%, not idle) for minutes +afterward, with zero FATAL lines and zero organic second recovery. What it +does **not** yet prove: an independent camera witness that the picture itself +came back, which the Cam Link capture instrumentation could not get running +this session (a macOS-side AVFoundation hang, not a player issue). See +PLAN.md's T7 entry for the full record, including exactly what remains open. + +## Reviewed 2026-08-15 + +Three adversarial reviews (Rust/unsafe soundness, libmpv API contract, gallery +operations). Fixed since: the event loop ignored `MPV_EVENT_END_FILE`, and +because `mpv_create` enables idle mode by default, ANY playback failure left the +process alive in idle **forever** on a black screen -- strictly worse than +crashing, and the most likely gallery failure (projector not awake at boot) hit +exactly that path. `user_data` pointed at a stack local; `read_fn` could return 0 +(= final EOF) for a zero-length request; `seek_fn`/`size_fn` used `-1` rather +than the documented `MPV_ERROR_UNSUPPORTED`; software decode could fall back +silently; mpv's diagnostics went nowhere. + +Still outstanding (see the story): asset+fps binding via an ingest sidecar, +read-only rootfs, and a frame-advance watchdog. + +**Second pass, same day**, after F3/F4/F7 landed and three more adversarial reviews +ran against that hardening: `--fps`/`--mode` with a missing value no longer evaporate +silently (they refused via `usage()`, matching `--opt`); a rejected mpv option now +exits 2, not 1 (it is a deterministic, operator-fixable bad invocation, not a runtime +failure the supervisor's restart loop could resolve); the event loop now treats +`MPV_EVENT_QUEUE_OVERFLOW` as fatal too, since mpv's internal event ring silently +drops events — potentially an END_FILE — once it chokes; sidecar strings support +`\uXXXX` escapes (surrogate pairs included), because a default-safe JSON serializer +escapes every non-ASCII byte that way, including inside informational keys this +player does not even interpret; the heartbeat's sub-zero temperature formatting no +longer drops the sign; and `build.rs` now also reruns on source changes (not just +`.git/HEAD`) and can take its git hash from a `.dex-build-id` stamp file, since the +Pi build is an rsync mirror, not a checkout, and `git rev-parse` there always failed. + +## Deployment + +Ship a **`.deb`**, don't build on the device. + +```bash +cargo deb # -> target/debian/dex-loop_0.1.0_arm64.deb +sudo apt install ./dex-loop_0.1.0_arm64.deb # apt, not dpkg -i: it resolves Depends +``` + +The package installs `dex-loop`, `dex-exhibit-apply` and `dex-wait-hdmi` to +`/usr/bin`, installs and enables the unit, creates the unprivileged `dex` user +with `video`/`render`, creates `/opt/dex`, and ships a stock +`/etc/dex/exhibit.json` (`display_mode: "auto"`, `asset: "/opt/dex/loop.265"`, +conffile — a hand edit survives a package upgrade). The display half is +deliberately inert; the `asset` is the path `ExecStart` used to hardcode, so a +stock install behaves exactly as it did before F6. It does **not** start the unit (that takes DRM +master, which an operator on an SSH console should time themselves) and it +ships **no asset** — the video and its sidecar are content, and baking one in +would mean rebuilding the software to change the artwork: + +```bash +scp loop.265 loop.265.json :/opt/dex/ +sudoedit /etc/dex/exhibit.json # F6: set display_mode (and kms_force if + # the sink needs one -- see README's + # "Exhibit config (F6)" section above), + # and `asset` if the file is not + # /opt/dex/loop.265 +sudo dex-exhibit-apply # reconciles cmdline.txt; reboot if it says to +sudo systemctl set-default multi-user.target # no desktop; nothing else may own DRM +sudo systemctl start dex-loop +``` + +**Why a package rather than `cargo build` on the Pi** (SPEC §5c): `Depends:` is +derived by `dpkg-shlibdeps` from the sonames the binary actually links, so a +libmpv ABI mismatch is refused by apt at install time, on a bench. Before, it +was checked nowhere — a black screen in a gallery was the first symptom. The +derived half cannot drift from reality because nobody writes it: + +``` +Depends: libc6 (>= 2.34), libmpv2 (>= 0.40.0) +``` + +The `0.40.0` is *not* derived, and that distinction is worth keeping straight. +Left to itself `dpkg-shlibdeps` says `libmpv2 (>= 0.19.0)` — the oldest libmpv +exporting the symbols we call. But we depend on mpv *behaviour*, not symbols: +`--gpu-hwdec-interop=drmprime-overlay` is a runtime option, and F1's recovery +rests on `END_FILE(reason=stop)` arriving for a `loadfile replace`. Both were +verified against 0.40. So `Cargo.toml` declares `$auto, libmpv2 (>= 0.40.0)`: +derive what can be derived, state what cannot, and raise the floor whenever a +fix is verified against a newer mpv. + +Two consequences worth knowing: + +* **Build environment must equal target environment.** CI + (`.github/workflows/dex-loop-deb.yml`) builds on an `ubuntu-24.04-arm` runner + *inside a `debian:trixie` container* — the runner for native arm64, the + container for the ABI. Linking against Ubuntu's libmpv and installing on + Debian would manufacture the exact mismatch the package prevents. +* **`/usr/bin`, not `/usr/local/bin`.** Debian policy reserves `/usr/local` for + the local administrator; the unit was repointed accordingly. + +`deploy/` carries the systemd unit, the HDMI connector wait, and +`dex-exhibit-apply` (F6). The unit's `StartLimitIntervalSec=0` is load-bearing: +the default rate limit would put the service into a permanent `failed` state +after a burst of crashes, which is the unattended failure this is meant to +prevent. A restart loop always beats a dead screen. + +The primary boot-order fix is at the KMS layer, not in the player: the exhibit +config's `kms_force` (with a trailing `D`) forces the connector to read +`connected` even before a sink is actually attached, so the Pi always believes +the intended mode is present rather than depending on boot order — see +"Exhibit config (F6)" above and `man dex-exhibit-apply`. `dex-wait-hdmi` is the +belt-and-braces fallback for connectors that have not been given a `kms_force`. + +**Watchdog (F10).** The unit also carries `WatchdogSec=180` + `NotifyAccess=main` +— an EXTERNAL actor for the one hazard tier 0/F9 cannot see: the event thread +hanging in our own code that is not an mpv call at all (e.g. `eprintln!` against +a wedged journald). `dex-loop` pings `WATCHDOG=1` once per ~10 s health-check +tick over a non-blocking `AF_UNIX` datagram socket (hand-written, `std` only — +zero new dependencies, see `src/watchdog.rs`); a dropped ping is counted, never +retried inline, and surfaced in the heartbeat (`watchdog=armed pings-dropped=N`). +`journalctl -u dex-loop` shows `watchdog: armed (window 180s, ping cadence 10s)` +at every start when running under the real unit, `watchdog: inert (...)` for any +manual/bench/CI invocation (no `$NOTIFY_SOCKET`). + +To PROVE it fires rather than trust the design (this crate's own standing bar — +see PLAN.md's F10 entry for the full three-part verification and exact +timestamps), on a bench, as an unprivileged user: + +```bash +sudo systemd-run --unit=wedge-test -p Type=simple -p NotifyAccess=main \ + -p WatchdogSec=15 -p Restart=on-failure -p RestartSec=2 \ + -p User=dex -p Group=dex -p SupplementaryGroups=video \ + /usr/bin/dex-loop /opt/dex/loop.265 --bench-no-sidecar --fps 30 \ + --bench-wedge-after-secs 0 --no-defaults --opt vo=null --opt vid=no --opt aid=no + +journalctl -u wedge-test -f # expect, ~15s later: +# wedge-test.service: Watchdog timeout (limit 15s)! +# wedge-test.service: Killing process NNNNN (dex-loop) with signal SIGABRT. +# wedge-test.service: Main process exited, code=killed, status=6/ABRT +# ...then it restarts and repeats. Stop it: systemctl stop wedge-test +``` + +A watchdog never seen to fire is indistinguishable from one wired to nothing. + +**Triage.** Every refusal and every runtime failure is journal-only today (see +PLAN.md's F8): on site this reads as a plain black rectangle, so start with +`journalctl -u dex-loop -n 20` for the last startup line, the gate that refused +(if exit 2), or the `playback ended (reason=..., error=...)` line (if exit 1). + +## Caveats + +* **Raw Annex-B only.** MP4 in, `.265` out, once at ingest (milestone 3). +* **Frame rate is metadata now.** A raw stream has no timestamps, so `--fps` + must accompany the asset. +* **No seeking, by construction.** `seek_fn` and `size_fn` deliberately report + unseekable/unknown, exactly like a pipe: an mpv that believes it can seek will + try to, and seeking is the operation that produces the seam. Fine for dex, + which only ever loops; disqualifying for a general-purpose player. + +## Licensing + +The dex project releases its works to the public domain as far as is legally and +practically possible. Public domain dedication is not uniformly recognised across +jurisdictions, and the instruments closest to it carry different trade-offs for +software and for content — so that single intent resolves to two licences: + +| | | +|---|---| +| **Content** (test cards, video masters, branding) | `CC0-1.0` | +| **Software** (this crate, dex code generally) | `MIT-0` | + +Both are maximally permissive — no attribution, no conditions. The split is the +same intent surviving two practical facts: CC0 expressly disclaims any patent +grant, and **Fedora has disallowed CC0 for code since 2022**, so using it for +software would forfeit a distribution channel for nothing; while `MIT-0` is MIT +minus the attribution clause, OSI-approved and accepted everywhere. GPL appears +in the project only where it was **inherited** (pi-gen → dex-os), never by choice. + +Licensing is machine-readable per the [REUSE](https://reuse.software) +specification — `REUSE.toml` plus `LICENSES/` — because the people who most need +a precise answer are distro packagers, and `reuse lint` gives them one that is +checked rather than asserted. It runs in CI. + +**The shipped `.deb` is a different question.** It links libmpv, which Debian +builds against GPL-3+ libsmbclient, so the binary is a combined work conveyed +under GPL-3+. Not a conflict and not a retreat: `MIT-0` is GPL-compatible, so +this source imposes nothing on anyone. The binary's terms simply are not ours to +set — they follow from how each distributor builds mpv, and an mpv without +libsmbclient is LGPL-2.1+. See [`LICENSE`](LICENSE), which the package embeds +verbatim as its copyright file so the distinction ships with it. diff --git a/packages/dexd/REUSE.toml b/packages/dexd/REUSE.toml new file mode 100644 index 0000000..2d3ff4b --- /dev/null +++ b/packages/dexd/REUSE.toml @@ -0,0 +1,23 @@ +# REUSE (reuse.software) — machine-readable licensing for dexd. +# +# Why bother: dex should eventually be installable from all the usual repos, and +# distro packagers are the people who have to answer "what is this licensed +# under, exactly, per file". REUSE makes that answer machine-checkable instead +# of a prose paragraph someone has to read and trust. `reuse lint` runs in CI. +# +# Why REUSE.toml rather than per-file SPDX headers: the whole crate is one +# author under one licence, so 28 headers would state the same fact 28 times +# above module docs that are already dense. A single declaration is easier to +# keep true — and being true is the entire value. Per-file headers become the +# better choice if this ever mixes licences or gains third-party code. +version = 1 + +SPDX-PackageName = "dexd" +SPDX-PackageSupplier = "Max Albrecht <1@178.is>" +SPDX-PackageDownloadLocation = "https://github.com/KTE/dex" + +[[annotations]] +path = "**" +precedence = "aggregate" +SPDX-FileCopyrightText = "Max Albrecht" +SPDX-License-Identifier = "MIT-0" diff --git a/packages/dexd/build.rs b/packages/dexd/build.rs new file mode 100644 index 0000000..8d7dabc --- /dev/null +++ b/packages/dexd/build.rs @@ -0,0 +1,91 @@ +//! Embed the git commit into the binary so the startup line identifies the +//! exact build. std only — zero dependencies. +//! +//! Three sources, in preference order: +//! 0. `DEX_BUILD_ID` — an environment variable. This exists because the file +//! in (1) is NOT cache-safe: `rerun-if-changed` on a path that did not +//! exist when the cached build ran is "simply never changed" (see the note +//! below), so a CI job that restores a `target/` cache and THEN writes the +//! stamp gets a stale binary reporting the old hash. Observed exactly that +//! on 2026-08-16: the stamp was written, the package still said `nogit`. +//! `rerun-if-env-changed` has no such hole — cargo compares the value. +//! 1. `.dex-build-id` — a one-line stamp file (gitignored) written by an +//! external sync step. This matters because the Pi build is NOT a git +//! checkout: it builds from an rsync mirror (`~/bench/dex-loop`), where +//! `git rev-parse` fails and the shipped binary would otherwise always +//! say "nogit" — every production binary permanently unidentifiable, +//! exactly the failure F7 exists to prevent. Once the sync step writes +//! the source commit hash here before each rsync, the Pi build picks it +//! up automatically. +//! 2. `git rev-parse` against this checkout, when one exists (the Mac side, +//! or any future build that IS a checkout). +//! +//! Falls back to "nogit" only when neither is available. + +use std::process::Command; + +fn main() { + let hash = env_hash() + .or_else(stamp_hash) + .or_else(git_hash_with_dirty) + .unwrap_or_else(|| "nogit".to_string()); + println!("cargo:rustc-env=DEX_GIT_HASH={hash}"); + + // Cargo compares the VALUE of this variable, so it invalidates correctly + // even from a warm cache — unlike a rerun-if-changed path that appears + // where none existed. + println!("cargo:rerun-if-env-changed=DEX_BUILD_ID"); + + // Re-run on anything that could change the identity above. A missing + // path is simply never "changed" -- fine, since not every build + // environment has all three (the Pi mirror has no .git; a build with no + // stamp file has no .dex-build-id). + println!("cargo:rerun-if-changed=.dex-build-id"); + // Source edits must invalidate the cached hash: emitting ANY + // rerun-if-changed replaces Cargo's default "rerun on any source + // change", so without this line, editing src/*.rs and rebuilding WITHOUT + // committing keeps reporting the previous (now stale) hash, with no + // +dirty -- misattributing whatever the bench measures to code that + // isn't in the binary. + println!("cargo:rerun-if-changed=src"); + println!("cargo:rerun-if-changed=Cargo.toml"); + // Re-run when HEAD moves (crate sits 2 levels below the repo root). The + // paths may not exist in a non-checkout build; that is fine. + println!("cargo:rerun-if-changed=../../.git/HEAD"); + println!("cargo:rerun-if-changed=../../.git/refs"); +} + +fn env_hash() -> Option { + let s = std::env::var("DEX_BUILD_ID").ok()?; + // Truncated to 12 to match the git path's --short=12: the startup line is + // read by a human on site, and one shape is easier to compare against a + // release note than "sometimes 12 chars, sometimes 40" depending on which + // source happened to win. + let s: String = s.trim().chars().take(12).collect(); + (!s.is_empty()).then_some(s) +} + +fn stamp_hash() -> Option { + let s = std::fs::read_to_string(".dex-build-id").ok()?; + let s = s.trim(); + (!s.is_empty()).then(|| s.to_string()) +} + +fn git_hash_with_dirty() -> Option { + let hash = Command::new("git") + .args(["rev-parse", "--short=12", "HEAD"]) + .output() + .ok() + .filter(|o| o.status.success()) + .and_then(|o| String::from_utf8(o.stdout).ok()) + .map(|s| s.trim().to_string()) + .filter(|s| !s.is_empty())?; + let dirty = Command::new("git") + .args(["status", "--porcelain"]) + .output() + .ok() + .filter(|o| o.status.success()) + .map(|o| !o.stdout.is_empty()) + .unwrap_or(false); + Some(format!("{hash}{}", if dirty { "+dirty" } else { "" })) +} diff --git a/packages/dexd/deny.toml b/packages/dexd/deny.toml new file mode 100644 index 0000000..db4e63e --- /dev/null +++ b/packages/dexd/deny.toml @@ -0,0 +1,86 @@ +# cargo-deny — turns SPEC §5c's dependency policy from prose into a gate. +# +# §5c says the requirement is not "no dependencies" but "a small, DECLARED, +# auditable dependency set". Declared is handled by Cargo.toml and +# dpkg-shlibdeps. This file handles the other two words: it fails the build if +# the set stops being small, stops being auditable, or acquires a known +# vulnerability. +# +# Run: cargo deny check + +[advisories] +# A gallery device runs unattended for weeks and is not network-connected, so +# the realistic threat is a malformed asset, not a remote attacker. Deny +# anyway: "not exploitable in our deployment" is a judgement that ages badly, +# and the whole point of a gate is that it does not require one. +yanked = "deny" +ignore = [] + +[licenses] +# An allow-list, not a deny-list. A new dependency arriving under something +# unexpected should fail loudly rather than be silently accepted because nobody +# thought to ban it. This also protects the licensing decision itself: dex aims +# to be packageable everywhere (see the licence open point in the story), and a +# copyleft dependency pulled in by accident would constrain that without anyone +# noticing until a distro said no. +# +# Deliberately trimmed to exactly what the graph contains today, rather than a +# generous list of "probably fine" licences. An unused allowance is not free: +# cargo-deny warns on each one, and a check that always prints warnings is a +# check people stop reading. A dependency arriving under BSD-3-Clause or +# Unicode-3.0 should therefore FAIL here and be waved through by an explicit +# one-line edit — which is the review this file exists to force. +# +# `MIT-0` is dexd's own licence. That entry changed from CC0-1.0 in the same +# commit as the licence decision — which is exactly the behaviour this list is +# for: it cannot silently disagree with what ships. +allow = [ + "Apache-2.0", # 14 crates + "MIT", # 16 crates + "MIT-0", # dexd itself + # Waved through deliberately, 2026-08-17 — which is exactly the one-line + # review this list exists to force, and the whole reason it is trimmed to + # what the graph actually contains. Arrived via + # yaml-rust2 -> hashlink -> hashbrown -> foldhash, and was the ONLY new + # licence in that entire subtree (yaml-rust2, arraydeque, hashlink and + # hashbrown are all already-allowed MIT/Apache-2.0). + # Zlib is permissive, OSI-approved and FSF Free/Libre, with no copyleft and + # no attribution obligation, so it constrains nothing about dex staying + # packageable everywhere — the property this allow-list protects. + "Zlib", # 1 crate (foldhash) +] +confidence-threshold = 0.9 + +[bans] +# Two versions of the same crate in one binary is the first symptom of a +# dependency set outgrowing "small enough to read". +multiple-versions = "warn" +wildcards = "deny" + +# THE PROC-MACRO BAN IS DELIBERATE AND LOAD-BEARING. +# +# serde is a dependency, but WITHOUT its `derive` feature: for a four-field +# struct, #[derive(Deserialize)] buys about fifteen lines of field extraction +# and costs the whole syn/quote/proc-macro2 toolchain on the critical path of +# every CI package build. That decision currently lives in a Cargo.toml comment, +# which is a rule; enabling the feature would quietly undo it and nothing would +# complain. Here it is a gate: turning on `derive` fails `cargo deny check`. +# +# If a future dependency genuinely needs proc macros, delete these entries +# deliberately and say why in the commit. That is the intended way through. +[[bans.deny]] +name = "syn" +[[bans.deny]] +name = "quote" +[[bans.deny]] +name = "proc-macro2" +[[bans.deny]] +name = "serde_derive" + +[sources] +unknown-registry = "deny" +unknown-git = "deny" +# crates.io only. A git dependency has no version, no audit trail and no +# guarantee of being there next year -- unacceptable for software that ships to +# devices expected to run untouched for the length of an exhibition. +allow-registry = ["https://github.com/rust-lang/crates.io-index"] diff --git a/packages/dexd/deploy/changelog b/packages/dexd/deploy/changelog new file mode 100644 index 0000000..91e74a4 --- /dev/null +++ b/packages/dexd/deploy/changelog @@ -0,0 +1,28 @@ +dexd (0.1.0-2) unstable; urgency=medium + + * F6: the display mode is now exhibit config, not a systemd-unit or + cmdline.txt hand-edit. New conffile /etc/dex/exhibit.json (binds + display_mode, kms_force, connector) and new privileged sibling binary + dex-exhibit-apply(1), which reconciles cmdline.txt's video= token with + it. dex-loop.service no longer passes --mode. Two new startup gates, + both before the asset is read: the exhibit config vs. running-kernel + cmdline.txt cross-check, and a sysfs pre-flight that refuses a + display_mode absent from the connector's own mode list. Fail-closed, + same contract as F3's sidecar: a missing or invalid exhibit config + refuses rather than guesses the display. + * postinst warns (non-fatal) if a leftover systemd drop-in still + overrides --mode. + + -- Max Albrecht <1@178.is> Mon, 17 Aug 2026 12:00:00 +0200 + +dexd (0.1.0-1) unstable; urgency=medium + + * Initial packaged release. Gapless 4K HEVC loop player for Raspberry Pi + gallery installations: libmpv fed by an endless in-process byte stream so + the decoder never sees EOF and never seeks, which removes the seam at the + loop point. + * Frame rate is bound to the asset by an ingest sidecar (fps + sha256); the + player refuses to start rather than guess a rate. + * Tier-0 self-healing: a bounded in-place recovery before escalating to exit. + + -- Max Albrecht <1@178.is> Sat, 15 Aug 2026 17:30:00 +0200 diff --git a/packages/dexd/deploy/dex-wait-hdmi b/packages/dexd/deploy/dex-wait-hdmi new file mode 100755 index 0000000..959b01b --- /dev/null +++ b/packages/dexd/deploy/dex-wait-hdmi @@ -0,0 +1,40 @@ +#!/usr/bin/env bash +# Wait for an HDMI connector to report `connected`, then exit 0. +# +# Install: /usr/local/bin/dex-wait-hdmi +# +# Why: a dex player is switched on at the mains, and a projector can take 30-90s +# to present EDID while the Pi boots in ~15. Starting before then means KMS picks +# a fallback mode (typically 1024x768) and STAYS there -- the display never +# corrects itself, and the installation shows the wrong thing all day. +# +# This is belt-and-braces. The primary fix is at the KMS layer, in +# /boot/firmware/cmdline.txt, so boot order stops mattering at all: +# +# # capture once, while the projector is awake and connected: +# cat /sys/class/drm/card*-HDMI-A-1/edid | sudo tee /lib/firmware/edid/dex.bin +# # then add to cmdline.txt (one line): +# drm.edid_firmware=HDMI-A-1:edid/dex.bin video=HDMI-A-1:3840x2160@30D +# +# The trailing `D` forces digital-enabled even when the connector reads +# disconnected. With that in place the Pi always believes a 4K30 display is +# attached and the projector simply locks on when it finishes warming up. +# This script then only covers the cases the baked EDID cannot: a swapped cable, +# or a replaced projector that invalidates it. +set -uo pipefail + +TIMEOUT="${DEX_HDMI_TIMEOUT:-120}" + +for _ in $(seq "$TIMEOUT"); do + # Glob the card number: vc4/v3d probe order makes card0/card1 unstable across + # kernel versions, so hardcoding it breaks on upgrade. + if grep -qx 'connected' /sys/class/drm/card*-HDMI-A-*/status 2>/dev/null; then + exit 0 + fi + sleep 1 +done + +echo "dex-wait-hdmi: no connected HDMI connector after ${TIMEOUT}s" >&2 +# Fail the unit so Restart=always retries the whole sequence, rather than +# starting the player into a display that is not there. +exit 1 diff --git a/packages/dexd/deploy/dexd.service b/packages/dexd/deploy/dexd.service new file mode 100644 index 0000000..d798b2e --- /dev/null +++ b/packages/dexd/deploy/dexd.service @@ -0,0 +1,162 @@ +# dex gapless loop player. +# +# Installed by the dexd .deb (SPEC 5c), which also creates the `dex` user +# and /opt/dex, and enables this unit. Remaining manual step: +# sudo systemctl set-default multi-user.target # no desktop; nothing else may own DRM +# +# Paths are /usr/bin, not /usr/local/bin: Debian policy reserves /usr/local for +# the local administrator, so a package may not write there. +# +# The player is useless without supervision: it runs unattended for weeks with +# no operator, and its only shutdown path is someone switching off the mains. +# Every design choice below follows from that. + +[Unit] +Description=dex gapless loop player +Documentation=https://github.com/KTE/dex +# A getty on tty1 holds DRM master and would prevent the player taking the +# display. Conflicting it away is what stops "device busy" ever happening, +# rather than handling it after the fact. +Conflicts=getty@tty1.service +After=multi-user.target + +# LOAD-BEARING. Without this, systemd's default rate limit (5 starts / 10 s) +# puts the unit into a permanent `failed` state after a burst of crashes -- +# exactly the unattended disaster the unit exists to prevent. A restart loop +# always beats a dead screen: never give up. +# +# MUST live in [Unit], not [Service]. It sat in [Service] until 2026-08-15, +# where systemd silently IGNORED it ("Unknown key ... ignoring") and the default +# limit applied the whole time -- so the single most safety-critical line in +# this file did nothing, while reading as though it did. `systemd-analyze +# verify` found it in one run; it is now in CI for that reason. +# (The key was accepted in [Service] before systemd 229 and moved to [Unit].) +StartLimitIntervalSec=0 + +[Service] +Type=simple +# F10 -- systemd watchdog (PLAN.md's F10 entry; implementation in +# src/watchdog.rs). Type stays `simple`, NOT `notify`: a notify unit that +# never sends READY=1 sits inactive forever, and this player has no natural +# "ready" moment before the endless stream starts (see src/watchdog.rs's +# module doc for why READY=1 is deliberately never sent). NotifyAccess=main +# is what makes WatchdogSec= work under Type=simple: it grants the main +# process the notify socket without requiring the READY handshake Type=notify +# would. +# +# LIVENESS CRITERION (see src/watchdog.rs's module doc for the full +# argument): dexd pings WATCHDOG=1 once per ~10s health-check tick, AFTER +# that tick's F1 evaluation has completed -- never from the 600s heartbeat. +# F1's in-place-recovery budget is cumulative and NEVER refills, so a +# display-wedged player's tick sequence is FORCED, by construction, through +# silence -> stall -> <=3 budgeted recoveries (~20s each) -> exit(1) -- +# bounding the pings a wedge can ever emit before either exiting (tier 1's +# exit-code restart already covers that) or, on the one path that can still +# hang (e.g. `eprintln!` against a wedged journald, INCLUDING on the exit(1) +# arm whose entire job is "let the supervisor take over" -- see that arm's +# own comment in main.rs), simply STOPPING. A wedged player therefore cannot +# keep this watchdog satisfied indefinitely -- there is no third state +# between "pings continue because the loop is genuinely progressing" and +# "pings stop because the loop hung". +# +# WatchdogSec=180 (PLAN.md's stated floor): >=17x the ping cadence, so a +# handful of missed/dropped pings (journald stutter, scheduling bursts; +# src/watchdog.rs's socket is non-blocking and a full receiver queue is +# treated as a DROPPED ping, never a retry or a block) cannot cause a +# spurious kill, while still bounding the worst-case black-wall time for +# this hazard class to a few minutes across a weeks-long unattended run. +# 180s also comfortably exceeds F1's own worst-case tier-0 episode (~2 min), +# so a watchdog kill can never preempt a recovery that tier 0 would have +# completed on its own -- see src/watchdog.rs's module doc, "gate placement". +# +# INVARIANT for future edits to this unit or to main.rs's event loop: do NOT +# add a "final ping" anywhere on an exit path (Escalate, END_FILE, +# QUEUE_OVERFLOW). Pings exist to prove the loop is STILL RUNNING, not that +# it once was; a ping sent right before the one hang this feature exists to +# catch would only reset the countdown and defeat it. +WatchdogSec=180 +NotifyAccess=main +# The display may not be awake when the Pi boots -- a projector can take 30-90s +# to present EDID while the Pi boots in ~15. Wait for the connector rather than +# starting into a fallback mode that never corrects itself. ExecStartPre runs +# BEFORE the main process is forked, i.e. before the watchdog timer starts +# (systemd starts counting WatchdogSec= at Type=simple's "start-up complete" +# point, which for `Type=simple` is exec of ExecStart) -- so this wait, however +# long, consumes none of the 180s budget above. Verified on the Pi, not just +# inferred: see PLAN.md's F10 entry for the bench unit that pins this. +# +# TimeoutStartSec MUST exceed dex-wait-hdmi's own bounded wait (DEX_HDMI_TIMEOUT, +# default 120s) with margin. Without this line, Debian's DefaultTimeoutStartSec= +# (90s) applies to ExecStartPre: a projector needing 91-120s would get the script +# SIGTERMed at ~90s with a generic "start-pre operation timed out" -- the script's +# own purpose-written "no connected HDMI connector after 120s" diagnostic would be +# dead code under this unit, and every retry would wait 30s less than the script +# intends. If DEX_HDMI_TIMEOUT is ever raised, raise this with it. +TimeoutStartSec=150 +ExecStartPre=/usr/bin/dex-wait-hdmi +# NO ARGUMENTS AT ALL, and each absence is deliberate. This unit states HOW to +# run the player; everything about WHAT it plays and WHERE lives in one +# operator-editable file, /etc/dex/exhibit.{yaml,json}. +# +# No asset path, as of 2026-08-17 (F6): WHICH artwork plays comes from the +# exhibit config's `asset` key. This line hardcoded /opt/dex/loop.265 until +# then, which meant changing the artwork required either overwriting that one +# path or editing a unit file the package owns -- and made "several assets in +# storage, the exhibit picks one" impossible, which was the stated reason for +# having an exhibit file at all. The shipped conffile's `asset` is that same +# /opt/dex/loop.265, so a stock install behaves exactly as it did. +# +# No --fps: the frame rate comes from .json (F3), written at ingest and +# hash-bound to the asset bytes. Deploy BOTH files. A missing or stale sidecar +# makes the player refuse (exit 2) rather than guess a rate -- a wrong guess +# would play slow forever with every metric green. +# +# No --mode (F6): the display mode is a property of the INSTALLATION (the +# venue's panel), not of this unit, so it lives in the same exhibit config an +# operator edits with `sudoedit /etc/dex/exhibit.json && sudo dex-exhibit-apply` +# (see man dex-exhibit-apply), rather than a value baked in here that every +# panel swap would require re-touching. A `--mode` here would additionally +# FIGHT the exhibit config: main.rs treats an explicit --mode as a cross-check +# against it and refuses to start on any disagreement. The same is now true of +# an asset path passed here. +ExecStart=/usr/bin/dexd + +Restart=always +RestartSec=2 +# NOTE: StartLimitIntervalSec lives in [Unit] above, not here. See the comment +# there -- putting it in this section is silently ignored. + +User=dex +SupplementaryGroups=video render + +# mpv resolves a shader-cache directory under $XDG_CACHE_HOME (falling back to +# $HOME/.cache) during VO init, and does so BEFORE consulting +# --gpu-shader-cache: verified by A/B on mpv 0.40, where setting that flag to +# `no` changes nothing. The `dex` user's home is /nonexistent, so without this +# the service logs +# Failed to create /nonexistent for shader cache (...)---disabling. +# on every single start. Harmless in itself, and exactly the kind of constant +# harmless error that trains an operator to stop reading the journal — which is +# the only diagnostic channel this player has on site. +# +# CacheDirectory= rather than a writable HOME: systemd creates +# /var/cache/dexd owned by User=, keeps it writable under +# ProtectSystem=strict, and cleans it up on purge. +CacheDirectory=dexd +Environment=XDG_CACHE_HOME=/var/cache/dexd + +# The player only ever reads one file. +ProtectSystem=strict +ProtectHome=yes +ReadOnlyPaths=/opt/dex +PrivateTmp=yes +NoNewPrivileges=yes + +# No graceful shutdown is implemented, deliberately: the mains switch IS the +# shutdown path, so abrupt death must be safe rather than avoidable. The kernel +# releases DRM master on exit, so the next start acquires it cleanly. +KillSignal=SIGTERM +TimeoutStopSec=5 + +[Install] +WantedBy=multi-user.target diff --git a/packages/dexd/deploy/exhibit.json.default b/packages/dexd/deploy/exhibit.json.default new file mode 100644 index 0000000..5e55e6f --- /dev/null +++ b/packages/dexd/deploy/exhibit.json.default @@ -0,0 +1,6 @@ +{ + "asset": "/opt/dex/loop.265", + "display_mode": "auto", + "kms_force": "none", + "note": "F6 stock default: inert on purpose (see man dex-exhibit-apply). auto asks mpv for the connector-preferred mode; none forces nothing at the KMS layer. Set display_mode explicitly for a real exhibit -- e.g. \"3840x2160@30\" -- and kms_force if the sink needs one (some panels, e.g. an Elgato Cam Link 4K, advertise 4K as their preferred mode and the driver still declines to build it unforced -- see PLAN.md F6). asset names WHICH file plays, so several assets can sit in /opt/dex and the exhibit picks one; the value here matches the path ExecStart used to hardcode, so a stock install behaves exactly as before. It is the one key with no safe default in the code: a config without it refuses to start rather than guess an artwork." +} diff --git a/packages/dexd/deploy/lintian-overrides b/packages/dexd/deploy/lintian-overrides new file mode 100644 index 0000000..9b23195 --- /dev/null +++ b/packages/dexd/deploy/lintian-overrides @@ -0,0 +1,53 @@ +# cargo-deb 2.12 hardcodes the systemd unit destination as lib/systemd/system. +# On a merged-usr system -- which trixie is, and every device we ship to -- that +# path IS usr/lib/systemd/system, the same directory by symlink. Verified: the +# package installs, enables, starts, stops, removes and purges correctly, and +# systemd resolves the unit as /usr/lib/systemd/system/dexd.service. Lintian +# is objecting to the spelling, not to the outcome. +# +# NOT worked around by dropping cargo-deb's systemd integration and hand-writing +# the enable/disable/mask lifecycle in maintainer scripts. That lifecycle was +# checked end-to-end on real hardware, including the subtle part (remove leaves +# the unit masked so a reinstall restores its enabled state; purge cleans it). +# Reimplementing deb-systemd-helper by hand to change a path spelling would +# trade a cosmetic tag for a real regression risk in the one code path that runs +# on every device. +# +# Fix properly by upgrading cargo-deb, which needs rustc >= 1.88; we are pinned +# to trixie's 1.85 on purpose (SPEC 5c, "Toolchain"). This override should be +# deleted the moment that pin moves. +dexd: aliased-location + +# Wants the changelog entry to close an ITP bug. That is a rule for packages +# being uploaded INTO the Debian archive; this one is built in CI and installed +# on our own devices, so there is no bug to close and never will be. +dexd: initial-upload-closes-no-bugs + +# FALSE POSITIVE, diagnosed rather than assumed -- an `embedded-library` tag is +# exactly the kind that should NOT be waved through on a hunch, since the real +# version of it means shipping unpatchable copies of someone else's CVEs. +# +# What lintian actually tests for (trixie, /usr/share/lintian/data/binaries/ +# embedded-libs line 86) is the single STRING: +# +# did not find expected +# +# That is one of libyaml's error messages -- and `yaml-rust2` is a from-scratch +# Rust port of libyaml's scanner/parser which reproduces its diagnostics +# verbatim, at src/parser.rs:366 and :525. So the fingerprint matches the +# message, not the code. +# +# Verified on the real arm64 binary (dexpi4, 2026-08-17), three ways: +# * `ldd` shows no libyaml among the dynamic dependencies +# * `strings` finds ZERO libyaml C symbols (no yaml_parser_*, yaml_emitter_*, +# yaml_document_*); every yaml symbol is a Rust-mangled yaml_rust2::* +# * no crate in the subtree (yaml-rust2, arraydeque, hashlink, hashbrown, +# foldhash) has a build.rs, a `links =` key, or is a *-sys crate, so there +# is no mechanism by which C could be compiled in at all +# +# There is therefore no embedded libyaml to update when libyaml has a CVE; the +# security-tracking concern the tag exists to serve does not apply. Re-check +# this if the yaml dependency is ever swapped for one with a C backend, which +# would make the same tag true. +dexd: embedded-library libyaml [usr/bin/dexd] +dexd: embedded-library libyaml [usr/bin/dex-exhibit-apply] diff --git a/packages/dexd/deploy/maintainer-scripts/postinst b/packages/dexd/deploy/maintainer-scripts/postinst new file mode 100755 index 0000000..02f5bcc --- /dev/null +++ b/packages/dexd/deploy/maintainer-scripts/postinst @@ -0,0 +1,102 @@ +#!/bin/sh +# Create the two things dexd.service requires but cannot create itself: +# the unprivileged `dex` user it runs as, and the asset directory it reads. +# +# Both were manual steps in the pre-package install recipe, which is precisely +# why they belong here: an unattended gallery device should not depend on +# whether whoever imaged the SD card remembered a chmod. +set -e + +case "$1" in +configure) + # A system user with no login, no home, and no password. It needs `video` + # and `render` only to open /dev/dri -- everything else is denied by the + # unit's own sandboxing (ProtectSystem=strict, ProtectHome, NoNewPrivileges). + if ! getent passwd dex >/dev/null; then + adduser --system --group --home /nonexistent --no-create-home \ + --shell /usr/sbin/nologin dex + fi + # Idempotent: adduser to a group the user is already in is a no-op, so this + # also repairs an install where the groups were dropped. + for g in video render; do + if getent group "$g" >/dev/null; then + adduser dex "$g" >/dev/null || true + fi + done + + # The asset directory. The package deliberately ships NO asset: the video and + # its sidecar are content, they change per installation, and baking one into + # the package would mean rebuilding the software to change the artwork. + # Root-owned and world-readable — the unit mounts it ReadOnlyPaths anyway, + # and the player only ever reads. + mkdir -p /opt/dex + chmod 0755 /opt/dex + + if [ ! -e /opt/dex/loop.265 ]; then + echo "dexd: installed, but /opt/dex/loop.265 is missing." >&2 + echo " Deploy BOTH the asset and its sidecar, then start the unit:" >&2 + echo " scp loop.265 loop.265.json :/opt/dex/" >&2 + echo " sudo systemctl start dexd" >&2 + echo " The player refuses to start without the sidecar, on purpose: a" >&2 + echo " guessed frame rate plays wrong forever with every metric green." >&2 + fi + + # F6 (PLAN.md): the display mode moved from a systemd drop-in's --mode into + # /etc/dex/exhibit.json. LOUD, non-fatal, because the package must not + # delete an admin-created /etc file itself -- and it need not: if the + # drop-in DISAGREES with the exhibit config, main.rs's own --mode + # cross-check refuses to start and names both values, so nothing can + # silently win. This warning exists for the case that still WOULD start: + # a drop-in that happens to agree, or one that has gone stale in some other + # way (see 2026-08-17's config.txt-vs-cmdline.txt drift on dexpi4) -- purely + # so the leftover gets noticed and removed rather than outliving its reason. + # F6, 2026-08-17: ExecStart no longer passes an asset path -- WHICH artwork + # plays now comes from the exhibit config's `asset` key. A FRESH install is + # fine (the shipped conffile has it), but dpkg conffile semantics keep an + # admin-edited /etc/dex/exhibit.json exactly as it was, so an upgraded device + # can end up with a config that predates the key. The player then refuses to + # start -- correctly and loudly, since guessing an artwork is the one guess + # this package will not make (see src/exhibit.rs's resolve_asset) -- but it + # would refuse at the next power cycle, in a gallery, with nobody watching. + # Warning here converts that into a message during the upgrade someone is + # sitting through. Non-fatal: the package must not edit an admin's /etc file. + # + # grep, not a JSON/YAML parser: a maintainer script may use only what the + # package depends on, and this is a "did you notice" prompt, not a gate. A + # false positive costs one glance at a file; the real gate is in the player. + # + # The key must start a line or follow a `{`/`,`, which covers pretty-printed + # JSON, single-line JSON, and YAML -- and, the reason for the anchor, does + # NOT match a commented-out `# asset:` in a YAML file. Driven against all + # eight shapes before shipping, not reasoned about: an unanchored pattern + # stayed silent on the commented case, which is precisely the config whose + # owner most needs the prompt. + for cfg in /etc/dex/exhibit.yaml /etc/dex/exhibit.json; do + [ -e "$cfg" ] || continue + if ! grep -qE '(^|[{,])[[:space:]]*"?asset"?[[:space:]]*:' "$cfg" 2>/dev/null; then + echo "dexd: NOTE: $cfg names no asset, and ExecStart no longer" >&2 + echo " passes one (F6). The player will refuse to start until it does --" >&2 + echo " deliberately: guessing which artwork to play is the one guess this" >&2 + echo " package will not make. Add the line, then restart:" >&2 + case "$cfg" in + *.yaml) echo " asset: /opt/dex/loop.265" >&2 ;; + *) echo " \"asset\": \"/opt/dex/loop.265\"," >&2 ;; + esac + echo " sudo systemctl restart dexd" >&2 + fi + done + + dropin_dir=/etc/systemd/system/dexd.service.d + if [ -d "$dropin_dir" ] && grep -q -- '--mode' "$dropin_dir"/*.conf 2>/dev/null; then + echo "dexd: NOTE: $dropin_dir still overrides --mode -- that setting" >&2 + echo " moved to /etc/dex/exhibit.json's display_mode (F6, PLAN.md). If the" >&2 + echo " drop-in's value disagrees with the exhibit config, the unit will" >&2 + echo " refuse to start (naming both). Remove the drop-in:" >&2 + echo " sudo rm $dropin_dir/*.conf && sudo systemctl daemon-reload" >&2 + fi + ;; +esac + +#DEBHELPER# + +exit 0 diff --git a/packages/dexd/deploy/maintainer-scripts/postrm b/packages/dexd/deploy/maintainer-scripts/postrm new file mode 100755 index 0000000..ed27ba8 --- /dev/null +++ b/packages/dexd/deploy/maintainer-scripts/postrm @@ -0,0 +1,13 @@ +#!/bin/sh +# Removal deliberately keeps BOTH the `dex` user and /opt/dex. +# +# The asset in /opt/dex is the artwork -- operator content the package never +# shipped and must not delete. And the user owns nothing but is referenced by +# anything an operator wrote themselves; reaping it on purge would silently +# break a hand-rolled unit or a cron job. Leaving a passwd entry behind is the +# cheaper mistake, and Debian policy permits keeping system users. +set -e + +#DEBHELPER# + +exit 0 diff --git a/packages/dexd/deploy/man/dex-exhibit-apply.1 b/packages/dexd/deploy/man/dex-exhibit-apply.1 new file mode 100644 index 0000000..1e0a156 --- /dev/null +++ b/packages/dexd/deploy/man/dex-exhibit-apply.1 @@ -0,0 +1,215 @@ +.TH DEX\-EXHIBIT\-APPLY 1 "2026-08-17" "dexd 0.1.0" "dex" +.SH NAME +dex\-exhibit\-apply \- reconcile the kernel cmdline with the F6 exhibit config +.SH SYNOPSIS +.B sudo dex\-exhibit\-apply +.RB [ \-\-exhibit\-config +.IR PATH ] +.RB [ \-\-cmdline\-path +.IR PATH ] +.SH DESCRIPTION +Rewrites +.IR /boot/firmware/cmdline.txt 's +.B video=: +token to match +.IR /etc/dex/exhibit.json 's +.B kms_force +\(em idempotently: every other token, its order, and every OTHER connector's +.B video= +token are preserved untouched. Writes one timestamped backup before any +change. +.PP +.B dexd (1) +itself never writes boot config \(em it runs as an unprivileged user under +.B ProtectSystem=strict +and that sandbox is a design feature, not an oversight to route around here. +This tool is the deliberate, separate, privileged step an operator runs by +hand after editing the exhibit config. Deploys in this project are manual +throughout (see README.md); this is no exception. +.SH WHY BOTH A KERNEL FORCE AND AN EXHIBIT CONFIG +Some sinks advertise a mode in their own EDID and the DRM driver still +declines to build it unforced \(em measured on an Elgato Cam Link 4K, which +lists 3840x2160@30 as its +.I preferred +detailed timing, yet +.B vc4 +built zero 3840x2160 modes from it unforced; forced, the identical timing +works. Other sinks (a 2560x1440 desktop monitor, in this project's own field +record) must +.B NOT +carry the force, because transmitting a mode the panel cannot show reads as a +player fault, not a config problem. The same +.I cmdline.txt +line is therefore correct for one venue and wrong for another, which is why it +is exhibit config \(em stated explicitly per install \(em rather than something +to remember in a comment. +.SH OPTIONS +.TP +.BI \-\-exhibit\-config " PATH" +Default: whichever of +.I /etc/dex/exhibit.yaml +and +.I /etc/dex/exhibit.json +exists \(em exactly one may, and both present is refused. This is deliberately +the same discovery +.BR dexd (1) +does, and shares its implementation: a tool whose job is to make the boot +cmdline agree with the config the player reads would be a drift GENERATOR if +it could read a different file. +.TP +.BI \-\-cmdline\-path " PATH" +Default: +.IR /boot/firmware/cmdline.txt . +Override exists for testing on a bench without touching real boot config. +.SH EXHIBIT CONFIG SCHEMA +The exhibit config comes in two formats and +.B the file extension decides which: +.B .json +is strict JSON (what the package ships, as a dpkg conffile so a hand edit +survives an upgrade), +.BR .yaml / .yml +is YAML (for hand editing on site \(em it takes comments). The schema is +identical; only the parser differs. YAML is a superset of JSON, so a +.B .json +file containing YAML is REFUSED rather than quietly accepted: the extension is +a promise to +.BR jq (1) +and every other consumer about what the bytes are. Strict JSON inside a +.B .yaml +file is fine, since it is valid YAML. +.PP +Either way the schema is STRICT \(em unlike the F3 asset sidecar, an unknown +key is refused rather than +ignored, because this file has no independent producer to stay compatible +with: a typo here (\fBkms_forse\fR for \fBkms_force\fR) must be a startup +refusal, not a silently dropped force. +.TP +.B asset +Absolute path to the file to play, e.g. +.IR /opt/dex/loop.265 . +This is the key that makes the file an EXHIBIT config rather than a display +config: an exhibit pairs a venue with an artwork, and several assets can sit in +.I /opt/dex +with this choosing one. +.B dex\-exhibit\-apply +itself does not use it \(em it reconciles the boot cmdline, which is a display +concern \(em but it is listed here because this page documents the whole schema. +Optional in the file; something must supply it, or +.BR dexd (1) +refuses to start rather than guess an artwork. +.TP +.B display_mode +(required) \fBauto\fR, or \fBWxH@R\fR with an INTEGER refresh, e.g. +\fB3840x2160@30\fR. What +.B dexd +asks mpv for via +.BR \-\-drm\-mode . +Non\-integer refresh forms are refused at config parse: bench\-verified +(2026\-08\-17), mpv rejects a rational refresh (\fB@30000/1001\fR) at option +parsing \(em a guaranteed restart loop \(em and silently rounds a decimal +(\fB@29.97\fR) to the integer vrefresh it names anyway. +.TP +.B kms_force +(default \fBnone\fR) \fBnone\fR, or \fBWxH@R\fR/\fBWxH@RD\fR with an INTEGER +refresh (the kernel's +.B video= +grammar has no fractional refresh \(em and since 2026\-08\-17 neither does +\fBdisplay_mode\fR, see above). The trailing +.B D +forces the connector to read \fBconnected\fR even before a sink is actually +attached \(em the boot\-order insurance for a projector that presents EDID +late. What +.B dex\-exhibit\-apply +writes into +.IR cmdline.txt . +.TP +.B connector +(default \fBHDMI\-A\-1\fR) e.g. \fBHDMI\-A\-2\fR. +.TP +.BR display ", " venue ", " note +Informational only, logged verbatim by +.B dexd +at every start \(em this is where the WHY that used to live in a +.I config.txt +comment block belongs now, since it travels with the config that is actually +enforced instead of a comment nobody updates. +.SH THE CMDLINE GATE +.B dexd +refuses to start (exit 2) if the exhibit config's +.B kms_force +disagrees with the token actually present in the RUNNING kernel's +.IR /proc/cmdline . +An edit to +.I exhibit.json +with no matching +.B dex\-exhibit\-apply ++ reboot is exactly this disagreement \(em caught at the next start rather +than black\-screening the venue for the run of the exhibition. The error names +both values and BOTH possible repairs: the gate cannot know whether the +config is the stale side or the cmdline is. In particular, on a machine whose +cmdline already carries a force the venue deliberately needs (some displays +build no 4K mode unforced), the repair is to update +.IR exhibit.json , +NOT to run this command \(em running it would delete the needed force. +.SH WHAT THE PRE\-FLIGHT DOES AND DOES NOT VALIDATE +.B dexd +checks the RESOLUTION half of +.B display_mode +against the connector's +.I /sys/class/drm/card*\-/modes +list before opening the asset. The +.B @R +refresh half is validated by GRAMMAR ONLY \(em the kernel's +.I modes +file has no refresh column, so no sysfs pre\-flight for it is possible. A +syntactically valid refresh the connector does not actually offer is settled +by mpv's +.B \-\-drm\-mode +matching at VO init instead: expect an mpv\-voiced VO error (and a restart +loop) in the journal rather than a message from this gate. If the mode you +configure came from EDID or +.B modetest +data, double\-check the refresh against what the connector really builds. +.SH EXAMPLE: SWAPPING THE PANEL +.nf +sudoedit /etc/dex/exhibit.json # display_mode/kms_force for the new panel +sudo dex\-exhibit\-apply +sudo reboot # only if it printed REBOOT REQUIRED +.fi +.SH MIGRATION FROM THE PRE\-F6 ARRANGEMENT +Before F6, the display mode lived in three places at once, none of them a +config file: a hand\-edited +.IR cmdline.txt , +a systemd drop\-in overriding +.B \-\-mode +in +.IR dexd.service , +and a comment block in +.I config.txt +carrying the reasoning \(em which drifted out of sync with the other two +within days of being written (see PLAN.md's F6 entry for the exact incident). +.B dexd.service +no longer passes +.BR \-\-mode , +so a leftover drop-in either agrees with the exhibit config (harmless until +removed; \fBdexd\fR's postinst warns if one still mentions +\fB\-\-mode\fR) or disagrees (refused loudly at the next start, naming both). +Remove it either way: +.nf +sudo rm /etc/systemd/system/dexd.service.d/*.conf +sudo systemctl daemon\-reload +.fi +.SH EXIT STATUS +.TP +.B 0 +cmdline.txt reconciled, or already correct ("no change"). +.TP +.B 1 +An I/O failure (cannot read or write a file) \(em not a bad config. +.TP +.B 2 +Refused: not running as root, an invalid exhibit config, or a rewrite that +would produce an empty or multi\-line cmdline.txt. +.SH SEE ALSO +.BR dexd (1), +.BR dex-wait-hdmi (1) diff --git a/packages/dexd/deploy/man/dex-wait-hdmi.1 b/packages/dexd/deploy/man/dex-wait-hdmi.1 new file mode 100644 index 0000000..677aca7 --- /dev/null +++ b/packages/dexd/deploy/man/dex-wait-hdmi.1 @@ -0,0 +1,30 @@ +.TH DEX\-WAIT\-HDMI 1 "2026-08-15" "dexd 0.1.0" "dex" +.SH NAME +dex\-wait\-hdmi \- block until a DRM connector reports a display +.SH SYNOPSIS +.B dex\-wait\-hdmi +.SH DESCRIPTION +Waits for an HDMI connector to become connected before returning. Run as +.B ExecStartPre +of +.BR dexd (1). +.PP +A projector can take 30\-90 seconds to present EDID while a Pi boots in about +15. Without this wait the player starts into a fallback mode that never +corrects itself, and the installation shows the wrong resolution until someone +notices. +.PP +The primary fix for boot ordering is at the KMS layer rather than here: bake +the display's EDID into +.I cmdline.txt +so the Pi always believes the intended mode is attached. This wait is the +fallback for devices where that has not been done. +.SH EXIT STATUS +.TP +.B 0 +A connector reported a display. +.TP +.B 1 +Timed out. The service still starts; the player will use whatever mode exists. +.SH SEE ALSO +.BR dexd (1) diff --git a/packages/dexd/deploy/man/dexd.1 b/packages/dexd/deploy/man/dexd.1 new file mode 100644 index 0000000..ce56e9a --- /dev/null +++ b/packages/dexd/deploy/man/dexd.1 @@ -0,0 +1,119 @@ +.TH DEXD 1 "2026-08-17" "dexd 0.1.0" "dex" +.SH NAME +dexd \- gapless HEVC loop player for unattended installations +.SH SYNOPSIS +.B dexd +.RI [ stream.265 ] +.RB [ \-\-fps +.IR F ] +.RB [ \-\-mode +.IR WxH@R ] +.RB [ \-\-exhibit\-config +.IR PATH ] +.RB [ \-\-bench\-no\-sidecar ] +.RB [ \-\-no\-defaults ] +.RB [ \-\-opt +.IR K=V ...] +.SH DESCRIPTION +Plays a raw Annex\-B HEVC elementary stream in an endless loop with no visible +seam. The stream is presented to libmpv through a callback that never reports +end\-of\-file, so the decoder never seeks and the wrap costs no frame. +.PP +.I stream.265 +is OPTIONAL: in a deployment the exhibit config's +.B asset +key names it, which is what lets several assets sit in +.I /opt/dex +with the exhibit choosing one \(em the packaged unit runs +.B dexd +with no arguments at all. Given here as well, the path must AGREE with the +config or startup is refused, naming both; it is REQUIRED with +.B \-\-bench\-no\-sidecar +(which consults no config). If neither names an asset, startup refuses rather +than guessing an artwork. +.PP +Intended for gallery devices that run for weeks unattended and are switched off +at the mains. There is no graceful shutdown path by design. +.SH OPTIONS +.TP +.BI \-\-fps " F" +Frame rate to bind, as \fB30\fR, \fB29.97\fR or \fB30000/1001\fR. A raw +elementary stream carries no timestamps, so this cannot be inferred. Normally +it comes from the sidecar; passing it here is only permitted if it agrees. +.TP +.BI \-\-mode " WxH@R" +Cross\-check against the exhibit config's \fBdisplay_mode\fR (F6); refused if +it disagrees, naming both. Under \fB\-\-bench\-no\-sidecar\fR this is the only +source (no exhibit config is consulted): omitted there, it defaults to +\fBauto\fR (the connector's preferred mode). +.TP +.BI \-\-exhibit\-config " PATH" +Path to the exhibit config (F6). Default: whichever of +\fI/etc/dex/exhibit.yaml\fR and \fI/etc/dex/exhibit.json\fR exists \(em exactly +one may, and both present is refused rather than resolved by precedence. The +file EXTENSION decides the parser: \fB.json\fR is strict JSON, +\fB.yaml\fR/\fB.yml\fR is YAML, same schema either way. See \fBFILES\fR and +\fBdex-exhibit-apply\fR(1). +.TP +.B \-\-bench\-no\-sidecar +Run without a sidecar AND without the exhibit config. Requires an explicit +\fB\-\-fps\fR. For bench use only: deployments must bind the rate to the asset +and the display to the exhibit config. +.TP +.B \-\-no\-defaults +Omit the built\-in mpv option set (see \fBNOTES\fR). +.TP +.BI \-\-opt " K=V" +Pass an additional mpv option. Repeatable. +.SH FILES +.TP +.I .json +Required sidecar beside the asset, binding \fBfps\fR and \fBsha256\fR to the +exact bytes. A missing, stale or mismatched sidecar refuses startup. +.TP +.I /opt/dex/loop.265 +The asset path the stock exhibit config names, and the one +.B ExecStart +hardcoded before F6. Nothing in the code defaults to it: it is a value in a +conffile the operator may change to any absolute path. +.TP +.I /etc/dex/exhibit.json \fRor\fI /etc/dex/exhibit.yaml +F6 exhibit config: binds \fBasset\fR (which artwork plays) and +\fBdisplay_mode\fR, and optionally \fBkms_force\fR and \fBconnector\fR. The \fB.json\fR form is what the package ships, as a dpkg +conffile so a hand edit survives an upgrade; the \fB.yaml\fR form is for hand +editing on site and takes comments. Exactly one may exist. Required +outside \fB\-\-bench\-no\-sidecar\fR \(em a missing +or invalid one refuses startup rather than guessing the display, the same +fail\-closed contract the sidecar has for frame rate. See +\fBdex-exhibit-apply\fR(1) for the schema and the migration from the +pre\-F6 \fB\-\-mode\fR\-in\-the\-unit arrangement. +.SH EXIT STATUS +.TP +.B 0 +Never, in normal operation: the player is designed not to end. +.TP +.B 1 +Playback failed or recovery was exhausted. +.TP +.B 2 +A startup gate refused: bad invocation, missing or invalid exhibit config, +exhibit config vs. kernel cmdline disagreement, a requested resolution absent +from the connector's own mode list, missing or mismatched sidecar, or an +asset whose first frames are not a closed GOP starting at an IDR. +.SH NOTES +By default the player selects the only zero\-copy path measured to reach +realtime 4K30 on a Pi 4: the decoder's SAND\-tiled frames go straight to a KMS +plane via \fB\-\-gpu\-hwdec\-interop=drmprime\-overlay\fR. Every other interop +detiles and misses realtime. +.PP +F6: the display mode is a property of the exhibit (the venue's panel), not of +the asset or the systemd unit \(em see \fI/etc/dex/exhibit.json\fR under +\fBFILES\fR and \fBdex-exhibit-apply\fR(1). +.SH SEE ALSO +.BR dex-exhibit-apply (1), +.BR dex-wait-hdmi (1), +.BR mpv (1), +.BR journalctl (1) +.PP +Diagnosis starts with +.BR "journalctl -u dexd -n 20" . diff --git a/packages/dexd/src/bin/dex-exhibit-apply.rs b/packages/dexd/src/bin/dex-exhibit-apply.rs new file mode 100644 index 0000000..75bb08b --- /dev/null +++ b/packages/dexd/src/bin/dex-exhibit-apply.rs @@ -0,0 +1,320 @@ +//! F6's privileged sibling: reconciles `/boot/firmware/cmdline.txt`'s +//! `video=:` token with `/etc/dex/exhibit.json`'s +//! `kms_force`, idempotently. +//! +//! WHY A SEPARATE BINARY. `dexd` runs as the unprivileged `dex` user +//! under `ProtectSystem=strict` and must never write boot config -- that +//! sandbox is a design feature (see deploy/dexd.service), not an +//! oversight to work around here. Deploys in this project are manual (see +//! README.md), so an operator runs this by hand, as root, after editing +//! `/etc/dex/exhibit.json` -- the same "remaining manual step" pattern the +//! systemd unit's own comment already documents for `set-default +//! multi-user.target`. +//! +//! All the actual logic -- the grammar, and the rewrite itself +//! (`reconcile_cmdline`) -- lives in `dexd::exhibit` and is Mac-testable +//! with no root and no real `/boot`. This binary is a thin, privileged shell +//! around it, exactly the split `src/main.rs` itself follows for the player. +//! +//! Usage: sudo dex-exhibit-apply [--exhibit-config PATH] [--cmdline-path PATH] +//! exit 0 -- cmdline.txt reconciled, or already correct ("no change"). +//! Prints REBOOT REQUIRED iff it actually changed. +//! exit 1 -- cannot read or write a file (I/O failure, not a bad config). +//! exit 2 -- refused: not running as root, an invalid exhibit config, or a +//! rewrite that would produce an empty/multi-line cmdline.txt. + +use dexd::exhibit::{ + load_exhibit_config, reconcile_cmdline, DEFAULT_EXHIBIT_CONFIG_PATH, + DEFAULT_EXHIBIT_CONFIG_PATHS, +}; + +use std::env; +use std::fs::{self, File}; +use std::io::Write; +use std::path::Path; +use std::process::ExitCode; +use std::time::{SystemTime, UNIX_EPOCH}; + +const DEFAULT_CMDLINE_PATH: &str = "/boot/firmware/cmdline.txt"; + +/// How many `cmdline.txt.bak-*` backups to keep. The boot partition is a +/// small FAT32 volume; unbounded accumulation would eventually fill it, and +/// a backup older than the last few applies has no recovery value anyway. +const BACKUPS_TO_KEEP: usize = 5; + +/// Write `contents` to `path` and fsync it before returning. A plain +/// `fs::write` leaves the data in the page cache with no durability +/// guarantee -- on the one file the Pi cannot boot without, for an +/// installation whose documented off-switch is the mains, "written" must +/// mean "on the card", not "scheduled". +fn write_synced(path: &str, contents: &[u8]) -> std::io::Result<()> { + let mut f = File::create(path)?; + f.write_all(contents)?; + f.sync_all() +} + +/// Best-effort fsync of `path`'s parent directory, so the rename that put +/// the file there is itself committed. Errors are deliberately ignored: +/// by this point the data blocks and the file are already synced, and some +/// filesystems refuse directory fsync -- failing the whole apply over the +/// least important of the three syncs would be worse than proceeding. +fn sync_parent_dir(path: &str) { + let parent = Path::new(path).parent().filter(|p| !p.as_os_str().is_empty()); + if let Some(dir) = parent { + if let Ok(d) = File::open(dir) { + let _ = d.sync_all(); + } + } +} + +/// Pick a backup path that does not already exist: two applies within the +/// same second must not silently truncate each other's backup. +fn fresh_backup_path(cmdline_path: &str, stamp: u64) -> String { + let base = format!("{cmdline_path}.bak-{stamp}"); + let mut candidate = base.clone(); + let mut n = 1u32; + while Path::new(&candidate).exists() { + candidate = format!("{base}.{n}"); + n += 1; + } + candidate +} + +/// Delete all but the newest [`BACKUPS_TO_KEEP`] `.bak-*` files. +/// Best-effort and loud about what it removes; a failure here never fails +/// the apply (the reconcile already succeeded), it only means one extra +/// backup survives until the next run. +/// +/// Ordered by modification time, NOT by name, and `just_written` is never a +/// prune candidate at all. Both matter for the same reason, caught live on +/// the bench (dexpi4, 2026-08-17): once pruning frees an unsuffixed +/// `bak-` name, a later same-second apply reuses it — and that name +/// sorts lexically BEFORE its older `.1`/`.2` siblings, so a name sort would +/// classify the NEWEST backup as oldest and delete the one backup that +/// still matches the file just replaced. FAT mtime granularity (2 s) can +/// still tie same-second backups, which the explicit `just_written` +/// exclusion makes harmless. +fn prune_old_backups(cmdline_path: &str, just_written: &str) { + let path = Path::new(cmdline_path); + let (Some(dir), Some(name)) = (path.parent(), path.file_name()) else { return }; + let prefix = format!("{}.bak-", name.to_string_lossy()); + let Ok(entries) = fs::read_dir(if dir.as_os_str().is_empty() { Path::new(".") } else { dir }) + else { + return; + }; + let mut backups: Vec<(std::time::SystemTime, String)> = entries + .flatten() + .filter_map(|e| { + let n = e.file_name().to_string_lossy().into_owned(); + if !n.starts_with(&prefix) { + return None; + } + let p = e.path().to_string_lossy().into_owned(); + if p == just_written { + return None; + } + let mtime = e.metadata().and_then(|m| m.modified()).unwrap_or(UNIX_EPOCH); + Some((mtime, p)) + }) + .collect(); + // just_written is excluded above but still counts toward the kept total. + let keep_others = BACKUPS_TO_KEEP.saturating_sub(1); + if backups.len() <= keep_others { + return; + } + backups.sort(); + let excess = backups.len() - keep_others; + for (_, old) in backups.into_iter().take(excess) { + match fs::remove_file(&old) { + Ok(()) => println!(" pruned old backup: {old}"), + Err(e) => eprintln!("warning: could not prune old backup {old}: {e}"), + } + } +} + +fn usage() -> ! { + eprintln!( + "usage: dex-exhibit-apply [--exhibit-config PATH] [--cmdline-path PATH] + +Reconciles the kernel cmdline's video=: token with the +exhibit config's kms_force (F6), idempotently -- every other token, its +order, and every OTHER connector's video= token are preserved untouched. +Writes one timestamped backup before any change. Must run as root. + + --exhibit-config PATH default: whichever of {DEFAULT_EXHIBIT_CONFIG_PATHS:?} + exists (exactly one may; the extension decides the + parser -- .json is strict JSON, .yaml is YAML) + --cmdline-path PATH default: {DEFAULT_CMDLINE_PATH} (test/bench override) + +Run this after editing the exhibit config. It prints REBOOT REQUIRED iff +cmdline.txt actually changed -- dexd binds the display from the RUNNING +kernel's /proc/cmdline, not from this file on disk, so an unrebooted change +has no effect yet and the next start's cmdline gate will say so." + ); + std::process::exit(2) +} + +fn main() -> ExitCode { + let args: Vec = env::args().skip(1).collect(); + // `None` means "discover it" -- not "use the JSON default" -- so this tool + // follows the same .json/.yaml discovery dexd does. + let mut exhibit_config_path: Option = None; + let mut cmdline_path = DEFAULT_CMDLINE_PATH.to_string(); + + let mut i = 0; + while i < args.len() { + match args[i].as_str() { + "--exhibit-config" => { + i += 1; + let Some(v) = args.get(i) else { usage() }; + exhibit_config_path = Some(v.clone()); + } + "--cmdline-path" => { + i += 1; + let Some(v) = args.get(i) else { usage() }; + cmdline_path = v.clone(); + } + "-h" | "--help" => usage(), + _ => usage(), + } + i += 1; + } + + // Refused before touching any file: "run this as root" is a clean, whole + // failure, rather than a confusing partial write that dies on the second + // fs::write with a permission error. + if !is_root() { + eprintln!("error: dex-exhibit-apply must run as root (sudo dex-exhibit-apply)"); + return ExitCode::from(2); + } + + // Deliberately the SAME loader dexd uses (exhibit::load_exhibit_config), + // not a local read: this tool's entire job is to make the boot cmdline + // agree with the config the player will read, so reading a different file + // than the player does would make it a drift GENERATOR. That includes the + // .json/.yaml discovery and the both-exist refusal. + let (config, exhibit_config_path) = + match load_exhibit_config(exhibit_config_path.as_deref(), &DEFAULT_EXHIBIT_CONFIG_PATHS) { + Ok(Some(found)) => found, + // Unlike the player, this tool has nothing useful to do without a + // config, so "none installed" is a plain refusal here rather than + // something deferred to a resolver. + Ok(None) => { + eprintln!( + "error: no exhibit config found (looked for {}). Create one — the .deb \ + ships {DEFAULT_EXHIBIT_CONFIG_PATH} — or name it with --exhibit-config", + DEFAULT_EXHIBIT_CONFIG_PATHS.join(", ") + ); + return ExitCode::from(2); + } + Err(e) => { + eprintln!("error: {e}"); + return ExitCode::from(2); + } + }; + + let current = match fs::read_to_string(&cmdline_path) { + Ok(t) => t, + Err(e) => { + eprintln!("error: cannot read {cmdline_path}: {e}"); + return ExitCode::from(1); + } + }; + + // Name the file this run is applying FROM, before saying anything about + // what it did. With two possible config names and a --exhibit-config + // override, "which file did that apply use?" is the first question anyone + // debugging a wrong mode asks, and the answer belongs in the output rather + // than in a reconstruction from argv. + println!( + "dex-exhibit-apply: applying {exhibit_config_path} (connector {}, kms_force {})", + config.connector, config.kms_force + ); + + let desired = match reconcile_cmdline(¤t, &config.connector, &config.kms_force) { + Ok(d) => d, + Err(e) => { + eprintln!("error: {e}"); + return ExitCode::from(2); + } + }; + + let current_trimmed = current.trim_end_matches(['\n', '\r']); + if current_trimmed == desired { + println!("dex-exhibit-apply: {cmdline_path} already matches the exhibit config -- no change"); + return ExitCode::SUCCESS; + } + + let stamp = SystemTime::now() + .duration_since(UNIX_EPOCH) + .map(|d| d.as_secs()) + .unwrap_or(0); + let backup_path = fresh_backup_path(&cmdline_path, stamp); + // Synced before the original is touched: a backup that is still only in + // the page cache when the mains go off is no backup at all. + if let Err(e) = write_synced(&backup_path, current.as_bytes()) { + eprintln!("error: cannot write backup {backup_path}: {e}"); + return ExitCode::from(1); + } + + // Preserve the file's own trailing-newline convention rather than impose + // one -- cmdline.txt is conventionally a single line with NO trailing + // newline, but writing back exactly what was there (minus the one token + // this tool changed) is the more conservative move regardless of which + // convention a given card image happens to use. + let to_write = if current.ends_with('\n') { + format!("{desired}\n") + } else { + desired.clone() + }; + // NEVER rewrite cmdline.txt in place. `fs::write` is open(O_TRUNC) + + // write: a mains cut between the truncate and the data commit leaves a + // zero-length or garbage cmdline.txt -- an unbootable Pi in a gallery + // with no operator, recoverable only by pulling the SD card on site. + // And this tool's next printed word is "REBOOT REQUIRED", i.e. it + // actively invites a power cycle while an unsynced write could still be + // sitting in the page cache. So: write a sibling temp file, fsync it, + // rename over the original, fsync the directory. Even where FAT32's + // rename atomicity is weak, temp+sync+rename strictly shrinks the + // corruption window versus in-place truncation. + let tmp_path = format!("{cmdline_path}.new"); + if let Err(e) = write_synced(&tmp_path, to_write.as_bytes()) { + eprintln!("error: cannot write {tmp_path}: {e}"); + return ExitCode::from(1); + } + if let Err(e) = fs::rename(&tmp_path, &cmdline_path) { + eprintln!("error: cannot rename {tmp_path} over {cmdline_path}: {e}"); + // Best-effort cleanup; the original is untouched either way. + let _ = fs::remove_file(&tmp_path); + return ExitCode::from(1); + } + sync_parent_dir(&cmdline_path); + + println!("dex-exhibit-apply: {cmdline_path}"); + println!(" old: {current_trimmed}"); + println!(" new: {desired}"); + println!(" backup: {backup_path}"); + prune_old_backups(&cmdline_path, &backup_path); + println!( + "REBOOT REQUIRED -- dexd binds the display from the RUNNING kernel's /proc/cmdline" + ); + + ExitCode::SUCCESS +} + +#[cfg(unix)] +fn is_root() -> bool { + // geteuid() has no safe std wrapper, and this crate deliberately keeps no + // libc dependency (Cargo.toml's SPEC §5c policy) for one syscall — the + // same reasoning that keeps main.rs's mpv bindings hand-written rather + // than pulled in via a crate. The unsafety is exactly this one call. + extern "C" { + fn geteuid() -> u32; + } + unsafe { geteuid() == 0 } +} + +#[cfg(not(unix))] +fn is_root() -> bool { + false +} diff --git a/packages/dexd/src/bin/dex-sidecar.rs b/packages/dexd/src/bin/dex-sidecar.rs new file mode 100644 index 0000000..f5fab62 --- /dev/null +++ b/packages/dexd/src/bin/dex-sidecar.rs @@ -0,0 +1,81 @@ +//! Round-trip checker for dexd's F3 sidecar contract. +//! +//! Validates a sidecar by running the player's OWN parser and hash +//! (`Sidecar::from_json`, `verify_payload`) rather than a bash/jq +//! reimplementation of the grammar that could silently drift from it. Used by +//! scripts/make-sidecar.sh on every sidecar it writes or checks. +//! +//! WHY THIS IS A CARGO BIN AND NOT A STANDALONE FILE: it used to `#[path]` +//! include ../dex-loop/src/{sidecar,sha256}.rs and build under plain +//! `rustc`. Those modules now use serde_json and sha2 (SPEC §5c), which +//! `rustc` alone cannot resolve — so the checker joined the crate rather than +//! give up the property that makes it worth having. It uses only the pure +//! library (no libmpv, no DRM), so it still builds and runs on the Mac exactly +//! as on the Pi. +//! +//! Build: cargo build --release --bin dex-sidecar +//! Usage: dex-sidecar +//! exit 0, "OK ..." on stdout -- sidecar parses AND its sha256 matches +//! the stream's exact on-disk bytes +//! exit 1, error on stderr -- whatever the real parser/verifier said +//! exit 2 -- usage error + +use dexd::sidecar; + +use std::env; +use std::fs; +use std::process::ExitCode; + +fn main() -> ExitCode { + let args: Vec = env::args().skip(1).collect(); + let [sidecar_path, stream_path] = args.as_slice() else { + eprintln!("usage: dex-sidecar "); + return ExitCode::from(2); + }; + + let text = match fs::read_to_string(sidecar_path) { + Ok(t) => t, + Err(e) => { + eprintln!("error: cannot read {sidecar_path}: {e}"); + return ExitCode::from(1); + } + }; + + let parsed = match sidecar::Sidecar::from_json(&text) { + Ok(s) => s, + Err(e) => { + eprintln!("error: {sidecar_path}: {e}"); + return ExitCode::from(1); + } + }; + + // The exact bytes the player reads: `fs::read(&path)` in main.rs, no + // filtering. Same call here, on purpose. + let payload = match fs::read(stream_path) { + Ok(p) => p, + Err(e) => { + eprintln!("error: cannot read {stream_path}: {e}"); + return ExitCode::from(1); + } + }; + + if let Err(e) = sidecar::verify_payload(&payload, &parsed) { + eprintln!("error: {e}"); + return ExitCode::from(1); + } + + println!( + "OK fps={} sha256={} width={} height={}", + parsed.fps, + parsed.sha256, + parsed + .width + .map(|w| w.to_string()) + .unwrap_or_else(|| "-".to_string()), + parsed + .height + .map(|h| h.to_string()) + .unwrap_or_else(|| "-".to_string()), + ); + ExitCode::SUCCESS +} diff --git a/packages/dexd/src/chunk.rs b/packages/dexd/src/chunk.rs new file mode 100644 index 0000000..28f746e --- /dev/null +++ b/packages/dexd/src/chunk.rs @@ -0,0 +1,185 @@ +//! T0 — the wrap arithmetic of the endless stream, extracted pure so it is +//! testable without libmpv. `read_fn` in main.rs is a thin unsafe shell over +//! `next_chunk`; the properties asserted here (never a zero-byte answer, the +//! concatenation property) ARE the program. + +/// One read request's answer: copy `n` bytes starting at payload offset +/// `start`; the reader's position afterwards is `next_pos`. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub struct Chunk { + /// Offset into the payload to copy from. Always < payload length. + pub start: usize, + /// Bytes to copy. Never 0. + pub n: usize, + /// Reader position after the copy. Always < payload length: the wrap + /// happens eagerly here, never lazily on the next call, so the invariant + /// holds between calls. + pub next_pos: usize, +} + +/// Decide the next chunk of the endless loop. +/// +/// `len` is the payload length, `pos` the reader position (any value is +/// tolerated; positions >= len wrap to 0 first), `want` the requested byte +/// count. +/// +/// Returns `None` exactly when no bytes can be produced without lying: a +/// zero-length request (`want == 0`), or the impossible-after-startup empty +/// payload (`len == 0`). The caller turns `None` into an mpv ERROR return +/// rather than 0 — NOT because mpv treats the two return values differently +/// (mpv 0.40's `stream_read_unbuffered` maps any `res <= 0` to EOF +/// uniformly, so a negative return would end the stream exactly like 0 +/// would) but because mpv never actually issues a zero-length read in the +/// first place (`stream.c` guards `len <= 0` before calling in at all) — the +/// distinction is for THIS CRATE'S OWN error ledger, so its diagnostics can +/// tell "asked for nothing" apart from "ran out of things to give". If mpv +/// ever did call with `want == 0` on some future version, the result would +/// be an ordinary END_FILE -> fatal exit -> supervisor restart, not a seam. +/// +/// A short read is legal (stream_cb.h), so the wrap is never stitched across +/// one call: the tail is returned now, the head on the next call. +pub fn next_chunk(len: usize, pos: usize, want: usize) -> Option { + if len == 0 || want == 0 { + return None; + } + let start = if pos >= len { 0 } else { pos }; + let avail = len - start; // >= 1, because start < len + let n = want.min(avail); // >= 1, because want >= 1 and avail >= 1 + let end = start + n; // <= len + let next_pos = if end == len { 0 } else { end }; + Some(Chunk { start, n, next_pos }) +} + +/// Clamp mpv's u64 request size to usize without ever turning a nonzero +/// request into 0. On a 32-bit target `as usize` truncates: an `nbytes` that +/// is an exact multiple of 2^32 would become a 0-byte request and therefore a +/// spurious final EOF. Saturating can only shrink the request, and short +/// reads are always legal. +pub fn clamp_want(nbytes: u64) -> usize { + usize::try_from(nbytes).unwrap_or(usize::MAX) +} + +#[cfg(test)] +mod tests { + use super::*; + + // T1: never returns n == 0 for any want >= 1 (a 0 return = final EOF to + // mpv, the exact event the design exists to prevent). + #[test] + fn never_zero_bytes_for_nonzero_want() { + for len in [1usize, 2, 3, 7, 64, 1000] { + for pos in [0usize, 1, len / 2, len.saturating_sub(1), len, len + 5] { + for want in [1usize, 2, len, len + 1, 10 * len, usize::MAX] { + let c = next_chunk(len, pos, want) + .expect("want >= 1 on a non-empty payload must produce a chunk"); + assert!(c.n >= 1, "n == 0 for len={len} pos={pos} want={want}"); + } + } + } + } + + // T1: want == 0 (and the impossible len == 0) are explicit, distinct + // outcomes — never conflated with a zero-byte "success" that mpv would + // read as EOF. + #[test] + fn zero_want_and_empty_payload_are_none_not_zero_chunks() { + assert_eq!(next_chunk(10, 0, 0), None); + assert_eq!(next_chunk(10, 9, 0), None); + assert_eq!(next_chunk(1, 0, 0), None); + assert_eq!(next_chunk(10, 10, 0), None); // even at the wrap point + assert_eq!(next_chunk(0, 0, 4096), None); // empty payload: error, not EOF + } + + // T1: wrap at the exact payload boundary. + #[test] + fn wraps_at_exact_boundary() { + // pos at end-of-payload (legacy lazy-caller state): wraps to 0 first. + let c = next_chunk(10, 10, 4).unwrap(); + assert_eq!((c.start, c.n, c.next_pos), (0, 4, 4)); + // tail shorter than want: short read of the tail, next_pos wraps to 0. + let c = next_chunk(10, 8, 4).unwrap(); + assert_eq!((c.start, c.n, c.next_pos), (8, 2, 0)); + // read ending exactly at len: next_pos is 0, not len (eager wrap). + let c = next_chunk(10, 6, 4).unwrap(); + assert_eq!((c.start, c.n, c.next_pos), (6, 4, 0)); + } + + // T1: payload smaller than the request. + #[test] + fn payload_smaller_than_want() { + // whole payload in one request: short read of everything, wrap to 0. + let c = next_chunk(3, 0, 4096).unwrap(); + assert_eq!((c.start, c.n, c.next_pos), (0, 3, 0)); + // single-byte payload: every read returns that byte, forever. + for _ in 0..5 { + let c = next_chunk(1, 0, 4096).unwrap(); + assert_eq!((c.start, c.n, c.next_pos), (0, 1, 0)); + } + } + + // T1: the position invariant read_fn's SAFETY comment relies on. + #[test] + fn next_pos_always_less_than_len() { + for len in [1usize, 2, 3, 5, 64, 4096] { + for pos in 0..=len + 2 { + for want in [1usize, 2, 3, len, len + 1, 3 * len] { + let c = next_chunk(len, pos, want).unwrap(); + assert!(c.next_pos < len, "next_pos={} len={len}", c.next_pos); + assert!(c.start < len); + assert!(c.start + c.n <= len); + } + } + } + } + + // T1: THE property — driving next_chunk repeatedly reproduces the payload + // repeated endlessly, byte for byte, for arbitrary request sizes. This is + // the only property that matters: the stream really is the loop. + #[test] + fn concatenation_reproduces_the_endless_loop() { + let payload: Vec = (0u8..=250).cycle().take(997).collect(); // prime length + // deterministic pseudo-random request sizes (LCG, no dependencies) + let mut rng: u64 = 0x853c49e6748fea9b; + let mut random_sizes = Vec::new(); + for _ in 0..2000 { + rng = rng + .wrapping_mul(6364136223846793005) + .wrapping_add(1442695040888963407); + random_sizes.push(((rng >> 33) % 300 + 1) as usize); // 1..=300 + } + let schedules: Vec> = vec![ + vec![1; 3000], // one byte at a time + vec![997; 8], // exactly the payload length + vec![996; 8], // one short of the payload + vec![998; 8], // one past the payload + vec![4096; 8], // far larger than the payload + random_sizes, // pseudo-random schedule + ]; + for schedule in schedules { + let mut pos = 0usize; + let mut out = Vec::new(); + for want in &schedule { + let c = next_chunk(payload.len(), pos, *want).unwrap(); + out.extend_from_slice(&payload[c.start..c.start + c.n]); + pos = c.next_pos; + } + let expected: Vec = payload.iter().copied().cycle().take(out.len()).collect(); + assert_eq!(out, expected, "stream diverged from the endless loop"); + } + } + + // T1: the u64 -> usize clamp saturates rather than truncating to 0. + // Honest limitation: on a 64-bit host try_from always succeeds, so the + // truncation branch only genuinely executes on a 32-bit target (the Pi 4 + // runs aarch64). This pins the contract against someone reintroducing + // `as usize`; a 32-bit CI target would then catch it. + #[test] + fn clamp_want_never_zero_for_nonzero_input() { + assert_eq!(clamp_want(0), 0); + assert_eq!(clamp_want(1), 1); + assert_eq!(clamp_want(4096), 4096); + assert!(clamp_want(1u64 << 32) != 0, "2^32 must not clamp to 0"); + assert!(clamp_want(u64::MAX) != 0); + assert_eq!(clamp_want(u64::MAX), usize::MAX); + } +} diff --git a/packages/dexd/src/exhibit.rs b/packages/dexd/src/exhibit.rs new file mode 100644 index 0000000..09eeb73 --- /dev/null +++ b/packages/dexd/src/exhibit.rs @@ -0,0 +1,2108 @@ +//! F6 — the exhibit display config: parse, validate, and decide the binding. +//! +//! Why: the display mode is a property of the INSTALLATION, not the asset. +//! One artwork may run on several panels over its life, and a panel may show +//! several artworks over a season — config whose lifetime differs from the +//! asset next to it eventually gets edited in the wrong copy. The +//! 2026-08-15 measurement is the concrete proof already on file: the same +//! `cmdline.txt` line is correct for one sink (Cam Link 4K, which vc4 +//! refuses to build a 4K mode for unforced even though the sink's own EDID +//! prefers it) and wrong for another (Dell U2719DC, which must NOT carry the +//! force or it transmits a signal the panel cannot show). So the mode +//! belongs to an `/etc/dex/exhibit.{json,yaml}` conffile — venue truth, not +//! asset truth — parsed with the F3 sidecar's own hardened, fail-closed flat +//! subset grammar. +//! +//! CORRECTION (2026-08-17): an earlier draft of this doc comment claimed "the +//! M5 soak played a 2160p30 asset on a 1440p Dell" as field evidence that one +//! asset runs on several panels. That run never happened — the M5 soak ran +//! 4K30 on the Cam Link with the Dell disconnected. Removed rather than left +//! to mislead a future reader; the Cam-Link-vs-Dell force disagreement above +//! is real, bench-verified evidence and stands on its own. +//! +//! TWO FORMATS, AND THE EXTENSION DECIDES WHICH — see [`ConfigFormat`] for +//! why that dispatch is a correctness rule and not a convenience. `.json` is +//! strict JSON (machine-writable, and what the `.deb` ships); `.yaml` is YAML +//! (comments, no quoting ceremony — this file gets hand-edited in a venue, +//! possibly on a phone over SSH). The two are the same schema: only the ~40 +//! lines that turn text into a flat key/value list differ, and +//! [`ExhibitConfig::from_pairs`] validates both. +//! +//! Format: +//! {"display_mode":"3840x2160@30","kms_force":"3840x2160@30", +//! "connector":"HDMI-A-1","display":"...","venue":"...","note":"..."} +//! or, identically: +//! display_mode: 3840x2160@30 # what mpv is asked for +//! kms_force: 3840x2160@30 # what the kernel cmdline must carry +//! connector: HDMI-A-1 +//! Required: `display_mode`. Optional: `kms_force` (default `"none"`), +//! `connector` (default `"HDMI-A-1"`), and the informational `display`, +//! `venue`, `note` strings, journal-logged at startup so the WHY that used to +//! live in a `config.txt` comment block travels with the config that is +//! actually enforced. +//! +//! UNKNOWN KEYS ARE REFUSED here — a deliberate divergence from the sidecar, +//! which tolerates them because ingest tooling evolves independently of +//! deployed players. This file has no such producer: it is hand-edited on the +//! same device the same `.deb` version reads it. Tolerating unknowns would +//! turn a typo (`"kms_forse"`) into a silently dropped force instead of a +//! caught one — a black gallery wall. Fail-closed IS the F-series semantics. +//! +//! Three further pure pieces live here, because none of them may live in +//! main.rs if they are to be testable without libmpv or real hardware +//! (§4 of the design): +//! +//! * [`resolve_display`] — mirrors `sidecar::resolve_fps`'s decision table: +//! the config binds; an agreeing `--mode` cross-checks it; a disagreeing +//! one is refused, naming both; an absent config (outside the bench escape +//! hatch) is refused rather than guessed. +//! * [`check_cmdline_matches`] / [`cmdline_video_token`] — the gate that +//! keeps the exhibit config and the KMS-layer `video=` token honest with +//! each other, so editing one without the other is caught at the next +//! start instead of black-screening a gallery for weeks. +//! * [`sysfs_modes_contains`] / [`mode_resolution`] — the pre-flight that +//! catches "the configured resolution is not even in this connector's mode +//! list" (the wrong-panel case) before asset/sidecar/NAL gates run. +//! * [`reconcile_cmdline`] — the pure rewrite `dex-exhibit-apply` (the +//! privileged sibling binary) uses to keep `cmdline.txt` in sync, +//! idempotently, preserving every other token untouched. + +use crate::sidecar::{parse_flat_json, Value}; +use yaml_rust2::{Event, Yaml, YamlLoader}; + +/// Default path for the exhibit config the `.deb` SHIPS, as a dpkg conffile. +/// JSON rather than YAML because a conffile is a fixed path installed by a +/// machine: the shipped artifact should be the machine-writable format, and +/// an operator who prefers YAML replaces it (see +/// [`DEFAULT_EXHIBIT_CONFIG_PATHS`]) rather than editing what dpkg tracks. +pub const DEFAULT_EXHIBIT_CONFIG_PATH: &str = "/etc/dex/exhibit.json"; + +/// Where the player looks when `--exhibit-config` is not given, in order. +/// +/// YAML first, so an operator who writes `exhibit.yaml` beside the shipped +/// `exhibit.json` gets what they wrote — the alternative (JSON wins, YAML +/// ignored) would let someone edit a file for an afternoon while the player +/// reads a different one, which is the exact "config drift" failure F6 exists +/// to end. `pick_default_config` refuses outright when both are present, so +/// "first wins" never silently decides anything: the order only fixes which +/// name the refusal calls the intended one. +pub const DEFAULT_EXHIBIT_CONFIG_PATHS: [&str; 2] = + ["/etc/dex/exhibit.yaml", "/etc/dex/exhibit.json"]; + +/// Choose the default config among those that actually exist on disk. +/// +/// `existing` is the subset of [`DEFAULT_EXHIBIT_CONFIG_PATHS`] that exists, +/// in that array's order; main.rs does the `Path::exists` calls so this stays +/// pure and testable without a filesystem. +/// +/// * exactly one → that one +/// * none → `None`, and the caller states the "no exhibit config" refusal +/// (one message, one place — `resolve_display`'s) +/// * both → **refuse**. Two configs for one player is the same class of fact +/// as a `--mode` that contradicts the config: there is a real, answerable +/// question about which the operator meant, and answering it by precedence +/// would hide it. Deleting the loser is one command; debugging a venue +/// running yesterday's mode is a day. +pub fn pick_default_config<'a>(existing: &[&'a str]) -> Result, String> { + match existing { + [] => Ok(None), + [only] => Ok(Some(only)), + several => { + // Name the LIKELY cause, not just the rule. The overwhelmingly + // common way to reach this is switching to YAML: writing + // exhibit.yaml leaves the .deb's own exhibit.json conffile sitting + // beside it, so the operator did one correct thing and got a + // refusal. A message that only restates the invariant would make + // that look like a bug in the player. + let shipped: Vec<&str> = several + .iter() + .copied() + .filter(|p| p.ends_with(".json")) + .collect(); + let hint = if shipped.is_empty() { + String::new() + } else { + format!( + " If you have just switched to YAML, the leftover is the one the package \ + installs: sudo rm {}.", + shipped.join(" ") + ) + }; + Err(format!( + "{} exhibit configs exist at once ({}) and nothing here can know which one you \ + meant. Keep exactly one — delete or rename the others — or name the intended \ + file explicitly with --exhibit-config.{hint}", + several.len(), + several.join(", ") + )) + } + } +} +pub const DEFAULT_CONNECTOR: &str = "HDMI-A-1"; +pub const DEFAULT_KMS_FORCE: &str = "none"; + +/// Is `t` a non-empty run of ASCII digits with at least one nonzero digit — +/// the same "positive integer" shape `sidecar::is_valid_fps` uses for its +/// numerator/denominator, duplicated here (rather than exposed from +/// `sidecar`) because it is a two-line primitive, not a shared contract. +fn positive_int(t: &str) -> bool { + !t.is_empty() && t.bytes().all(|b| b.is_ascii_digit()) && t.bytes().any(|b| b != b'0') +} + +/// Is `t` a non-empty run of ASCII digits (leading zeros and an all-zero +/// value both allowed — connector numbering is not a rate, "0" is a +/// legitimate enumeration index on some drivers). +fn digits_only(t: &str) -> bool { + !t.is_empty() && t.bytes().all(|b| b.is_ascii_digit()) +} + +/// Split `"WxH@R"` into its three fields. Neither half is validated here — +/// callers apply their own grammar to `w`/`h`/`r`. +fn split_mode(s: &str) -> Option<(&str, &str, &str)> { + let (wh, r) = s.split_once('@')?; + let (w, h) = wh.split_once('x')?; + Some((w, h, r)) +} + +/// `display_mode` grammar: `"auto"`, or `"WxH@R"` with W, H, R positive +/// INTEGERS — the same refresh rule as `kms_force`, and deliberately NOT the +/// `--fps` grammar an earlier revision borrowed. That revision reasoned +/// "mpv's `--drm-mode` accepts a fractional refresh, so `@29.97` and +/// `@30000/1001` are representable"; bench-driving both through the real +/// deploy path (dexpi4, mpv 0.40, 2026-08-17) disproved it twice over: +/// +/// * `@30000/1001` fails mpv's OPTION PARSER outright (`set +/// drm-mode=3840x2160@30000/1001: error setting option (-7)`) — a config +/// value this grammar accepted could NEVER play, only produce a 2 s-cadence +/// restart loop. Fail-closed belongs at config parse, not at VO init. +/// * `@29.97` parses and PLAYS — because mpv matches DRM modes by integer +/// `vrefresh` rounding, i.e. it silently drove the same 30 Hz mode that +/// `@30` names honestly. A decimal buys nothing over its rounded integer +/// (the kernel mode's timing is what it is) while implying a precision +/// that does not exist — the silent-wrongness class this crate refuses. +/// +/// (An INTEGER refresh the connector does not offer is caught loudly at VO +/// init — `Could not find mode matching 3840x2160@60`, same bench — since +/// the sysfs pre-flight can only validate the resolution half; see +/// `mode_resolution`.) The `@R` part is mandatory: `"3840x2160"` alone would +/// let mpv pick among same-resolution timings by list order, which is again +/// silent wrongness. +pub fn is_valid_display_mode(s: &str) -> bool { + if s == "auto" { + return true; + } + match split_mode(s) { + Some((w, h, r)) => positive_int(w) && positive_int(h) && positive_int(r), + None => false, + } +} + +/// `kms_force` grammar: `"none"`, or `"WxH@R"`/`"WxH@RD"` with W, H, R +/// positive integers — R is an integer here, unlike `display_mode`, because +/// the kernel's `video=` cmdline grammar has no fractional refresh. The +/// trailing `D` forces the connector to read `connected` even when nothing is +/// attached yet (the boot-order insurance F6 replaces the old comment block +/// with); it is a suffix on the whole token, not part of the refresh number. +pub fn is_valid_kms_force(s: &str) -> bool { + if s == "none" { + return true; + } + let core = s.strip_suffix('D').unwrap_or(s); + match split_mode(core) { + Some((w, h, r)) => positive_int(w) && positive_int(h) && positive_int(r), + None => false, + } +} + +/// `asset` grammar: an ABSOLUTE path to a file, with no trailing slash and no +/// ASCII control characters. +/// +/// Absolute because the player runs as a systemd service whose working +/// directory is `/`: a relative path would resolve somewhere the operator did +/// not type, and — being a *plausible* path — would fail with "no such file" +/// pointing at a name that looks right. Control characters because this string +/// is printed into the journal at every start, and the journal is the only +/// diagnostic channel a gallery device has; a newline inside it would forge a +/// second log line. +/// +/// Deliberately NOT checked here: whether the file exists, or what extension +/// it has. Existence is main.rs's job (it produces the read error, which is +/// more informative than a grammar refusal), and this crate has no business +/// deciding that an artwork must be called `.265`. +pub fn is_valid_asset(s: &str) -> bool { + s.starts_with('/') + && !s.ends_with('/') + && !s.chars().any(|c| c.is_ascii_control()) +} + +/// `connector` grammar: `"HDMI-A-"`, n a plain digit run. Also governs the +/// `video=:...` token dex-exhibit-apply writes and the +/// `/sys/class/drm/card*-` glob the sysfs pre-flight reads, so +/// accepting garbage here would surface far from where it was typed. +pub fn is_valid_connector(s: &str) -> bool { + match s.strip_prefix("HDMI-A-") { + Some(n) => digits_only(n), + None => false, + } +} + +/// Which parser reads an exhibit config — decided by the FILE EXTENSION, never +/// by sniffing the bytes. +/// +/// YAML is a superset of JSON, so "parse everything as YAML" would pass every +/// test this file could write and still be wrong: it would accept comments, +/// anchors and unquoted keys inside a file named `.json`, and that file would +/// then break `jq`, `python -m json.tool`, and any other consumer that trusts +/// the name. **An extension is a promise to the rest of the world about what +/// the bytes are.** Honouring it is the entire reason for offering two formats +/// instead of one, so the dispatch is here, at the outermost layer, where it +/// cannot be bypassed by a convenience helper. +/// +/// **Why not YAML's own JSON schema, which exists for exactly this?** Because +/// it solves a different problem than the one here. YAML 1.2 defines three +/// schemas (failsafe, JSON, core), and all three govern only how an *untagged +/// scalar* resolves to a type — they say nothing about SYNTAX. A YAML parser +/// set to the JSON schema still accepts `#` comments, block style, unquoted +/// keys, anchors and `---` document markers; it would simply resolve +/// `3840x2160` differently. So "parse `.json` with a YAML parser in JSON-schema +/// mode" would NOT deliver the promise a `.json` name makes, which is precisely +/// the promise this type exists to keep. That is why the JSON path runs on +/// `serde_json`, a real JSON parser, rather than on a configured YAML one. +/// +/// It would also be the wrong choice for the `.yaml` path, in the opposite +/// direction: under the JSON schema a plain scalar matching none of +/// null/bool/int/float is an ERROR, so `display_mode: 3840x2160@30` — an +/// unquoted string, and the entire ergonomic point of offering YAML — would +/// fail to resolve. (Moot in practice: `yaml-rust2` hardwires core-ish +/// resolution and exposes no schema selection at all. See +/// [`parse_flat_yaml`], which documents the one place it departs from 1.2 +/// core.) +/// +/// Everything after tree-building is shared: both parsers produce the same +/// `Vec<(String, Value)>` flat map that [`crate::sidecar::parse_flat_json`] +/// already produces, and [`ExhibitConfig::from_pairs`] does 100% of the +/// mapping and validation for both. Only the ~40 lines that turn text into +/// pairs differ. The subset that flat map enforces — strings and non-negative +/// integers, one level deep — is stricter in node types than any of YAML's +/// three schemas, but it applies AFTER resolution, so resolution decisions +/// remain observable: `venue: 2026` resolves to an integer and is then refused +/// as "must be a string", exactly as `"venue": 2026` is on the JSON side. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum ConfigFormat { + Json, + Yaml, +} + +impl ConfigFormat { + /// Decide the format from a path's extension, case-insensitively (an + /// operator's editor may have written `EXHIBIT.YAML`, and refusing that + /// would be pedantry rather than safety — the promise the extension makes + /// is the same either way). + /// + /// An unrecognised or absent extension REFUSES rather than defaulting to + /// either parser. Defaulting is how a `.txt` full of YAML ends up being + /// read as JSON, or worse, the reverse: silently accepting YAML in a file + /// the rest of the toolchain will read as JSON is precisely the failure + /// this type exists to make impossible. + /// + /// Uses `Path::extension` rather than splitting on the last `.`, so a + /// DOTFILE named `.json` has no extension and refuses — which is what the + /// rest of the world thinks too, and the point of this function is to + /// agree with the rest of the world about what a file name means. + pub fn from_path(path: &str) -> Result { + let ext = std::path::Path::new(path) + .extension() + .map(|e| e.to_string_lossy().to_ascii_lowercase()); + match ext.as_deref() { + Some("json") => Ok(ConfigFormat::Json), + Some("yaml") | Some("yml") => Ok(ConfigFormat::Yaml), + _ => Err(format!( + "exhibit config {path:?}: cannot tell the format from the file name — name it \ + .json (strict JSON) or .yaml/.yml (YAML). The extension decides the parser, so \ + that whatever else reads this file gets the format its name promises" + )), + } + } +} + +/// Parse `text` as a single flat YAML mapping under the SAME subset grammar +/// [`crate::sidecar::parse_flat_json`] enforces for JSON — strings and +/// non-negative integers only, one level deep, no duplicate keys — returning +/// its key/value pairs in source order. +/// +/// The subset is not a limitation grudgingly inherited from the JSON side; it +/// is what keeps the two formats interchangeable. A YAML feature this refuses +/// (a nested mapping, a list, an anchor) is one that could not be written in +/// the `.json` form of the same config, and a config whose meaning depends on +/// which extension it was saved under would defeat the point of supporting +/// both. Anchors and aliases are refused BEFORE the load, by +/// [`refuse_anchors_and_aliases`] — see there for why the node-type check +/// below cannot do it. +/// +/// Three YAML-specific hazards, all refused rather than coerced: +/// +/// * **Booleans.** `display_mode: true` is a boolean, not the string +/// `"true"`, and refuses with a message naming the quoting fix. +/// +/// MEASURED, not assumed, because the received wisdom here is wrong for +/// this library: `yaml-rust2` 0.11 resolves scalars close to the **YAML 1.2 +/// core schema**, where ONLY `true`/`false` (any of three casings) are +/// booleans. The famous "Norway problem" — `no` silently becoming `false` — +/// is a YAML **1.1** behaviour and does not occur here, so `kms_force: no` +/// arrives as the string `"no"` and is refused a step later by +/// [`is_valid_kms_force`]'s grammar, naming the valid values. Both paths +/// refuse; only the message differs. Pinned by test in both directions, so +/// a version bump that adopted 1.1 resolution could not slip through. +/// +/// "Close to", not "is": its null resolution (`yaml.rs`'s `from_str`) is +/// `"" | "~" | "null"`, where the 1.2 core schema also lists `Null` and +/// `NULL`. Driven against the real binary — `note: Null` and `note: NULL` +/// are accepted as STRINGS, while `null`, `~` and an empty value resolve to +/// null and are refused. Harmless for this schema (only the informational +/// keys could receive such a value, and taking it literally is the friendlier +/// of the two readings) but stated exactly, because "it implements the core +/// schema" is the kind of nearly-true sentence a future reader would rely on. +/// +/// **There is no schema knob to reach for**: `yaml-rust2` hardwires this +/// resolution and exposes no way to select the failsafe or JSON schema. See +/// [`ConfigFormat`] for why the JSON schema would not have been the right +/// tool for the `.json` path anyway. +/// * **Floats.** `29.97` is `Yaml::Real`. The JSON side already refuses +/// non-integers, and the bench evidence in [`is_valid_display_mode`] is why: +/// a decimal refresh silently means its rounded integer. +/// * **Multiple documents.** A `---`-separated stream has no single answer to +/// "what is the config", so it refuses instead of taking the first. +/// +/// Duplicate keys are rejected by `yaml-rust2` itself (its loader errors on +/// insert rather than last-wins, unlike most YAML libraries) — the one rule +/// the JSON side had to hand-roll a serde visitor for. Pinned by test, not +/// assumed: see `yaml_duplicate_key_refused`. +pub fn parse_flat_yaml(text: &str) -> Result, String> { + refuse_anchors_and_aliases(text)?; + let docs = YamlLoader::load_from_str(text).map_err(|e| format!("exhibit config YAML: {e}"))?; + let doc = match docs.len() { + // load_from_str returns zero documents for empty or comment-only + // input. Treated as a parse failure, not an empty config: a config + // that parses to "no keys at all" would then fail on the missing + // required key with a message about `display_mode`, burying the real + // problem (the file is blank -- truncated write, wrong path, editor + // that saved nothing). + 0 => { + return Err( + "exhibit config YAML: no document — the file is empty or contains only comments" + .into(), + ) + } + 1 => &docs[0], + n => { + return Err(format!( + "exhibit config YAML: {n} documents in one file (`---` separators) — an exhibit \ + config must be exactly one mapping, since nothing here could say which document \ + is the authoritative one" + )) + } + }; + let map = match doc { + Yaml::Hash(h) => h, + other => { + return Err(format!( + "exhibit config YAML: top level is a {}, expected a mapping of key: value pairs", + yaml_type_name(other) + )) + } + }; + let mut out = Vec::with_capacity(map.len()); + for (k, v) in map { + let key = match k { + Yaml::String(s) => s.clone(), + other => { + return Err(format!( + "exhibit config YAML: key is a {}, expected a string", + yaml_type_name(other) + )) + } + }; + let value = match v { + Yaml::String(s) => Value::Str(s.clone()), + // Non-negative only, matching the JSON subset's u64. A negative + // number is not merely out of range for these keys, it is out of + // range for the FORMAT -- so it refuses here, in the same voice a + // `.json` file's `-1` would. + Yaml::Integer(i) if *i >= 0 => Value::Num(*i as u64), + Yaml::Integer(i) => { + return Err(format!( + "exhibit config YAML: key {key:?} has negative value {i} — this format \ + carries strings and non-negative integers only" + )) + } + Yaml::Boolean(_) => { + return Err(format!( + "exhibit config YAML: key {key:?} resolved to a BOOLEAN — YAML reads bare \ + true/false as booleans, not as the text \"true\"/\"false\". Quote the value \ + if you meant a string" + )) + } + other => { + return Err(format!( + "exhibit config YAML: key {key:?} has a {} value — this format carries \ + strings and non-negative integers only, one level deep", + yaml_type_name(other) + )) + } + }; + out.push((key, value)); + } + Ok(out) +} + +/// Refuse a document that declares an anchor (`&name`) or uses an alias +/// (`*name`), BEFORE it is loaded. +/// +/// Must happen at the event level, because by the time `YamlLoader` hands back +/// a tree the aliases are gone: it resolves `*name` to a *copy* of the anchored +/// node, so a config using them arrives looking exactly like one that spelled +/// the value out. [`Yaml::Alias`] therefore never appears in a loaded document, +/// and the `Alias` arm in [`yaml_type_name`] was unreachable — this function is +/// what makes that documented refusal real. (Found 2026-08-17 by driving the +/// classic YAML footguns through the shipped parser rather than reasoning about +/// them: `note: &a hello` / `venue: *a` was silently ACCEPTED, with `venue` +/// carrying a value the file never assigns to it.) +/// +/// Refused for the reason the whole subset exists: **a config must not mean +/// something different from what it appears to say, and must not mean something +/// different depending on which extension it was saved under.** JSON has no +/// anchors, so a `.yaml` file using them could not be expressed as the `.json` +/// form of the same config — which is this module's stated test for whether a +/// YAML feature belongs in the subset. +/// +/// It also removes the one unbounded cost in this parser. Alias expansion is +/// what makes "billion laughs" possible: nested aliases expand exponentially +/// during LOADING, before any of this crate's node-type checks can run. The +/// exposure here is small (a root-owned local file on a device) — but a player +/// whose entire design is refusing to guess should not have a startup path that +/// can be made to allocate without bound by a config typo. +fn refuse_anchors_and_aliases(text: &str) -> Result<(), String> { + let mut parser = yaml_rust2::parser::Parser::new_from_str(text); + loop { + // A syntax error is not this function's business -- return Ok and let + // YamlLoader produce the real, marked parse error a line later, so the + // operator gets one good message instead of two half-ones. + let Ok((event, _marker)) = parser.next_token() else { + return Ok(()); + }; + let anchor_id = match &event { + Event::Alias(_) => { + return Err( + "exhibit config YAML: uses an alias (`*name`), which this format refuses. An \ + alias makes the file mean something it does not say -- the loader replaces \ + it with a copy of the anchored value -- and it has no JSON equivalent, so \ + the same config could not be written in the .json form. Write the value out." + .into(), + ) + } + Event::Scalar(_, _, id, _) => *id, + Event::MappingStart(id, _) | Event::SequenceStart(id, _) => *id, + Event::StreamEnd => return Ok(()), + _ => 0, + }; + // Anchor ids start at 1; 0 means "no anchor". An anchor with no alias + // is harmless in itself, but it is the half of the feature that makes + // the other half possible, and leaving it accepted would mean the + // refusal above depends on how far the operator got. + if anchor_id > 0 { + return Err( + "exhibit config YAML: declares an anchor (`&name`), which this format refuses. \ + Anchors exist to be referenced by aliases, which make a file mean something it \ + does not say and have no JSON equivalent. Write the value out." + .into(), + ); + } + } +} + +/// Name a `Yaml` node's kind for an error message, in the vocabulary an +/// operator editing YAML would recognise — not the Rust variant name. +/// +/// Returns the noun WITHOUT an article, so call sites choose their own +/// ("top level is a list" vs "has a list value"). An earlier revision baked +/// "a " into these and produced "has a a nested mapping value" at one of the +/// three call sites. +fn yaml_type_name(y: &Yaml) -> &'static str { + match y { + Yaml::Real(_) => "decimal number (this format takes integers only)", + Yaml::Integer(_) => "integer", + Yaml::String(_) => "string", + Yaml::Boolean(_) => "boolean", + Yaml::Array(_) => "list", + Yaml::Hash(_) => "nested mapping", + // Unreachable in a loaded document -- see refuse_anchors_and_aliases, + // which rejects both halves of the feature before the load. Kept so + // the match stays exhaustive without a catch-all that would silently + // absorb a future variant. + Yaml::Alias(_) => "alias (`*anchor`)", + Yaml::Null => "null (an empty value)", + Yaml::BadValue => "unreadable value", + } +} + +/// Does `p` exist, distinguishing "not there" from "cannot tell"? +/// +/// `Path::exists()` would be one line, but it maps EVERY error to `false` — +/// including EACCES on a parent directory. That is the same misreport the F6 +/// review already caught once on the read path (an unreadable config +/// producing "create this file" for a file the operator can see exists), and +/// the convenient call would quietly reintroduce it. +fn config_exists(p: &str) -> Result { + match std::fs::metadata(p) { + Ok(_) => Ok(true), + Err(e) if e.kind() == std::io::ErrorKind::NotFound => Ok(false), + Err(e) => Err(format!( + "cannot stat {p}: {e} -- this is not \"missing\", it is \"cannot tell\"; check the \ + permissions on {p} and on every directory above it (the dex user must be able to \ + traverse them)" + )), + } +} + +/// Locate, read and parse the exhibit config, returning it with the path it +/// actually came from. `Ok(None)` means no config exists. +/// +/// **The one function in this module that touches the filesystem**, and it +/// earns the exception: WHICH FILE IS THE CONFIG is a policy, and both +/// binaries that answer it — `dexd`, which enforces the config, and +/// `dex-exhibit-apply`, which reconciles the boot cmdline *to* the config — +/// must answer it identically. Two copies of this logic would drift, and the +/// specific way they would drift is that the apply tool writes a cmdline for +/// one file while the player reads another: F6's own failure mode, produced by +/// F6's own implementation. (The module's testability rule is about libmpv and +/// real DRM, not about `stat` — `defaults` is injectable precisely so this is +/// testable against a temp directory.) +/// +/// `Ok(None)` is deliberately NOT an error: [`resolve_display`] is the single +/// place that states the "no exhibit config" refusal, mirroring how +/// `sidecar::resolve_fps` states the analogous "no sidecar" one. Every OTHER +/// failure is stated here, because each is a distinct operational fact with a +/// distinct repair. +pub fn load_exhibit_config( + override_path: Option<&str>, + defaults: &[&str], +) -> Result, String> { + let path = match override_path { + // An EXPLICITLY NAMED file that is not there is not the same fact as + // "no config was ever installed", and must not borrow the latter's + // message: "create /etc/dex/exhibit.json" is actively wrong advice for + // an operator who just told us to read something else. + Some(p) => { + if !config_exists(p)? { + return Err(format!( + "--exhibit-config {p:?}: no such file. (This is the explicitly named path; \ + drop --exhibit-config to use the installed default.)" + )); + } + p.to_string() + } + None => { + let mut existing = Vec::new(); + for p in defaults { + if config_exists(p)? { + existing.push(*p); + } + } + match pick_default_config(&existing)? { + Some(p) => p.to_string(), + None => return Ok(None), + } + } + }; + let format = ConfigFormat::from_path(&path)?; + // Reaching this read means the file existed a moment ago, so NotFound is + // no longer the expected miss -- every error here is a real one. + let text = std::fs::read_to_string(&path).map_err(|e| { + format!( + "cannot read {path}: {e} -- the file exists but is not readable; check its \ + owner/permissions (the dex user must be able to read it)" + ) + })?; + let cfg = ExhibitConfig::parse(&text, format).map_err(|e| format!("{path}: {e}"))?; + Ok(Some((cfg, path))) +} + +/// A parsed, validated exhibit config. +#[derive(Debug, Clone, PartialEq, Eq)] +pub struct ExhibitConfig { + /// WHICH artwork this exhibit plays. Optional in the file, but something + /// must supply it -- see [`resolve_asset`]. + pub asset: Option, + pub display_mode: String, + pub kms_force: String, + pub connector: String, + pub display: Option, + pub venue: Option, + pub note: Option, +} + +impl ExhibitConfig { + /// Parse and validate an exhibit config, with `format` deciding the + /// parser. Fail-closed, same theory as the F3 sidecar: an unparseable + /// exhibit config and a missing one are the same operational fact — both + /// refuse rather than run on a display nobody stated. + /// + /// Callers get `format` from [`ConfigFormat::from_path`], never from the + /// bytes: see that type's docs for why sniffing would be wrong even though + /// it would always work. + pub fn parse(text: &str, format: ConfigFormat) -> Result { + match format { + ConfigFormat::Json => Self::from_json(text), + ConfigFormat::Yaml => Self::from_yaml(text), + } + } + + /// Parse and validate an exhibit config's **strict JSON** text. + pub fn from_json(text: &str) -> Result { + Self::from_pairs(parse_flat_json(text).map_err(|e| format!("exhibit config JSON: {e}"))?) + } + + /// Parse and validate an exhibit config's **YAML** text. + pub fn from_yaml(text: &str) -> Result { + Self::from_pairs(parse_flat_yaml(text)?) + } + + /// Map a parsed flat key/value tree onto the struct, and validate it. + /// + /// **This is the whole schema, and both formats reach it unchanged.** The + /// two parsers above differ only in how text becomes `kv`; every key name, + /// default, grammar check and error message lives here exactly once, so a + /// `.json` and a `.yaml` file expressing the same config cannot diverge in + /// meaning or in what they refuse. + /// + /// UNLIKE the sidecar, unknown keys are refused (module docs above) — + /// this is the one place this parser's contract deliberately differs + /// from `Sidecar::from_json`'s. + fn from_pairs(kv: Vec<(String, Value)>) -> Result { + let mut asset = None; + let mut display_mode = None; + let mut kms_force = None; + let mut connector = None; + let mut display = None; + let mut venue = None; + let mut note = None; + for (k, v) in kv { + match (k.as_str(), v) { + ("asset", Value::Str(s)) => asset = Some(s), + ("asset", Value::Num(_)) => { + return Err("exhibit config: asset must be a string".into()) + } + ("display_mode", Value::Str(s)) => display_mode = Some(s), + ("display_mode", Value::Num(_)) => { + return Err("exhibit config: display_mode must be a string".into()) + } + ("kms_force", Value::Str(s)) => kms_force = Some(s), + ("kms_force", Value::Num(_)) => { + return Err("exhibit config: kms_force must be a string".into()) + } + ("connector", Value::Str(s)) => connector = Some(s), + ("connector", Value::Num(_)) => { + return Err("exhibit config: connector must be a string".into()) + } + ("display", Value::Str(s)) => display = Some(s), + ("display", Value::Num(_)) => { + return Err("exhibit config: display must be a string".into()) + } + ("venue", Value::Str(s)) => venue = Some(s), + ("venue", Value::Num(_)) => { + return Err("exhibit config: venue must be a string".into()) + } + ("note", Value::Str(s)) => note = Some(s), + ("note", Value::Num(_)) => { + return Err("exhibit config: note must be a string".into()) + } + (other, _) => { + return Err(format!( + "exhibit config: unknown key {other:?} (strict schema, unlike the \ + sidecar — known keys: asset, display_mode, kms_force, connector, \ + display, venue, note; a typo here must not silently drop a force)" + )) + } + } + } + let display_mode = + display_mode.ok_or("exhibit config: missing required key \"display_mode\"")?; + if !is_valid_display_mode(&display_mode) { + return Err(format!( + "exhibit config: invalid display_mode {display_mode:?} (expect \"auto\" or \ + \"WxH@R\", e.g. \"3840x2160@30\")" + )); + } + let kms_force = kms_force.unwrap_or_else(|| DEFAULT_KMS_FORCE.to_string()); + if !is_valid_kms_force(&kms_force) { + return Err(format!( + "exhibit config: invalid kms_force {kms_force:?} (expect \"none\" or \ + \"WxH@R\"/\"WxH@RD\" with an integer refresh)" + )); + } + let connector = connector.unwrap_or_else(|| DEFAULT_CONNECTOR.to_string()); + if !is_valid_connector(&connector) { + return Err(format!( + "exhibit config: invalid connector {connector:?} (expect \"HDMI-A-\")" + )); + } + if let Some(a) = &asset { + if !is_valid_asset(a) { + return Err(format!( + "exhibit config: invalid asset {a:?} (expect an absolute path to the file \ + to play, e.g. \"/opt/dex/loop.265\" -- no trailing slash, no control \ + characters)" + )); + } + } + Ok(ExhibitConfig { + asset, + display_mode, + kms_force, + connector, + display, + venue, + note, + }) + } +} + +/// Where a bound display config came from: the exhibit config, or the bench +/// escape hatch (`--bench-no-sidecar` [`--mode `]). +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum DisplaySource { + Config, + Bench, +} + +/// The resolved display binding: what the player asks mpv for, what the +/// kernel cmdline is expected to carry, and which connector. +#[derive(Debug, Clone, PartialEq, Eq)] +pub struct ResolvedDisplay { + pub display_mode: String, + pub kms_force: String, + pub connector: String, + pub source: DisplaySource, +} + +/// Decide the display binding, mirroring `sidecar::resolve_fps`'s table. +/// +/// | config | `--mode` | bench | result | +/// |---|---|---|---| +/// | present | absent | no | config binds | +/// | present | == config | no | config binds (cross-check) | +/// | present | != config | no | refuse, naming both | +/// | absent | any | no | refuse: no exhibit config | +/// | any | any | yes | CLI wins (`--mode` or `"auto"`); config ignored | +/// +/// Under the bench flag the cmdline gate is ALSO skipped (main.rs), since a +/// `~/bench` build runs on hand-managed boot state by definition — that is +/// why this function reports `kms_force: "none"` and the default connector +/// for the bench branch rather than anything derived from a config that is, +/// by definition, not consulted. +pub fn resolve_display( + config: Option<&ExhibitConfig>, + cli_mode: Option<&str>, + bench_no_sidecar: bool, +) -> Result { + if bench_no_sidecar { + return match cli_mode { + Some(m) if is_valid_display_mode(m) => Ok(ResolvedDisplay { + display_mode: m.to_string(), + kms_force: DEFAULT_KMS_FORCE.to_string(), + connector: DEFAULT_CONNECTOR.to_string(), + source: DisplaySource::Bench, + }), + Some(m) => Err(format!( + "--mode {m:?} is not a valid display mode (expect \"auto\" or \"WxH@R\", e.g. \ + \"3840x2160@30\")" + )), + None => Ok(ResolvedDisplay { + display_mode: "auto".to_string(), + kms_force: DEFAULT_KMS_FORCE.to_string(), + connector: DEFAULT_CONNECTOR.to_string(), + source: DisplaySource::Bench, + }), + }; + } + match config { + None => Err( + "no exhibit config; refusing to guess the display. Create /etc/dex/exhibit.json \ + (the .deb ships one) or use --bench-no-sidecar on a bench" + .into(), + ), + Some(cfg) => match cli_mode { + None => Ok(ResolvedDisplay { + display_mode: cfg.display_mode.clone(), + kms_force: cfg.kms_force.clone(), + connector: cfg.connector.clone(), + source: DisplaySource::Config, + }), + Some(m) if m == cfg.display_mode => Ok(ResolvedDisplay { + display_mode: cfg.display_mode.clone(), + kms_force: cfg.kms_force.clone(), + connector: cfg.connector.clone(), + source: DisplaySource::Config, + }), + Some(m) => Err(format!( + "--mode {m:?} contradicts exhibit config display_mode {:?}; drop --mode (the \ + exhibit config is authoritative) or fix /etc/dex/exhibit.json", + cfg.display_mode + )), + }, + } +} + +/// Where the asset path came from. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum AssetSource { + /// The exhibit config's `asset` key — the deployment path. + Config, + /// A path given on the command line. Legitimate on a bench, and as a + /// one-off on a device without editing `/etc`; never how a show runs. + Cli, +} + +/// The resolved asset: which file to play, and who said so. +#[derive(Debug, Clone, PartialEq, Eq)] +pub struct ResolvedAsset { + pub path: String, + pub source: AssetSource, +} + +/// Decide WHICH ASSET to play — the capability the exhibit file exists for. +/// +/// Before this, `dexd.service`'s `ExecStart` hardcoded +/// `/opt/dex/loop.265`, so changing the artwork meant overwriting that one +/// path or editing a systemd unit the package owns. Several assets could not +/// sit in storage with the exhibit choosing one, which was Max's stated reason +/// for wanting an exhibit file at all: an exhibit is a pairing of a venue with +/// an artwork, and it could previously express only the venue half. +/// +/// The table mirrors [`resolve_display`] and `sidecar::resolve_fps` exactly, +/// because a third decision surface with its own idiom is how invariants get +/// forgotten: +/// +/// | config `asset` | CLI path | bench | result | +/// |---|---|---|---| +/// | present | absent | no | config binds | +/// | present | == config | no | config binds (cross-check) | +/// | present | != config | no | refuse, naming both | +/// | absent | present | no | CLI binds (one-off, logged as such) | +/// | absent | absent | no | refuse: nothing names an asset | +/// | any | present | yes | CLI binds; config ignored | +/// | any | absent | yes | refuse: bench must be explicit | +/// +/// The row that could have gone the other way is the fifth. Defaulting to +/// `/opt/dex/loop.265` would carry every existing deployment through untouched +/// — and would be the worst guess in the program: it would silently play LAST +/// season's artwork for an operator who edited the config but mistyped the +/// key, with every metric green. Every other guess this crate refuses (frame +/// rate, display mode) is refused for a weaker version of that reason. So: +/// refuse, name the exact line to add, and warn at package-install time (see +/// `deploy/maintainer-scripts/postinst`) so an upgrade surfaces it while +/// someone is watching rather than at the next power cycle. +pub fn resolve_asset( + config: Option<&ExhibitConfig>, + cli_path: Option<&str>, + bench_no_sidecar: bool, +) -> Result { + if bench_no_sidecar { + return match cli_path { + Some(p) => Ok(ResolvedAsset { + path: p.to_string(), + source: AssetSource::Cli, + }), + None => Err( + "--bench-no-sidecar consults no exhibit config, so the asset must be given on \ + the command line: dexd --bench-no-sidecar --fps " + .into(), + ), + }; + } + match (config.and_then(|c| c.asset.as_deref()), cli_path) { + (Some(a), None) => Ok(ResolvedAsset { + path: a.to_string(), + source: AssetSource::Config, + }), + (Some(a), Some(p)) if a == p => Ok(ResolvedAsset { + path: a.to_string(), + source: AssetSource::Config, + }), + (Some(a), Some(p)) => Err(format!( + "command-line asset {p:?} contradicts the exhibit config's asset {a:?}; drop the \ + path (the exhibit config is authoritative) or fix the config" + )), + (None, Some(p)) => Ok(ResolvedAsset { + path: p.to_string(), + source: AssetSource::Cli, + }), + (None, None) => Err( + "no asset: the exhibit config does not name one and none was given on the command \ + line. Add it to the exhibit config -- `asset: /opt/dex/loop.265` (YAML) or \ + `\"asset\": \"/opt/dex/loop.265\"` (JSON) -- which is what lets several assets sit \ + in /opt/dex with the exhibit choosing one" + .into(), + ), + } +} + +/// Find the `video=:` token for exactly `connector` inside a +/// kernel cmdline string (space-separated tokens, as `/proc/cmdline` and +/// `cmdline.txt` both are). Other connectors' `video=` tokens, and every +/// other kind of token, are ignored. Returns the `` half only. +pub fn cmdline_video_token(cmdline: &str, connector: &str) -> Option { + cmdline.split_whitespace().find_map(|tok| { + let rest = tok.strip_prefix("video=")?; + let (conn, mode) = rest.split_once(':')?; + (conn == connector).then(|| mode.to_string()) + }) +} + +/// Find a CONNECTORLESS `video=` token (`video=WxH@R`, no `conn:` prefix) — +/// the kernel's grammar also accepts this shape, and it forces ALL +/// connectors. Neither this module's gate nor its reconciler can reason +/// about one (which connector does it bind? how does the kernel arbitrate it +/// against a per-connector token?), so callers REFUSE when one is present +/// rather than reporting "no token" (the gate) or appending a second, +/// overlapping force (the reconciler) — refusing beats guessing at kernel +/// arbitration order for a token shape no dex install is supposed to carry. +/// Returns the whole token, for the refusal message to name verbatim. +pub fn connectorless_video_token(cmdline: &str) -> Option { + cmdline.split_whitespace().find_map(|tok| { + let rest = tok.strip_prefix("video=")?; + (!rest.contains(':')).then(|| tok.to_string()) + }) +} + +/// The gate that keeps the exhibit config and the KMS-layer `video=` token +/// honest with each other (§3.3 of the design): an operator who edits one and +/// forgets the other must be refused, loudly, naming what to run — not +/// black-screen the gallery on whichever value the kernel happens to have +/// booted with. `kms_force == "none"` means "expect NO token for this +/// connector"; anything else means "expect this EXACT token". +/// Every refusal names BOTH possible repairs, because this gate cannot know +/// which side is stale: the config may be right and the cmdline leftover — +/// but equally the cmdline may carry a force this venue deliberately needs +/// (the Cam Link builds no 4K mode unforced) that a freshly installed +/// default config simply does not know about yet. A message prescribing only +/// `dex-exhibit-apply` in that second case would instruct the operator to +/// DELETE the needed force, after which `display_mode=auto` silently plays +/// at whatever the connector negotiates — the exact wrongness class this +/// gate exists to close, reached by following its own instructions. +pub fn check_cmdline_matches(cmdline: &str, connector: &str, kms_force: &str) -> Result<(), String> { + if let Some(tok) = connectorless_video_token(cmdline) { + return Err(format!( + "the kernel cmdline carries a connectorless token {tok:?}, which forces ALL \ + connectors -- this gate cannot reconcile it with the exhibit config's \ + per-connector kms_force. Qualify it with a connector \ + (video={connector}:) or remove it from /boot/firmware/cmdline.txt, \ + then reboot" + )); + } + let found = cmdline_video_token(cmdline, connector); + match (kms_force, found) { + (DEFAULT_KMS_FORCE, None) => Ok(()), + (DEFAULT_KMS_FORCE, Some(found)) => Err(format!( + "exhibit config says kms_force=none for {connector}, but the kernel cmdline \ + carries video={connector}:{found} -- EITHER the config is stale (a force this \ + venue deliberately needs, e.g. a display that builds no 4K mode unforced: set \ + kms_force={found:?} and a matching display_mode in /etc/dex/exhibit.json to \ + keep it) OR the cmdline is (run 'sudo dex-exhibit-apply' and reboot to remove \ + the force)" + )), + (want, Some(found)) if want == found => Ok(()), + (want, Some(found)) => Err(format!( + "exhibit config says kms_force={want}, but the kernel cmdline carries \ + video={connector}:{found} -- EITHER the config is stale (fix kms_force in \ + /etc/dex/exhibit.json to match the venue) OR the cmdline is (run 'sudo \ + dex-exhibit-apply' and reboot)" + )), + (want, None) => Err(format!( + "exhibit config says kms_force={want}, but the kernel cmdline carries no video= \ + token for {connector} -- EITHER the config is stale (set kms_force=none in \ + /etc/dex/exhibit.json if this venue needs no force) OR the cmdline is (run \ + 'sudo dex-exhibit-apply' and reboot to add the token)" + )), + } +} + +/// The `WxH` half of a `display_mode` (`"auto"` has none — sysfs has no +/// "auto" entry to check against, so callers must skip the pre-flight for it). +pub fn mode_resolution(display_mode: &str) -> Option<&str> { + if display_mode == "auto" { + return None; + } + display_mode.split_once('@').map(|(wh, _)| wh) +} + +/// Does `modes_text` (the verbatim contents of +/// `/sys/class/drm/card*-/modes` — one `WxH` per line, no refresh +/// column) list `want_wh`? Pure over fixture text so the wrong-panel case +/// (Dell asked for 2160) is testable without real DRM. +pub fn sysfs_modes_contains(modes_text: &str, want_wh: &str) -> bool { + modes_text.lines().map(str::trim).any(|l| l == want_wh) +} + +/// Rewrite a single-line `cmdline.txt`'s `video=:...` token to +/// match `kms_force`, preserving every other token, its position, and every +/// OTHER connector's `video=` token untouched. `kms_force == "none"` means +/// "remove the token for this connector, if present". Idempotent: running +/// this twice on its own output is a no-op, because a token that already +/// exists is replaced in place rather than moved to the end — the second run +/// finds the same token already correct and changes nothing. +pub fn reconcile_cmdline(cmdline_text: &str, connector: &str, kms_force: &str) -> Result { + let trimmed = cmdline_text.trim_end_matches(['\n', '\r']); + if trimmed.contains('\n') || trimmed.contains('\r') { + return Err("cmdline.txt must be a single line".into()); + } + if !is_valid_kms_force(kms_force) { + return Err(format!("reconcile_cmdline: invalid kms_force {kms_force:?}")); + } + // A connectorless `video=` token forces ALL connectors; rewriting around + // it would leave two overlapping forces for the kernel to arbitrate -- + // see connectorless_video_token's doc for why refusing beats guessing. + if let Some(tok) = connectorless_video_token(trimmed) { + return Err(format!( + "cmdline.txt carries a connectorless token {tok:?}, which forces ALL connectors \ + -- refusing to reconcile per-connector video={connector}:... tokens around it \ + (the kernel would arbitrate two overlapping forces). Qualify or remove {tok:?} \ + by hand first" + )); + } + let desired_token = if kms_force == DEFAULT_KMS_FORCE { + None + } else { + Some(format!("video={connector}:{kms_force}")) + }; + + let mut out: Vec = Vec::new(); + let mut replaced = false; + for tok in trimmed.split_whitespace() { + let is_ours = tok + .strip_prefix("video=") + .and_then(|rest| rest.split_once(':')) + .is_some_and(|(conn, _)| conn == connector); + if is_ours { + replaced = true; + if let Some(d) = &desired_token { + out.push(d.clone()); + } + // else: kms_force == "none" -- drop this token. + continue; + } + out.push(tok.to_string()); + } + if !replaced { + if let Some(d) = desired_token { + out.push(d); + } + } + + let result = out.join(" "); + if result.trim().is_empty() { + return Err( + "reconcile_cmdline: result would be an empty cmdline.txt -- refusing to write it" + .into(), + ); + } + Ok(result) +} + +#[cfg(test)] +mod tests { + use super::*; + + // ---- is_valid_display_mode / is_valid_kms_force / is_valid_connector -- + + #[test] + fn display_mode_grammar() { + for ok in [ + "auto", + "3840x2160@30", + "1x1@1", + // Leading zeros are accepted -- positive_int only requires "all + // digits, at least one nonzero", the same rule sidecar::is_valid_fps + // applies to its numerator/denominator. Documented here rather than + // left to be discovered by a future reader of the bad list below. + "03840x2160@30", + ] { + assert!(is_valid_display_mode(ok), "{ok} should be valid"); + } + for bad in [ + "", + "Auto", + "3840x2160", // no @R -- mandatory, see docs + "3840x2160@", // empty R + "x2160@30", // empty W + "3840x@30", // empty H + "3840X2160@30", // uppercase X separator + "3840x2160@0", // R must be positive + "3840x2160@-30", // negative R + "3840x2160@30D", // D suffix is a kms_force thing, not display_mode + "auto@30", + // Bench-disproven forms an earlier revision accepted (dexpi4, + // mpv 0.40, 2026-08-17 -- see is_valid_display_mode's docs): + // rational fails mpv's option parser (-7, guaranteed restart + // loop), decimal silently rounds to the integer vrefresh. + "3840x2160@30000/1001", + "2560x1440@59.95", + ] { + assert!(!is_valid_display_mode(bad), "{bad:?} should be invalid"); + } + } + + #[test] + fn kms_force_grammar() { + for ok in ["none", "3840x2160@30", "3840x2160@30D", "1x1@1D"] { + assert!(is_valid_kms_force(ok), "{ok} should be valid"); + } + for bad in [ + "", + "None", + "auto", // auto is a display_mode concept, not kms_force + "3840x2160@29.97", // fractional refresh: kernel video= grammar has none + "3840x2160", + "3840x2160@", + "3840x2160@30d", // lowercase d is not the force suffix + "3840x2160@30DD", + "3840x2160@0", + "3840x2160@0D", + ] { + assert!(!is_valid_kms_force(bad), "{bad:?} should be invalid"); + } + } + + #[test] + fn connector_grammar() { + for ok in ["HDMI-A-1", "HDMI-A-2", "HDMI-A-0", "HDMI-A-10"] { + assert!(is_valid_connector(ok), "{ok} should be valid"); + } + for bad in ["", "HDMI-A-", "HDMI-A", "DP-1", "hdmi-a-1", "HDMI-A-1 "] { + assert!(!is_valid_connector(bad), "{bad:?} should be invalid"); + } + } + + // ---- ExhibitConfig::from_json ------------------------------------------ + + #[test] + fn parses_the_canonical_config() { + let text = r#"{"display_mode":"3840x2160@30","kms_force":"3840x2160@30", + "connector":"HDMI-A-1","display":"Cam Link 4K", + "venue":"bench","note":"see PLAN.md F6"}"#; + let c = ExhibitConfig::from_json(text).unwrap(); + assert_eq!(c.display_mode, "3840x2160@30"); + assert_eq!(c.kms_force, "3840x2160@30"); + assert_eq!(c.connector, "HDMI-A-1"); + assert_eq!(c.display.as_deref(), Some("Cam Link 4K")); + assert_eq!(c.venue.as_deref(), Some("bench")); + assert_eq!(c.note.as_deref(), Some("see PLAN.md F6")); + } + + /// The .deb ships this exact file as `/etc/dex/exhibit.json`'s stock + /// conffile content (Cargo.toml's `assets`). If a future grammar change + /// ever made the shipped default itself invalid, every fresh install + /// would refuse to start with no asset ever having been touched -- + /// pinned here rather than discovered on a bench. + #[test] + fn the_shipped_default_config_parses() { + let text = include_str!("../deploy/exhibit.json.default"); + let c = ExhibitConfig::from_json(text).expect("shipped default must parse"); + assert_eq!(c.display_mode, "auto"); + assert_eq!(c.kms_force, "none"); + // The shipped default MUST name an asset. Without one, a stock install + // now refuses at startup -- ExecStart no longer passes a path, and the + // code deliberately has no fallback (see resolve_asset). This assert is + // the gate on that: the conffile is the only thing standing between a + // fresh install and "no asset" on first start. + assert_eq!(c.asset.as_deref(), Some("/opt/dex/loop.265")); + // ...and it must actually resolve, not merely parse. + assert_eq!( + resolve_asset(Some(&c), None, false).unwrap().path, + "/opt/dex/loop.265" + ); + } + + #[test] + fn minimal_config_gets_defaults() { + let c = ExhibitConfig::from_json(r#"{"display_mode":"auto"}"#).unwrap(); + assert_eq!(c.display_mode, "auto"); + assert_eq!(c.kms_force, DEFAULT_KMS_FORCE); + assert_eq!(c.connector, DEFAULT_CONNECTOR); + assert_eq!(c.display, None); + } + + #[test] + fn unknown_keys_are_refused_unlike_the_sidecar() { + let e = ExhibitConfig::from_json(r#"{"display_mode":"auto","future_key":"x"}"#) + .unwrap_err(); + assert!(e.contains("unknown key") && e.contains("future_key"), "{e}"); + } + + #[test] + fn typo_in_kms_force_is_caught_not_silently_dropped() { + // The exact field scenario the strict-schema decision defends against: + // "kms_forse" (typo) must be a parse error, not an ignored key that + // leaves the intended force unset. + let e = ExhibitConfig::from_json( + r#"{"display_mode":"3840x2160@30","kms_forse":"3840x2160@30"}"#, + ) + .unwrap_err(); + assert!(e.contains("kms_forse"), "{e}"); + } + + #[test] + fn missing_display_mode_is_refused_by_name() { + let e = ExhibitConfig::from_json(r#"{"kms_force":"none"}"#).unwrap_err(); + assert!(e.contains("display_mode"), "{e}"); + } + + #[test] + fn invalid_display_mode_and_kms_force_and_connector_are_refused() { + assert!(ExhibitConfig::from_json(r#"{"display_mode":"nope"}"#).is_err()); + assert!(ExhibitConfig::from_json( + r#"{"display_mode":"auto","kms_force":"nope"}"# + ) + .is_err()); + assert!(ExhibitConfig::from_json( + r#"{"display_mode":"auto","connector":"DP-1"}"# + ) + .is_err()); + } + + #[test] + fn display_mode_as_number_is_refused() { + let e = ExhibitConfig::from_json(r#"{"display_mode":30}"#).unwrap_err(); + assert!(e.contains("string"), "{e}"); + } + + #[test] + fn duplicate_keys_are_refused_inherited_from_the_sidecar_grammar() { + let e = ExhibitConfig::from_json( + r#"{"display_mode":"auto","display_mode":"3840x2160@30"}"#, + ) + .unwrap_err(); + assert!(e.contains("duplicate"), "{e}"); + } + + // ---- resolve_display ---------------------------------------------------- + + fn cfg(display_mode: &str, kms_force: &str) -> ExhibitConfig { + ExhibitConfig { + asset: None, + display_mode: display_mode.to_string(), + kms_force: kms_force.to_string(), + connector: DEFAULT_CONNECTOR.to_string(), + display: None, + venue: None, + note: None, + } + } + + /// `cfg`, plus an asset — for the `resolve_asset` table. + fn cfg_with_asset(asset: &str) -> ExhibitConfig { + ExhibitConfig { + asset: Some(asset.to_string()), + ..cfg("auto", "none") + } + } + + #[test] + fn deploy_path_takes_display_from_config() { + let c = cfg("3840x2160@30", "3840x2160@30"); + let r = resolve_display(Some(&c), None, false).unwrap(); + assert_eq!(r.display_mode, "3840x2160@30"); + assert_eq!(r.kms_force, "3840x2160@30"); + assert_eq!(r.source, DisplaySource::Config); + } + + #[test] + fn agreeing_cli_mode_allowed_disagreeing_refused_naming_both() { + let c = cfg("3840x2160@30", "none"); + assert!(resolve_display(Some(&c), Some("3840x2160@30"), false).is_ok()); + let e = resolve_display(Some(&c), Some("2560x1440@60"), false).unwrap_err(); + assert!(e.contains("3840x2160@30") && e.contains("2560x1440@60"), "{e}"); + } + + #[test] + fn missing_config_is_refused_without_the_bench_flag() { + let e = resolve_display(None, Some("3840x2160@30"), false).unwrap_err(); + assert!(e.contains("exhibit config"), "{e}"); + assert!(resolve_display(None, None, false).is_err()); + } + + #[test] + fn bench_escape_hatch_ignores_config_entirely() { + let c = cfg("2560x1440@60", "2560x1440@60"); + // CLI --mode wins even though it contradicts the config. + let r = resolve_display(Some(&c), Some("3840x2160@30"), true).unwrap(); + assert_eq!(r.display_mode, "3840x2160@30"); + assert_eq!(r.source, DisplaySource::Bench); + assert_eq!(r.kms_force, DEFAULT_KMS_FORCE); + } + + #[test] + fn bench_with_no_mode_defaults_to_auto_unlike_fps() { + // Unlike resolve_fps, which REQUIRES --fps under the bench flag, + // resolve_display treats a missing --mode as "auto" -- see the design + // table: "no --mode means auto". + let r = resolve_display(None, None, true).unwrap(); + assert_eq!(r.display_mode, "auto"); + assert_eq!(r.source, DisplaySource::Bench); + } + + #[test] + fn bench_with_invalid_mode_is_refused() { + let e = resolve_display(None, Some("banana"), true).unwrap_err(); + assert!(e.contains("banana"), "{e}"); + } + + // ---- cmdline_video_token / check_cmdline_matches ------------------------- + + #[test] + fn cmdline_video_token_finds_only_the_named_connector() { + let cl = "console=ttyS0 video=HDMI-A-1:3840x2160@30D quiet video=HDMI-A-2:1920x1080@60"; + assert_eq!( + cmdline_video_token(cl, "HDMI-A-1"), + Some("3840x2160@30D".to_string()) + ); + assert_eq!( + cmdline_video_token(cl, "HDMI-A-2"), + Some("1920x1080@60".to_string()) + ); + assert_eq!(cmdline_video_token(cl, "HDMI-A-3"), None); + } + + #[test] + fn cmdline_gate_none_expected_none_present_matches() { + assert!(check_cmdline_matches("console=ttyS0 quiet", "HDMI-A-1", "none").is_ok()); + } + + #[test] + fn cmdline_gate_none_expected_but_present_refuses_naming_the_fix() { + let e = check_cmdline_matches( + "video=HDMI-A-1:3840x2160@30", + "HDMI-A-1", + "none", + ) + .unwrap_err(); + assert!(e.contains("dex-exhibit-apply") && e.contains("3840x2160@30"), "{e}"); + } + + #[test] + fn cmdline_gate_wrong_mode_refuses_naming_both() { + let e = check_cmdline_matches( + "video=HDMI-A-1:3840x2160@30", + "HDMI-A-1", + "2560x1440@60", + ) + .unwrap_err(); + assert!(e.contains("2560x1440@60") && e.contains("3840x2160@30"), "{e}"); + } + + #[test] + fn cmdline_gate_wrong_connector_is_treated_as_absent() { + // The configured connector's token is missing even though a DIFFERENT + // connector's token is present -- must refuse "no token", not match. + let e = check_cmdline_matches( + "video=HDMI-A-2:1920x1080@60", + "HDMI-A-1", + "3840x2160@30", + ) + .unwrap_err(); + assert!(e.contains("no video=") || e.contains("carries no"), "{e}"); + } + + #[test] + fn cmdline_gate_refuses_a_connectorless_video_token() { + // Kernel grammar also accepts `video=WxH@R` with no connector, which + // forces ALL connectors. Previously invisible to the gate (reported + // as "no video= token") -- it must refuse, naming the token. + let e = check_cmdline_matches( + "console=ttyS0 video=1920x1080@60 quiet", + "HDMI-A-1", + "3840x2160@30", + ) + .unwrap_err(); + assert!(e.contains("video=1920x1080@60") && e.contains("ALL connectors"), "{e}"); + // Even when the per-connector expectation is "none": the global + // force still binds our connector, so "matches" would be a lie. + let e = check_cmdline_matches("video=1920x1080@60", "HDMI-A-1", "none").unwrap_err(); + assert!(e.contains("connectorless"), "{e}"); + } + + #[test] + fn connectorless_video_token_ignores_per_connector_tokens() { + assert_eq!(connectorless_video_token("video=HDMI-A-1:3840x2160@30 quiet"), None); + assert_eq!( + connectorless_video_token("quiet video=1024x768"), + Some("video=1024x768".to_string()) + ); + assert_eq!(connectorless_video_token("console=ttyS0 quiet"), None); + } + + #[test] + fn cmdline_gate_expected_present_matches() { + assert!(check_cmdline_matches( + "root=/dev/mmcblk0p2 video=HDMI-A-1:3840x2160@30D rootwait", + "HDMI-A-1", + "3840x2160@30D", + ) + .is_ok()); + } + + // ---- mode_resolution / sysfs_modes_contains ------------------------------ + + #[test] + fn mode_resolution_strips_the_refresh_auto_has_none() { + assert_eq!(mode_resolution("3840x2160@30"), Some("3840x2160")); + assert_eq!(mode_resolution("auto"), None); + } + + #[test] + fn sysfs_modes_contains_checks_whole_line_matches() { + let modes = "3840x2160\n3840x2160\n1920x1080\n1024x768\n"; + assert!(sysfs_modes_contains(modes, "3840x2160")); + assert!(sysfs_modes_contains(modes, "1024x768")); + assert!(!sysfs_modes_contains(modes, "2560x1440")); + // Multi-card glob output concatenated is still just lines. + let two_cards = "3840x2160\n1920x1080\n2560x1440\n1920x1080\n"; + assert!(sysfs_modes_contains(two_cards, "2560x1440")); + assert!(!sysfs_modes_contains(two_cards, "7680x4320")); + } + + #[test] + fn sysfs_modes_contains_trims_whitespace() { + assert!(sysfs_modes_contains(" 3840x2160 \r\n", "3840x2160")); + } + + // ---- reconcile_cmdline --------------------------------------------------- + + #[test] + fn reconcile_adds_a_token_when_none_existed() { + let out = reconcile_cmdline( + "console=ttyS0 root=/dev/mmcblk0p2 rootwait quiet", + "HDMI-A-1", + "3840x2160@30D", + ) + .unwrap(); + assert_eq!( + out, + "console=ttyS0 root=/dev/mmcblk0p2 rootwait quiet video=HDMI-A-1:3840x2160@30D" + ); + } + + #[test] + fn reconcile_replaces_in_place_preserving_position_and_order() { + let out = reconcile_cmdline( + "console=ttyS0 video=HDMI-A-1:3840x2160@30 rootwait quiet", + "HDMI-A-1", + "2560x1440@60", + ) + .unwrap(); + assert_eq!( + out, + "console=ttyS0 video=HDMI-A-1:2560x1440@60 rootwait quiet" + ); + } + + #[test] + fn reconcile_removes_the_token_for_none() { + let out = reconcile_cmdline( + "console=ttyS0 video=HDMI-A-1:3840x2160@30 rootwait", + "HDMI-A-1", + "none", + ) + .unwrap(); + assert_eq!(out, "console=ttyS0 rootwait"); + } + + #[test] + fn reconcile_leaves_other_connectors_video_tokens_untouched() { + let out = reconcile_cmdline( + "video=HDMI-A-2:1920x1080@60 video=HDMI-A-1:3840x2160@30 quiet", + "HDMI-A-1", + "none", + ) + .unwrap(); + assert_eq!(out, "video=HDMI-A-2:1920x1080@60 quiet"); + } + + #[test] + fn reconcile_is_idempotent() { + let once = reconcile_cmdline( + "console=ttyS0 rootwait", + "HDMI-A-1", + "3840x2160@30D", + ) + .unwrap(); + let twice = reconcile_cmdline(&once, "HDMI-A-1", "3840x2160@30D").unwrap(); + assert_eq!(once, twice); + + // And the "no change" case: apply the SAME force that is already + // present -- dex-exhibit-apply's real first-run-on-dexpi4 scenario. + let already = "console=ttyS0 video=HDMI-A-1:3840x2160@30 rootwait"; + let out = reconcile_cmdline(already, "HDMI-A-1", "3840x2160@30").unwrap(); + assert_eq!(out, already); + } + + #[test] + fn reconcile_refuses_a_connectorless_video_token() { + // Previously the reconciler would leave the global force in place and + // append its own per-connector token -- two overlapping forces for + // the kernel to arbitrate. Refuse instead, naming the token. + let e = reconcile_cmdline("console=ttyS0 video=1024x768 rootwait", "HDMI-A-1", "3840x2160@30") + .unwrap_err(); + assert!(e.contains("video=1024x768") && e.contains("by hand"), "{e}"); + // Same refusal on the removal direction (kms_force=none). + assert!(reconcile_cmdline("video=1024x768", "HDMI-A-1", "none").is_err()); + } + + #[test] + fn reconcile_no_op_when_none_requested_and_none_present() { + let out = reconcile_cmdline("console=ttyS0 rootwait", "HDMI-A-1", "none").unwrap(); + assert_eq!(out, "console=ttyS0 rootwait"); + } + + #[test] + fn reconcile_rejects_multi_line_input() { + assert!(reconcile_cmdline("a b\nc d", "HDMI-A-1", "none").is_err()); + } + + #[test] + fn reconcile_tolerates_a_trailing_newline() { + // cmdline.txt conventionally has no trailing newline, but a file an + // editor "fixed" by adding one must not be treated as multi-line. + let out = reconcile_cmdline("console=ttyS0\n", "HDMI-A-1", "none").unwrap(); + assert_eq!(out, "console=ttyS0"); + } + + #[test] + fn reconcile_refuses_an_empty_result() { + let e = reconcile_cmdline("video=HDMI-A-1:3840x2160@30", "HDMI-A-1", "none").unwrap_err(); + assert!(e.contains("empty"), "{e}"); + } + + #[test] + fn reconcile_rejects_an_invalid_kms_force() { + assert!(reconcile_cmdline("console=ttyS0", "HDMI-A-1", "banana").is_err()); + } + + // ---- the format dispatch ------------------------------------------- + // + // The rule under test is "the extension decides", so these check the + // DISPATCH, not the parsers. The parsers get their own sections below. + + #[test] + fn extension_decides_the_parser() { + assert_eq!( + ConfigFormat::from_path("/etc/dex/exhibit.json").unwrap(), + ConfigFormat::Json + ); + assert_eq!( + ConfigFormat::from_path("/etc/dex/exhibit.yaml").unwrap(), + ConfigFormat::Yaml + ); + assert_eq!( + ConfigFormat::from_path("/etc/dex/exhibit.yml").unwrap(), + ConfigFormat::Yaml + ); + } + + #[test] + fn extension_match_is_case_insensitive() { + assert_eq!( + ConfigFormat::from_path("/tmp/EXHIBIT.JSON").unwrap(), + ConfigFormat::Json + ); + assert_eq!( + ConfigFormat::from_path("/tmp/Exhibit.Yaml").unwrap(), + ConfigFormat::Yaml + ); + } + + #[test] + fn unknown_or_absent_extension_refuses_rather_than_defaulting() { + for p in [ + "/etc/dex/exhibit", // no extension at all + "/etc/dex/exhibit.txt", // an extension, but not one of ours + "/etc/dex/.json", // a DOTFILE named .json -- no extension + "/etc/dex/exhibit.json.bak", // the backup, not the config + ] { + let e = ConfigFormat::from_path(p).unwrap_err(); + assert!(e.contains(".json") && e.contains(".yaml"), "{p}: {e}"); + } + } + + /// THE point of the whole dispatch: YAML syntax inside a file named + /// `.json` must be refused, even though a YAML parser would accept it + /// happily. A `.json` file that only `dexd` can read is a broken + /// promise to `jq` and everything else downstream. + #[test] + fn yaml_syntax_in_a_json_file_is_refused() { + let yaml_text = "display_mode: 3840x2160@30\nkms_force: none\n"; + // The YAML parser accepts it, proving the input is valid YAML... + assert!(ExhibitConfig::parse(yaml_text, ConfigFormat::Yaml).is_ok()); + // ...and the JSON parser must still refuse it under a .json name. + assert!(ExhibitConfig::parse(yaml_text, ConfigFormat::Json).is_err()); + } + + /// The converse, which must NOT be an error: JSON is a subset of YAML, so + /// a machine that writes strict JSON into a `.yaml` file still parses. + /// This is what lets an ingest tool emit one format for both names. + #[test] + fn json_text_parses_under_the_yaml_parser_too() { + let json_text = r#"{"display_mode":"3840x2160@30","kms_force":"none"}"#; + let as_json = ExhibitConfig::parse(json_text, ConfigFormat::Json).unwrap(); + let as_yaml = ExhibitConfig::parse(json_text, ConfigFormat::Yaml).unwrap(); + assert_eq!(as_json, as_yaml); + } + + /// Equivalence: the same config in either format produces the same + /// struct, byte for byte. This is the test that would fail first if the + /// two paths ever stopped sharing `from_pairs`. + #[test] + fn both_formats_agree_on_a_full_config() { + let json = ExhibitConfig::from_json( + r#"{"display_mode":"3840x2160@30","kms_force":"3840x2160@30D", + "connector":"HDMI-A-2","display":"Elgato Cam Link 4K", + "venue":"gallery east wall","note":"vc4 builds no 4K mode unforced"}"#, + ) + .unwrap(); + let yaml = ExhibitConfig::from_yaml( + "# the same thing, with the comments JSON cannot carry\n\ + display_mode: 3840x2160@30\n\ + kms_force: 3840x2160@30D # trailing D: force `connected`\n\ + connector: HDMI-A-2\n\ + display: Elgato Cam Link 4K\n\ + venue: gallery east wall\n\ + note: vc4 builds no 4K mode unforced\n", + ) + .unwrap(); + assert_eq!(json, yaml); + } + + /// Validation is shared, so a YAML file gets the JSON path's messages -- + /// including the strict-schema unknown-key refusal that a typo'd + /// `kms_forse` must produce in either format. + #[test] + fn yaml_inherits_the_strict_schema_and_the_grammars() { + let e = ExhibitConfig::from_yaml("display_mode: auto\nkms_forse: none\n").unwrap_err(); + assert!(e.contains("kms_forse") && e.contains("unknown key"), "{e}"); + + let e = ExhibitConfig::from_yaml("display_mode: 3840x2160@29.97\n").unwrap_err(); + assert!(e.contains("invalid display_mode"), "{e}"); + + let e = ExhibitConfig::from_yaml("kms_force: none\n").unwrap_err(); + assert!(e.contains("display_mode"), "{e}"); + } + + // ---- the YAML subset ------------------------------------------------ + + /// yaml-rust2's loader errors on a duplicate key rather than last-wins. + /// That is the ONE rule the JSON side needed a hand-written serde visitor + /// for, so it is load-bearing that the YAML side gets it for free -- + /// pinned by test, because it is a property of the dependency and would + /// otherwise silently regress on a version bump. + #[test] + fn yaml_duplicate_key_refused() { + let e = + ExhibitConfig::from_yaml("display_mode: auto\ndisplay_mode: 3840x2160@30\n").unwrap_err(); + assert!(e.contains("duplicate"), "{e}"); + } + + /// A bare `true`/`false` is a boolean, refused with the quoting fix. + #[test] + fn yaml_bare_true_false_are_booleans_and_refused() { + for text in [ + "display_mode: true\n", + "display_mode: auto\nnote: FALSE\n", + ] { + let e = ExhibitConfig::from_yaml(text).unwrap_err(); + assert!(e.contains("BOOLEAN"), "{text:?}: {e}"); + assert!(e.contains("Quote the value"), "{text:?}: {e}"); + } + } + + /// The OTHER half, and the one that is easy to get wrong from memory: + /// yaml-rust2 0.11 resolves close to the YAML **1.2 core schema**, so the + /// "Norway problem" does NOT apply -- `no` stays the string `"no"` and is + /// refused by the kms_force GRAMMAR, not by the boolean branch. Asserted + /// on the message so that a future version adopting 1.1 resolution (which + /// would make `kms_force: no` mean `false`) fails here rather than + /// changing what a deployed config means. + #[test] + fn yaml_bare_no_stays_a_string_under_the_1_2_core_schema() { + let e = ExhibitConfig::from_yaml("display_mode: auto\nkms_force: no\n").unwrap_err(); + assert!( + e.contains("invalid kms_force \"no\""), + "expected the grammar refusal for the STRING \"no\", not a boolean one: {e}" + ); + assert!(!e.contains("BOOLEAN"), "{e}"); + // ...and stating it properly is the fix. + let ok = ExhibitConfig::from_yaml("display_mode: auto\nkms_force: none\n").unwrap(); + assert_eq!(ok.kms_force, "none"); + } + + /// Where yaml-rust2 DEPARTS from the 1.2 core schema, pinned so the docs + /// cannot quietly become wrong: core lists `null | Null | NULL | ~ | empty` + /// as null, but this library's `from_str` matches only `""`, `"~"` and + /// `"null"` — case-sensitively. So the capitalised spellings arrive as + /// ordinary strings. + /// + /// Harmless here (only the informational keys could carry such a value, + /// and reading it literally is the friendlier of the two options), but + /// "it implements the core schema" is a nearly-true sentence a future + /// reader would rely on, and this is the test that keeps it honest. + #[test] + fn yaml_null_resolution_is_case_sensitive_unlike_the_1_2_core_schema() { + for null_spelling in ["null", "~", ""] { + let e = ExhibitConfig::from_yaml(&format!("display_mode: auto\nnote: {null_spelling}\n")) + .unwrap_err(); + assert!(e.contains("null"), "{null_spelling:?}: {e}"); + } + for string_spelling in ["Null", "NULL"] { + let c = ExhibitConfig::from_yaml(&format!( + "display_mode: auto\nnote: {string_spelling}\n" + )) + .expect("core would call this null; yaml-rust2 does not"); + assert_eq!(c.note.as_deref(), Some(string_spelling)); + } + } + + /// Scalar resolution happens BEFORE this crate's subset check, so it stays + /// observable — and both formats must land in the same place. A bare + /// number resolves to an integer and is then refused for a string-only + /// key, identically to the JSON spelling of the same thing. + #[test] + fn a_bare_number_is_refused_the_same_way_in_both_formats() { + let y = ExhibitConfig::from_yaml("display_mode: auto\nvenue: 2026\n").unwrap_err(); + let j = ExhibitConfig::from_json(r#"{"display_mode":"auto","venue":2026}"#).unwrap_err(); + assert_eq!(y, j, "the two formats must refuse identically"); + assert!(y.contains("venue must be a string"), "{y}"); + // ...and quoting is the fix in both. + assert_eq!( + ExhibitConfig::from_yaml("display_mode: auto\nvenue: \"2026\"\n").unwrap(), + ExhibitConfig::from_json(r#"{"display_mode":"auto","venue":"2026"}"#).unwrap() + ); + } + + /// Anchors and aliases are refused, and this test exists because the + /// original implementation only *documented* that it refused them. + /// + /// `YamlLoader` resolves `*name` into a copy of the anchored node, so + /// `Yaml::Alias` never reaches the node-type check and the config below was + /// silently ACCEPTED -- with `venue` carrying "hello", a value the file + /// never assigns to it. Exactly the "means something other than it says" + /// failure the subset exists to prevent, hidden by the fact that the + /// refusal had been written down. + #[test] + fn yaml_anchors_and_aliases_are_refused_not_silently_expanded() { + // The case that used to pass: venue is never assigned in the text. + let e = + ExhibitConfig::from_yaml("display_mode: auto\nnote: &a hello\nvenue: *a\n").unwrap_err(); + assert!(e.contains("anchor") || e.contains("alias"), "{e}"); + + // An anchor with no alias is refused too -- it is the half that makes + // the other half possible, and accepting it would make the refusal + // depend on how far the operator got. + let e = ExhibitConfig::from_yaml("display_mode: auto\nnote: &unused hello\n").unwrap_err(); + assert!(e.contains("anchor"), "{e}"); + + // ...while the spelled-out equivalent is fine, which is the point: the + // refusal costs the operator one retyped value, not a capability. + let ok = + ExhibitConfig::from_yaml("display_mode: auto\nnote: hello\nvenue: hello\n").unwrap(); + assert_eq!(ok.note.as_deref(), Some("hello")); + assert_eq!(ok.venue.as_deref(), Some("hello")); + } + + /// A syntax error must still produce YamlLoader's own marked message, not + /// a vaguer one from the anchor pre-scan that now runs first. + #[test] + fn the_anchor_prescan_does_not_swallow_real_syntax_errors() { + let e = ExhibitConfig::from_yaml("display_mode: auto\nnote:\n\tx: 1\n").unwrap_err(); + assert!(e.contains("tab"), "expected the scanner's own diagnostic: {e}"); + } + + /// Error messages must read as English. `yaml_type_name` returns bare + /// nouns and each call site supplies its own article -- an earlier + /// revision baked "a " into the names and emitted "has a a nested mapping + /// value" at one of the three sites. + #[test] + fn type_names_do_not_double_their_article() { + for text in [ + "display_mode: auto\nnote:\n a: b\n", + "display_mode: auto\nnote:\n - a\n", + "- a\n- b\n", + "just a string\n", + ] { + let e = ExhibitConfig::from_yaml(text).unwrap_err(); + assert!(!e.contains(" a a "), "doubled article: {e}"); + assert!(!e.contains(" a an "), "doubled article: {e}"); + } + } + + #[test] + fn yaml_nesting_lists_and_null_are_refused() { + let nested = ExhibitConfig::from_yaml("display_mode: auto\nnote:\n a: b\n").unwrap_err(); + assert!(nested.contains("nested mapping"), "{nested}"); + let list = ExhibitConfig::from_yaml("display_mode: auto\nnote:\n - a\n").unwrap_err(); + assert!(list.contains("a list"), "{list}"); + let null = ExhibitConfig::from_yaml("display_mode: auto\nnote:\n").unwrap_err(); + assert!(null.contains("null"), "{null}"); + } + + #[test] + fn yaml_multiple_documents_refused_rather_than_taking_the_first() { + let e = ExhibitConfig::from_yaml( + "display_mode: auto\n---\ndisplay_mode: 3840x2160@30\n", + ) + .unwrap_err(); + assert!(e.contains("2 documents"), "{e}"); + } + + #[test] + fn yaml_empty_or_comment_only_says_so_rather_than_blaming_display_mode() { + for text in ["", " \n", "# just a comment\n"] { + let e = ExhibitConfig::from_yaml(text).unwrap_err(); + assert!(e.contains("no document"), "{text:?}: {e}"); + } + } + + #[test] + fn yaml_top_level_scalar_or_list_refused() { + let e = ExhibitConfig::from_yaml("just a string\n").unwrap_err(); + assert!(e.contains("top level"), "{e}"); + let e = ExhibitConfig::from_yaml("- a\n- b\n").unwrap_err(); + assert!(e.contains("top level") && e.contains("a list"), "{e}"); + } + + /// An integer VALUE parses (the shared mapper then refuses it for these + /// particular keys, in the same words the JSON path uses) -- but a + /// NEGATIVE one is out of range for the format itself. + #[test] + fn yaml_integers_follow_the_json_subset() { + let e = ExhibitConfig::from_yaml("display_mode: 30\n").unwrap_err(); + assert!(e.contains("display_mode must be a string"), "{e}"); + let e = ExhibitConfig::from_yaml("display_mode: auto\nnote: -1\n").unwrap_err(); + assert!(e.contains("negative"), "{e}"); + } + + // ---- default-config discovery --------------------------------------- + + #[test] + fn one_default_config_is_picked_none_is_none() { + assert_eq!(pick_default_config(&[]).unwrap(), None); + assert_eq!( + pick_default_config(&["/etc/dex/exhibit.yaml"]).unwrap(), + Some("/etc/dex/exhibit.yaml") + ); + assert_eq!( + pick_default_config(&["/etc/dex/exhibit.json"]).unwrap(), + Some("/etc/dex/exhibit.json") + ); + } + + /// Two configs at once refuses, naming both -- never "YAML wins". + /// Precedence here would let an operator edit one file all afternoon + /// while the player reads the other. + #[test] + fn two_default_configs_refuse_naming_both() { + let e = pick_default_config(&DEFAULT_EXHIBIT_CONFIG_PATHS).unwrap_err(); + assert!(e.contains("exhibit.yaml") && e.contains("exhibit.json"), "{e}"); + assert!(e.contains("--exhibit-config"), "{e}"); + // The likely cause, named: writing exhibit.yaml leaves the shipped + // conffile beside it, so the operator did the documented thing and + // still got refused. Without this the message reads like a bug. + assert!( + e.contains("sudo rm /etc/dex/exhibit.json"), + "must name the leftover conffile as the fix: {e}" + ); + } + + /// ...and that hint is CONDITIONAL, not glued on: a collision between two + /// non-shipped files must not tell the operator to remove a package file + /// that has nothing to do with it. + #[test] + fn the_leftover_conffile_hint_only_appears_when_a_json_is_involved() { + let e = pick_default_config(&["/srv/a.yaml", "/srv/b.yml"]).unwrap_err(); + assert!(!e.contains("sudo rm"), "{e}"); + } + + // ---- resolve_asset: the whole decision table ------------------------- + + #[test] + fn deploy_path_takes_the_asset_from_the_config() { + let c = cfg_with_asset("/opt/dex/spring.265"); + let r = resolve_asset(Some(&c), None, false).unwrap(); + assert_eq!(r.path, "/opt/dex/spring.265"); + assert_eq!(r.source, AssetSource::Config); + } + + #[test] + fn an_agreeing_cli_path_cross_checks_and_the_config_still_binds() { + let c = cfg_with_asset("/opt/dex/spring.265"); + let r = resolve_asset(Some(&c), Some("/opt/dex/spring.265"), false).unwrap(); + assert_eq!(r.source, AssetSource::Config); + } + + #[test] + fn a_contradicting_cli_path_is_refused_naming_both() { + let c = cfg_with_asset("/opt/dex/spring.265"); + let e = resolve_asset(Some(&c), Some("/opt/dex/autumn.265"), false).unwrap_err(); + assert!(e.contains("spring.265") && e.contains("autumn.265"), "{e}"); + } + + /// A config with no `asset` key still accepts a hand-given path -- the + /// one-off case (try another file on a deployed device without editing + /// /etc), reported as CLI-sourced so the journal cannot be misread. + #[test] + fn a_cli_path_works_when_the_config_names_no_asset() { + let c = cfg("auto", "none"); + let r = resolve_asset(Some(&c), Some("/opt/dex/try.265"), false).unwrap(); + assert_eq!(r.source, AssetSource::Cli); + } + + /// THE fail-closed row: nothing names an asset, so nothing is guessed -- + /// specifically NOT /opt/dex/loop.265, which would silently play last + /// season's artwork for someone who mistyped the key. + #[test] + fn no_asset_anywhere_refuses_rather_than_defaulting_to_loop_265() { + let c = cfg("auto", "none"); + let e = resolve_asset(Some(&c), None, false).unwrap_err(); + assert!(e.contains("no asset"), "{e}"); + // ...and it names the exact line to add, in both formats. + assert!(e.contains("asset: /opt/dex/loop.265"), "{e}"); + assert!(e.contains(r#""asset": "/opt/dex/loop.265""#), "{e}"); + // Also with NO config at all (that case refuses earlier, at + // resolve_display -- but this function must not invent a path either). + assert!(resolve_asset(None, None, false).is_err()); + } + + #[test] + fn bench_takes_the_cli_path_and_ignores_the_config() { + let c = cfg_with_asset("/opt/dex/spring.265"); + let r = resolve_asset(Some(&c), Some("/tmp/bench.265"), true).unwrap(); + assert_eq!(r.path, "/tmp/bench.265"); + assert_eq!(r.source, AssetSource::Cli); + } + + #[test] + fn bench_without_a_path_refuses_rather_than_falling_back_to_the_config() { + let c = cfg_with_asset("/opt/dex/spring.265"); + let e = resolve_asset(Some(&c), None, true).unwrap_err(); + assert!(e.contains("command line"), "{e}"); + } + + // ---- the asset grammar ---------------------------------------------- + + #[test] + fn asset_must_be_an_absolute_path() { + assert!(is_valid_asset("/opt/dex/loop.265")); + assert!(is_valid_asset("/srv/art/Karte–Süd.265")); // non-ASCII is fine + // A relative path would resolve against the service's working + // directory (`/`), i.e. somewhere the operator did not type. + assert!(!is_valid_asset("loop.265")); + assert!(!is_valid_asset("./loop.265")); + assert!(!is_valid_asset("")); + assert!(!is_valid_asset("/opt/dex/")); // a directory, not a file + } + + /// The asset path is printed into the journal at every start, and the + /// journal is the only diagnostic channel a gallery device has -- a + /// newline in it would forge a second log line. + #[test] + fn asset_with_a_control_character_is_refused() { + assert!(!is_valid_asset("/opt/dex/loop.265\ndexd: all fine here")); + assert!(!is_valid_asset("/opt/dex/loop\t.265")); + } + + #[test] + fn asset_parses_from_both_formats_and_is_grammar_checked() { + let j = ExhibitConfig::from_json( + r#"{"asset":"/opt/dex/spring.265","display_mode":"auto"}"#, + ) + .unwrap(); + let y = ExhibitConfig::from_yaml("asset: /opt/dex/spring.265\ndisplay_mode: auto\n").unwrap(); + assert_eq!(j, y); + assert_eq!(j.asset.as_deref(), Some("/opt/dex/spring.265")); + + let e = ExhibitConfig::from_yaml("asset: loop.265\ndisplay_mode: auto\n").unwrap_err(); + assert!(e.contains("invalid asset"), "{e}"); + } + + // ---- load_exhibit_config, against a real temp directory -------------- + // + // `defaults` is injectable, so the discovery policy both binaries share is + // testable here rather than only through the CLI on a machine that happens + // to have /etc/dex. + + /// A unique temp path per call site, so tests never collide with each + /// other or with a previous run's leftovers. + fn tmp(name: &str) -> String { + let dir = std::env::temp_dir().join(format!("dex-exhibit-test-{}", std::process::id())); + std::fs::create_dir_all(&dir).unwrap(); + dir.join(name).to_string_lossy().into_owned() + } + + #[test] + fn load_finds_the_yaml_default_and_parses_it_as_yaml() { + let y = tmp("load-yaml.yaml"); + std::fs::write(&y, "display_mode: 3840x2160@30 # venue panel\n").unwrap(); + let (cfg, path) = load_exhibit_config(None, &[&y, &tmp("load-yaml-absent.json")]) + .unwrap() + .unwrap(); + assert_eq!(cfg.display_mode, "3840x2160@30"); + assert_eq!(path, y); + let _ = std::fs::remove_file(&y); + } + + #[test] + fn load_returns_none_when_no_default_exists() { + let found = load_exhibit_config( + None, + &[&tmp("load-absent.yaml"), &tmp("load-absent.json")], + ) + .unwrap(); + assert!(found.is_none(), "{found:?}"); + } + + /// The whole reason both binaries call this: with both names present it + /// refuses instead of picking, so `dex-exhibit-apply` can never reconcile + /// the cmdline against a file `dexd` will not read. + #[test] + fn load_refuses_when_both_defaults_exist() { + let (j, y) = (tmp("load-both.json"), tmp("load-both.yaml")); + std::fs::write(&j, r#"{"display_mode":"auto"}"#).unwrap(); + std::fs::write(&y, "display_mode: auto\n").unwrap(); + let e = load_exhibit_config(None, &[&y, &j]).unwrap_err(); + assert!(e.contains("2 exhibit configs"), "{e}"); + let _ = std::fs::remove_file(&j); + let _ = std::fs::remove_file(&y); + } + + /// An explicitly named missing file must NOT borrow the "no exhibit + /// config, create the default" message -- the operator named a different + /// path, and telling them to create /etc/dex/exhibit.json is wrong advice. + #[test] + fn load_names_the_explicit_path_when_it_is_missing() { + let p = tmp("load-explicitly-absent.json"); + let _ = std::fs::remove_file(&p); + let e = load_exhibit_config(Some(&p), &DEFAULT_EXHIBIT_CONFIG_PATHS).unwrap_err(); + assert!(e.contains(&p), "{e}"); + assert!(!e.contains("Create /etc/dex"), "{e}"); + } + + /// A config whose name promises neither format refuses at the dispatch, + /// before any parse is attempted -- even though its CONTENTS would parse + /// perfectly well as either. + #[test] + fn load_refuses_an_unrecognised_extension_even_with_valid_contents() { + let p = tmp("load-nameless.conf"); + std::fs::write(&p, r#"{"display_mode":"auto"}"#).unwrap(); + let e = load_exhibit_config(Some(&p), &DEFAULT_EXHIBIT_CONFIG_PATHS).unwrap_err(); + assert!(e.contains("cannot tell the format"), "{e}"); + let _ = std::fs::remove_file(&p); + } +} diff --git a/packages/dexd/src/ffi_consts.rs b/packages/dexd/src/ffi_consts.rs new file mode 100644 index 0000000..70b7712 --- /dev/null +++ b/packages/dexd/src/ffi_consts.rs @@ -0,0 +1,99 @@ +//! libmpv ABI constants, transcribed from mpv/client.h (mpv v0.40.0) and — +//! more importantly — verified against the LIVE library at test time by +//! tests/ffi_constants.rs via mpv_event_name()/mpv_error_string(). +//! +//! Why the paranoia: an earlier revision transcribed MPV_EVENT_LOG_MESSAGE as +//! 6. It is 2; 6 is MPV_EVENT_START_FILE. The event handler then cast a +//! start-file payload to a log-message struct and dereferenced garbage +//! pointers — a segfault on the first frame. A transcription can silently +//! rot; the live library cannot. These are NOT sequential-by-category. + +use std::ffi::c_int; + +pub const MPV_EVENT_NONE: c_int = 0; +pub const MPV_EVENT_SHUTDOWN: c_int = 1; +pub const MPV_EVENT_LOG_MESSAGE: c_int = 2; +/// Delivered in reply to `mpv_command_async`. F1's tier-0 in-place recovery +/// issues its `loadfile` through the async command API specifically so the +/// event thread can never block on it (see src/health.rs's module doc); +/// this event is used only to log whether the queued command was accepted, +/// nothing gates on it. +pub const MPV_EVENT_COMMAND_REPLY: c_int = 5; +pub const MPV_EVENT_START_FILE: c_int = 6; +pub const MPV_EVENT_END_FILE: c_int = 7; +/// Delivered when a property registered via `mpv_observe_property` changes +/// value (or, for some properties, at mpv's own internal polling cadence). +/// F1's health check subscribes to `time-pos` through this rather than +/// polling `mpv_get_property_string` synchronously -- see src/health.rs's +/// module doc for why that distinction matters. +pub const MPV_EVENT_PROPERTY_CHANGE: c_int = 22; +/// "At least one event had to be dropped." Delivered once the internal +/// 1000-slot event ring chokes and starts silently discarding events -- +/// including, potentially, an END_FILE. Treated as fatal: see the event +/// loop in main.rs. +pub const MPV_EVENT_QUEUE_OVERFLOW: c_int = 24; + +/// The documented "not supported" sentinel for stream callbacks. `-1` is +/// MPV_ERROR_EVENT_QUEUE_FULL, which happens to work only because mpv 0.40 +/// tests the sign rather than the value. +pub const MPV_ERROR_UNSUPPORTED: c_int = -18; + +/// `mpv_format` tag for a plain floating-point property value (client.h's +/// `mpv_format` enum: NONE=0, STRING=1, OSD_STRING=2, FLAG=3, INT64=4, +/// DOUBLE=5, NODE=6, ...). Unlike the event ids above, mpv exposes no +/// `mpv_format_name()` to verify this against the live library the same +/// way, so main.rs checks this tag on every `MPV_EVENT_PROPERTY_CHANGE` +/// payload before ever reading it as an `f64`, rather than trusting the +/// transcription blindly -- the same class of bug that made +/// MPV_EVENT_LOG_MESSAGE's mistranscription a segfault instead of a caught +/// error. +pub const MPV_FORMAT_DOUBLE: c_int = 5; + +/// `mpv_format` tag for a 64-bit integer property value (client.h's +/// `mpv_format` enum: NONE=0, STRING=1, OSD_STRING=2, FLAG=3, INT64=4, +/// DOUBLE=5, ...). F9's two observed drop counters use it. Their NATIVE +/// type is `int` (`m_property_int_ro`, player/command.c:763-781), but the +/// client API converts on the way out -- getproperty_fn routes INT64 +/// through M_PROPERTY_GET_NODE -> conv_node_to_format (player/client.c +/// :1417-1442), and m_property_do synthesizes GET_NODE from GET for the +/// int type (options/m_property.c:171-188) -- so INT64 is a legal request +/// for them and the payload is an i64. +/// +/// Like MPV_FORMAT_DOUBLE this cannot be checked against the live library +/// (mpv exposes no mpv_format_name()), so main.rs checks the tag on every +/// payload before dereferencing it. A wrong value here fails in the SAFE +/// direction, exactly like MPV_END_FILE_REASON_STOP: the tag never +/// matches, the counters read "n/a" forever, and nothing is misread. +pub const MPV_FORMAT_INT64: c_int = 4; + +/// `mpv_format` tag mpv substitutes for a property's real format whenever +/// the property is unavailable or a getter errored -- client.h: "Warning: if +/// a property is unavailable or retrieving it caused an error, +/// MPV_FORMAT_NONE will be set in mpv_event_property, even if the format +/// parameter was set to a different value. In this case, the +/// mpv_event_property.data field is invalid." F9's two observed drop +/// counters see this instead of MPV_FORMAT_INT64 whenever no vo_chain +/// exists (`M_PROPERTY_UNAVAILABLE`, player/command.c:763-781) -- at startup, +/// and during a tier-0 recovery's teardown/rebuild. main.rs's +/// property-change handler acts on it (see `ObservedCounter::mark_unavailable`) +/// rather than silently dropping it, so a reset that mpv's event coalescing +/// hides cannot silently under-count the heartbeat's totals. +pub const MPV_FORMAT_NONE: c_int = 0; + +/// `mpv_end_file_reason` tag for MPV_EVENT_END_FILE's `reason` field: +/// "Playback was stopped by an external action" (client.h). Bench-confirmed +/// live on the Pi's mpv 0.40.0 (three independent adversarial reviews, +/// 2026-08-15, IPC/ctypes probes) that `loadfile replace` delivers +/// exactly this reason for the file being replaced. F1's in-place recovery +/// issues that exact command, so main.rs must recognize and absorb this one +/// reason for precisely the one command it issued itself -- every other +/// reason, and every STOP with no recovery in flight, stays fatal. See +/// PLAN.md's F1 addendum for why treating ANY end-file as fatal made +/// tier-0 recovery unreachable before this constant existed. +/// +/// Unlike the event ids above, mpv exposes no runtime name lookup for +/// END_FILE reasons, so this is transcription-only -- but a wrong value +/// here fails in the SAFE direction: main.rs's absorption check simply +/// never matches, so an unmatched STOP falls through to the pre-existing +/// fatal path exactly as it did before this feature existed. +pub const MPV_END_FILE_REASON_STOP: c_int = 2; diff --git a/packages/dexd/src/health.rs b/packages/dexd/src/health.rs new file mode 100644 index 0000000..0c28f46 --- /dev/null +++ b/packages/dexd/src/health.rs @@ -0,0 +1,604 @@ +//! F1 — tier-0 self-healing: the escalation policy for the periodic health +//! check, extracted pure so the policy that decides "recover in place" vs. +//! "give up and let the supervisor restart" is testable without libmpv, a +//! display, or even the crate's FFI half. `main.rs` is a thin driver over +//! this: it samples a playback-position property on a fixed cadence and +//! feeds the sample to [`HealthMonitor::tick`]; every other decision is made +//! here. +//! +//! # Why this exists (PLAN.md F1) +//! +//! The event loop already treats a fatal mpv event (END_FILE, +//! QUEUE_OVERFLOW) as tier 1: exit non-zero, let `Restart=always` recover. +//! That covers a display absent at BOOT. It does not cover an INTERMITTENT +//! loss mid-show — the projector's HDMI blinking, a sink waking up late — +//! where mpv itself never emits a fatal event at all: it just silently stops +//! making progress while the process stays alive and every existing signal +//! (heartbeat, supervisor) reads green. Killing a player that LOOKS healthy +//! over one glitch is heavy-handed and loses seconds of picture in front of +//! the public; never checking at all is bug #1's exact shape (alive, green, +//! screen black) with a different trigger. Tier 0 is the middle path: try a +//! cheap in-place repair first, and escalate only if that keeps failing. +//! +//! # The escalation policy, precisely +//! +//! * A "tick" happens on a fixed wall-clock cadence (main.rs: ~10 s, off the +//! decode path — see that module's comment on the event-wait timeout for +//! why the cadence has to be enforced by a TIMEOUT, not just by events). +//! * Each tick reports the LATEST known playback position (e.g. `time-pos`). +//! `None` means no update has arrived since the monitor was created (or +//! since it was last reset by a recovery) — treated identically to "the +//! value is unchanged", because a wedged core stops producing property +//! updates entirely; the absence of an update IS the stall symptom, not a +//! distinct case needing its own handling. +//! * The position must strictly increase between two ticks to count as +//! progress. Two CONSECUTIVE non-advancing ticks (not one) are required +//! before acting, to absorb ordinary jitter around a check boundary — see +//! PLAN.md's F1 text ("has not advanced across two consecutive checks"). +//! * Once that bar is met, the caller is told to attempt an in-place +//! recovery (re-issue `loadfile ... replace`, which opens a brand-new +//! `loop://` stream, restarting demux+decode and forcing a `vo_reconfig`). +//! In mpv v0.40, a playlist replace tears down and rebuilds the demuxer +//! and decoder chain but does NOT tear down the video output itself +//! (`uninit_video_out` runs only on process termination) -- so this +//! plausibly repairs a decode-side wedge, but a fault in the DRM/GPU +//! context surviving the replace may still need a full process restart +//! (tier 1) to clear. Unverified against a physical HDMI-loss bench test +//! as of 2026-08-15; see PLAN.md's F1 addendum. +//! * Recovery attempts are drawn from a budget fixed at construction and +//! NEVER replenished for the life of the process — see "why the budget +//! never resets" below. Once exhausted, the next qualifying stall +//! escalates instead of attempting another recovery. +//! +//! # Why the budget never resets on a temporary recovery +//! +//! The obvious alternative — give the counter back after some period of +//! sustained health following a recovery — reopens exactly the hole the cap +//! exists to close. A fault that FLAPS (heals for a while, stalls again, +//! repeat) would keep resetting the counter before it ever reached the cap, +//! producing a total number of in-place retries that is unbounded across the +//! process's lifetime even though each individual episode looks bounded. +//! That is "an unbounded in-place retry loop... wearing a different hat" — +//! exactly the failure mode the task brief warns against, just spread out in +//! time instead of packed into one burst. A cumulative, never-replenished +//! budget is the only shape that bounds the WORST case, not just the common +//! one. And leaning on tier 1 sooner than strictly necessary is cheap and +//! safe here: `Restart=always` with `StartLimitIntervalSec=0` +//! (`deploy/dexd.service`) never gives up either, so the process comes +//! back regardless of which tier does the healing — tier 0 running out of +//! budget means tier 1 takes over, not that the show stops. +//! +//! # Why the sample that drives this arrives asynchronously (F9's concern) +//! +//! A prior review (F9, SUSPECTED, recorded in PLAN.md) flagged that a +//! synchronous `mpv_get_property_string` call made from a supervisor thread +//! is itself a core-wedge risk: if the mpv core is ever stuck (e.g. the VO +//! thread blocked in a DRM ioctl against a dying projector, holding whatever +//! the core needs), a synchronous property read could block that thread +//! forever — reproducing bug #1's exact shape (alive, supervisor green, +//! screen black) through a new door instead of closing it. F1 takes that +//! seriously rather than reproducing it: +//! +//! * `main.rs` never polls `mpv_get_property_string` (or any synchronous +//! property read) for this feature. It registers exactly ONE +//! `mpv_observe_property` call at startup — documented in client.h as +//! non-blocking, queuing a subscription and returning immediately — and +//! thereafter only ever learns the position from `MPV_EVENT_PROPERTY_CHANGE` +//! events delivered through the SAME `mpv_wait_event` loop that already +//! proves, by the very fact that this program correctly detects END_FILE +//! and QUEUE_OVERFLOW today, that it cannot block indefinitely (the wait +//! call takes a caller-supplied timeout). +//! * If the core wedges, no new property-change events arrive at all. That +//! surfaces here as `tick(None)` (or an unchanging value) — exactly the +//! stall signal this module already exists to notice, produced by the same +//! mechanism that already detects every other fault, rather than by a new +//! blocking call that could itself hang. +//! * Recovery is issued the same way: `mpv_command_async`, not the +//! synchronous `mpv_command` this program already uses once at startup +//! (safe there — nothing has had a chance to wedge before the first +//! frame). Using the async variant for recovery means the one new +//! synchronous-shaped risk this feature could have introduced — blocking +//! on `loadfile` from the event thread while trying to fix a wedged core — +//! does not exist either. +//! +//! In short: every new mpv-facing call this feature adds is either +//! documented non-blocking (`mpv_observe_property`, `mpv_command_async`) or +//! not a new call at all (`mpv_wait_event`, already relied on). No new +//! synchronous call is introduced, so F1 does not add a new way to hang — +//! the exact requirement the task brief states. + +/// What the caller should do after feeding one health-check sample to +/// [`HealthMonitor::tick`]. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum HealthAction { + /// Progress since the last tick, or too early to judge yet (the very + /// first sample, or only one non-advancing tick so far). Nothing to do. + Healthy, + /// No progress across >= 2 consecutive ticks, and the recovery budget is + /// not yet exhausted. The caller should force VO reconfiguration + /// (re-issue `loadfile ... replace`, asynchronously) and keep running. + AttemptRecovery { + /// This attempt's number (1-based) against `max`. + attempt: u32, + max: u32, + }, + /// No progress across >= 2 consecutive ticks, and the recovery budget is + /// exhausted. The caller should escalate to tier 1: exit non-zero so the + /// supervisor restarts the whole process. + Escalate, +} + +/// The tier-0 escalation state machine. See the module doc for the policy +/// and the reasoning behind it; this struct only holds the state that +/// policy needs. +pub struct HealthMonitor { + last_position: Option, + consecutive_stalls: u32, + recovery_attempts_used: u32, + max_recovery_attempts: u32, +} + +impl HealthMonitor { + /// `max_recovery_attempts` is the CUMULATIVE, process-lifetime budget of + /// in-place recoveries — see the module doc for why it never refills. + pub fn new(max_recovery_attempts: u32) -> Self { + Self { + last_position: None, + consecutive_stalls: 0, + recovery_attempts_used: 0, + max_recovery_attempts, + } + } + + /// Feed one health-check tick. `position` is the latest known playback + /// position (e.g. `time-pos` seconds), or `None` if no update has been + /// observed since the monitor was created, or since it was last reset by + /// a recovery attempt — see the module doc for why that is treated the + /// same as "value unchanged". + pub fn tick(&mut self, position: Option) -> HealthAction { + let advanced = match (position, self.last_position) { + (Some(p), Some(prev)) => p > prev, + // First sample since start (or since the last recovery reset + // the baseline below): nothing to compare against yet. Treating + // this as progress, not a stall, avoids counting mpv's own + // startup/reload latency -- a few hundred ms to a few seconds + // for a 4K decode to begin -- as a fault. A REAL startup hang + // (no sample ever arrives) is still caught: two consecutive + // `None` ticks below is exactly that case. + (Some(_), None) => true, + (None, _) => false, + }; + + if let Some(p) = position { + self.last_position = Some(p); + } + + if advanced { + self.consecutive_stalls = 0; + return HealthAction::Healthy; + } + + self.consecutive_stalls += 1; + if self.consecutive_stalls < 2 { + return HealthAction::Healthy; + } + + // Two consecutive non-advancing ticks: a qualifying stall episode. + // Reset the LOCAL jitter counter so the next attempt (if any) or the + // eventual escalation gets its own clean 2-tick window rather than + // re-triggering on the very next tick -- but do NOT reset the + // budget itself; see the module doc. + self.consecutive_stalls = 0; + self.issue_recovery_or_escalate() + } + + /// T7 -- bench-only live-fire probe (PLAN.md): force EXACTLY the + /// `AttemptRecovery`/`Escalate` decision an organic stall would produce, + /// without waiting for one. Draws from the SAME cumulative budget and + /// performs the SAME position-baseline reset as `tick` -- see "why the + /// budget never resets" above -- so a forced probe exercises the + /// IDENTICAL mpv-facing code path (main.rs's `loadfile ... replace`, + /// absorbing the resulting `END_FILE(reason=stop)`) that a real stall + /// would, rather than a look-alike that could pass while the real path + /// stays broken. That identity is the whole point: C1 (every recovery + /// attempt killing the process on its own first step) shipped and + /// reached the bench without ever having been exercised against a live + /// mpv, because nothing -- test or otherwise -- had ever driven this + /// path for real. See main.rs's `--force-recovery-after-secs` for what + /// decides WHEN to call this. + /// + /// Not a stall: whatever jitter `consecutive_stalls` was accumulating + /// before this call is stale the moment a recovery actually issues (the + /// same reset `tick` performs on a qualifying stall), so it is cleared + /// here too rather than left to bleed into the next organic judgment. + pub fn force_recovery(&mut self) -> HealthAction { + self.consecutive_stalls = 0; + self.issue_recovery_or_escalate() + } + + /// The budget check + attempt/escalate decision shared by `tick`'s + /// qualifying-stall arm and `force_recovery` -- the only two places + /// allowed to spend the recovery budget. Callers are responsible for + /// whatever precedes "a recovery decision is due now" (stall counting + /// for `tick`, nothing for `force_recovery`); this is only the part + /// that must stay identical between them. + fn issue_recovery_or_escalate(&mut self) -> HealthAction { + if self.recovery_attempts_used >= self.max_recovery_attempts { + return HealthAction::Escalate; + } + self.recovery_attempts_used += 1; + // A successful in-place recovery re-opens the stream from byte 0 + // (main.rs issues `loadfile ... replace` on the SAME loop:// URL, + // which creates a brand-new stream via open_fn), so it legitimately + // restarts mpv's own position counter near zero. Forget the + // pre-recovery baseline so the very next real sample -- however + // small -- reads as progress rather than "still less than the old + // high-water mark", which would otherwise misread a recovery that + // WORKED as a continuing stall. + self.last_position = None; + HealthAction::AttemptRecovery { + attempt: self.recovery_attempts_used, + max: self.max_recovery_attempts, + } + } +} + +/// T7 -- bench-only live-fire probe (PLAN.md): decides WHEN to fire a single +/// forced tier-0 recovery, entirely independent of whether anything has +/// actually stalled. Pure so the "fires exactly once, at or after N seconds +/// of uptime, never before, never twice" contract is testable without mpv. +/// +/// Fires ONCE, not repeatedly: T7 exists to prove the recovery mechanism +/// survives contact with a live mpv at all (the C1 regression), not to run +/// an ongoing chaos-monkey campaign against it -- a single forced episode, +/// same as the reviewer's manual probe that caught C1, is the minimal thing +/// that closes the gap PLAN.md describes. See `main.rs`'s +/// `--force-recovery-after-secs` for how a run arms this, and +/// `HealthMonitor::force_recovery` for what firing does once armed. +pub struct ForceRecoveryTrigger { + after_secs: u64, + fired: bool, +} + +impl ForceRecoveryTrigger { + pub fn new(after_secs: u64) -> Self { + Self { + after_secs, + fired: false, + } + } + + /// Feed the current uptime. Returns `true` exactly once -- the first + /// call where `uptime_secs >= after_secs` -- and `false` on every call + /// before or after that. + pub fn should_fire(&mut self, uptime_secs: u64) -> bool { + if self.fired || uptime_secs < self.after_secs { + return false; + } + self.fired = true; + true + } +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn first_ever_tick_is_healthy_even_with_no_prior_position() { + let mut m = HealthMonitor::new(3); + assert_eq!(m.tick(Some(0.0)), HealthAction::Healthy); + } + + #[test] + fn steady_advancement_never_acts() { + let mut m = HealthMonitor::new(3); + let mut pos = 0.0; + for _ in 0..50 { + pos += 0.033; // roughly one 30fps frame's worth per tick + assert_eq!(m.tick(Some(pos)), HealthAction::Healthy); + } + } + + #[test] + fn a_single_non_advancing_tick_is_not_enough_to_act() { + // PLAN.md: "if it has NOT ADVANCED ACROSS TWO CONSECUTIVE checks" -- + // one stalled tick must not trigger anything, or ordinary jitter + // around a check boundary would cause spurious recoveries. + let mut m = HealthMonitor::new(3); + assert_eq!(m.tick(Some(1.0)), HealthAction::Healthy); + assert_eq!(m.tick(Some(1.0)), HealthAction::Healthy); // stall #1 + assert_eq!(m.tick(Some(2.0)), HealthAction::Healthy); // recovers before 2 + } + + #[test] + fn two_consecutive_stalls_trigger_the_first_recovery_attempt() { + let mut m = HealthMonitor::new(3); + assert_eq!(m.tick(Some(1.0)), HealthAction::Healthy); + assert_eq!(m.tick(Some(1.0)), HealthAction::Healthy); // stall #1 + assert_eq!( + m.tick(Some(1.0)), // stall #2 -- qualifies + HealthAction::AttemptRecovery { attempt: 1, max: 3 } + ); + } + + #[test] + fn none_samples_count_as_stalls_exactly_like_an_unchanged_value() { + // No property-change event arriving at all is the same symptom as + // the value not changing -- see the module doc's F9 discussion. + // Unlike a genuine `Some` value, `None` gets NO "first sample" grace + // period: the forgiveness for `Some` exists to absorb ordinary + // startup LATENCY (a value arrives a moment late), but a string of + // `None`s is exactly the real-startup-hang case the module doc + // calls out ("a value that never arrives at all") -- so it must not + // be forgiven for an extra tick the way a merely-late value is. + let mut m = HealthMonitor::new(3); + assert_eq!(m.tick(None), HealthAction::Healthy); // stall #1 + assert_eq!( + m.tick(None), // stall #2 -- qualifies, no grace period for None + HealthAction::AttemptRecovery { attempt: 1, max: 3 } + ); + } + + #[test] + fn recovery_gets_a_fresh_two_tick_window_before_being_judged_again() { + let mut m = HealthMonitor::new(3); + m.tick(Some(1.0)); + m.tick(Some(1.0)); // stall #1 + assert_eq!( + m.tick(Some(1.0)), // stall #2: attempt 1 + HealthAction::AttemptRecovery { attempt: 1, max: 3 } + ); + // Immediately after an attempt, ONE more sample must NOT + // re-trigger -- the attempt needs its own 2-tick window, otherwise + // a recovery that has not even had time to take effect gets judged + // as having already failed. + assert_eq!(m.tick(Some(0.05)), HealthAction::Healthy); + } + + #[test] + fn recovery_resets_the_position_baseline_so_a_reload_restart_is_not_misread_as_still_stalled( + ) { + // loadfile replace opens a BRAND NEW loop:// stream, so a real + // recovery legitimately restarts mpv's own position counter near + // zero. Without resetting the monitor's baseline too, the very next + // real sample (small) would compare as "less than" the + // pre-recovery high-water mark and register as ANOTHER stall -- + // punishing the exact recovery that worked. + let mut m = HealthMonitor::new(3); + m.tick(Some(500.0)); + m.tick(Some(500.0)); // stall #1 + assert_eq!( + m.tick(Some(500.0)), // stall #2: recovery issued, baseline dropped + HealthAction::AttemptRecovery { attempt: 1, max: 3 } + ); + // The reload starts over near zero -- this must read as progress, + // not as "0.1 < 500, another stall". + assert_eq!(m.tick(Some(0.1)), HealthAction::Healthy); + assert_eq!(m.tick(Some(0.2)), HealthAction::Healthy); + } + + #[test] + fn a_healed_stream_stays_healthy_and_does_not_consume_further_budget() { + let mut m = HealthMonitor::new(3); + m.tick(Some(1.0)); + m.tick(Some(1.0)); + assert_eq!( + m.tick(Some(1.0)), + HealthAction::AttemptRecovery { attempt: 1, max: 3 } + ); + // The stream advances again (recovery worked): healthy indefinitely. + let mut pos = 0.0; + for _ in 0..20 { + pos += 0.1; + assert_eq!(m.tick(Some(pos)), HealthAction::Healthy); + } + } + + #[test] + fn budget_is_cumulative_and_never_refills_across_separate_episodes() { + // The behaviour the module doc's "why the budget never resets" + // section defends: a fault that heals in between episodes must NOT + // get its budget back, or a flapping fault produces an unbounded + // total number of recovery attempts spread across many separate + // episodes -- the same failure the cap exists to prevent, just + // spread out in time instead of packed into one burst. + let mut m = HealthMonitor::new(2); + + // Episode 1: stalls, recovers (attempt 1/2). + m.tick(Some(1.0)); + m.tick(Some(1.0)); + assert_eq!( + m.tick(Some(1.0)), + HealthAction::AttemptRecovery { attempt: 1, max: 2 } + ); + // Fully healthy for a long stretch afterwards. + let mut pos = 0.0; + for _ in 0..100 { + pos += 0.1; + assert_eq!(m.tick(Some(pos)), HealthAction::Healthy); + } + + // Episode 2: stalls again, recovers (attempt 2/2 -- budget exhausted). + let stalled_at = pos; + m.tick(Some(stalled_at)); + assert_eq!( + m.tick(Some(stalled_at)), + HealthAction::AttemptRecovery { attempt: 2, max: 2 } + ); + // Healthy again for a long stretch. + pos = 0.0; + for _ in 0..100 { + pos += 0.1; + assert_eq!(m.tick(Some(pos)), HealthAction::Healthy); + } + + // Episode 3: budget is gone -- this one must escalate, not attempt + // a third recovery, even though the stream has been perfectly + // healthy for hundreds of ticks since the last stall. + let stalled_at = pos; + m.tick(Some(stalled_at)); + assert_eq!(m.tick(Some(stalled_at)), HealthAction::Escalate); + } + + #[test] + fn exhausting_the_budget_escalates_every_qualifying_stall_after() { + let mut m = HealthMonitor::new(1); + m.tick(Some(1.0)); + m.tick(Some(1.0)); + assert_eq!( + m.tick(Some(1.0)), + HealthAction::AttemptRecovery { attempt: 1, max: 1 } + ); + // Budget is now exhausted (1/1 used). The next qualifying stall + // must escalate, not attempt a second recovery. + m.tick(Some(0.1)); + m.tick(Some(0.1)); + assert_eq!(m.tick(Some(0.1)), HealthAction::Escalate); + } + + #[test] + fn zero_budget_escalates_on_the_first_qualifying_stall() { + // Edge case sanity: a monitor configured with NO recovery budget at + // all still requires 2 consecutive stalls (the jitter guard is + // unconditional) but then escalates immediately rather than ever + // attempting an in-place recovery. + let mut m = HealthMonitor::new(0); + m.tick(Some(1.0)); + m.tick(Some(1.0)); + assert_eq!(m.tick(Some(1.0)), HealthAction::Escalate); + } + + #[test] + fn position_going_backwards_counts_as_a_stall_not_progress() { + // Should never happen for a monotonic, clock-driven time-pos, but + // the policy must not treat it as advancement if it ever does (e.g. + // a property glitch during a VO reconfigure) -- strictly `>`, not + // `!=`. + let mut m = HealthMonitor::new(3); + m.tick(Some(5.0)); + m.tick(Some(4.0)); // stall #1 (went backwards) + assert_eq!( + m.tick(Some(4.0)), // stall #2 + HealthAction::AttemptRecovery { attempt: 1, max: 3 } + ); + } + + // ---- T7: force_recovery / ForceRecoveryTrigger -------------------- + + #[test] + fn force_recovery_attempts_immediately_with_no_stall_observed() { + // The whole point of T7: unlike `tick`, this needs no stall history + // at all -- perfectly healthy, freshly-advancing playback still + // gets forced into a recovery attempt. + let mut m = HealthMonitor::new(3); + m.tick(Some(1.0)); + m.tick(Some(2.0)); + m.tick(Some(3.0)); // steady progress, nowhere near a stall + assert_eq!( + m.force_recovery(), + HealthAction::AttemptRecovery { attempt: 1, max: 3 } + ); + } + + #[test] + fn force_recovery_shares_the_same_cumulative_budget_as_tick() { + // A forced probe must spend from the SAME budget an organic stall + // would -- two independent counters would let the total number of + // in-place retries exceed max_recovery_attempts, exactly the + // unbounded-worst-case hole the budget exists to close (see the + // module doc). + let mut m = HealthMonitor::new(2); + assert_eq!( + m.force_recovery(), + HealthAction::AttemptRecovery { attempt: 1, max: 2 } + ); + m.tick(Some(10.0)); + m.tick(Some(10.0)); // stall #1 + assert_eq!( + m.tick(Some(10.0)), // stall #2 -- second and LAST budgeted attempt + HealthAction::AttemptRecovery { attempt: 2, max: 2 } + ); + // Budget exhausted by one forced + one organic attempt: a third + // request of EITHER kind must escalate, not attempt again. + assert_eq!(m.force_recovery(), HealthAction::Escalate); + } + + #[test] + fn force_recovery_escalates_once_the_budget_is_already_exhausted() { + let mut m = HealthMonitor::new(0); + assert_eq!(m.force_recovery(), HealthAction::Escalate); + } + + #[test] + fn force_recovery_resets_the_position_baseline_like_a_real_recovery() { + // Same reasoning as tick's own reset test: the forced recovery's + // loadfile-replace restarts mpv's position counter near zero, so + // the monitor's baseline must drop too or the next real sample + // reads as "still less than the old high-water mark" -- a stall + // that never happened. + let mut m = HealthMonitor::new(3); + m.tick(Some(500.0)); + assert_eq!( + m.force_recovery(), + HealthAction::AttemptRecovery { attempt: 1, max: 3 } + ); + assert_eq!(m.tick(Some(0.1)), HealthAction::Healthy); + } + + #[test] + fn force_recovery_clears_accumulated_jitter_so_the_next_stall_needs_its_own_two_ticks( + ) { + // One non-advancing tick (not yet qualifying) followed by a forced + // probe must not leave that lone stall "banked" -- the next + // organic judgment needs its own fresh two-tick window, same as + // after any other recovery. + let mut m = HealthMonitor::new(3); + m.tick(Some(1.0)); + m.tick(Some(1.0)); // stall #1 -- does not yet qualify + assert_eq!( + m.force_recovery(), + HealthAction::AttemptRecovery { attempt: 1, max: 3 } + ); + // If the pre-existing stall had survived, this single non-advancing + // tick would immediately qualify as "stall #2". It must not. + assert_eq!(m.tick(Some(0.1)), HealthAction::Healthy); + } + + #[test] + fn force_recovery_trigger_does_not_fire_before_its_deadline() { + let mut t = ForceRecoveryTrigger::new(10); + assert!(!t.should_fire(0)); + assert!(!t.should_fire(9)); + } + + #[test] + fn force_recovery_trigger_fires_exactly_once_at_the_deadline() { + let mut t = ForceRecoveryTrigger::new(10); + assert!(!t.should_fire(9)); + assert!(t.should_fire(10)); + // Same call again (a later tick at the same or a later uptime) must + // not re-fire -- T7 is a single live-fire probe, not a repeating one. + assert!(!t.should_fire(10)); + assert!(!t.should_fire(11)); + assert!(!t.should_fire(1_000_000)); + } + + #[test] + fn force_recovery_trigger_fires_late_if_polled_late_but_still_only_once() { + // The driver in main.rs polls this on a cadence, not continuously -- + // a poll that lands after the deadline must still fire (once), not + // wait for an exact match. + let mut t = ForceRecoveryTrigger::new(10); + assert!(!t.should_fire(3)); + assert!(t.should_fire(47)); + assert!(!t.should_fire(48)); + } + + #[test] + fn force_recovery_trigger_zero_secs_fires_on_the_first_poll() { + let mut t = ForceRecoveryTrigger::new(0); + assert!(t.should_fire(0)); + assert!(!t.should_fire(0)); + } +} diff --git a/packages/dexd/src/heartbeat.rs b/packages/dexd/src/heartbeat.rs new file mode 100644 index 0000000..4a4eeca --- /dev/null +++ b/packages/dexd/src/heartbeat.rs @@ -0,0 +1,540 @@ +//! F7/F9 — the heartbeat line: one log line every ~10 min so a weeks-later +//! field failure is diagnosable from the journal after the fact (was it +//! degrading? hot? dropping frames? did the core stop answering?). +//! Formatting is pure and unit-tested here; scheduling and the mpv/sysfs +//! plumbing live in main.rs. +//! +//! # F9 — why every field here comes from an event, never a synchronous read +//! +//! Before F9, `emit_heartbeat` called `mpv_get_property_string` directly from +//! the event thread. Verified against mpv v0.40.0 source: +//! `mpv_get_property_string` -> `run_locked` -> `mp_dispatch_lock` +//! (misc/dispatch.c:364-394) spins on `mp_cond_wait` with **no timeout** +//! until the core thread is trapped inside `mp_dispatch_queue_process()`. A +//! core thread wedged in a DRM ioctl mid-playloop never reaches that trap +//! point, so the caller -- this event thread, the same one `mpv_wait_event` +//! runs on -- blocks forever. That is bug #1's exact shape (process alive, +//! supervisor green, screen black) entered through the one diagnostic that +//! was supposed to help detect it. +//! +//! The fix is not "read less often", it is "never call in". Every value a +//! heartbeat line prints now arrives ONLY via `MPV_EVENT_PROPERTY_CHANGE`, +//! the same door F1's `time-pos` subscription already uses. Three source-backed +//! properties make that strictly better than the synchronous read it +//! replaces, not merely no-worse: +//! +//! 1. **The read migrates off our thread.** Observed-property getters run on +//! the core thread via `send_client_property_changes()`, which explicitly +//! drops the client lock around the getter call (player/client.c +//! :1694-1699, "property getters can do whatever they want"). A wedged +//! getter blocks neither our event thread nor `mpv_wait_event`. If the +//! core is wedged we get silence, not a hang -- and silence is exactly +//! the signal `pos-age=` below exists to surface. +//! 2. **Steady-state cost is a handful of events per counter, ever, never +//! queued.** Registration forces one initial notification per counter +//! (client.h) that arrives promptly as `format=NONE`/`data=NULL` because +//! no VO chain exists yet (player/command.c:763-781); a second event +//! carries the first real `INT64` value once the VO chain comes up +//! (availability flipping counts as a change independent of value +//! equality, player/client.c:1715). After that, change events fire only +//! when the value actually changes (`equal_mpv_value`, +//! player/client.c:1715-1717). A gallery run with no drops emits two +//! events per counter around startup, then silence for three weeks. +//! 3. **They cannot contribute to `QUEUE_OVERFLOW`.** Property-change events +//! are "never queued" (player/client.c:942-943) -- generated inside +//! `mpv_wait_event` only once the queue has drained. Observing more +//! properties cannot push this program toward the one event it treats as +//! fatal. +//! +//! `mpv_get_property_string` and the `mpv_free` it requires are gone from +//! main.rs's FFI surface entirely (not merely unused) -- see the comment +//! left in their place there. With no `*mut MpvHandle` in scope, +//! `emit_heartbeat` cannot call into mpv even by accident. That is a gate, +//! not a rule: per this repo's gates-over-rules principle, "don't add a +//! blocking read here" as a comment can be forgotten; a missing parameter +//! cannot. + +/// One mpv counter as the heartbeat knows it: a value learned ONLY from +/// `MPV_EVENT_PROPERTY_CHANGE`, never from a synchronous read (see the +/// module doc for why that distinction is the whole of F9). +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub struct ObservedCounter { + observed: bool, + last_raw: Option, + /// Whether ANY real sample has ever arrived, via `sample`. Deliberately + /// separate from `last_raw`: `mark_unavailable` clears `last_raw` (the + /// diffing baseline) but must NOT make an already-earned `total` render + /// as "n/a" again -- that would trade one false-all-clear risk (a raw + /// counter reading 0 right after a reset) for another (a real, + /// already-known total flickering back to "no value yet" for the + /// ordinary few-second VO-chain gap every recovery causes). + has_sample: bool, + total: u64, +} + +impl ObservedCounter { + /// `mpv_observe_property` succeeded: this run can produce values. + pub const fn observed() -> Self { + Self { observed: true, last_raw: None, has_sample: false, total: 0 } + } + + /// `mpv_observe_property` FAILED: this run never will. Renders `off`, + /// which is a different fact from "no value yet" -- see the module doc + /// (principle 2, "distinguish the two no-value cases"). + pub const fn unobserved() -> Self { + Self { observed: false, last_raw: None, has_sample: false, total: 0 } + } + + /// Feed one `MPV_EVENT_PROPERTY_CHANGE` payload (already unwrapped from + /// `MPV_FORMAT_INT64`). Accumulates across per-file counter resets -- + /// see the body comment for why that is required, not merely tidy. + pub fn sample(&mut self, raw: i64) { + // mpv's drop counters are PER PLAYBACK SESSION: F1's tier-0 recovery + // (`loadfile ... replace`) rebuilds the VO chain and restarts both + // at 0. Printed raw, the last heartbeat before a field failure could + // read `frame-drops=0` purely because a recovery reset it four + // minutes earlier -- a false all-clear in the ONE line that is the + // only diagnostic channel on site. So accumulate: add forward + // deltas, and treat any decrease as a reset whose post-reset value + // is itself new. + // + // `max(0)`: mpv never reports a negative drop count, but this value + // arrives through a hand-transcribed FFI tag check, and a clamp is + // cheaper than a cast that could wrap into billions. + let raw = raw.max(0) as u64; + let delta = match self.last_raw { + Some(prev) if raw >= prev => raw - prev, + Some(_) => raw, // counter reset: everything visible now is new + None => raw, // first sample of the process + }; + self.total = self.total.saturating_add(delta); + self.last_raw = Some(raw); + self.has_sample = true; + } + + /// Cumulative count since process start, or `None` if no sample has + /// EVER arrived. Test/caller accessor; rendering goes through `Display`. + /// Gated on `has_sample`, not `last_raw`: `mark_unavailable` clears the + /// latter (a diffing-baseline reset) without un-earning a total this + /// counter has already accumulated -- see `has_sample`'s doc comment. + pub fn total(&self) -> Option { + self.has_sample.then_some(self.total) + } + + /// Record that mpv reported this property as currently UNAVAILABLE (an + /// `MPV_EVENT_PROPERTY_CHANGE` with `format=MPV_FORMAT_NONE` arrived -- + /// no VO chain, at startup or mid-recovery). `total` (and whether + /// `total()` renders it at all -- see `has_sample`) is untouched: this + /// is not itself a counter reset, and a total already earned must not + /// flicker back to "n/a" for the ordinary few-second gap every recovery + /// causes. `last_raw` IS cleared, so the next real sample -- however + /// small -- is read as a fresh first sample (added in full) rather than + /// diffed against a value that may already belong to a dead playback + /// session. + /// + /// Why this exists (MINOR, three adversarial reviews, 2026-08-15): + /// `sample`'s only reset signal is a numeric DECREASE, but mpv coalesces + /// property-change events -- only the latest state per changed property + /// survives to the next drain (client.h). If a recovery's teardown + /// (unavailable), the new session's restart at 0, and a climb past the + /// old session's total all happen before this program's event thread + /// drains -- plausible exactly during a recovery, when that thread is + /// busy absorbing the END_FILE/START_FILE burst -- the visible sequence + /// can be e.g. 5 -> 7 with no decrease at all, and `sample` would count + /// a delta of 2 when 7 new drops actually occurred. Marking the + /// unavailable state narrows that window (the NONE event itself could + /// still be coalesced away) rather than closing it -- full closure is + /// not possible from the client side, and isn't worth more machinery + /// for a diagnostic line. + pub fn mark_unavailable(&mut self) { + self.last_raw = None; + } +} + +impl std::fmt::Display for ObservedCounter { + fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + if !self.observed { + // Subscription never registered: this number is not evidence of + // anything, look at the startup lines instead. + write!(f, "off") + } else if let Some(total) = self.total() { + write!(f, "{total}") + } else { + // Registered, but no real sample has EVER arrived (`total()` is + // gated on `has_sample`, not on the diffing baseline + // `mark_unavailable` clears -- so this branch, unlike a plain + // "no current value" check, is NOT re-entered by an ordinary + // few-second VO-chain gap once at least one sample has already + // landed). Two distinct causes render identically, both + // correctly: (1) no VO chain has EVER come up in this run's + // whole life -- expected only for heartbeat #0, unusual after; + // (2) the property NAME is unknown to this mpv build. + // `mpv_observe_property` never validates the name -- it fails + // only for a bad FORMAT or OOM (client.h: "Observing a property + // that doesn't exist is allowed") -- so a future mpv renaming + // `frame-drop-count`/`vo-delayed-frame-count` would not fail + // `observe()` (which would render "off" instead, see above); it + // would silently subscribe successfully and sit here forever. + // Whichever cause, a run stuck here for HOURS (not seconds) IS + // the diagnosis: the core asked and never usefully answered. + write!(f, "n/a") + } + } +} + +/// The last known playback position and how stale it is. One struct, not +/// two `Option`s, because "position known but age unknown" cannot occur -- +/// both come from the same `MPV_EVENT_PROPERTY_CHANGE`. +#[derive(Debug, Clone, Copy, PartialEq)] +pub struct PositionSample { + pub secs: f64, + pub age_secs: u64, +} + +/// Everything one heartbeat line prints. A struct with NAMED fields, not a +/// multi-argument function: two of the fields are `ObservedCounter` and two +/// are integers, so a positional call site could silently transpose +/// frame_drops/vo_delayed (or wraps/uptime) and no test in the crate would +/// catch it -- main.rs is the one file the Mac cannot run. +#[derive(Debug, Clone, Copy)] +pub struct HeartbeatSnapshot { + pub wraps: u64, + pub uptime_secs: u64, + pub temp_millicelsius: Option, + pub position: Option, + pub frame_drops: ObservedCounter, + pub vo_delayed: ObservedCounter, + /// F10: `None` when the systemd watchdog is inert for this run (no + /// `$NOTIFY_SOCKET` -- every Mac/bench/CI run); `Some(n)` when armed, + /// `n` being the cumulative count of pings that could not be sent + /// (`dexd::watchdog::PingOutcome::Dropped`) since process start. + /// Kept as a plain `Option` rather than importing + /// `dexd::watchdog`'s own types here -- this module already prints + /// two other subsystems' state (F1's position, F9's counters) as plain + /// values, and pulling in a third module's enum just to render one + /// number would be the odd one out. + pub watchdog_pings_dropped: Option, +} + +impl HeartbeatSnapshot { + /// Render one heartbeat line. + pub fn render(&self) -> String { + let temp = match self.temp_millicelsius { + Some(m) => { + // Extract the sign before dividing: `m / 1000` truncates + // toward zero, so for m in -999..=-1 it is 0 -- an unheated + // venue on a winter cold-boot (e.g. -250 m°C) would + // otherwise render as a plausible-looking POSITIVE "0.2C" in + // the journal. `unsigned_abs` sidesteps the one panic hazard + // in this shape (`i64::MIN.abs()`) entirely, though sysfs + // never reports a value near it. + let (sign, mag) = if m < 0 { ("-", m.unsigned_abs()) } else { ("", m as u64) }; + format!("{sign}{}.{}C", mag / 1000, (mag % 1000) / 100) + } + None => "n/a".to_string(), + }; + let (pos, pos_age) = match self.position { + Some(PositionSample { secs, age_secs }) => (format!("{secs:.1}s"), format!("{age_secs}s")), + None => ("n/a".to_string(), "n/a".to_string()), + }; + // F10: "inert" and "armed pings-dropped=0" are different facts (no + // systemd watchdog at all, vs. one that is armed and has never lost + // a ping yet) -- both worth telling apart from the journal after + // the fact, same reasoning as `ObservedCounter`'s off/n/a split. + let watchdog = match self.watchdog_pings_dropped { + Some(n) => format!("armed pings-dropped={n}"), + None => "inert".to_string(), + }; + format!( + "dexd: heartbeat wraps={} uptime={}s temp={temp} frame-drops={} \ + vo-delayed={} pos={pos} pos-age={pos_age} watchdog={watchdog}", + self.wraps, self.uptime_secs, self.frame_drops, self.vo_delayed, + ) + } +} + +#[cfg(test)] +mod tests { + use super::*; + + // --- ObservedCounter ---------------------------------------------- + + #[test] + fn unobserved_renders_off_even_after_a_stray_sample() { + let mut c = ObservedCounter::unobserved(); + assert_eq!(c.to_string(), "off"); + // A stray sample must not happen in practice (nothing feeds an + // unobserved counter), but the render must stay "off" regardless -- + // "off" means "the subscription never registered", which a sample + // arriving cannot retroactively change. + c.sample(5); + assert_eq!(c.to_string(), "off"); + } + + #[test] + fn observed_with_no_sample_renders_na() { + assert_eq!(ObservedCounter::observed().to_string(), "n/a"); + assert_eq!(ObservedCounter::observed().total(), None); + } + + #[test] + fn first_sample_becomes_the_total() { + let mut c = ObservedCounter::observed(); + c.sample(7); + assert_eq!(c.total(), Some(7)); + assert_eq!(c.to_string(), "7"); + } + + #[test] + fn monotone_samples_accumulate_by_delta() { + let mut c = ObservedCounter::observed(); + c.sample(2); + c.sample(5); + c.sample(5); // no change: delta 0 + c.sample(9); + assert_eq!(c.total(), Some(9)); + } + + #[test] + fn a_decrease_is_a_tier_0_recovery_reset_and_accumulates_the_post_reset_value() { + // The scenario this pins: F1's in-place recovery (`loadfile ... + // replace`) rebuilds the VO chain, and mpv restarts frame-drop-count + // at 0 for the new playback session. Printed raw, that reset would + // make `frame-drops=0` look like a false all-clear right after a + // recovery. The heartbeat must instead keep counting forward from + // where it was. + let mut c = ObservedCounter::observed(); + c.sample(40); // pre-recovery: 40 drops accumulated + c.sample(3); // recovery reset the session counter to 3 + assert_eq!(c.total(), Some(43), "the post-reset value must be added, not replace the total"); + c.sample(3); // steady after the reset: no further change + assert_eq!(c.total(), Some(43)); + } + + #[test] + fn mark_unavailable_keeps_total_but_makes_the_next_sample_a_fresh_first_sample() { + // The scenario `mark_unavailable` exists for: mpv coalesces + // property-change events, so an unavailability blip that lands + // between two drains can hide an ENTIRE recovery episode's session + // reset from `sample`'s only reset signal (a numeric decrease) -- + // see `mark_unavailable`'s doc comment. `total` must survive + // untouched; `last_raw` must NOT, so the next sample is read as a + // fresh first sample (added in full) rather than diffed against a + // value that may belong to a dead session. + let mut c = ObservedCounter::observed(); + c.sample(40); + assert_eq!(c.total(), Some(40)); + + c.mark_unavailable(); + assert_eq!( + c.total(), + Some(40), + "an already-earned total must not flicker back to n/a for an \ + ordinary unavailability blip -- total() is gated on has_sample, \ + not on the diffing baseline mark_unavailable clears" + ); + + c.sample(7); // a coalesced sequence would otherwise read as delta=7 + assert_eq!( + c.total(), + Some(47), + "post-unavailability sample must be added in full as a fresh \ + first sample, not diffed against the stale pre-unavailability \ + last_raw" + ); + } + + #[test] + fn negative_raw_is_clamped_to_zero() { + let mut c = ObservedCounter::observed(); + c.sample(-1); + assert_eq!(c.total(), Some(0)); + } + + #[test] + fn total_saturates_at_u64_max() { + // Three `i64::MAX` contributions overflow u64 (u64::MAX = 2 * + // i64::MAX + 1, so a third addition of i64::MAX always overflows). + // Each contribution arrives via a "reset to 0, then jump back to + // i64::MAX" pair so every sample stays a legal i64 input -- this + // tests `saturating_add`, not a way to smuggle an out-of-range raw + // value past the FFI boundary. + let mut c = ObservedCounter::observed(); + c.sample(i64::MAX); + c.sample(0); // reset + c.sample(i64::MAX); + c.sample(0); // reset + c.sample(i64::MAX); + assert_eq!(c.total(), Some(u64::MAX)); + } + + // --- HeartbeatSnapshot::render -------------------------------------- + + #[test] + fn renders_the_full_line() { + let mut frame_drops = ObservedCounter::observed(); + frame_drops.sample(0); + let mut vo_delayed = ObservedCounter::observed(); + vo_delayed.sample(2); + let snap = HeartbeatSnapshot { + wraps: 143, + uptime_secs: 3600, + temp_millicelsius: Some(48_250), + position: Some(PositionSample { secs: 3599.4, age_secs: 0 }), + frame_drops, + vo_delayed, + watchdog_pings_dropped: Some(0), + }; + assert_eq!( + snap.render(), + "dexd: heartbeat wraps=143 uptime=3600s temp=48.2C frame-drops=0 \ + vo-delayed=2 pos=3599.4s pos-age=0s watchdog=armed pings-dropped=0" + ); + } + + #[test] + fn all_missing_sources_degrade_to_na_not_errors() { + let snap = HeartbeatSnapshot { + wraps: 0, + uptime_secs: 0, + temp_millicelsius: None, + position: None, + frame_drops: ObservedCounter::observed(), + vo_delayed: ObservedCounter::observed(), + watchdog_pings_dropped: None, + }; + assert_eq!( + snap.render(), + "dexd: heartbeat wraps=0 uptime=0s temp=n/a frame-drops=n/a \ + vo-delayed=n/a pos=n/a pos-age=n/a watchdog=inert" + ); + } + + #[test] + fn unregistered_counters_render_off() { + let snap = HeartbeatSnapshot { + wraps: 0, + uptime_secs: 0, + temp_millicelsius: None, + position: None, + frame_drops: ObservedCounter::unobserved(), + vo_delayed: ObservedCounter::unobserved(), + watchdog_pings_dropped: None, + }; + assert_eq!( + snap.render(), + "dexd: heartbeat wraps=0 uptime=0s temp=n/a frame-drops=off \ + vo-delayed=off pos=n/a pos-age=n/a watchdog=inert" + ); + } + + #[test] + fn armed_watchdog_prints_the_dropped_ping_count() { + let snap = HeartbeatSnapshot { + wraps: 0, + uptime_secs: 0, + temp_millicelsius: None, + position: None, + frame_drops: ObservedCounter::observed(), + vo_delayed: ObservedCounter::observed(), + watchdog_pings_dropped: Some(7), + }; + assert!( + snap.render().ends_with("watchdog=armed pings-dropped=7"), + "{}", + snap.render() + ); + } + + #[test] + fn inert_watchdog_never_prints_a_dropped_count() { + // `None` means "no systemd watchdog at all" -- distinct from "armed, + // zero drops so far" (Some(0), pinned by `renders_the_full_line` + // above). A dropped-count number here would misleadingly suggest a + // watchdog exists when it does not. + let snap = HeartbeatSnapshot { + wraps: 0, + uptime_secs: 0, + temp_millicelsius: None, + position: None, + frame_drops: ObservedCounter::observed(), + vo_delayed: ObservedCounter::observed(), + watchdog_pings_dropped: None, + }; + let s = snap.render(); + assert!(s.ends_with("watchdog=inert"), "{s}"); + assert!(!s.contains("pings-dropped"), "{s}"); + } + + #[test] + fn sub_zero_temperatures_keep_their_sign() { + // -250 m°C: an unheated venue on a winter cold-boot. Before the + // original fix, `m / 1000 == 0` for any m in -999..=-1, so this + // silently rendered as the POSITIVE "0.2C" -- a plausible-looking + // wrong reading in exactly the diagnostic line this module exists + // to make trustworthy. Carried over verbatim from F7. + let base = HeartbeatSnapshot { + wraps: 0, + uptime_secs: 0, + temp_millicelsius: Some(-250), + position: None, + frame_drops: ObservedCounter::observed(), + vo_delayed: ObservedCounter::observed(), + watchdog_pings_dropped: None, + }; + assert_eq!( + base.render(), + "dexd: heartbeat wraps=0 uptime=0s temp=-0.2C frame-drops=n/a \ + vo-delayed=n/a pos=n/a pos-age=n/a watchdog=inert" + ); + let mut base = base; + base.temp_millicelsius = Some(-1_500); + assert_eq!( + base.render(), + "dexd: heartbeat wraps=0 uptime=0s temp=-1.5C frame-drops=n/a \ + vo-delayed=n/a pos=n/a pos-age=n/a watchdog=inert" + ); + // i64::MIN: the one value where a naive `.abs()` would panic. + // `unsigned_abs()` does not. + base.temp_millicelsius = Some(i64::MIN); + let s = base.render(); + assert!(s.contains("temp=-"), "{s}"); + } + + #[test] + fn position_formats_to_one_decimal() { + let snap = HeartbeatSnapshot { + wraps: 0, + uptime_secs: 0, + temp_millicelsius: None, + position: Some(PositionSample { secs: 12.049, age_secs: 4 }), + frame_drops: ObservedCounter::observed(), + vo_delayed: ObservedCounter::observed(), + watchdog_pings_dropped: None, + }; + assert!(snap.render().contains("pos=12.0s pos-age=4s"), "{}", snap.render()); + } + + #[test] + fn a_large_position_from_weeks_of_uptime_does_not_go_exponential() { + // Three weeks at ~1x realtime: well past the point where `{}` would + // switch a float to scientific notation if this used a bare + // Display impl instead of a fixed `{:.1}` format. + let three_weeks_secs = 3600.0 * 24.0 * 21.0; + let snap = HeartbeatSnapshot { + wraps: 0, + uptime_secs: 0, + temp_millicelsius: None, + position: Some(PositionSample { secs: three_weeks_secs, age_secs: 0 }), + frame_drops: ObservedCounter::observed(), + vo_delayed: ObservedCounter::observed(), + watchdog_pings_dropped: None, + }; + let s = snap.render(); + assert!(s.contains("pos=1814400.0s"), "{s}"); + // Not a whole-line check: "heartbeat" itself contains 'e'. Scope the + // scientific-notation check to the pos field's rendered text alone. + let pos_field = s.split("pos=").nth(1).unwrap().split(' ').next().unwrap(); + assert!(!pos_field.contains('e') && !pos_field.contains('E'), "{pos_field}"); + } +} diff --git a/packages/dexd/src/lib.rs b/packages/dexd/src/lib.rs new file mode 100644 index 0000000..0894fb3 --- /dev/null +++ b/packages/dexd/src/lib.rs @@ -0,0 +1,20 @@ +//! dexd's pure core: every piece of logic that can exist without libmpv. +//! +//! The binary (src/main.rs) is deliberately a thin unsafe shell over this +//! library: FFI structs, callbacks, and the event loop. Everything decidable — +//! wrap arithmetic, hashing, sidecar binding, NAL validation, log formatting — +//! lives here, testable on any machine with `cargo test --lib`, no libmpv and +//! no display required. That split is T0 of the hardening plan: every serious +//! bug so far lived where testing could not reach. + +#![forbid(unsafe_code)] + +pub mod chunk; +pub mod exhibit; +pub mod ffi_consts; +pub mod health; +pub mod heartbeat; +pub mod nal; +pub mod sha256; +pub mod sidecar; +pub mod watchdog; diff --git a/packages/dexd/src/main.rs b/packages/dexd/src/main.rs new file mode 100644 index 0000000..a025034 --- /dev/null +++ b/packages/dexd/src/main.rs @@ -0,0 +1,1938 @@ +//! Gapless HEVC looper for the Raspberry Pi: libmpv fed by an endless byte stream. +//! +//! # Why this exists +//! +//! Every mpv looping mechanism stalls at the wrap, because each one makes the +//! decoder re-enter the file. Measured on a Pi 4 at 4K30 and 1080p60 against an +//! HDMI capture card (2026-08-13/14): +//! +//! | mechanism | hold at the wrap | +//! |----------------------------------|-------------------------| +//! | `--loop-file=inf` | 83 ms, every loop | +//! | `--ab-loop-a/b` | 83 ms, identical | +//! | `--playlist` + `--prefetch` | 117-133 ms, worse | +//! | `ffmpeg -stream_loop` + vout_drm | 67-217 ms, 3 per loop | +//! | `--loop-file=inf` on raw `.265` | freezes on the last frame | +//! +//! The only configuration with **zero** held frames is one where the decoder +//! never reaches EOF. That works because the asset has a closed GOP with an IDR +//! at frame 0, so presenting byte 0 straight after the last byte is an ordinary +//! mid-stream IDR rather than a seek. +//! +//! `while true; do cat loop.265; done | mpv -` proves it (verified seamless: +//! zero held frames across 19 wraps at 4K30, and a 3.5 h soak with flat memory) +//! but is not shippable: a process per loop (~29k/day for a 3 s card), and a +//! SIGPIPE hot-spin burning a core if mpv ever exits. +//! +//! A Python version of this same design displayed correctly but ran at **0.6x +//! realtime**, with frames held at random points across the loop -- the +//! signature of a starved feed. The bytes have to arrive at ~5 MB/s against hard +//! per-frame deadlines, and a ctypes callback under the GIL cannot promise that. +//! libmpv's `stream_cb` is a C API, so in Rust the same design is a memcpy. +//! +//! # The whole trick +//! +//! [`read_fn`] never returns 0. Returning 0 means EOF to mpv; wrapping the +//! offset back to the start instead means the stream simply never ends. +//! +//! # Requirements +//! +//! * A raw Annex-B HEVC elementary stream, not MP4: +//! `ffmpeg -i card.mp4 -c:v copy -bsf:v hevc_mp4toannexb -f hevc loop.265` +//! * The real frame rate, because a raw stream carries no timestamps. This is +//! why frame rate has to become ingest metadata (milestone 3). + +#![deny(unsafe_op_in_unsafe_fn)] + +use dexd::chunk::{clamp_want, next_chunk}; +use dexd::exhibit::{ + check_cmdline_matches, load_exhibit_config, mode_resolution, resolve_asset, resolve_display, + sysfs_modes_contains, AssetSource, DisplaySource, DEFAULT_EXHIBIT_CONFIG_PATH, + DEFAULT_EXHIBIT_CONFIG_PATHS, +}; +use dexd::ffi_consts::{ + MPV_END_FILE_REASON_STOP, MPV_ERROR_UNSUPPORTED, MPV_EVENT_COMMAND_REPLY, MPV_EVENT_END_FILE, + MPV_EVENT_LOG_MESSAGE, MPV_EVENT_NONE, MPV_EVENT_PROPERTY_CHANGE, MPV_EVENT_QUEUE_OVERFLOW, + MPV_EVENT_SHUTDOWN, MPV_EVENT_START_FILE, MPV_FORMAT_DOUBLE, MPV_FORMAT_INT64, + MPV_FORMAT_NONE, +}; +use dexd::health::{ForceRecoveryTrigger, HealthAction, HealthMonitor}; +use dexd::heartbeat::{HeartbeatSnapshot, ObservedCounter, PositionSample}; +use dexd::nal::validate_leading_nals; +use dexd::sidecar::{resolve_fps, verify_payload, FpsSource, Sidecar}; +use dexd::watchdog::{self, PingOutcome, WatchdogDecision, WatchdogEnv}; +use std::env; +use std::ffi::{c_char, c_int, c_void, CStr, CString}; +use std::fs; +use std::os::unix::net::{SocketAddr as UnixSocketAddr, UnixDatagram}; +use std::process::ExitCode; +use std::sync::atomic::{AtomicU64, Ordering}; +use std::time::Instant; + +// --------------------------------------------------------------------------- +// libmpv FFI. Only the handful of entry points this program needs, transcribed +// from mpv/client.h and mpv/stream_cb.h. Hand-written rather than bindgen: the +// surface is small, stable, and a build-time codegen dependency would outweigh +// it on a device that has to build offline. +// --------------------------------------------------------------------------- + +#[repr(C)] +struct MpvHandle { + _private: [u8; 0], +} + +/// `mpv_stream_cb_info` — the callbacks mpv will use for one opened stream. +#[repr(C)] +struct MpvStreamCbInfo { + cookie: *mut c_void, + read_fn: Option i64>, + seek_fn: Option i64>, + size_fn: Option i64>, + close_fn: Option, + cancel_fn: Option, +} + +#[repr(C)] +struct MpvEvent { + event_id: c_int, + error: c_int, + reply_userdata: u64, + data: *mut c_void, +} + +/// `mpv_event_end_file`. Only the first two fields are read; the trailing +/// playlist fields exist in the C struct but are irrelevant to a single-file +/// appliance, and reading a prefix of a #[repr(C)] struct is well-defined. +#[repr(C)] +struct MpvEventEndFile { + reason: c_int, + error: c_int, +} + +#[repr(C)] +struct MpvEventLogMessage { + prefix: *const c_char, + level: *const c_char, + text: *const c_char, + log_level: c_int, +} + +/// `mpv_event_property`. Only read when `format == MPV_FORMAT_DOUBLE` (F1's +/// `time-pos` subscription); `data` is a tagged union whose true type +/// depends on `format`, so the tag MUST be checked before `data` is ever +/// dereferenced -- see the MPV_FORMAT_DOUBLE doc comment in ffi_consts.rs. +#[repr(C)] +struct MpvEventProperty { + name: *const c_char, + format: c_int, + data: *mut c_void, +} + +#[link(name = "mpv")] +extern "C" { + fn mpv_create() -> *mut MpvHandle; + fn mpv_initialize(ctx: *mut MpvHandle) -> c_int; + fn mpv_terminate_destroy(ctx: *mut MpvHandle); + fn mpv_set_option_string( + ctx: *mut MpvHandle, + name: *const c_char, + data: *const c_char, + ) -> c_int; + fn mpv_command(ctx: *mut MpvHandle, args: *const *const c_char) -> c_int; + // Async counterpart of mpv_command: queues the command and returns + // immediately (client.h), replying later via MPV_EVENT_COMMAND_REPLY. + // F1's in-place recovery uses this, never the synchronous mpv_command, + // specifically so the event thread cannot block on it -- see + // src/health.rs's module doc. + fn mpv_command_async( + ctx: *mut MpvHandle, + reply_userdata: u64, + args: *const *const c_char, + ) -> c_int; + fn mpv_wait_event(ctx: *mut MpvHandle, timeout: f64) -> *mut MpvEvent; + fn mpv_error_string(error: c_int) -> *const c_char; + fn mpv_request_log_messages(ctx: *mut MpvHandle, min_level: *const c_char) -> c_int; + fn mpv_stream_cb_add_ro( + ctx: *mut MpvHandle, + protocol: *const c_char, + user_data: *mut c_void, + open_fn: Option c_int>, + ) -> c_int; + // Non-blocking (client.h): queues a subscription and returns + // immediately. F1's health check calls this exactly ONCE at startup and + // thereafter only ever reads the position out of the resulting + // MPV_EVENT_PROPERTY_CHANGE events -- never a synchronous property + // read -- see src/health.rs's module doc for why that distinction is + // the whole point. + fn mpv_observe_property( + ctx: *mut MpvHandle, + reply_userdata: u64, + name: *const c_char, + format: c_int, + ) -> c_int; + // DELIBERATELY ABSENT: mpv_get_property_string / mpv_get_property (and + // the mpv_free they require). Every synchronous property read goes + // through run_locked -> mp_dispatch_lock (player/client.c:1059, + // misc/dispatch.c:364), which waits WITHOUT A TIMEOUT until the core + // thread is trapped in its dispatch loop -- a core wedged in a DRM + // ioctl never gets there, so the caller blocks forever. That is F9: + // the heartbeat's own diagnostic becoming bug #1 (alive, supervisor + // green, screen black). Values come from mpv_observe_property + + // MPV_EVENT_PROPERTY_CHANGE instead. Do not re-add these bindings. +} + +// --------------------------------------------------------------------------- +// The endless stream +// --------------------------------------------------------------------------- + +/// One reader's position within the looping payload. +/// +/// The payload is `&'static [u8]` because it is leaked once at startup and must +/// outlive every mpv thread; there is no meaningful point at which freeing it +/// would be correct while the player runs. +struct LoopStream { + data: &'static [u8], + pos: usize, +} + +// The cookie is created on one mpv thread, read on the demux thread, and freed +// on whichever thread closes the stream. That requires Send. It is Send today, +// but the raw-pointer laundering through `cookie` means the compiler never +// checks -- so assert it, and a future field (Rc, *mut, an mmap guard) becomes +// a compile error rather than a data race. +const _: () = { + const fn assert_send() {} + assert_send::(); +}; + +/// Completed passes over the payload — incremented by `read_fn` on the demux +/// thread each time the position wraps to 0, read by the heartbeat on the +/// event thread. This counts DEMUXER passes, which run ~1 s (readahead) +/// ahead of what is on screen. Relaxed ordering: a monotonic diagnostic +/// counter, not a synchronization point. +static WRAP_COUNT: AtomicU64 = AtomicU64::new(0); + +/// Heartbeat cadence: frequent enough to bound "when did it die" to a useful +/// journal window, rare enough to cost nothing. +const HEARTBEAT_SECS: u64 = 600; + +/// F1 tier-0 health-check cadence: "order 10 s" per PLAN.md -- low enough +/// frequency, off the decode path, to cost nothing; the escalation policy +/// itself (dexd::health::HealthMonitor) requires TWO consecutive +/// non-advancing checks before acting, so real detection latency for an +/// actual stall is roughly 2x this. +const HEALTH_CHECK_SECS: u64 = 10; + +/// F1's cumulative, process-lifetime in-place-recovery budget -- see +/// dexd::health's module doc ("why the budget never resets") for the +/// full reasoning. 3 is small enough to guarantee the worst case is bounded +/// (at most 3 loadfile-reload attempts, ever, before conceding to tier 1) +/// yet large enough to absorb a handful of isolated transient glitches +/// (HDMI blink, a sink waking late) over a multi-week unattended run +/// without needlessly forcing a full process restart for something tier 0 +/// could fix in place. +const MAX_RECOVERY_ATTEMPTS: u32 = 3; + +/// `reply_userdata` tag for F1's `time-pos` subscription, so its +/// MPV_EVENT_PROPERTY_CHANGE events are told apart from any property this +/// program observes in the future. +const HEALTH_CHECK_USERDATA: u64 = 1; + +/// `reply_userdata` tag for F1's in-place-recovery `loadfile` command, so +/// its MPV_EVENT_COMMAND_REPLY is told apart from any async command this +/// program issues in the future. +const RECOVERY_COMMAND_USERDATA: u64 = 2; + +/// `reply_userdata` tags for F9's two observed drop counters, told apart +/// from HEALTH_CHECK_USERDATA/RECOVERY_COMMAND_USERDATA and from each other +/// so the MPV_EVENT_PROPERTY_CHANGE handler never has to `CStr`-compare +/// `p.name` to know which counter a payload belongs to. +const FRAME_DROPS_USERDATA: u64 = 3; +const VO_DELAYED_USERDATA: u64 = 4; + +/// The stream read callback: a thin unsafe shell over +/// [`dexd::chunk::next_chunk`], which owns (and tests) every rule that +/// matters — never return 0 (to mpv, 0 is final EOF, the one event this +/// program exists to prevent), wrap eagerly, saturate the u64 request size. +/// This function only performs the memcpy the pure core cannot. +/// +/// The `None` (zero-length request) branch below is unreachable in mpv 0.40 +/// (`stream.c` guards `len <= 0` before ever calling in) and, if it ever did +/// fire, a negative return is treated identically to 0 by +/// `stream_read_unbuffered` (`res <= 0` -> EOF either way) -- so returning an +/// error here is not a mechanism mpv honors specially. It is kept as a +/// defensive sentinel so THIS crate's own diagnostics can tell "asked for +/// nothing" apart from "ran out of things to give"; if the path ever does +/// fire on a future mpv, the result is an ordinary END_FILE -> fatal exit -> +/// supervisor restart, not a seam. +extern "C" fn read_fn(cookie: *mut c_void, buf: *mut c_char, nbytes: u64) -> i64 { + // SAFETY: `cookie` is the Box leaked in `open_fn`, and mpv + // guarantees it is passed back unmodified for the life of the stream. + // The `&mut` additionally requires exclusivity. That comes from mpv's + // STREAM LAYER, not from any serialization guarantee in stream_cb.h -- + // verified against mpv v0.40.0's include/mpv/stream_cb.h: it documents + // no such thing, and the one callback whose threading it DOES document + // (cancel_fn) is explicitly cross-thread ("will be called from a + // separate thread than the demux thread", stream_cb.h:154-155). The + // real mechanism: a stream_t is single-owner and driven by one thread at + // a time -- open_cb runs the seek_fn(cookie, 0) probe during open, + // before fill_buffer/close are ever installed (stream/stream_cb.c), so + // open-time access happens-before every read, and close (after demux + // teardown) happens-after the last one. `cancel_fn` is `None` today + // (see `open_fn` below) specifically so nothing can touch this cookie + // from that documented second thread; if a future change ever needs + // `cancel_fn`, this exclusivity argument breaks and the cookie needs a + // redesign (e.g. an atomic/lock) before it is wired up. + let s = unsafe { &mut *(cookie as *mut LoopStream) }; + + let Some(c) = next_chunk(s.data.len(), s.pos, clamp_want(nbytes)) else { + // Zero-length request (or an impossible empty payload). Report an + // error, never 0 -- see the doc comment above for why this is + // belt-and-braces rather than load-bearing against mpv itself. + return i64::from(MPV_ERROR_UNSUPPORTED); + }; + + // SAFETY: mpv guarantees `buf` is writable for `nbytes` bytes; next_chunk + // guarantees c.n >= 1, c.n <= nbytes (the request is clamped, never + // grown) and c.start + c.n <= data.len(), and the ranges cannot overlap. + // + // `.cast::()` rather than `as *mut u8`, because `c_char` is NOT the same + // type on both machines this crate is built on: it is `i8` on macOS/aarch64 + // and `u8` on Linux/aarch64. So `buf as *mut u8` is a real conversion on the + // dev Mac and a no-op on the Pi -- where clippy then rejects it as an + // unnecessary cast. `.cast()` is correct and lint-clean on both. Do not + // "simplify" it to `buf`: that only compiles on the Pi. + unsafe { + std::ptr::copy_nonoverlapping(s.data.as_ptr().add(c.start), buf.cast::(), c.n); + } + s.pos = c.next_pos; + if c.next_pos == 0 { + // The copy reached the payload's end: one full pass completed. + WRAP_COUNT.fetch_add(1, Ordering::Relaxed); + } + c.n as i64 +} + +/// Report the stream as unseekable, exactly like a pipe. +/// +/// Deliberate: an mpv that believes it can seek will try to, and seeking is the +/// operation that produces the seam. Refusing here keeps the only available +/// behaviour "keep reading forwards". +extern "C" fn seek_fn(_cookie: *mut c_void, _offset: i64) -> i64 { + i64::from(MPV_ERROR_UNSUPPORTED) +} + +/// Report the size as unknown, again like a pipe. +/// +/// Returning the payload length would let mpv compute a duration and a progress +/// position for a stream that has neither, and would invite it to treat the end +/// of the buffer as the end of the media. +extern "C" fn size_fn(_cookie: *mut c_void) -> i64 { + i64::from(MPV_ERROR_UNSUPPORTED) +} + +extern "C" fn close_fn(cookie: *mut c_void) { + // SAFETY: reclaims the Box leaked in `open_fn`; mpv calls this exactly once. + unsafe { drop(Box::from_raw(cookie as *mut LoopStream)) }; +} + +/// Open callback for the `loop://` protocol. The URI is ignored: the payload is +/// fixed at startup, so there is nothing to parse and nothing that can fail here. +extern "C" fn open_fn( + user_data: *mut c_void, + _uri: *mut c_char, + info: *mut MpvStreamCbInfo, +) -> c_int { + // SAFETY: `user_data` is the &'static [u8] passed to mpv_stream_cb_add_ro. + let data: &'static [u8] = unsafe { *(user_data as *mut &'static [u8]) }; + let stream = Box::new(LoopStream { data, pos: 0 }); + + // SAFETY: mpv provides a valid, writable info struct for us to fill. + unsafe { + (*info).cookie = Box::into_raw(stream) as *mut c_void; + (*info).read_fn = Some(read_fn); + (*info).seek_fn = Some(seek_fn); + (*info).size_fn = Some(size_fn); + (*info).close_fn = Some(close_fn); + (*info).cancel_fn = None; + } + 0 +} + +// --------------------------------------------------------------------------- + +fn err(ctx: *mut MpvHandle, what: &str, code: c_int) -> String { + let _ = ctx; + // SAFETY: mpv_error_string returns a static NUL-terminated string. + let msg = unsafe { CStr::from_ptr(mpv_error_string(code)) }; + format!("{what}: {} ({code})", msg.to_string_lossy()) +} + +fn set_opt(ctx: *mut MpvHandle, name: &str, value: &str) -> Result<(), String> { + let n = CString::new(name).map_err(|e| e.to_string())?; + let v = CString::new(value).map_err(|e| e.to_string())?; + let rc = unsafe { mpv_set_option_string(ctx, n.as_ptr(), v.as_ptr()) }; + if rc < 0 { + return Err(err(ctx, &format!("set {name}={value}"), rc)); + } + Ok(()) +} + +/// Pi SoC temperature in millidegrees C, if the kernel exposes it. Absent on +/// non-Linux and never an error: the heartbeat degrades to n/a. +fn read_temp_millicelsius() -> Option { + std::fs::read_to_string("/sys/class/thermal/thermal_zone0/temp") + .ok()? + .trim() + .parse() + .ok() +} + +/// Register one property observer. Non-blocking by construction: +/// `mpv_observe_property` only takes the client-local lock and wakes the +/// core (player/client.c:1536-1572). Returns whether the subscription is +/// live for this run; a failure is never fatal, but per design principle 2 +/// it is always loud. +fn observe(ctx: *mut MpvHandle, userdata: u64, name: &str, format: c_int) -> bool { + let Ok(n) = CString::new(name) else { return false }; + let rc = unsafe { mpv_observe_property(ctx, userdata, n.as_ptr(), format) }; + if rc < 0 { + eprintln!("warning: {}", err(ctx, &format!("mpv_observe_property({name})"), rc)); + } + rc >= 0 +} + +/// One heartbeat line to stderr. Runs on the event thread, off the decode +/// path. Takes no `*mut MpvHandle` -- see the module doc on +/// `dexd::heartbeat` for why that absence is the point of F9: with no +/// mpv handle in scope, it is type-level impossible for this function to +/// call into mpv. Every value it prints was already learned from an event. +fn emit_heartbeat( + started: Instant, + last_position: Option<(f64, Instant)>, + frame_drops: ObservedCounter, + vo_delayed: ObservedCounter, + watchdog_pings_dropped: Option, +) { + eprintln!( + "{}", + HeartbeatSnapshot { + wraps: WRAP_COUNT.load(Ordering::Relaxed), + uptime_secs: started.elapsed().as_secs(), + temp_millicelsius: read_temp_millicelsius(), + position: last_position.map(|(secs, at)| PositionSample { + secs, + age_secs: at.elapsed().as_secs(), + }), + frame_drops, + vo_delayed, + watchdog_pings_dropped, + } + .render() + ); +} + +/// F10: the live systemd-watchdog resources held across the whole run -- +/// the non-blocking socket (opened once, never recreated) plus the address +/// it pings and a running count of drops for the heartbeat line. `None` +/// (via [`setup_watchdog`]) for every non-systemd run: Mac dev, bench tmux, +/// CI, any manual invocation off systemd -- see `dexd::watchdog`'s +/// module doc. +/// +/// This struct -- and the actual `send_to_addr` call site -- deliberately +/// live HERE, in the driver, not in the library: `dexd::watchdog` +/// already owns everything decidable (the handshake, the address +/// construction, the non-blocking send wrapper, all unit-tested with real +/// sockets under `cargo test --lib`, see that module's tests). What is left +/// is process-lifetime resource ownership and the "when do we call this" +/// wiring -- exactly what `main.rs` already does for every other piece of +/// mpv-facing state in this program (`health`, `frame_drops`, `vo_delayed`), +/// per `lib.rs`'s own "main.rs is a thin driver" principle. +struct WatchdogRuntime { + socket: UnixDatagram, + addr: UnixSocketAddr, + pings_dropped: u64, +} + +/// One policy for every "resolve() said Armed but the ping socket cannot be +/// established" arm of [`setup_watchdog`]. Which way it goes depends on +/// whether systemd's kill timer is demonstrably running (`$WATCHDOG_USEC` +/// present -- systemd exports it exactly when `WatchdogSec=` is configured): +/// +/// - **Timer armed: `exit(1)` now.** Nothing this process logs can disarm +/// the timer on systemd's side -- "running without pings" under an armed +/// `WatchdogSec=` means being SIGABRT-killed every window, forever (a ~2 s +/// black hiccup every 3 min under the shipped unit), while a +/// "watchdog DISABLED" journal line actively hides the cause of every one +/// of those kills. `Restart=always` + `RestartSec=2` retries in seconds, +/// and this failure class (address resolution, fd exhaustion, +/// `set_nonblocking`) is transient -- a clean fast retry strictly beats a +/// `WatchdogSec`-cadence kill loop with a lying journal line. +/// - **No timer: run without pings.** Nothing will kill us for not pinging, +/// so pings are genuinely pointless; be loud once and play the asset. +fn watchdog_setup_failed( + cause: &str, + kill_timer_armed: bool, + window_secs: Option, +) -> Option { + if kill_timer_armed { + let window = + window_secs.map(|w| format!("{w}s")).unwrap_or_else(|| "WatchdogSec".to_string()); + eprintln!( + "error: dexd: watchdog: {cause} -- systemd's WatchdogSec timer IS armed \ + ($WATCHDOG_USEC is set) and cannot be disarmed from inside this process: without \ + pings, systemd would SIGABRT this process every {window} while it plays normally. \ + Exiting now so Restart= retries cleanly instead." + ); + std::process::exit(1); + } + eprintln!( + "warning: dexd: watchdog: {cause} -- no WatchdogSec timer is armed ($WATCHDOG_USEC \ + absent), so pings would prove nothing; running WITHOUT watchdog pings for this run" + ); + None +} + +/// Resolve the systemd watchdog handshake and, if armed, open the +/// non-blocking socket it needs. Failures along the way (address resolution, +/// socket creation, `set_nonblocking`) follow [`watchdog_setup_failed`]'s +/// policy: they only downgrade to "no pings" when systemd's own kill timer +/// is NOT running -- when it is, no in-process downgrade exists (the timer +/// keeps counting regardless of what we log), so the process exits for a +/// clean fast retry rather than limping into a guaranteed +/// `WatchdogSec`-cadence kill loop. +fn setup_watchdog(tick_secs: u64) -> Option { + let env = WatchdogEnv::from_process_env(); + // systemd exports $WATCHDOG_USEC exactly when WatchdogSec= is configured + // on the unit -- its presence means a kill timer is counting RIGHT NOW, + // no matter what this process does or logs about its own pings. + let kill_timer_armed = env.watchdog_usec.is_some(); + let decision = watchdog::resolve(&env, std::process::id(), tick_secs); + let (addr, window_secs, warning) = match decision { + WatchdogDecision::Inert(reason) => { + eprintln!("dexd: watchdog: inert ({reason})"); + // The pid-mismatch arm is a NORMAL inert condition when no timer + // is armed -- but under an armed WatchdogSec= it is a death + // sentence on a schedule: systemd expects pings from the unit's + // main pid, this process (rightly) refuses to ping under someone + // else's identity, and the timer fires every window regardless. + // Say so, so the journal explains the SIGABRT kills that follow + // (e.g. a future edit wrapping ExecStart in a shell). + if kill_timer_armed + && matches!(reason, watchdog::InertReason::WatchdogPidMismatch { .. }) + { + eprintln!( + "warning: dexd: watchdog: $WATCHDOG_USEC is set, so systemd's WatchdogSec \ + timer IS armed and expects pings from the unit's MAIN pid -- with none \ + arriving, systemd will kill this process every watchdog window; expect a \ + kill/restart loop until the unit is fixed (is ExecStart wrapped in a shell?)" + ); + } + return None; + } + WatchdogDecision::Armed { addr, window_secs, warning } => (addr, window_secs, warning), + }; + + let sockaddr = match watchdog::socket_addr(&addr) { + Ok(a) => a, + Err(e) => { + return watchdog_setup_failed( + &format!("cannot resolve $NOTIFY_SOCKET address ({addr:?}): {e}"), + kill_timer_armed, + window_secs, + ); + } + }; + let socket = match UnixDatagram::unbound() { + Ok(s) => s, + Err(e) => { + return watchdog_setup_failed( + &format!("UnixDatagram::unbound failed: {e}"), + kill_timer_armed, + window_secs, + ); + } + }; + // LOAD-BEARING: see dexd::watchdog's module doc §"no new blocking + // call" -- a blocking send against a full receiver queue would be a new + // way for THIS feature to hang the event thread, exactly the class of + // bug F9 already had to fix once for the heartbeat. + if let Err(e) = socket.set_nonblocking(true) { + return watchdog_setup_failed( + &format!("set_nonblocking failed: {e} (refusing a socket that could block the event thread)"), + kill_timer_armed, + window_secs, + ); + } + + let window = window_secs.map(|w| format!("{w}s")).unwrap_or_else(|| "unknown".to_string()); + eprintln!("dexd: watchdog: armed (window {window}, ping cadence {tick_secs}s)"); + if let Some(w) = warning { + eprintln!("warning: dexd: watchdog: {w}"); + } + + Some(WatchdogRuntime { socket, addr: sockaddr, pings_dropped: 0 }) +} + +/// Whether an `MPV_EVENT_END_FILE` with this `reason` is an expected +/// teardown half of F1's in-place recovery -- its own `loadfile ... +/// replace` command(s) -- rather than a real failure. Pure and separately +/// tested (unlike the rest of the event loop, which needs libmpv) so this +/// one condition -- the entire fix for the regression where every recovery +/// attempt killed the process on its own first step -- cannot silently +/// break again without a failing test. See `MPV_END_FILE_REASON_STOP`'s doc +/// comment for the mpv behaviour this encodes. +/// +/// `recovery_stops_pending` is a COUNT, not a bool: the organic health-check +/// tick and T7's forced probe can both issue a recovery in the same loop +/// iteration (main.rs's tick block runs before the trigger block), and +/// `mpv_command_async` only QUEUES a loadfile against a core that may still +/// be busy from a still-wedged episode, so a second recovery can be issued +/// (and accepted) before the first attempt's stop has been observed. Two +/// in-flight recoveries produce two END_FILE(reason=stop) events; a bool can +/// absorb only the first and would treat the second -- an entirely expected +/// teardown -- as fatal. See PLAN.md's F1 addendum ("overlapping tier-0 +/// recoveries") for the confirmed scenario this fixes. +fn is_expected_recovery_stop(recovery_stops_pending: u32, reason: c_int) -> bool { + recovery_stops_pending > 0 && reason == MPV_END_FILE_REASON_STOP +} + +/// Act on a [`HealthAction`], whatever produced it. Shared by the organic +/// health-check tick and T7's `--force-recovery-after-secs` bench probe +/// (`HealthMonitor::force_recovery`) specifically so a forced probe drives +/// the EXACT SAME mpv-facing mechanics -- `mpv_command_async(loadfile ... +/// replace)`, then incrementing `recovery_stops_pending` so the resulting +/// `END_FILE(reason=stop)` is absorbed rather than treated as fatal (see +/// `is_expected_recovery_stop`) -- that a real stall would. That identity is +/// the point of T7: it is what lets a bench probe stand in for a real +/// stall's recovery path at all. `reason` is only the situational log +/// prefix; the two callers differ in WHY a recovery is due, not in what +/// happens once it is. +fn act_on_health_action( + ctx: *mut MpvHandle, + action: HealthAction, + reason: &str, + last_position: &mut Option<(f64, Instant)>, + recovery_stops_pending: &mut u32, +) { + match action { + HealthAction::Healthy => {} + HealthAction::AttemptRecovery { attempt, max } => { + eprintln!( + "dexd: health check: {reason} -- attempting in-place recovery \ + {attempt}/{max} (re-issuing loadfile: restarts demux+decode and forces \ + a VO reconfigure; does not tear down/re-init the DRM/GPU context itself \ + -- see PLAN.md's F1 addendum)" + ); + // Mirror HealthMonitor's own baseline reset (health.rs: + // `self.last_position = None`, shared by tick's AttemptRecovery + // arm and force_recovery via issue_recovery_or_escalate). + // Without this, the driver would keep feeding the STALE + // pre-recovery position back into the next tick, which the + // monitor -- now comparing against its own `None` baseline -- + // would misread as a fresh first sample (i.e. progress), buying + // a spurious "Healthy" tick that stretches the real escalation + // timeline. Since F9, `last_position` also carries the sample's + // `Instant` in the same tuple, so this one line clears + // `pos-age=`'s staleness clock too. + *last_position = None; + // mpv_command_async, not mpv_command: this call runs on the + // SAME event thread that also has to keep detecting every fatal + // event, and a synchronous command could block that thread + // against a wedged core -- see dexd::health's module doc. + let cmd_loadfile = CString::new("loadfile").unwrap(); + let cmd_url = CString::new("loop://endless").unwrap(); + let cmd_replace = CString::new("replace").unwrap(); + let argv: [*const c_char; 4] = [ + cmd_loadfile.as_ptr(), + cmd_url.as_ptr(), + cmd_replace.as_ptr(), + std::ptr::null(), + ]; + let rc = unsafe { mpv_command_async(ctx, RECOVERY_COMMAND_USERDATA, argv.as_ptr()) }; + if rc < 0 { + eprintln!("warning: {}", err(ctx, "mpv_command_async(loadfile)", rc)); + } else { + // Only now: the command was actually queued, so mpv WILL + // deliver an END_FILE(reason=stop) for the file being + // replaced (see MPV_END_FILE_REASON_STOP's doc comment) -- + // that event must be absorbed, not treated as the fatal + // failure it would otherwise look like. If mpv_command_async + // itself failed (above), no such event is coming, so the + // count must NOT be incremented. Saturating: bounded in + // practice by MAX_RECOVERY_ATTEMPTS (the cumulative budget + // this same command draws from), so saturation never + // actually engages -- it is here so a future change to that + // relationship fails safe (an undercount that stays fatal) + // rather than wrapping into a silent lie. + *recovery_stops_pending = recovery_stops_pending.saturating_add(1); + } + } + HealthAction::Escalate => { + eprintln!( + "dexd: FATAL: tier-0 self-healing exhausted its recovery budget \ + ({MAX_RECOVERY_ATTEMPTS} attempt(s)) with no progress ({reason}) -- \ + exiting so the supervisor restarts (tier 1)" + ); + // std::process::exit, not `break` into the shared + // mpv_terminate_destroy() teardown at the end of main: Escalate + // fires precisely because the core looks wedged (or, for a + // forced probe, to prove the SAME exit path an organic + // escalation would take), and mpv_terminate_destroy + // synchronously joins mpv's own threads -- a core that cannot + // advance time-pos may not be able to complete that join + // either, which would block the one exit path whose entire job + // is to let the supervisor take over. Per design principle 4 + // (the mains switch IS the shutdown path), exiting abruptly + // here is not a shortcut, it is correct. + // + // F10 INVARIANT: this `eprintln!` line above -- or anything + // else added to this arm before `exit(1)` -- can in principle + // block (e.g. against a wedged journald), and that possibility + // is EXACTLY what the systemd watchdog (`WatchdogSec=` in + // deploy/dexd.service, dexd::watchdog) exists to catch: + // if this arm hangs here, the tick loop never completes another + // iteration, so no further WATCHDOG=1 pings are sent, and + // systemd's timer fires. Do NOT add a "final ping" to this arm + // to try to look more alive on the way out -- that would reset + // the watchdog's countdown right before the one hang this + // feature exists to catch, defeating it. See + // dexd::watchdog's module doc for the full argument. + std::process::exit(1); + } + } +} + +fn usage() -> ! { + eprintln!( + "usage: dexd [] [--fps ] [--mode WxH@R] [--bench-no-sidecar] [--no-defaults] [--opt K=V ...] + + raw Annex-B HEVC elementary stream, looped endlessly. + OPTIONAL in a deployment: the exhibit config's `asset` + key names it, which is what lets several assets sit in + /opt/dex with the exhibit choosing one. Given here too, + it must AGREE with the config or startup refuses, naming + both. REQUIRED with --bench-no-sidecar, which consults no + config. If neither names an asset, startup refuses rather + than guessing an artwork. + .json ingest sidecar, REQUIRED: {{\"fps\":\"30\",\"sha256\":\"<64 hex>\"}} + fps comes from it; the sha256 must match the asset bytes + --fps F optional cross-check; must equal the sidecar fps exactly + --mode WxH@R cross-check against the exhibit config's display_mode (F6); + optional alongside --bench-no-sidecar, where it is the only + source instead (defaults to auto there) + --exhibit-config PATH + F6: path to the exhibit config. Default: whichever of + {DEFAULT_EXHIBIT_CONFIG_PATHS:?} exists -- exactly one may, + and two at once is refused rather than resolved by + precedence. THE EXTENSION DECIDES THE PARSER: .json is + strict JSON, .yaml/.yml is YAML; same schema either way. + Binds the display mode, the expected KMS force, and the + connector -- see man dex-exhibit-apply. + --bench-no-sidecar BENCH ONLY: skip the sidecar AND the exhibit config, take + --fps/--mode as given + --force-recovery-after-secs N + T7 BENCH ONLY: force a tier-0 in-place recovery N seconds after + the loadfile request (NOT N seconds of confirmed playback -- + decode startup takes time too), whether or not anything has + stalled. For a live-fire run meant to catch mid-playback issues + rather than startup ones, pick N with margin over real decode + startup latency. REQUIRES --bench-no-sidecar (refused otherwise) + so it can never fire against a real, sidecar-bound deployment + asset. + --bench-wedge-after-secs N + F10 BENCH ONLY: N seconds after startup, deliberately hang the + event thread FOREVER -- simulates the one hazard class F1/F9 + cannot see (an event-thread hang outside any mpv call), to prove + whether a systemd watchdog (WatchdogSec=) actually fires and + restarts this process. The process never recovers on its own once + this fires; only an external actor (systemd, or a test harness's + own kill) can end it. REQUIRES --bench-no-sidecar (refused + otherwise) so it can never fire against a real, sidecar-bound + deployment asset. + --opt K=V pass an extra mpv option (repeatable) + --no-defaults omit the built-in Pi 4 zero-copy option set + +exit codes: 2 = refused before playback (bad invocation/asset/sidecar/display; fix and redeploy) + 1 = playback/runtime failure (the supervisor restarts)" + ); + std::process::exit(2) +} + +/// Locate `/sys/class/drm/card-/modes`. The card number is not +/// hardcoded: vc4/v3d probe order makes it unstable across kernel versions +/// (the same reason `deploy/dex-wait-hdmi` globs it rather than assuming +/// `card1`). Returns `None` if no such entry exists -- e.g. no DRM at all (a +/// CI container), or a mistyped connector name. +fn find_sysfs_modes_path(connector: &str) -> Option { + let suffix = format!("-{connector}"); + let entries = fs::read_dir("/sys/class/drm").ok()?; + for entry in entries.flatten() { + let name = entry.file_name(); + let name = name.to_string_lossy(); + if name.starts_with("card") && name.ends_with(suffix.as_str()) { + let modes_path = entry.path().join("modes"); + if modes_path.is_file() { + return Some(modes_path.to_string_lossy().into_owned()); + } + } + } + None +} + +fn main() -> ExitCode { + // Identify the build before anything can fail: a field journal that + // starts with an unidentifiable process is undebuggable weeks later. + eprintln!( + "dexd {} ({})", + env!("CARGO_PKG_VERSION"), + env!("DEX_GIT_HASH") + ); + + let args: Vec = env::args().skip(1).collect(); + // NO `if args.is_empty() { usage() }`. Since F6 moved the asset into the + // exhibit config, an EMPTY argv is the normal deployment invocation -- + // `ExecStart=/usr/bin/dexd`, everything else in + // /etc/dex/exhibit.{yaml,json}. That guard survived the asset change for + // about ten minutes and would have put the shipped unit into a permanent + // exit-2 restart loop on the device while every Mac-side test passed, + // because every test passes arguments. A bare `dexd` now proceeds to + // the config, and refuses there if the config cannot answer. + + let mut path: Option = None; + let mut cli_fps: Option = None; + let mut mode: Option = None; + let mut extra: Vec<(String, String)> = Vec::new(); + let mut defaults = true; + let mut bench_no_sidecar = false; + let mut force_recovery_after_secs: Option = None; + let mut bench_wedge_after_secs: Option = None; + let mut exhibit_config_path: Option = None; + // Test-only override so integration tests can supply a synthetic kernel + // cmdline instead of depending on the actual host's /proc/cmdline, which + // varies by machine (a CI container has no video= token at all; the real + // bench Pi, once F6 is deployed there, always does). Never printed in + // usage(): a real deployment always reads the real /proc/cmdline. + let mut proc_cmdline_path: Option = None; + + let mut i = 0; + while i < args.len() { + match args[i].as_str() { + "--fps" => { + i += 1; + // A missing value here (flag is the last token -- an edited + // systemd unit, a line-continuation typo) must refuse loudly, + // not evaporate: a silently-dropped --fps falls through to + // "no cross-check", and a silently-dropped --mode falls + // through to the connector-preferred mode -- wrong cadence + // or wrong resolution, forever, with no error. + let Some(v) = args.get(i) else { usage() }; + cli_fps = Some(v.clone()); + } + "--mode" => { + i += 1; + let Some(v) = args.get(i) else { usage() }; + mode = Some(v.clone()); + } + "--exhibit-config" => { + i += 1; + let Some(v) = args.get(i) else { usage() }; + exhibit_config_path = Some(v.clone()); + } + "--proc-cmdline" => { + i += 1; + let Some(v) = args.get(i) else { usage() }; + proc_cmdline_path = Some(v.clone()); + } + "--no-defaults" => defaults = false, + "--bench-no-sidecar" => bench_no_sidecar = true, + "--force-recovery-after-secs" => { + i += 1; + // Same "refuse loudly, never evaporate" discipline as + // --fps/--mode above: a dropped value here would silently + // leave T7 disarmed, which is harmless, but a MALFORMED + // value (e.g. a typo'd flag argument) must not be read as + // "flag absent" either -- usage() either way. + let Some(v) = args.get(i) else { usage() }; + let Ok(n) = v.parse::() else { usage() }; + force_recovery_after_secs = Some(n); + } + "--bench-wedge-after-secs" => { + i += 1; + // Same "refuse loudly, never evaporate" discipline as + // --force-recovery-after-secs above. + let Some(v) = args.get(i) else { usage() }; + let Ok(n) = v.parse::() else { usage() }; + bench_wedge_after_secs = Some(n); + } + "--opt" => { + i += 1; + let kv = args.get(i).cloned().unwrap_or_default(); + match kv.split_once('=') { + Some((k, v)) => extra.push((k.to_string(), v.to_string())), + None => usage(), + } + } + "-h" | "--help" => usage(), + s if !s.starts_with('-') && path.is_none() => path = Some(s.to_string()), + _ => usage(), + } + i += 1; + } + + // No `let Some(path) = path else { usage() }` any more: since F6 the asset + // may come from the exhibit config instead, so "which asset" is a + // RESOLUTION (exhibit::resolve_asset, below, after the config is read) and + // not an argv shape. usage() here would refuse the normal deployment + // invocation -- `ExecStart=/usr/bin/dexd` with no path at all. + + // T7 (PLAN.md) -- the "impossible to enable accidentally in a + // deployment" requirement, enforced as a gate rather than left to + // operator discipline. `--force-recovery-after-secs` REQUIRES + // `--bench-no-sidecar`. This is not an arbitrary pairing: it ties the + // bench-only recovery probe to the SAME escape hatch that already keeps + // `--bench-no-sidecar` out of every real deployment (deploy/dexd.service + // never passes it -- a real asset is bound to its ingest sidecar, full + // stop), so a live-fire probe can never end up armed against a gallery + // show by an operator pasting a bench command line into the wrong + // place. Checked here, before the asset is even read, so the refusal is + // unconditional on CLI shape alone -- it does not depend on whether a + // sidecar happens to exist on disk. + if force_recovery_after_secs.is_some() && !bench_no_sidecar { + eprintln!( + "error: --force-recovery-after-secs requires --bench-no-sidecar -- it is a \ + T7 BENCH-ONLY live-fire probe (PLAN.md) that forces a tier-0 in-place \ + recovery on a timer, whether or not anything has actually stalled, and \ + must never be armed against what could be a real, sidecar-bound \ + deployment asset. Add --bench-no-sidecar --fps to run it on a bench, \ + or drop --force-recovery-after-secs to run normally." + ); + return ExitCode::from(2); + } + + // F10 (PLAN.md) -- the identical "impossible to enable accidentally in + // a deployment" gate as T7 above, for the same reason: a probe that + // deliberately hangs the event thread forever must never be reachable + // against a real, sidecar-bound show, however it got pasted into a + // command line. + if bench_wedge_after_secs.is_some() && !bench_no_sidecar { + eprintln!( + "error: --bench-wedge-after-secs requires --bench-no-sidecar -- it is an F10 \ + BENCH-ONLY probe (PLAN.md) that deliberately hangs the event thread forever to \ + prove whether a systemd watchdog actually fires, and must never be armed against \ + what could be a real, sidecar-bound deployment asset. Add --bench-no-sidecar \ + --fps to run it on a bench, or drop --bench-wedge-after-secs to run normally." + ); + return ExitCode::from(2); + } + + // F6 -- the exhibit display config. Runs BEFORE the asset is read: these + // are the cheapest gates in the program and must not depend on asset + // presence (a wrong-panel install is worth catching even if the asset + // path is also wrong). Order: config parse -> cmdline gate -> sysfs mode + // pre-flight. See src/exhibit.rs module docs and PLAN.md's F6 entry. + let found = if bench_no_sidecar { + None + } else { + match load_exhibit_config(exhibit_config_path.as_deref(), &DEFAULT_EXHIBIT_CONFIG_PATHS) { + Ok(found) => found, + Err(e) => { + eprintln!("error: {e}"); + return ExitCode::from(2); + } + } + }; + // When nothing was found there is no path to name, so the startup line + // falls back to the installed default -- which is also the file + // resolve_display's refusal tells the operator to create. + let (exhibit_config, exhibit_config_path) = match found { + Some((cfg, path)) => (Some(cfg), path), + None => (None, DEFAULT_EXHIBIT_CONFIG_PATH.to_string()), + }; + + let display = match resolve_display(exhibit_config.as_ref(), mode.as_deref(), bench_no_sidecar) + { + Ok(d) => d, + Err(e) => { + eprintln!("error: {e}"); + return ExitCode::from(2); + } + }; + + // The cmdline gate: keeps the exhibit config and the KMS-layer `video=` + // token honest with each other. Skipped under the bench flag, since a + // ~/bench build runs on hand-managed boot state by definition (§2.3/§3.3 + // of the F6 design) -- the whole POINT of --bench-no-sidecar is running + // outside the deployment config's authority. + if !bench_no_sidecar { + let proc_cmdline_path = proc_cmdline_path + .clone() + .unwrap_or_else(|| "/proc/cmdline".to_string()); + match fs::read_to_string(&proc_cmdline_path) { + Ok(cmdline) => { + if let Err(e) = check_cmdline_matches(&cmdline, &display.connector, &display.kms_force) + { + eprintln!("error: {e}"); + return ExitCode::from(2); + } + } + Err(e) => { + eprintln!("error: cannot read {proc_cmdline_path}: {e}"); + return ExitCode::from(2); + } + } + } + + // The sysfs mode pre-flight: catches "the configured resolution is not + // even in this connector's mode list" (the wrong-panel case) before any + // asset/sidecar/NAL gate runs. Plain-text sysfs, no privilege, no libmpv + // -- `/sys/class/drm/card*-/modes`, one WxH per line. Runs + // whenever a specific mode is requested, bench or not: it is a real + // hardware-agreement check that reads nothing from /etc or /proc, so + // nothing about "bench" exempts it (unlike the cmdline gate above, which + // is specifically about /etc vs /boot agreement). + if let Some(want_wh) = mode_resolution(&display.display_mode) { + match find_sysfs_modes_path(&display.connector) { + Some(modes_path) => match fs::read_to_string(&modes_path) { + Ok(modes_text) => { + if !sysfs_modes_contains(&modes_text, want_wh) { + let offered: Vec<&str> = modes_text.lines().map(str::trim).collect(); + eprintln!( + "error: display_mode {want_wh:?} is not among the modes {} \ + ({modes_path}) offers: {offered:?} -- wrong panel, or the \ + cmdline force (if any) has not taken effect yet (reboot?)", + display.connector + ); + return ExitCode::from(2); + } + } + Err(e) => { + eprintln!("error: cannot read {modes_path}: {e}"); + return ExitCode::from(2); + } + }, + None => { + eprintln!( + "error: display_mode {want_wh:?} requested for connector {}, but no sysfs \ + modes list was found for it under /sys/class/drm -- is the connector name \ + correct, and is DRM available?", + display.connector + ); + return ExitCode::from(2); + } + } + } + + eprintln!( + "dexd: display {} ({}), connector {}, kms-force {}", + display.display_mode, + match display.source { + DisplaySource::Config => format!("exhibit config {exhibit_config_path}"), + DisplaySource::Bench => "BENCH OVERRIDE, unbound".to_string(), + }, + display.connector, + if bench_no_sidecar { + "not checked (bench)".to_string() + } else { + display.kms_force.clone() + }, + ); + + // F6 -- WHICH asset. Resolved from the exhibit config and/or the command + // line by the same decision table shape as the display and the frame rate + // (exhibit::resolve_asset). This is what lets several assets sit in + // /opt/dex with the exhibit choosing one, instead of ExecStart naming a + // single hardcoded path. + let asset = match resolve_asset(exhibit_config.as_ref(), path.as_deref(), bench_no_sidecar) { + Ok(a) => a, + Err(e) => { + eprintln!("error: {e}"); + return ExitCode::from(2); + } + }; + let path = asset.path; + eprintln!( + "dexd: asset {path} ({})", + match asset.source { + AssetSource::Config => format!("exhibit config {exhibit_config_path}"), + // Loud on purpose. A show must run off the config; a CLI path is + // either a bench or a hand-started one-off, and either way the + // journal should say so rather than let someone read a + // hand-invoked run as evidence about the deployed one. + AssetSource::Cli => "COMMAND LINE, not the exhibit config".to_string(), + } + ); + + // Read the loop once. These are small (1.3 MB at 1080p, 14.8 MB at 4K for a + // 3 s card) and holding it in memory removes the filesystem from the hot + // path: no re-open, no page-cache dependency, no I/O stall at the wrap. + let payload = match fs::read(&path) { + Ok(p) if !p.is_empty() => p, + Ok(_) => { + eprintln!("error: {path} is empty"); + return ExitCode::from(2); + } + Err(e) => { + eprintln!("error: cannot read {path}: {e}"); + return ExitCode::from(2); + } + }; + let leaked: &'static [u8] = Box::leak(payload.into_boxed_slice()); + + // F3 — bind the asset to its ingest sidecar. A raw Annex-B stream has no + // timestamps: a WRONG --fps plays slow/fast forever with zero errors and + // every metric nominal — the one failure that is undetectable by + // construction. So the frame rate travels WITH the asset, bound by a + // sha256, and an unbound asset is refused. `--bench-no-sidecar --fps F` + // is the deliberate two-flag bench escape hatch. + let sidecar_path = format!("{path}.json"); + let sidecar: Option = if bench_no_sidecar { + None + } else { + match fs::read_to_string(&sidecar_path) { + Ok(text) => match Sidecar::from_json(&text) { + Ok(s) => Some(s), + Err(e) => { + eprintln!("error: {sidecar_path}: {e}"); + return ExitCode::from(2); + } + }, + Err(e) => { + eprintln!( + "error: cannot read sidecar {sidecar_path}: {e}\n\ + an asset without its ingest sidecar is unbound (fps would be a \ + guess); re-ingest to produce it, or use --bench-no-sidecar \ + --fps on a bench" + ); + return ExitCode::from(2); + } + } + }; + + let (fps, fps_source) = match resolve_fps( + sidecar.as_ref().map(|s| s.fps.as_str()), + cli_fps.as_deref(), + bench_no_sidecar, + ) { + Ok(r) => r, + Err(e) => { + eprintln!("error: {e}"); + return ExitCode::from(2); + } + }; + + if let Some(s) = &sidecar { + if let Err(e) = verify_payload(leaked, s) { + eprintln!("error: {path}: {e}"); + return ExitCode::from(2); + } + } + + // F6 addendum (2026-08-17 review): display_mode "auto" skips the sysfs + // pre-flight by construction (no mode to check against), which converts + // fail-closed into fail-silent on the one config the .deb ships by + // default -- a forgotten /etc/dex/exhibit.json edit on hardware that + // builds no 4K mode unforced (the Cam Link case) plays the artwork at + // whatever the connector negotiates, for weeks, with every metric green. + // The sidecar is hash-bound to the asset and already names its + // resolution, so at least SAY SO when the connector cannot even offer + // it. A warning, not a refusal: whether "auto" should stay the factory + // default at all (vs an explicit "unset" that refuses like a missing + // config) is one of the parked F6 design questions -- see PLAN.md. + if display.display_mode == "auto" { + if let Some((w, h)) = sidecar.as_ref().and_then(|s| s.width.zip(s.height)) { + let want = format!("{w}x{h}"); + if let Some(modes_path) = find_sysfs_modes_path(&display.connector) { + if let Ok(modes_text) = fs::read_to_string(&modes_path) { + if !sysfs_modes_contains(&modes_text, &want) { + let offered: Vec<&str> = modes_text.lines().map(str::trim).collect(); + eprintln!( + "warning: display_mode is \"auto\" and the asset is {want} (per its \ + sidecar), but connector {} offers only {offered:?} ({modes_path}) \ + -- KMS will drive whatever fallback it negotiates and the artwork \ + will play at the WRONG resolution with every metric green. If this \ + display needs a forced mode to build {want} (e.g. the Cam Link \ + builds no 4K mode unforced), set display_mode and kms_force in \ + /etc/dex/exhibit.json, run 'sudo dex-exhibit-apply', and reboot", + display.connector + ); + } + } + } + } + } + + // F4 — validate the leading NALs. The wrap is only seamless because byte + // 0 begins VPS/SPS/PPS + IDR; a wrong-but-intact asset (open GOP, no + // leading IDR, not Annex-B at all) would glitch at every wrap, silently, + // ~29k times/day. Truncation is caught by the F3 hash above; this catches + // shape. Runs in bench mode too — the premise holds there as well. + if let Err(e) = validate_leading_nals(leaked) { + eprintln!("error: {path}: {e}"); + return ExitCode::from(2); + } + + eprintln!( + "dexd: {} bytes, fps {fps} ({}), looping endlessly", + leaked.len(), + match fps_source { + FpsSource::Sidecar => "sidecar", + FpsSource::BenchOverride => "BENCH OVERRIDE, unbound", + } + ); + + if let Some(n) = force_recovery_after_secs { + // Loud on purpose (principle 2): the gate above makes this + // impossible to reach without --bench-no-sidecar already having + // been accepted, but a run that silently, quietly forces its own + // recovery mid-show is exactly the kind of surprise that belongs in + // the journal in giant letters, not inferred later from a + // recovery log line with no explanation of why it fired. + eprintln!( + "warning: BENCH ONLY (T7): --force-recovery-after-secs={n} is ARMED -- this \ + run will FORCE a tier-0 in-place recovery {n}s after the loadfile request was \ + queued (NOT {n}s of confirmed playback -- decode startup can itself take a \ + few seconds, so a small N can fire during startup rather than steady \ + playback), whether or not anything has actually stalled. Never pass this \ + flag on a real deployment asset (see PLAN.md's T7 entry)." + ); + } + + if let Some(n) = bench_wedge_after_secs { + // Same "loud on purpose" discipline as T7 above. + eprintln!( + "warning: BENCH ONLY (F10 wedge probe): --bench-wedge-after-secs={n} is ARMED -- \ + this run will deliberately hang the event thread FOREVER {n}s after startup, \ + simulating F10's hazard class (an event-thread hang outside any mpv call). \ + Never pass this flag on a real deployment asset (see PLAN.md's F10 entry)." + ); + } + + let ctx = unsafe { mpv_create() }; + if ctx.is_null() { + eprintln!("error: mpv_create failed"); + return ExitCode::FAILURE; + } + + let mut opts: Vec<(String, String)> = Vec::new(); + if defaults { + // Measured on Pi 4 / trixie. The decoder emits Broadcom SAND-tiled NV12 + // and the display can scan SAND out natively, but ONLY straight onto a + // KMS plane -- every other path detiles (CPU: 14.3 fps, GL: 5 fps). + // `drmprime-overlay` is the interop that puts the frame on a plane; + // plain `drmprime` imports into GL and is 2x slower. + for (k, v) in [ + ("vo", "gpu"), + ("hwdec", "drm"), + ("gpu-context", "drm"), + ("gpu-api", "opengl"), + ("gpu-hwdec-interop", "drmprime-overlay"), + // Video on the primary plane, mpv's GL/OSD surface on the overlay -- + // SWAPPED from mpv's defaults, deliberately: it keeps the 4K video + // off the V3D render path entirely. Caveat: mpv sets ZPOS only on + // the video plane, so video-under-GL visibility relies on vc4's + // default plane ordering rather than anything mpv guarantees. + // Verified on this Pi 4 + kernel; re-verify after a kernel upgrade + // or on any other DRM driver. + ("drm-draw-plane", "overlay"), + ("drm-drmprime-video-plane", "primary"), + ("video-sync", "display-resample"), + // Without this, a decoder that cannot use the hardware path falls + // back to software SILENTLY and plays 4K30 at ~14 fps. Making it + // fatal turns an invisible performance collapse into an END_FILE + // error, which is handled and restartable. + ("hwdec-software-fallback", "no"), + ("fullscreen", "yes"), + ("osc", "no"), + ("input-default-bindings", "no"), + ("terminal", "no"), + // A raw elementary stream has no timestamps; mpv must generate them. + ("correct-pts", "no"), + // NOTE: this is belt-and-braces, not the load-bearing bound. mpv + // only runs its aggressive cache for streams flagged as network, + // and stream_cb streams are not; readahead here is governed by + // demuxer-readahead-secs instead. The flat memory measured over + // 3.5 h is due to that, not to this cap. + ("demuxer-max-bytes", "64MiB"), + // The actual prefetch depth. One second of decoded-ahead insurance + // across the wrap, where the whole gaplessness claim is decided. + ("demuxer-readahead-secs", "1.0"), + ] { + opts.push((k.to_string(), v.to_string())); + } + } + opts.push(("container-fps-override".into(), fps)); + // F6 -- the resolved exhibit display binding. "auto" means: pass no + // drm-mode at all, which is mpv's own documented default + // (drm-mode=preferred) -- see exhibit::resolve_display's docs on what + // "auto" precisely means. drm-connector is passed unconditionally: every + // ResolvedDisplay carries a connector (defaulted if not configured), so + // this is always explicit rather than relying on mpv's own connector + // pick, which the pre-F6 code never stated either way. + if display.display_mode != "auto" { + opts.push(("drm-mode".into(), display.display_mode.clone())); + } + opts.push(("drm-connector".into(), display.connector.clone())); + opts.extend(extra); + + for (k, v) in &opts { + if let Err(e) = set_opt(ctx, k, v) { + eprintln!("error: {e}"); + unsafe { mpv_terminate_destroy(ctx) }; + // A rejected option is deterministic given these inputs: the + // same asset + flags fail identically on every restart, so per + // the exit-code contract this is "bad invocation" (2) -- fix and + // redeploy -- not a runtime failure (1) that the supervisor's + // restart loop could ever resolve on its own. + return ExitCode::from(2); + } + } + + // Register `loop://` BEFORE initialize, so the protocol exists by the time + // the play command is issued. + // + // `user_data` is a LEAKED Box, not a pointer to a local. mpv keeps this + // pointer until mpv_terminate_destroy returns and may dereference it from + // its own threads at any point; aiming it at a stack slot in main() worked + // only because every exit path happens to tear mpv down first. That is UB + // the moment anything unwinds (a dev build does), and one refactor away + // from UB even in release. 16 bytes, leaked once, removes the hazard. + let proto = CString::new("loop").unwrap(); + let user_data = Box::into_raw(Box::new(leaked)) as *mut c_void; + let rc = unsafe { mpv_stream_cb_add_ro(ctx, proto.as_ptr(), user_data, Some(open_fn)) }; + if rc < 0 { + eprintln!("error: {}", err(ctx, "mpv_stream_cb_add_ro", rc)); + unsafe { mpv_terminate_destroy(ctx) }; + return ExitCode::FAILURE; + } + + // Without this, libmpv discards every diagnostic it produces: `terminal=no` + // is the libmpv default, so log output goes nowhere unless it is requested + // as events. For an appliance whose value is a performance property that + // functional testing cannot see, this is the difference between a field + // failure being diagnosable and being a mystery. + let lvl = CString::new("warn").unwrap(); + let rc = unsafe { mpv_request_log_messages(ctx, lvl.as_ptr()) }; + if rc < 0 { + eprintln!("warning: {}", err(ctx, "mpv_request_log_messages", rc)); + } + + let rc = unsafe { mpv_initialize(ctx) }; + if rc < 0 { + eprintln!("error: {}", err(ctx, "mpv_initialize", rc)); + unsafe { mpv_terminate_destroy(ctx) }; + return ExitCode::FAILURE; + } + + // F1 tier-0 health check: subscribe to time-pos so the event loop can + // later detect a stalled decode without ever polling the core + // synchronously -- see dexd::health's module doc for the full + // reasoning. This call is itself documented non-blocking; only the + // SUBSEQUENT samples, delivered as ordinary MPV_EVENT_PROPERTY_CHANGE + // events through the same wait loop already proven (by END_FILE and + // QUEUE_OVERFLOW handling) never to hang, are load-bearing. A failure + // to register is NOT fatal -- the health check is a best-effort safety + // net on top of a working player, not a gate the show depends on -- but + // per principle 2 it must be loud, never silent. + let health_check_registered = observe(ctx, HEALTH_CHECK_USERDATA, "time-pos", MPV_FORMAT_DOUBLE); + // Whether the health check is actually usable this run. Gates the tick + // loop below (`health: Option`) -- without this gate, a + // failed registration would leave `last_position` permanently `None`, + // and every tick would read as a stall forever, eventually issuing + // recovery commands (and, once the budget is exhausted, exiting) + // against a perfectly healthy player: the exact opposite of the + // "DISABLED" warning below. `observe()` already logged the underlying + // mpv error; this is the feature-specific consequence of that failure. + if !health_check_registered { + eprintln!( + "warning: tier-0 self-healing is DISABLED for this run (time-pos \ + subscription failed above); tier 1 (process restart on a fatal \ + event) still applies" + ); + } + + // F9: the heartbeat's two drop counters, subscribed the same way as + // time-pos above and for the same reason -- see dexd::heartbeat's + // module doc. Registered here, before `loadfile`, simply to group all of + // this program's mpv_observe_property calls at startup rather than + // scattering them -- registration order relative to loadfile does not + // affect completeness either way: mpv forces an initial notification for + // EVERY observer at the moment it registers (client.h), regardless of + // when that happens, so nothing is "missed" by a later subscription. + // That forced initial event for these two properties arrives promptly + // as format=NONE/data=NULL (no VO chain yet -- see the property-change + // handler below); the first real INT64 value is a second, later event + // once the VO chain exists. A failed registration is not fatal -- + // `observe()` already warned loudly -- it just means this run's + // heartbeat prints "off" for that counter for its whole life; no + // separate warning is needed beyond `observe`'s own, since the + // heartbeat repeats the fact every 10 minutes anyway. + let mut frame_drops = if observe(ctx, FRAME_DROPS_USERDATA, "frame-drop-count", MPV_FORMAT_INT64) { + ObservedCounter::observed() + } else { + ObservedCounter::unobserved() + }; + let mut vo_delayed = if observe(ctx, VO_DELAYED_USERDATA, "vo-delayed-frame-count", MPV_FORMAT_INT64) { + ObservedCounter::observed() + } else { + ObservedCounter::unobserved() + }; + + let cmd_loadfile = CString::new("loadfile").unwrap(); + let cmd_url = CString::new("loop://endless").unwrap(); + let argv: [*const c_char; 3] = [cmd_loadfile.as_ptr(), cmd_url.as_ptr(), std::ptr::null()]; + let rc = unsafe { mpv_command(ctx, argv.as_ptr()) }; + if rc < 0 { + eprintln!("error: {}", err(ctx, "loadfile", rc)); + unsafe { mpv_terminate_destroy(ctx) }; + return ExitCode::FAILURE; + } + + // Run until mpv stops playing. The stream is infinite by construction, so + // there is no benign way for playback to end: END_FILE means something + // failed, and it must be FATAL here. + // + // This is the single most important behaviour in the program, and it is not + // obvious. libmpv is not the CLI player: `mpv_create` enables idle mode by + // default (client.h), so a failed playback emits END_FILE and then sits in + // idle FOREVER -- it never emits SHUTDOWN. A loop that waits only for + // SHUTDOWN therefore blocks forever with the process alive, healthy to any + // supervisor, and the wall black. That is strictly worse than crashing. + // + // Compounding it: with vo=gpu the video output is created during file load, + // NOT during mpv_initialize. So every display-side failure -- projector not + // awake, no EDID, DRM master held by a getty -- passes both the initialize + // and loadfile return codes and lands here. Which is exactly the most + // likely failure in a gallery. + // + // So: exit non-zero and let the supervisor restart us. + let started = Instant::now(); + let mut last_heartbeat = Instant::now(); + // F1 tier-0 health check state. `last_position` is updated ONLY by the + // MPV_EVENT_PROPERTY_CHANGE handler below -- never read synchronously + // from mpv -- and fed to `health` on a fixed cadence. See + // dexd::health's module doc for the full policy and reasoning. The + // Instant travels WITH the position (one tuple, not two separately + // updated locals) so they cannot desync -- see dexd::heartbeat's + // `PositionSample` doc for why that is a struct-shape decision, not + // just a style one. + let mut last_position: Option<(f64, Instant)> = None; + // F10: resolve the systemd watchdog handshake and open its socket (if + // armed) BEFORE heartbeat #0, so that very first line already reports + // the real watchdog state instead of a stale default -- see + // dexd::watchdog's module doc and `setup_watchdog` above. + let mut watchdog_runtime = setup_watchdog(HEALTH_CHECK_SECS); + // Heartbeat #0: proves temperature reading and line formatting on every + // boot, and anchors the journal. It is NOT proof of the mpv property + // subscriptions (F9 removed the only call that could prove that + // synchronously): frame-drops/vo-delayed/pos all read "n/a" here by + // design, since nothing has been decoded yet. The on-device check that + // the subscriptions actually work belongs to the deploy checklist, not + // this line -- see PLAN.md's F9 entry ("must show numbers, not n/a"). + emit_heartbeat( + started, + last_position, + frame_drops, + vo_delayed, + watchdog_runtime.as_ref().map(|w| w.pings_dropped), + ); + let mut health: Option = if health_check_registered { + Some(HealthMonitor::new(MAX_RECOVERY_ATTEMPTS)) + } else { + None + }; + let mut last_health_check = Instant::now(); + // T7 (PLAN.md): armed only when the CLI gate above accepted + // --force-recovery-after-secs (which itself required --bench-no-sidecar). + // `None` here is the overwhelmingly common case -- every real deployment + // run -- and costs one `Option` check per loop iteration. + let mut force_recovery_trigger: Option = + force_recovery_after_secs.map(ForceRecoveryTrigger::new); + // F10 (PLAN.md): the bench-only wedge probe, armed only when the CLI + // gate below accepted --bench-wedge-after-secs (which itself requires + // --bench-no-sidecar, same escape hatch as T7). See its firing site + // below for what it proves and why. + let mut bench_wedge_trigger: Option = + bench_wedge_after_secs.map(ForceRecoveryTrigger::new); + // Counts recoveries issued whose matching END_FILE(reason=stop) has not + // yet been observed (bench-confirmed live, three independent reviews, + // 2026-08-15) -- see the MPV_EVENT_END_FILE handler below and PLAN.md's + // F1 addendum. A COUNT, not a single flag: the organic tick and T7's + // forced probe can both issue a recovery in the same loop iteration + // (mpv_command_async only queues against a core that may still be busy + // from a prior attempt), so more than one can be in flight at once, and + // each produces its own END_FILE(stop) to absorb. Nothing else in this + // program ever issues a command that produces a STOP-reason end-file, so + // "count > 0" is what tells "one of our own recoveries' expected + // teardowns" apart from an actual failure that happens to carry the same + // reason code. + let mut recovery_stops_pending: u32 = 0; + + let exit = ExitCode::SUCCESS; + loop { + // Wake at least every HEALTH_CHECK_SECS. In the healthy steady + // state mpv delivers frequent time-pos property-change events on + // its own, waking this thread without any help from the timeout -- + // but during an actual STALL, by definition NO such events arrive + // (that absence IS the stall signal; see dexd::health), so the + // timeout is what guarantees the health check still gets evaluated + // on schedule in precisely the one case that matters. A wake this + // cheap (drain one event, compare two numbers) on the event thread, + // separate from the decode/VO threads, costs nothing on the decode + // path. + let ev = unsafe { mpv_wait_event(ctx, HEALTH_CHECK_SECS as f64) }; + let id = unsafe { (*ev).event_id }; + let reply_userdata = unsafe { (*ev).reply_userdata }; + + if last_heartbeat.elapsed().as_secs() >= HEARTBEAT_SECS { + emit_heartbeat( + started, + last_position, + frame_drops, + vo_delayed, + watchdog_runtime.as_ref().map(|w| w.pings_dropped), + ); + last_heartbeat = Instant::now(); + } + + // F1's tick AND F10's ping share this cadence gate, but the ping is + // NOT nested inside `if let Some(h) = health` below -- see + // dexd::watchdog's module doc, "gate placement": if the + // time-pos subscription itself failed to register (health is + // `None`, near-zero probability), F1 is disabled but the player may + // still be perfectly healthy, and stopping pings in that mode would + // convert a merely-degraded run into a guaranteed watchdog kill + // loop. The ping is emitted AFTER the tick/act_on_health_action + // pair for this iteration has fully completed, per PLAN.md F10 §1 -- + // that ordering, plus F1's cumulative never-refilling recovery + // budget (dexd::health, "why the budget never resets"), is what + // makes this a real liveness criterion rather than "the process + // runs": a display-wedged player's tick sequence is FORCED, by + // construction, through silence -> stall -> <=3 budgeted recoveries + // -> Escalate -> process exit, so it can only ever emit a BOUNDED + // number of pings before either exiting (tier 1 already handles + // that) or -- on the one path that can still hang, e.g. `eprintln!` + // against a wedged journald on the Escalate arm itself, see that + // arm's own comment below -- simply stopping, which is exactly what + // the watchdog is here to catch. The loop cannot both hang and keep + // pinging. (Scope: that covers wedged-core/wedged-thread failures + // only -- a signal-level failure where time-pos advances with no + // photons on the wall pings forever, out of F10's scope by design; + // see dexd::watchdog's module doc, "Scope, stated precisely".) + if last_health_check.elapsed().as_secs() >= HEALTH_CHECK_SECS { + last_health_check = Instant::now(); + if let Some(h) = health.as_mut() { + let action = h.tick(last_position.map(|(secs, _)| secs)); + act_on_health_action( + ctx, + action, + &format!( + "no progress across 2 consecutive checks ({HEALTH_CHECK_SECS}s apart)" + ), + &mut last_position, + &mut recovery_stops_pending, + ); + } + if let Some(wd) = watchdog_runtime.as_mut() { + match watchdog::send_ping(&wd.socket, &wd.addr) { + PingOutcome::Sent => {} + PingOutcome::Dropped => wd.pings_dropped = wd.pings_dropped.saturating_add(1), + } + } + } + + // T7 (PLAN.md): the bench-only live-fire probe. Checked every + // iteration (two integer comparisons; `ForceRecoveryTrigger` is + // `None` and this whole block skipped on every real deployment run) + // rather than gated on the HEALTH_CHECK_SECS cadence above, so it + // does not additionally wait out however much of that ~10s window + // was already elapsed when `--force-recovery-after-secs` was + // reached. It can still be delayed up to HEALTH_CHECK_SECS in the + // worst case (mpv_wait_event's timeout bounds how often this loop + // body runs at all when nothing else is waking it) -- acceptable + // for a bench diagnostic whose job is to prove the mechanism works + // at all, not to fire at a precise instant. + if let Some(trigger) = force_recovery_trigger.as_mut() { + if trigger.should_fire(started.elapsed().as_secs()) { + match health.as_mut() { + Some(h) => { + let action = h.force_recovery(); + act_on_health_action( + ctx, + action, + "T7 bench probe: --force-recovery-after-secs elapsed", + &mut last_position, + &mut recovery_stops_pending, + ); + } + None => { + // health_check_registered was false (the time-pos + // subscription itself failed at startup, already + // warned loudly there) -- there is no HealthMonitor + // to spend a recovery attempt from, so the probe is + // inert. Loud, not silent: a bench operator staring + // at a run that never fires needs to know why. + eprintln!( + "warning: --force-recovery-after-secs elapsed but cannot fire: \ + tier-0 health check is DISABLED for this run (time-pos \ + subscription failed above)" + ); + } + } + } + } + + // F10: the bench-only wedge probe -- deliberately parks THIS event + // thread forever, simulating the one hazard class F1/F9 cannot see + // (a hang in our own code that is not an mpv call at all -- see + // dexd::watchdog's module doc "Framing"). Checked every + // iteration, same reasoning as T7's trigger above, and placed + // BEFORE the event-id dispatch below for the same reason that + // matters here even more than it does for T7: on a run where mpv + // reaches END_FILE/QUEUE_OVERFLOW almost immediately (e.g. this + // file's own `--opt vid=no --opt aid=no` test convention), the + // dispatch's own `std::process::exit(1)` would otherwise win the + // race and this probe would never get a chance to fire at all. + // Firing hangs the thread PERMANENTLY (a real `loop`, not a single + // long sleep) -- there is no "and then it resumes"; the whole point + // is that only an external actor (systemd's watchdog, if armed; the + // test harness's own deadline-kill otherwise) can end this process + // from here on. If nothing ever un-hangs it, that IS the pass + // condition -- see PLAN.md's F10 entry and README.md for how this + // is used on the Pi to prove the watchdog fires. + if let Some(trigger) = bench_wedge_trigger.as_mut() { + if trigger.should_fire(started.elapsed().as_secs()) { + eprintln!( + "warning: BENCH ONLY (F10 wedge probe): --bench-wedge-after-secs elapsed -- \ + deliberately parking the event thread forever to simulate F10's hazard \ + class (an event-thread hang outside any mpv call). If a systemd watchdog \ + is armed above, it should fire and this process should be restarted by \ + the supervisor; if this process is still alive well past WatchdogSec, the \ + watchdog did not fire and F10 has a gap." + ); + loop { + std::thread::sleep(std::time::Duration::from_secs(3600)); + } + } + } + + if id == MPV_EVENT_NONE { + continue; + } + if id == MPV_EVENT_SHUTDOWN { + break; + } + if id == MPV_EVENT_START_FILE { + continue; + } + if id == MPV_EVENT_PROPERTY_CHANGE { + // SAFETY: `data` is an mpv_event_property for this event id, + // valid until the next mpv_wait_event call. + let p = unsafe { &*((*ev).data as *const MpvEventProperty) }; + // Check BOTH the reply_userdata tag and the format tag before + // ever touching `data` -- `data`'s true type is decided by + // `format` (a tagged union), and trusting a hand-transcribed + // format constant without checking it is exactly the class of + // bug that made MPV_EVENT_LOG_MESSAGE's mistranscription a + // segfault instead of a caught error. A mismatch on either tag + // is not acted on: the next health-check tick simply sees no + // new sample, which HealthMonitor already treats identically to + // a genuine stall (see dexd::health's module doc) -- so + // failing to interpret an unexpected payload here fails toward + // "the health check is slightly more eager", never toward + // reading garbage. F9's two drop-counter observers follow the + // exact same discipline for MPV_FORMAT_INT64 -- and DO also act + // on MPV_FORMAT_NONE, deliberately, see below. + if reply_userdata == HEALTH_CHECK_USERDATA + && p.format == MPV_FORMAT_DOUBLE + && !p.data.is_null() + { + // SAFETY: the format tag confirms `data` points to an f64, + // per client.h's mpv_event_property contract for + // MPV_FORMAT_DOUBLE. `.cast::()` rather than + // `as *const f64`: this crate's convention (see read_fn's + // SAFETY comment above, where a platform-dependent `as` + // cast already caused a real bug) is `.cast()` for every raw + // pointer conversion, so a pointee-type mismatch is always a + // compile error instead of a silent reinterpretation. + last_position = Some((unsafe { *p.data.cast::() }, Instant::now())); + } else if p.format == MPV_FORMAT_INT64 && !p.data.is_null() { + // SAFETY: the format tag confirms `data` points to an i64, + // per MPV_FORMAT_INT64's doc comment in ffi_consts.rs. + let raw = unsafe { *p.data.cast::() }; + match reply_userdata { + FRAME_DROPS_USERDATA => frame_drops.sample(raw), + VO_DELAYED_USERDATA => vo_delayed.sample(raw), + _ => {} + } + } else if p.format == MPV_FORMAT_NONE + && matches!(reply_userdata, FRAME_DROPS_USERDATA | VO_DELAYED_USERDATA) + { + // A property that is momentarily UNAVAILABLE (no vo_chain: + // before the first frame, and during a recovery's teardown) + // arrives as format=MPV_FORMAT_NONE with data=NULL + // (player/client.c:1810-1816). `total` must survive this + // untouched -- it is not a value and not itself a counter + // reset -- but `last_raw` must NOT: mpv coalesces property + // events (client.h: "only once the event queue becomes + // empty ... one event per changed property"), so if a + // recovery's teardown (unavailable), the new session's + // restart at 0, and a climb past the old session's total all + // happen before this event thread next drains -- plausible + // exactly then, since the thread is busy absorbing the + // recovery's END_FILE/START_FILE burst -- `sample()` would + // see e.g. 5 -> 7 with no visible decrease and under-count + // by however many drops actually occurred (MINOR, three + // adversarial reviews, 2026-08-15). Recording the + // unavailability here means the next delivered sample, + // however small, is read as a fresh first sample rather + // than diffed against a `last_raw` that may already belong + // to a dead session -- narrows the window (this NONE event + // itself could still be coalesced away) rather than closing + // it; full closure isn't possible from the client side and + // isn't worth more machinery for a diagnostic line. See + // `ObservedCounter::mark_unavailable`'s doc comment. + match reply_userdata { + FRAME_DROPS_USERDATA => frame_drops.mark_unavailable(), + VO_DELAYED_USERDATA => vo_delayed.mark_unavailable(), + _ => {} + } + } + continue; + } + if id == MPV_EVENT_COMMAND_REPLY { + // Diagnostic only -- nothing gates on this. The next + // health-check tick judges the recovery by its actual effect + // (did time-pos start advancing again), not by whether mpv + // accepted the command; logging a rejection just makes that + // judgment call diagnosable from the journal afterwards. + if reply_userdata == RECOVERY_COMMAND_USERDATA { + let error = unsafe { (*ev).error }; + if error < 0 { + eprintln!( + "dexd: health check: in-place recovery's loadfile command \ + was rejected: {}", + err(ctx, "mpv_command_async(loadfile) reply", error) + ); + // Rejected AFTER being queued (mpv_command_async itself + // returned success) but before taking effect: no + // matching END_FILE(reason=stop) will ever arrive for + // THIS attempt, so its slot must not sit waiting for one + // -- a stale count left too high here could otherwise + // mask a real, later, unrelated END_FILE(stop) as an + // attempt's expected teardown. Decrement by one, not + // reset to zero: RECOVERY_COMMAND_USERDATA is shared by + // every recovery command, so a second attempt may + // legitimately still be in flight and its own stop is + // still owed. + recovery_stops_pending = recovery_stops_pending.saturating_sub(1); + } + } + continue; + } + if id == MPV_EVENT_LOG_MESSAGE { + // SAFETY: mpv guarantees `data` is an mpv_event_log_message for + // this event id, with NUL-terminated strings valid until the next + // mpv_wait_event call. + let m = unsafe { &*((*ev).data as *const MpvEventLogMessage) }; + let pfx = unsafe { CStr::from_ptr(m.prefix) }.to_string_lossy(); + let txt = unsafe { CStr::from_ptr(m.text) }.to_string_lossy(); + eprint!("mpv/{pfx}: {txt}"); + continue; + } + if id == MPV_EVENT_END_FILE { + // SAFETY: `data` is an mpv_event_end_file for this event id. + let ef = unsafe { &*((*ev).data as *const MpvEventEndFile) }; + if is_expected_recovery_stop(recovery_stops_pending, ef.reason) { + // One of the in-place recovery's `loadfile ... replace` + // calls just produced the END_FILE(reason=stop) mpv always + // emits for the file being replaced -- the expected teardown + // half of a recovery that is still in progress, not a + // failure. See MPV_END_FILE_REASON_STOP's doc comment and + // PLAN.md's F1 addendum: before this check existed, EVERY + // in-place recovery attempt killed the process on its own + // first step, making tier 0 unreachable. Absorb exactly one + // -- not reset to zero -- because the organic tick and T7's + // forced probe can both have issued a recovery in the same + // loop iteration, in which case a SECOND matching stop is + // still owed and must not be treated as fatal either. + recovery_stops_pending -= 1; + eprintln!( + "dexd: health check: in-place recovery's loadfile replaced the \ + stream; absorbing the expected END_FILE(reason=stop) for the file \ + it replaced ({recovery_stops_pending} more still outstanding), not \ + treating it as a failure" + ); + continue; + } + let why = unsafe { CStr::from_ptr(mpv_error_string(ef.error)) }; + eprintln!( + "dexd: FATAL: playback ended (reason={}, error={}) -- an endless \ + stream must never end; exiting so the supervisor restarts", + ef.reason, + why.to_string_lossy() + ); + // std::process::exit, not `break`: see the Escalate arm's + // comment above for why the shared mpv_terminate_destroy() + // teardown is the wrong tool on a fatal exit path -- the same + // reasoning applies here (and to QUEUE_OVERFLOW below). + std::process::exit(1); + } + if id == MPV_EVENT_QUEUE_OVERFLOW { + // mpv's internal event ring chokes at 1000 pending events and + // silently drops every event after that -- including END_FILE -- + // until the client drains back to empty (client.c send_event). + // There is no reservation for fatal events, so a dropped + // END_FILE would otherwise leave this program in mpv's default + // idle mode forever: exactly bug #1's failure, entered through a + // different door. We cannot know what was lost, so treat this + // exactly like END_FILE: exit and let the supervisor restart. + eprintln!( + "dexd: FATAL: mpv event queue overflowed -- at least one event was \ + dropped and may have been the one that mattered; exiting so the \ + supervisor restarts" + ); + std::process::exit(1); + } + } + + unsafe { mpv_terminate_destroy(ctx) }; + exit +} + +// --------------------------------------------------------------------------- +// This is the `dexd` BIN target, so these tests link libmpv (Pi only: +// `cargo test`) even though they call no mpv function -- `cargo check +// --all-targets` type-checks them on the Mac without linking. `cargo test +// --lib` (the Mac-safe command) does not run this module; it only runs +// tests under the `dexd` LIB target (src/lib.rs and its submodules). +// --------------------------------------------------------------------------- +#[cfg(test)] +mod tests { + use super::*; + + // Historical bug #3 lived exactly at this line: `read_fn` returning 0 + // for a zero-length request, which mpv reads as final EOF. `chunk.rs`'s + // `next_chunk(_, _, 0) == None` test pins the pure boundary; this test + // pins the shell around it -- reintroduce `else { return 0; }` here and + // every test in the crate stays green except this one (mpv itself + // essentially never issues a zero-length read, so the Pi integration + // suite can't see it either). + #[test] + fn read_fn_reports_mpv_error_not_zero_for_a_zero_length_request() { + let data: &'static [u8] = Box::leak(vec![1u8, 2, 3].into_boxed_slice()); + let cookie = Box::into_raw(Box::new(LoopStream { data, pos: 0 })) as *mut c_void; + let mut buf = [0u8; 8]; + let r = read_fn(cookie, buf.as_mut_ptr().cast::(), 0); + assert_eq!( + r, + i64::from(MPV_ERROR_UNSUPPORTED), + "must be an mpv error, never 0 -- to mpv, 0 means final EOF" + ); + // Reclaim what open_fn would normally leave leaked for the stream's + // lifetime, via the same path close_fn uses. + unsafe { drop(Box::from_raw(cookie as *mut LoopStream)) }; + } + + #[test] + fn read_fn_copies_bytes_and_advances_the_shared_position() { + let data: &'static [u8] = Box::leak(vec![10u8, 20, 30, 40, 50].into_boxed_slice()); + let cookie = Box::into_raw(Box::new(LoopStream { data, pos: 0 })) as *mut c_void; + let mut buf = [0u8; 8]; + + let r = read_fn(cookie, buf.as_mut_ptr().cast::(), 3); + assert_eq!(r, 3); + assert_eq!(&buf[..3], &[10, 20, 30]); + // SAFETY: single-threaded test; no other call is touching `cookie`. + let pos_after_first = unsafe { &*(cookie as *mut LoopStream) }.pos; + assert_eq!(pos_after_first, 3, "the callback must thread position through the same cookie, not reset per call"); + + let r = read_fn(cookie, buf.as_mut_ptr().cast::(), 4); + assert_eq!(r, 2, "short read: only 2 bytes remain before the wrap"); + assert_eq!(&buf[..2], &[40, 50]); + let pos_after_second = unsafe { &*(cookie as *mut LoopStream) }.pos; + assert_eq!(pos_after_second, 0, "eager wrap: position must land back at 0, not at len"); + + unsafe { drop(Box::from_raw(cookie as *mut LoopStream)) }; + } + + // The C1 regression this pins: `loadfile ... replace` (F1's in-place + // recovery) makes mpv emit END_FILE(reason=stop) for the file being + // replaced. Before `is_expected_recovery_stop` existed, the event loop + // treated ANY end-file as fatal, so the recovery's own first step + // always killed the process -- tier 0 was unreachable through the real + // event loop. These three cases are the whole fix. + + #[test] + fn a_pending_recovery_absorbs_its_own_stop_reason_end_file() { + assert!(is_expected_recovery_stop(1, MPV_END_FILE_REASON_STOP)); + } + + #[test] + fn a_stop_reason_end_file_with_no_recovery_pending_stays_fatal() { + // Nothing else in this program issues a command that produces a + // stop-reason end-file -- but if one ever did, it must not be + // silently swallowed just because the reason code matches. + assert!(!is_expected_recovery_stop(0, MPV_END_FILE_REASON_STOP)); + } + + #[test] + fn any_other_reason_stays_fatal_even_while_a_recovery_is_pending() { + // eof=0, quit=3, error=4, redirect=5 (2 is stop, tested above). + for reason in [0, 3, 4, 5] { + assert!( + !is_expected_recovery_stop(1, reason), + "reason {reason} must not be absorbed" + ); + } + } + + // The follow-up regression these two pin (three independent adversarial + // reviews, 2026-08-15, MAJOR): the organic health-check tick and T7's + // forced probe can both issue a recovery before either one's END_FILE + // arrives (mpv_command_async only QUEUES against a core that may still + // be busy from the first attempt), producing TWO END_FILE(reason=stop) + // events for one episode. A single bool could absorb only the first and + // treated the second -- an entirely expected teardown for a recovery + // that just worked -- as fatal, killing a healthy-again process with a + // journal line indistinguishable from C1. See main.rs's + // `recovery_stops_pending` doc comment and PLAN.md's F1 addendum. + + #[test] + fn two_overlapping_recoveries_both_absorb_their_own_stop() { + let mut pending: u32 = 0; + pending += 1; // organic tick issues attempt 1 + pending += 1; // T7's forced probe issues attempt 2, same iteration + assert_eq!(pending, 2); + + assert!(is_expected_recovery_stop(pending, MPV_END_FILE_REASON_STOP)); + pending -= 1; // first END_FILE(stop) absorbed + assert_eq!(pending, 1, "one recovery's stop is still owed"); + + assert!( + is_expected_recovery_stop(pending, MPV_END_FILE_REASON_STOP), + "the second stop must NOT be treated as fatal just because the \ + first one already cleared the flag -- this is the exact bug a \ + plain bool could not represent" + ); + pending -= 1; + assert_eq!(pending, 0); + + // Budget is exhausted now (both attempts absorbed): a THIRD + // stop-reason end-file with nothing outstanding is a real failure. + assert!(!is_expected_recovery_stop(pending, MPV_END_FILE_REASON_STOP)); + } + + #[test] + fn a_rejected_command_reply_decrements_by_one_not_to_zero() { + // Mirrors the MPV_EVENT_COMMAND_REPLY handler: a rejection after + // queueing means THAT attempt's stop will never arrive, but + // RECOVERY_COMMAND_USERDATA is shared by every recovery command, so + // a second, still-legitimate attempt may be in flight and its stop + // must remain expected. + let mut pending: u32 = 2; + pending = pending.saturating_sub(1); // one attempt's reply was rejected + assert_eq!(pending, 1); + assert!(is_expected_recovery_stop(pending, MPV_END_FILE_REASON_STOP)); + } +} diff --git a/packages/dexd/src/nal.rs b/packages/dexd/src/nal.rs new file mode 100644 index 0000000..741b44c --- /dev/null +++ b/packages/dexd/src/nal.rs @@ -0,0 +1,201 @@ +//! F4 — the asset validation gate: refuse an asset whose leading NALs cannot +//! support the gaplessness premise. +//! +//! The endless-stream design only wraps seamlessly because byte 0 begins a +//! closed GOP: parameter sets (VPS/SPS/PPS) then an IDR, so re-entering at +//! byte 0 mid-stream is an ordinary keyframe, not a seek. An asset that +//! starts with anything else — an open-GOP CRA, a trailing slice, no +//! parameter sets — would "play" and then glitch at EVERY wrap (~29k visible +//! artefacts/day for a 3 s loop), silently. The hash (F3) proves the bytes +//! are the ingested bytes; this gate proves the ingested bytes have the +//! required SHAPE. Truncation is F3's job: a truncated copy has intact +//! leading NALs and passes this gate by design. +//! +//! Why IDR only (19/20), not any IRAP (16-23): CRA (21) admits RASL leading +//! pictures whose wrap-join correctness depends on the content; BLA (16-18) +//! never comes from a sane ingest; 22/23 are reserved. The premise stated +//! everywhere in this crate is "IDR at frame 0" — so that is what the gate +//! enforces. Relax knowingly if an asset ever justifies it. + +/// HEVC nal_unit_type values (ITU-T H.265 Table 7-1) this gate names. +pub const NAL_VPS: u8 = 32; +pub const NAL_SPS: u8 = 33; +pub const NAL_PPS: u8 = 34; +pub const NAL_IDR_W_RADL: u8 = 19; +pub const NAL_IDR_N_LP: u8 = 20; + +/// Validate the leading NAL units of a raw Annex-B HEVC stream. +/// +/// Passes iff, before the first VCL NAL (type 0-31), all of VPS/SPS/PPS have +/// appeared, and that first VCL NAL is an IDR (19 or 20). Everything after +/// the first VCL NAL is out of scope — the F3 hash covers byte-level +/// integrity of the whole asset. +/// +/// Scanning is a plain 00 00 01 search (3- and 4-byte start codes both +/// resolve to it): encoders insert emulation-prevention bytes precisely so +/// that pattern never occurs inside a NAL payload, so the search cannot +/// false-positive on a well-formed stream. +pub fn validate_leading_nals(data: &[u8]) -> Result<(), String> { + let mut vps = false; + let mut sps = false; + let mut pps = false; + let mut found_any = false; + let mut iter = StartCodeIter { data, i: 0 }; + while let Some((b0, _b1)) = iter.next_nal_header() { + found_any = true; + if b0 & 0x80 != 0 { + return Err("corrupt NAL header (forbidden_zero_bit set)".into()); + } + let nal_type = (b0 >> 1) & 0x3f; + match nal_type { + NAL_VPS => vps = true, + NAL_SPS => sps = true, + NAL_PPS => pps = true, + 0..=31 => { + // First VCL NAL: the gate's decision point. + let missing: Vec<&str> = [(!vps, "VPS"), (!sps, "SPS"), (!pps, "PPS")] + .iter() + .filter(|(m, _)| *m) + .map(|(_, n)| *n) + .collect(); + if !missing.is_empty() { + return Err(format!( + "first slice appears before parameter sets ({} missing); not a \ + valid loop asset — re-ingest with a closed-GOP encode", + missing.join("/") + )); + } + return match nal_type { + NAL_IDR_W_RADL | NAL_IDR_N_LP => Ok(()), + 21 => Err( + "leading keyframe is CRA (open GOP), not IDR; the wrap would \ + splice mid-GOP — re-ingest with a closed-GOP encode (IDR at \ + frame 0)" + .into(), + ), + t => Err(format!( + "first slice NAL is type {t}, not an IDR (19/20); the stream does \ + not start on a clean keyframe — re-ingest with a closed-GOP encode" + )), + }; + } + _ => {} // other non-VCL (AUD, SEI, ...): fine before the IDR + } + } + if !found_any { + return Err( + "no Annex-B start code found; this is not a raw HEVC elementary stream (MP4? \ + use: ffmpeg -i in.mp4 -c:v copy -bsf:v hevc_mp4toannexb -f hevc out.265)" + .into(), + ); + } + Err("parameter sets but no slice found in the asset".into()) +} + +/// Finds each 00 00 01 start code (the 4-byte form contains it) and yields +/// the two NAL header bytes that follow. +struct StartCodeIter<'a> { + data: &'a [u8], + i: usize, +} + +impl StartCodeIter<'_> { + fn next_nal_header(&mut self) -> Option<(u8, u8)> { + let d = self.data; + let mut i = self.i; + while i + 2 < d.len() { + if d[i] == 0 && d[i + 1] == 0 && d[i + 2] == 1 { + let h = i + 3; + self.i = h + 1; // keep searching after this start code + if h + 1 < d.len() { + return Some((d[h], d[h + 1])); + } + return None; // start code at EOF, no room for a header + } + i += 1; + } + self.i = i; + None + } +} + +#[cfg(test)] +mod tests { + use super::*; + + /// One NAL: 4-byte start code + 2-byte header (layer 0, tid+1 = 1). + fn nal(nal_type: u8, payload: &[u8]) -> Vec { + let mut v = vec![0, 0, 0, 1, nal_type << 1, 0x01]; + v.extend_from_slice(payload); + v + } + + fn stream(types: &[u8]) -> Vec { + let mut v = Vec::new(); + for &t in types { + v.extend(nal(t, &[0x2a; 8])); + } + v + } + + #[test] + fn valid_closed_gop_asset_passes() { + assert!(validate_leading_nals(&stream(&[32, 33, 34, 19])).is_ok()); // IDR_W_RADL + assert!(validate_leading_nals(&stream(&[32, 33, 34, 20])).is_ok()); // IDR_N_LP + // non-VCL noise before/among parameter sets is fine (AUD=35, SEI=39) + assert!(validate_leading_nals(&stream(&[35, 32, 39, 33, 34, 19])).is_ok()); + // trailing slices after the IDR are out of scope for the gate + assert!(validate_leading_nals(&stream(&[32, 33, 34, 19, 1, 0, 1])).is_ok()); + } + + #[test] + fn three_byte_start_codes_pass_too() { + let mut v = Vec::new(); + for t in [32u8, 33, 34, 19] { + v.extend([0, 0, 1, t << 1, 0x01]); + v.extend([0x2a; 8]); + } + assert!(validate_leading_nals(&v).is_ok()); + } + + #[test] + fn missing_parameter_sets_are_refused_and_named() { + let e = validate_leading_nals(&stream(&[32, 33, 19])).unwrap_err(); + assert!(e.contains("PPS"), "{e}"); + let e = validate_leading_nals(&stream(&[34, 19])).unwrap_err(); + assert!(e.contains("VPS") && e.contains("SPS"), "{e}"); + } + + #[test] + fn non_idr_first_slice_is_refused() { + // TRAIL_R (1) first: a copy that lost its head, or a cut mid-GOP. + // Would glitch at EVERY wrap. + let e = validate_leading_nals(&stream(&[32, 33, 34, 1])).unwrap_err(); + assert!(e.contains("not an IDR"), "{e}"); + } + + #[test] + fn cra_open_gop_is_refused_by_name() { + let e = validate_leading_nals(&stream(&[32, 33, 34, 21])).unwrap_err(); + assert!(e.contains("CRA"), "{e}"); + } + + #[test] + fn parameter_sets_after_the_slice_do_not_count() { + let e = validate_leading_nals(&stream(&[32, 33, 19, 34])).unwrap_err(); + assert!(e.contains("PPS"), "{e}"); + } + + #[test] + fn garbage_and_degenerate_inputs_are_refused() { + assert!(validate_leading_nals(&[]).is_err()); + let e = validate_leading_nals(&[0x47; 4096]).unwrap_err(); + assert!(e.contains("start code"), "{e}"); // no Annex-B start code at all + // parameter sets only, no slice ever + assert!(validate_leading_nals(&stream(&[32, 33, 34])).is_err()); + // forbidden_zero_bit set on the first NAL header + let mut v = vec![0, 0, 0, 1, 0x80 | (32 << 1), 0x01]; + v.extend([0x2a; 8]); + assert!(validate_leading_nals(&v).is_err()); + } +} diff --git a/packages/dexd/src/sha256.rs b/packages/dexd/src/sha256.rs new file mode 100644 index 0000000..7ec36c9 --- /dev/null +++ b/packages/dexd/src/sha256.rs @@ -0,0 +1,110 @@ +//! F3 prerequisite — one-shot SHA-256 (FIPS 180-4) for the asset<->sidecar +//! binding, over the `sha2` crate. +//! +//! This module was hand-rolled (185 lines, FIPS 180-4 from scratch) under the +//! crate's former zero-dependency rule. That rule is withdrawn — SPEC §5c — +//! so the implementation is now RustCrypto's and this file is a thin wrapper. +//! +//! **The wrapper is deliberate: the module keeps its API so the vectors below +//! keep testing the thing the player actually calls.** They now pin `sha2` +//! rather than our own compression function, which is the point — the tests +//! were always a statement about `sha256_hex`'s output, never about who +//! computed it. +//! +//! Before the swap, the hand-rolled implementation was confirmed to agree with +//! coreutils `sha256sum` on the real bench assets (loop4k.265, loop.265). That +//! check is why existing sidecars remain valid across this change: had it +//! disagreed, every sidecar in the field would have encoded a wrong digest and +//! this "refactor" would have turned the F3 startup gate into a fleet-wide +//! boot loop. A hash change is a data-format change. + +use sha2::{Digest, Sha256}; + +/// SHA-256 of `data`, one shot. +pub fn sha256(data: &[u8]) -> [u8; 32] { + let mut out = [0u8; 32]; + out.copy_from_slice(&Sha256::digest(data)); + out +} + +/// SHA-256 of `data` as 64 lowercase hex characters — the sidecar format. +pub fn sha256_hex(data: &[u8]) -> String { + const HEX: &[u8; 16] = b"0123456789abcdef"; + let d = sha256(data); + let mut s = String::with_capacity(64); + for b in d { + s.push(HEX[(b >> 4) as usize] as char); + s.push(HEX[(b & 0xf) as usize] as char); + } + s +} + +#[cfg(test)] +mod tests { + use super::*; + + fn a_times(n: usize) -> Vec { + vec![b'a'; n] + } + + #[test] + fn nist_vectors() { + assert_eq!( + sha256_hex(b""), + "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855" + ); + assert_eq!( + sha256_hex(b"abc"), + "ba7816bf8f01cfea414140de5dae2223b00361a396177a9cb410ff61f20015ad" + ); + assert_eq!( + sha256_hex(b"abcdbcdecdefdefgefghfghighijhijkijkljklmklmnlmnomnopnopq"), + "248d6a61d20638b8e5c026930c3e6039a33ce45964ff2167f6ecedd419db06c1" + ); + } + + // The NIST long vector: one million 'a'. Exercises many blocks and the + // rem == 0 padding path on a large input. Milliseconds even in debug. + #[test] + fn nist_million_a() { + assert_eq!( + sha256_hex(&a_times(1_000_000)), + "cdc76e5c9914fb9281a1c7e284d73e67f1809a48a497200e046d39ccc7112cd0" + ); + } + + // Padding boundaries: 55 is the last 1-padding-block length, 56 the first + // 2-block one; 63/64/65 straddle the block size; 112 covers a mid-size + // rem. Reference digests generated with `shasum -a 256` on 2026-08-15. + #[test] + fn padding_boundaries() { + for (n, want) in [ + ( + 55, + "9f4390f8d30c2dd92ec9f095b65e2b9ae9b0a925a5258e241c9f1e910f734318", + ), + ( + 56, + "b35439a4ac6f0948b6d6f9e3c6af0f5f590ce20f1bde7090ef7970686ec6738a", + ), + ( + 63, + "7d3e74a05d7db15bce4ad9ec0658ea98e3f06eeecf16b4c6fff2da457ddc2f34", + ), + ( + 64, + "ffe054fe7ae0cb6dc65c3af9b61d5209f439851db43d0ba5997337df154668eb", + ), + ( + 65, + "635361c48bb9eab14198e76ea8ab7f1a41685d6ad62aa9146d301d4f17eb0ae0", + ), + ( + 112, + "f54353008a2553262ecdc4a34749563ba0950e8b0fc8652780b0a614b99683c1", + ), + ] { + assert_eq!(sha256_hex(&a_times(n)), want, "length {n}"); + } + } +} diff --git a/packages/dexd/src/sidecar.rs b/packages/dexd/src/sidecar.rs new file mode 100644 index 0000000..56d38a0 --- /dev/null +++ b/packages/dexd/src/sidecar.rs @@ -0,0 +1,549 @@ +//! F3 — the asset+fps sidecar: parse, validate, and decide the binding. +//! +//! Why: the one failure that is undetectable BY CONSTRUCTION. A raw Annex-B +//! stream has no timestamps, so `--fps 25` on a 30 fps asset plays 20% slow, +//! forever, with zero errors and every metric nominal. The fix: the frame +//! rate travels WITH the asset (a sidecar written at ingest), bound by a +//! sha256 so a stale/wrong/truncated asset is refused at startup. +//! +//! Format — `.json` next to the asset (`loop.265` -> `loop.265.json`): +//! {"fps":"30","sha256":"<64 hex>","width":3840,"height":2160, +//! "source":"card.mp4","encoder_cmd":"ffmpeg ..."} +//! `fps` is a STRING, not a JSON number: "30000/1001" must survive exactly, +//! and 29.97 as a float invites drift. It is passed verbatim to mpv's +//! container-fps-override after grammar validation. Required: fps, sha256. +//! Optional, informational: width, height (integers), source, encoder_cmd +//! (strings). Unknown keys are ignored so ingest can add metadata without +//! breaking deployed players. +//! +//! The parser accepts a STRICT SUBSET of JSON — one flat object, string and +//! unsigned-integer values only (this applies uniformly to every key, known +//! or not: an ignored key's value must still be a string or unsigned +//! integer, never an array/bool/null/nested object). Anything else is a +//! parse error, and a parse error refuses startup. Fail-closed IS the F3 +//! semantics: an unparseable sidecar and a missing one are the same +//! operational fact. +//! +//! The subset is enforced by a serde visitor over `serde_json` (SPEC §5c); it +//! was hand-rolled while the crate had a zero-dependency rule. Exactly ONE +//! rule survives as our own code, because it is the one a `Map` cannot state: +//! **duplicate keys are rejected** rather than silently last-wins. +//! +//! String escapes (\" \\ \/ \n \r \t, and \uXXXX including UTF-16 surrogate +//! pairs) are serde_json's problem now — but they are recorded here as a +//! REQUIREMENT, because the tempting "simplification" is to ban non-ASCII and +//! it would be wrong. \uXXXX is what every "safe by default" JSON serializer +//! reaches for: Python's `json.dumps` (default `ensure_ascii=True`) turns ANY +//! non-ASCII character into \uXXXX, and Go's `encoding/json` does the same for +//! `<`, `>`, `&`. That includes informational keys like `source`/`encoder_cmd` +//! which are never interpreted here — a `source` filename may legitimately be +//! "Karte–Süd.mp4". Refusing escapes would refuse byte-perfect, correctly +//! hashed assets over nothing but an ingest tool's serializer settings. +//! (`fps` and `sha256` are ASCII by their own grammars, validated below.) + +/// A parsed JSON value, restricted to the sidecar subset: strings and +/// unsigned integers only. +#[derive(Debug, Clone, PartialEq, Eq)] +pub enum Value { + Str(String), + Num(u64), +} + +/// Parse `text` as a single flat JSON object under the sidecar subset grammar +/// (see module docs), returning its key/value pairs in source order. +pub fn parse_flat_json(text: &str) -> Result, String> { + serde_json::from_str::(text) + .map(|f| f.0) + .map_err(|e| e.to_string()) +} + +/// A newtype whose `Deserialize` impl *is* the subset grammar. +/// +/// Hand-written visitor rather than `#[derive]` or `serde_json::Map`, for one +/// reason: **the grammar rejects duplicate keys and a Map cannot express +/// that** — it silently keeps the last. `{"fps":"30","fps":"25"}` has to be an +/// error rather than a coin flip decided by which parser reads it, because F3 +/// is a fail-closed gate: an ambiguous sidecar and a missing one are the same +/// operational fact. Everything else the old hand-rolled parser did — lexing, +/// escapes, surrogate pairs, trailing-data and structural errors — is +/// serde_json's now. +struct FlatObject(Vec<(String, Value)>); + +impl<'de> serde::Deserialize<'de> for FlatObject { + fn deserialize>(d: D) -> Result { + // deserialize_map, not deserialize_any: a top-level array or scalar is + // rejected by serde_json with a type error before we see it. + d.deserialize_map(FlatObjectVisitor) + } +} + +struct FlatObjectVisitor; + +/// Name a rejected value's type for the error message. Kept exhaustive rather +/// than `_ =>` so a future serde_json variant is a compile error here. +fn type_name(v: &serde_json::Value) -> &'static str { + match v { + serde_json::Value::Null => "null", + serde_json::Value::Bool(_) => "a boolean", + serde_json::Value::Array(_) => "an array", + serde_json::Value::Object(_) => "a nested object", + serde_json::Value::String(_) => "a string", + serde_json::Value::Number(_) => "a number", + } +} + +impl<'de> serde::de::Visitor<'de> for FlatObjectVisitor { + type Value = FlatObject; + + fn expecting(&self, f: &mut std::fmt::Formatter) -> std::fmt::Result { + f.write_str("a flat JSON object whose values are strings or unsigned integers") + } + + fn visit_map(self, mut map: A) -> Result + where + A: serde::de::MapAccess<'de>, + { + use serde::de::Error as _; + let mut out: Vec<(String, Value)> = Vec::new(); + while let Some(key) = map.next_key::()? { + if out.iter().any(|(k, _)| *k == key) { + return Err(A::Error::custom(format!("duplicate key {key:?}"))); + } + // The subset applies UNIFORMLY, to unknown keys too: an ignored + // key's value must still be a string or unsigned integer. Reading + // into serde_json::Value first is what lets us say so — and lets + // an informational key like `source` hold any text, escapes and + // all, which is the behaviour the format actually needs. + let value = match map.next_value::()? { + serde_json::Value::String(s) => Value::Str(s), + serde_json::Value::Number(n) => Value::Num(n.as_u64().ok_or_else(|| { + A::Error::custom(format!( + "value for {key:?} must be an unsigned integer, got {n} \ + (floats and negatives are outside the sidecar subset; \ + write fps as a string)" + )) + })?), + other => { + return Err(A::Error::custom(format!( + "unsupported value for {key:?}: {} (subset: strings and \ + unsigned integers only)", + type_name(&other) + ))) + } + }; + out.push((key, value)); + } + Ok(FlatObject(out)) + } +} + + +/// Is `s` a well-formed frame rate string: a positive integer ("30"), a +/// positive decimal ("29.97"), or a positive rational ("30000/1001")? +pub fn is_valid_fps(s: &str) -> bool { + fn positive_int(t: &str) -> bool { + !t.is_empty() && t.bytes().all(|b| b.is_ascii_digit()) && t.bytes().any(|b| b != b'0') + } + if let Some((num, den)) = s.split_once('/') { + return positive_int(num) && positive_int(den); + } + if let Some((int, frac)) = s.split_once('.') { + let digits_ok = !int.is_empty() + && int.bytes().all(|b| b.is_ascii_digit()) + && !frac.is_empty() + && frac.bytes().all(|b| b.is_ascii_digit()); + let nonzero = int.bytes().chain(frac.bytes()).any(|b| b != b'0'); + return digits_ok && nonzero; + } + positive_int(s) +} + +/// A parsed, validated sidecar: the frame rate and the sha256 the asset must +/// match, plus optional informational dimensions. +#[derive(Debug, Clone, PartialEq, Eq)] +pub struct Sidecar { + pub fps: String, + pub sha256: String, + pub width: Option, + pub height: Option, +} + +impl Sidecar { + /// Parse and validate a sidecar's JSON text. Fail-closed: any grammar + /// violation, missing required key, invalid fps, or malformed sha256 is + /// an error, on the theory that an unparseable sidecar and a missing one + /// are the same operational fact. + pub fn from_json(text: &str) -> Result { + let kv = parse_flat_json(text).map_err(|e| format!("sidecar JSON: {e}"))?; + let mut fps = None; + let mut sha = None; + let mut width = None; + let mut height = None; + for (k, v) in kv { + match (k.as_str(), v) { + ("fps", Value::Str(s)) => fps = Some(s), + ("fps", Value::Num(_)) => { + return Err( + "sidecar: fps must be a JSON STRING (\"30\", \"30000/1001\") \ + so rational rates survive exactly" + .into(), + ) + } + ("sha256", Value::Str(s)) => sha = Some(s), + ("sha256", Value::Num(_)) => return Err("sidecar: sha256 must be a string".into()), + ("width", Value::Num(n)) => width = Some(n), + ("height", Value::Num(n)) => height = Some(n), + ("width", Value::Str(_)) | ("height", Value::Str(_)) => { + return Err("sidecar: width/height must be integers".into()) + } + // Unknown keys and informational strings: ignored, so ingest + // can add metadata without breaking deployed players. + _ => {} + } + } + let fps = fps.ok_or("sidecar: missing required key \"fps\"")?; + if !is_valid_fps(&fps) { + return Err(format!( + "sidecar: invalid fps {fps:?} (expect \"30\", \"29.97\" or \"30000/1001\")" + )); + } + let sha = sha.ok_or("sidecar: missing required key \"sha256\"")?; + let sha = sha.to_ascii_lowercase(); + if sha.len() != 64 || !sha.bytes().all(|b| b.is_ascii_hexdigit()) { + return Err(format!( + "sidecar: sha256 must be 64 hex digits, got {sha:?}" + )); + } + Ok(Sidecar { + fps, + sha256: sha, + width, + height, + }) + } +} + +/// Where a bound fps value came from: the asset's sidecar, or the bench +/// escape hatch (`--bench-no-sidecar --fps `). +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum FpsSource { + Sidecar, + BenchOverride, +} + +/// Decide the fps to bind to, and where it came from. +/// +/// Rules: with a sidecar present, its fps wins; an explicit `--fps` is +/// allowed only if it agrees with the sidecar (a mismatch is refused, naming +/// both values, since the sidecar is authoritative). With no sidecar, the +/// deploy path refuses to guess — the bench escape hatch is a deliberate, +/// two-flag act (`--bench-no-sidecar` AND `--fps`), never a silent fallback. +pub fn resolve_fps( + sidecar_fps: Option<&str>, + cli_fps: Option<&str>, + bench_no_sidecar: bool, +) -> Result<(String, FpsSource), String> { + if bench_no_sidecar { + return match cli_fps { + Some(f) if is_valid_fps(f) => Ok((f.to_string(), FpsSource::BenchOverride)), + Some(f) => Err(format!("--fps {f:?} is not a valid frame rate")), + None => Err("--bench-no-sidecar requires an explicit --fps".into()), + }; + } + match (sidecar_fps, cli_fps) { + (Some(s), None) => Ok((s.to_string(), FpsSource::Sidecar)), + (Some(s), Some(c)) if s == c => Ok((s.to_string(), FpsSource::Sidecar)), + (Some(s), Some(c)) => Err(format!( + "--fps {c} contradicts sidecar fps {s}; drop --fps (the sidecar is authoritative) \ + or fix the sidecar" + )), + (None, _) => Err( + "no sidecar found; refusing to guess the frame rate. Re-ingest the asset to \ + produce .json, or use --bench-no-sidecar --fps on a bench" + .into(), + ), + } +} + +/// Verify `payload`'s sha256 matches the sidecar's — the binding that refuses +/// a stale, wrong, or truncated asset at startup instead of playing it wrong +/// forever. +pub fn verify_payload(payload: &[u8], sidecar: &Sidecar) -> Result<(), String> { + let actual = crate::sha256::sha256_hex(payload); + if actual != sidecar.sha256 { + return Err(format!( + "asset does not match its sidecar: sha256 {actual} != sidecar {}; the asset or \ + sidecar is stale, wrong, or truncated — re-ingest", + sidecar.sha256 + )); + } + Ok(()) +} + +#[cfg(test)] +mod tests { + use super::*; + + const GOOD_SHA: &str = "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855"; + + #[test] + fn parses_the_canonical_sidecar() { + let text = format!( + r#"{{"fps":"30","sha256":"{GOOD_SHA}","width":3840,"height":2160,"source":"card.mp4","encoder_cmd":"ffmpeg -i card.mp4 -c:v copy"}}"# + ); + let s = Sidecar::from_json(&text).unwrap(); + assert_eq!(s.fps, "30"); + assert_eq!(s.sha256, GOOD_SHA); + assert_eq!(s.width, Some(3840)); + assert_eq!(s.height, Some(2160)); + } + + #[test] + fn unknown_keys_are_ignored() { + let text = format!(r#"{{"fps":"30","sha256":"{GOOD_SHA}","future_key":"whatever"}}"#); + assert!(Sidecar::from_json(&text).is_ok()); + } + + #[test] + fn fps_as_number_is_refused_with_guidance() { + let text = format!(r#"{{"fps":30,"sha256":"{GOOD_SHA}"}}"#); + let e = Sidecar::from_json(&text).unwrap_err(); + assert!(e.contains("STRING"), "{e}"); + } + + #[test] + fn missing_required_keys_are_refused_by_name() { + let e = Sidecar::from_json(&format!(r#"{{"sha256":"{GOOD_SHA}"}}"#)).unwrap_err(); + assert!(e.contains("fps"), "{e}"); + let e = Sidecar::from_json(r#"{"fps":"30"}"#).unwrap_err(); + assert!(e.contains("sha256"), "{e}"); + } + + #[test] + fn bad_sha256_is_refused_uppercase_is_normalized() { + for sha in ["", "abc", &"g".repeat(64), &"a".repeat(63), &"a".repeat(65)] { + let text = format!(r#"{{"fps":"30","sha256":"{sha}"}}"#); + assert!(Sidecar::from_json(&text).is_err(), "sha {sha:?} accepted"); + } + let text = format!(r#"{{"fps":"30","sha256":"{}"}}"#, GOOD_SHA.to_uppercase()); + assert_eq!(Sidecar::from_json(&text).unwrap().sha256, GOOD_SHA); + } + + #[test] + fn fps_grammar() { + for ok in ["30", "25", "29.97", "23.976", "30000/1001", "60"] { + assert!(is_valid_fps(ok), "{ok} should be valid"); + } + for bad in [ + "", "0", "00", "0/30", "30/0", "-30", "+30", " 30", "30 ", "30/", "/1001", "1e3", + "29.97.5", "29.", ".97", "banana", "0.0", + ] { + assert!(!is_valid_fps(bad), "{bad:?} should be invalid"); + } + } + + #[test] + fn parser_rejects_everything_outside_the_subset() { + for bad in [ + "", // no object + "[1,2]", // array at top level + r#"{"a":{"b":1}}"#, // nested object + r#"{"a":[1]}"#, // array value + r#"{"a":true}"#, // boolean + r#"{"a":null}"#, // null + r#"{"a":-1}"#, // negative number + r#"{"a":1.5}"#, // float + r#"{"a":1e3}"#, // exponent + r#"{"a":"x"}"trailing"#, // trailing data + r#"{"a":"x""b":"y"}"#, // missing comma + r#"{"a":"unterminated}"#, // unterminated string + r#"{"a":"x","a":"y"}"#, // duplicate key + ] { + assert!(parse_flat_json(bad).is_err(), "accepted: {bad}"); + } + } + + #[test] + fn unicode_escapes_are_decoded() { + // These are raw string literals: the JSON *source text* the parser + // receives contains the literal four characters `\`, `u`, and four + // hex digits -- the parser itself must turn that into a code point. + // The expected side uses Rust's OWN (unrelated) `\u{...}` syntax + // purely so this file's source stays plain ASCII. + + // ASCII code point spelled via the escape A. + let json = format!(r#"{{"a":"\{}0041"}}"#, 'u'); + let kv = parse_flat_json(&json).unwrap(); + assert_eq!(kv[0].1, Value::Str("A".to_string())); + + // The actual field bug: json.dumps({"source": "Zürich.mp4"}) with + // Python's DEFAULT ensure_ascii=True produces the escape ü for + // "ü". Built with `format!` so this file's source stays plain ASCII; + // `\u{fc}` on the expected side is Rust's own (unrelated) escape. + let json = format!(r#"{{"a":"Z\{}00fcrich.mp4"}}"#, 'u'); + let kv = parse_flat_json(&json).unwrap(); + assert_eq!(kv[0].1, Value::Str("Z\u{fc}rich.mp4".to_string())); + + // Outside the BMP: a UTF-16 surrogate pair, standard JSON (not a + // sidecar-specific extension) -- U+1F389 PARTY POPPER is + // 🎉. + let json = format!(r#"{{"a":"\{0}d83c\{0}df89"}}"#, 'u'); + let kv = parse_flat_json(&json).unwrap(); + assert_eq!(kv[0].1, Value::Str("\u{1F389}".to_string())); + + // A \u escape next to a plain escape in the same string, to prove + // they compose: é (é) then a plain \n. + let json = format!(r#"{{"a":"caf\{}00e9\nmore"}}"#, 'u'); + let kv = parse_flat_json(&json).unwrap(); + assert_eq!(kv[0].1, Value::Str("caf\u{e9}\nmore".to_string())); + } + + // The other half of the escape requirement, and the half that is easier to + // lose: a serializer that does NOT escape (jq, Rust's own serde_json, + // Python with ensure_ascii=False) writes the character as raw UTF-8 bytes. + // Both spellings mean the same sidecar and both must parse. + // + // This exists because "no sidecar value can hold a non-ASCII byte, so + // reject non-ASCII at the door" was proposed during the SPEC §5c work and + // is WRONG: it is true of `fps` and `sha256`, and false of `source` -- + // asset filenames are routinely not ASCII. Narrowing the contract there + // would have refused correctly-hashed assets over their filename. + #[test] + fn raw_utf8_in_an_informational_key_parses_like_its_escaped_form() { + // Built with char escapes so this file's source stays plain ASCII. + let raw = "Z\u{fc}rich.mp4".to_string(); + let json = format!(r#"{{"source":"{raw}"}}"#); + let kv = parse_flat_json(&json).unwrap(); + assert_eq!(kv[0].1, Value::Str(raw.clone())); + + // ... and is indistinguishable from the \u-escaped spelling. + let escaped = format!(r#"{{"source":"Z\{}00fcrich.mp4"}}"#, 'u'); + assert_eq!(parse_flat_json(&escaped).unwrap(), kv); + + // A full sidecar with a non-ASCII source must still bind normally -- + // `source` is informational and never interpreted. + let text = format!( + r#"{{"fps":"30","sha256":"{GOOD_SHA}","source":"{raw}"}}"# + ); + let s = Sidecar::from_json(&text).unwrap(); + assert_eq!(s.fps, "30"); + } + + // Duplicate rejection is the ONE grammar rule still implemented by hand + // (SPEC §5c): serde_json's Map silently keeps the last value, so the + // visitor has to say so itself. Tested by name rather than only inside the + // reject-everything batch, because a refactor that dropped the visitor for + // a plain Map would still pass every other test in this file. + #[test] + fn duplicate_keys_are_refused_and_named() { + let e = parse_flat_json(r#"{"fps":"30","fps":"25"}"#).unwrap_err(); + assert!(e.contains("duplicate"), "{e}"); + assert!(e.contains("fps"), "{e}"); + + // Including a duplicated key the sidecar does not interpret: the rule + // is about the document being unambiguous, not about which keys matter. + let e = parse_flat_json(r#"{"source":"a","source":"b"}"#).unwrap_err(); + assert!(e.contains("duplicate"), "{e}"); + } + + #[test] + fn malformed_unicode_escapes_are_refused() { + for bad in [ + r#"{"a":"\u12"}"#, // truncated: only 2 hex digits + r#"{"a":"\u12zz"}"#, // non-hex digits + r#"{"a":"\ud800"}"#, // lone high surrogate, no pair follows + r#"{"a":"\udc00"}"#, // lone low surrogate + r#"{"a":"\ud800A"}"#, // high surrogate followed by a non-surrogate + r#"{"a":"\ud800\udbff"}"#, // high surrogate followed by ANOTHER high surrogate + ] { + assert!(parse_flat_json(bad).is_err(), "accepted: {bad}"); + } + } + + #[test] + fn unicode_escape_in_an_ignored_sidecar_key_no_longer_breaks_startup() { + // The concrete field scenario: an ingest tool's default-safe JSON + // serializer \u-escapes a non-ASCII byte inside "source", a key this + // player does not even interpret -- and startup used to refuse + // anyway, on an otherwise byte-perfect, correctly-hashed asset. Built + // with `format!` so the JSON text itself contains the literal escape + // ü, not an already-decoded byte -- that is the actual bug. + let text = format!( + r#"{{"fps":"30","sha256":"{GOOD_SHA}","source":"Z\{}00fcrich.mp4"}}"#, + 'u' + ); + let s = Sidecar::from_json(&text).unwrap(); + assert_eq!(s.fps, "30"); + } + + #[test] + fn parser_accepts_the_subset() { + let kv = parse_flat_json(r#" { "a" : "x\n\"q\"" , "n" : 42 } "#).unwrap(); + assert_eq!( + kv, + vec![ + ("a".to_string(), Value::Str("x\n\"q\"".to_string())), + ("n".to_string(), Value::Num(42)), + ] + ); + assert_eq!(parse_flat_json("{}").unwrap(), vec![]); + } + + #[test] + fn deploy_path_takes_fps_from_sidecar() { + assert_eq!( + resolve_fps(Some("30"), None, false).unwrap(), + ("30".to_string(), FpsSource::Sidecar) + ); + } + + #[test] + fn agreeing_cli_fps_allowed_disagreeing_refused_naming_both() { + assert!(resolve_fps(Some("30"), Some("30"), false).is_ok()); + let e = resolve_fps(Some("30"), Some("25"), false).unwrap_err(); + assert!(e.contains("30") && e.contains("25"), "{e}"); + } + + #[test] + fn missing_sidecar_is_refused_without_the_bench_flag() { + let e = resolve_fps(None, Some("30"), false).unwrap_err(); + assert!(e.contains("bench"), "{e}"); + assert!(resolve_fps(None, None, false).is_err()); + } + + #[test] + fn bench_escape_hatch_requires_both_flags_and_a_valid_rate() { + assert_eq!( + resolve_fps(None, Some("30"), true).unwrap(), + ("30".to_string(), FpsSource::BenchOverride) + ); + // bench flag with a sidecar present: the sidecar is IGNORED — that is + // what "bench" means — and the CLI value wins. + assert_eq!( + resolve_fps(Some("25"), Some("30"), true).unwrap(), + ("30".to_string(), FpsSource::BenchOverride) + ); + assert!(resolve_fps(None, None, true).is_err()); + assert!(resolve_fps(None, Some("banana"), true).is_err()); + } + + #[test] + fn verify_payload_binds_bytes_to_sidecar() { + let payload = b"the asset bytes"; + let s = Sidecar { + fps: "30".into(), + sha256: crate::sha256::sha256_hex(payload), + width: None, + height: None, + }; + assert!(verify_payload(payload, &s).is_ok()); + let bad = Sidecar { + fps: "30".into(), + sha256: "a".repeat(64), + width: None, + height: None, + }; + let e = verify_payload(payload, &bad).unwrap_err(); + assert!(e.contains("sha256"), "{e}"); + } +} diff --git a/packages/dexd/src/watchdog.rs b/packages/dexd/src/watchdog.rs new file mode 100644 index 0000000..318dae3 --- /dev/null +++ b/packages/dexd/src/watchdog.rs @@ -0,0 +1,641 @@ +//! F10 — systemd `WatchdogSec=` + a hand-rolled `sd_notify`: the actor of +//! last resort for the one hazard F1/F9 cannot see. See PLAN.md's F10 entry +//! for the full design record (the "F9/F10 distinction" framing, the socket +//! analysis, the dependency decision, the timing math); this doc condenses +//! the parts an implementer reading this file needs, plus the parts a +//! reviewer of *this file specifically* needs. +//! +//! # Framing: what F10 is for, and why it is a separate feature from F1/F9 +//! +//! F9 proved a wedged mpv core produces SILENCE on this program's event +//! thread, not a block (`mpv_observe_property` getters run on the core +//! thread with the client lock dropped — verified against mpv 0.40 source, +//! see `dexd::health`'s module doc). F1 acts on that silence in-process, +//! with a bounded budget. Neither covers the one thing left: **our own event +//! thread hanging in code that is not an mpv call at all** — the canonical +//! case being `eprintln!` blocking against a wedged journald, including on +//! the `Escalate` arm whose entire job is "exit so tier 1 can take over" (see +//! `main.rs::act_on_health_action`'s `Escalate` arm). Nothing in-process can +//! detect that; it needs an external actor. That is systemd, driven by +//! `WatchdogSec=` in the unit (`deploy/dexd.service`) plus this module. +//! +//! # The liveness criterion — and why a WEDGED PLAYER FAILS it +//! +//! `main.rs` pings `WATCHDOG=1` once per `HEALTH_CHECK_SECS` tick, **after** +//! that tick's `HealthMonitor::tick` evaluation (and any resulting recovery +//! command) has completed — never from anywhere else, and specifically never +//! from the 600s heartbeat (longer than any sane `WatchdogSec`). "The process +//! is alive" alone would be worthless (PLAN.md's own words) — the reason +//! this criterion is NOT that is a property F1 already guarantees: its +//! recovery budget (`MAX_RECOVERY_ATTEMPTS`, `main.rs`) is cumulative and +//! NEVER refills (see `dexd::health`'s "why the budget never resets"), +//! so a display-wedged player's tick sequence is forced, by construction, +//! through: silence → stall detected at 2 ticks (~20s) → ≤3 budgeted +//! recoveries (~20s each) → `Escalate` → `std::process::exit(1)`. A wedged +//! display therefore emits a BOUNDED number of pings and then either exits +//! (moot — tier 1's exit-code path already handles it) or, on the one path +//! that can still hang (`Escalate`'s own `eprintln!` before `exit(1)`, or a +//! hang anywhere else in the loop body), STOPS pinging entirely — because +//! the ping sits at the very end of a duty cycle that a genuine hang, by +//! definition, never completes again. Within this hazard class — a wedged +//! mpv core or a hung event thread — there is no third state: either +//! progress continues (pings continue, correctly, nothing to do), or the +//! duty cycle stops completing (pings stop, watchdog fires). The loop +//! cannot both hang and ping. +//! +//! Scope, stated precisely: that guarantee covers wedged-core/wedged-thread +//! failures ONLY. A player whose `time-pos` keeps advancing while no photons +//! reach the wall — HDMI signal lost mid-run, panel powered off, the plane +//! presenting to a disconnected sink — reads as healthy to F1 (decode and +//! present proceed internally) and therefore pings forever. No in-process +//! liveness criterion can see that; it is a signal-level failure, explicitly +//! out of F10's scope — see PLAN.md's F10 entry ("residual states") for the +//! record and a possible future closure (DRM connector-status polling). +//! +//! # Gate placement: the ping is emitted OUTSIDE the health `Option` gate +//! +//! `main.rs` pings regardless of whether `health: Option` is +//! `Some` — i.e. even in the near-zero-probability case where the `time-pos` +//! `mpv_observe_property` registration itself failed and F1 is DISABLED for +//! the run. In that mode the ping only certifies "the event loop completed +//! an iteration," a strictly weaker claim — but stopping pings there instead +//! would convert a degraded-but-otherwise-fine run into a guaranteed +//! `WatchdogSec`-later kill loop, which is a worse outcome than the +//! diagnostically-honest weaker claim. The DISABLED warning at startup gains +//! a clause saying so. See `main.rs`'s tick site for where this is wired. +//! +//! # No new blocking call — the socket analysis +//! +//! `sd_notify` is `sendto(2)` on an `AF_UNIX SOCK_DGRAM` socket to +//! `$NOTIFY_SOCKET`. Unlike UDP, Unix datagram sockets have flow control: if +//! the receiver's (PID 1's) queue is full, a *blocking* `sendto` blocks +//! rather than dropping. PID 1's queue is generally enormous relative to a +//! 10-byte payload, but "generally" was exactly the standard of proof F9's +//! lesson raised the bar past — so this module does not rely on it: the +//! socket this crate opens is put into **non-blocking mode** +//! (`UnixDatagram::set_nonblocking(true)`, done once in `main.rs`, the +//! socket's owner), and `EAGAIN`/`EWOULDBLOCK` is treated as a **dropped +//! ping**, never retried synchronously, never panicked on — see +//! [`interpret_send_result`] and [`send_ping`]. With a 10s ping cadence and +//! `WatchdogSec=180` (`deploy/dexd.service`), systemd needs ~17 +//! consecutive drops before a spurious kill; a run that sick is not a wrong +//! restart. The dropped-ping count is surfaced in the 600s heartbeat line +//! (`watchdog=armed pings-dropped=N`, see `heartbeat.rs`) so a run degraded +//! this way is diagnosable after the fact, per this crate's "loud, never +//! silent" principle. +//! +//! # Dependency: zero. Hand-written, `std` only +//! +//! The entire protocol this crate needs is: read `$NOTIFY_SOCKET`, open an +//! unbound `AF_UNIX SOCK_DGRAM` socket, send the literal bytes `WATCHDOG=1`. +//! `std::os::unix::net::UnixDatagram` covers all of it, including the +//! abstract-namespace case (`std::os::linux::net::SocketAddrExt`, stable +//! since 1.70 — comfortably inside this crate's `rust-version = "1.85"`). +//! No `libsystemd` binding (would add a shared-object link, growing the +//! `.deb`'s `$auto`-derived `Depends`, to avoid ~150 lines) and no +//! `sd-notify` crate (removes less than it appears to — the ping policy, +//! the env handshake, and an audit of its socket handling for the +//! non-blocking guarantee this feature requires would all still be owned +//! here). See PLAN.md's F10 entry, §3, for the full comparison this crate's +//! SPEC §5c ("a dependency earns its place by removing code we would +//! otherwise own") is weighed against. `Cargo.toml`'s dependency list and +//! `$auto`-derived `.deb` `Depends` are both UNCHANGED by this feature. +//! +//! Deliberately **not** sent: `READY=1` (a no-op under `Type=simple` + +//! `NotifyAccess=main`, and sending it would invite a future switch to +//! `Type=notify` that PLAN.md explicitly forbids — a notify unit that never +//! sends `READY=1` sits inactive forever, and this program has no natural +//! "ready" moment before the endless stream starts). Also not sent: +//! `STOPPING=1` (no graceful shutdown exists, by design — the mains switch +//! IS the shutdown path). +//! +//! # A deliberate deviation from PLAN.md's literal test recipe, and why +//! +//! PLAN.md's §6 test plan calls for the non-blocking-guarantee test to +//! `shrink [the receiver's] SO_RCVBUF` via `setsockopt` before flooding it. +//! `std` exposes no safe way to do that, and this crate is +//! `#![forbid(unsafe_code)]` crate-wide (`lib.rs`) — a `forbid`, not a +//! `deny`, so it cannot be locally overridden even inside a `#[cfg(test)]` +//! module; that is the entire point of choosing `forbid` there. Rather than +//! reach for a `libc`/raw-FFI dependency (itself a SPEC §5c question this +//! feature's whole point is to avoid needing) just to shrink a buffer for a +//! test, [`tests::send_ping_eventually_reports_dropped_once_the_receiver_queue_fills`] +//! instead floods the OS's UNMODIFIED default receive buffer within a +//! bounded iteration count. A 10-byte datagram still costs real per-message +//! kernel accounting overhead, so even a generous default buffer fills +//! within at most a few thousand sends — the bound used has two orders of +//! magnitude of headroom over that. This is a real behavioural test (an +//! actual unread socket, actually exhausted) of the exact property this +//! module depends on, just without artificially engineering the buffer size +//! first. Flagged here per this crate's practice of stating a deviation +//! rather than silently reinterpreting the brief. + +use std::io; +use std::os::unix::net::{SocketAddr, UnixDatagram}; + +/// The entire wire payload this program ever sends. No trailing newline: +/// sd_notify's protocol is newline-separated `KEY=VALUE` lines, and this is +/// the only line this crate ever emits (see the module doc for why +/// `READY=1`/`STOPPING=1` are deliberately never sent) — a trailing newline +/// would be decoration systemd does not require (`sd-daemon` sources treat a +/// single-line unterminated datagram as a complete, valid notification). +pub const WATCHDOG_PING_PAYLOAD: &[u8] = b"WATCHDOG=1"; + +/// Raw environment inputs the handshake decision needs, as explicit values +/// rather than three direct `std::env::var` calls inside [`resolve`] — that +/// keeps `resolve` pure (no I/O, no global state) and testable with +/// synthetic inputs, without mutating the real process environment (which +/// races every other test in the same test binary; see `std::env::set_var`'s +/// own safety caveats). [`WatchdogEnv::from_process_env`] is the one place +/// that actually reads the real environment, for `main.rs` to call once at +/// startup. +#[derive(Debug, Clone, Default, PartialEq, Eq)] +pub struct WatchdogEnv { + pub notify_socket: Option, + pub watchdog_pid: Option, + pub watchdog_usec: Option, +} + +impl WatchdogEnv { + /// Read the three systemd notify-protocol variables from the real + /// process environment. The only non-pure function in this module; + /// everything downstream of it ([`resolve`], [`socket_addr`], + /// [`send_ping`]) takes its inputs explicitly instead. + pub fn from_process_env() -> Self { + Self { + notify_socket: std::env::var("NOTIFY_SOCKET").ok(), + watchdog_pid: std::env::var("WATCHDOG_PID").ok(), + watchdog_usec: std::env::var("WATCHDOG_USEC").ok(), + } + } +} + +/// Where to send pings, in a form that survives being decided on ANY +/// platform. Kept as this crate's own enum rather than a real +/// `std::os::unix::net::SocketAddr` — turning `Abstract` into a concrete +/// address needs `std::os::linux::net::SocketAddrExt`, which does not exist +/// outside Linux/Android, so [`resolve`] (which must stay compilable and +/// testable on macOS, the dev machine) cannot construct one directly. See +/// [`socket_addr`] for the platform-gated conversion. +#[derive(Debug, Clone, PartialEq, Eq)] +pub enum NotifySocketAddr { + /// A filesystem path — the common case (containers, VMs, most systemd + /// configurations). + Path(String), + /// `$NOTIFY_SOCKET` began with `@`: an abstract-namespace name (the + /// leading `@` already stripped, matching `sd_notify`'s own convention — + /// see systemd's `sd-daemon.c`). + Abstract(String), +} + +/// Why the watchdog is inert for this run. Every arm here describes a +/// NORMAL run, not an error: the Mac dev machine, a bench tmux session, CI, +/// and any manual invocation off systemd all land in +/// [`InertReason::NoNotifySocket`] — the overwhelmingly common case. Logged +/// once at startup, not warned. +#[derive(Debug, Clone, PartialEq, Eq)] +pub enum InertReason { + /// No `$NOTIFY_SOCKET` at all: not running under a systemd unit with + /// `NotifyAccess=` set. + NoNotifySocket, + /// `$WATCHDOG_PID` is set and does not identify this process. The + /// watchdog handshake belongs to a DIFFERENT process (e.g. a wrapper + /// script systemd also tracks under the same unit) — pinging under + /// someone else's identity would be actively wrong, not merely useless, + /// so this is inert rather than armed-anyway. A value that fails to + /// parse as a pid at all is treated identically to "set and different": + /// it cannot possibly equal our own pid either way. + WatchdogPidMismatch { ours: u32, unit: String }, +} + +impl std::fmt::Display for InertReason { + fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + match self { + InertReason::NoNotifySocket => write!(f, "no NOTIFY_SOCKET"), + InertReason::WatchdogPidMismatch { ours, unit } => write!( + f, + "WATCHDOG_PID={unit} does not match our pid {ours} -- the watchdog handshake \ + belongs to a different process" + ), + } + } +} + +/// A startup condition worth a loud warning even though the watchdog IS (or +/// becomes) armed — distinct from [`InertReason`], which is never a problem. +#[derive(Debug, Clone, PartialEq, Eq)] +pub enum ArmedWarning { + /// `$WATCHDOG_USEC` is set but small enough that this program's fixed + /// ping cadence does not clear systemd's own "ping at least twice per + /// window" guidance with any margin (PLAN.md F10 §4). Pings are still + /// sent every tick regardless — there is nothing better to do from + /// inside the process — this is purely diagnostic, flagging a likely + /// unit misconfiguration rather than something this code can fix. + WindowTooShortForCadence { window_secs: u64, tick_secs: u64 }, +} + +impl std::fmt::Display for ArmedWarning { + fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + match self { + ArmedWarning::WindowTooShortForCadence { window_secs, tick_secs } => write!( + f, + "WatchdogSec window ({window_secs}s) gives our {tick_secs}s ping cadence little \ + margin (systemd recommends pinging at least twice per window) -- likely a \ + misconfigured unit; pings are still sent every tick regardless" + ), + } + } +} + +/// The one decision this module exists to make, taken once at startup from +/// [`WatchdogEnv`] alone. +#[derive(Debug, Clone, PartialEq, Eq)] +pub enum WatchdogDecision { + Inert(InertReason), + Armed { + addr: NotifySocketAddr, + /// `$WATCHDOG_USEC`, converted to whole seconds, when systemd + /// exported it (it does whenever `WatchdogSec=` is configured on + /// the unit — absent only if `NotifyAccess=` was granted without + /// `WatchdogSec=`, an unusual configuration). Carried through so + /// `main.rs`'s startup log line can state the real window rather + /// than the value this crate merely hopes the unit file has. + window_secs: Option, + warning: Option, + }, +} + +/// Decide whether — and where — to ping, from raw env strings alone. Pure: +/// no I/O, no socket, no env access (see [`WatchdogEnv::from_process_env`] +/// for the one caller that supplies real values). `tick_secs` is +/// `HEALTH_CHECK_SECS` (`main.rs`), passed in explicitly rather than +/// imported, so this module has zero dependency on `main`'s constants. +pub fn resolve(env: &WatchdogEnv, our_pid: u32, tick_secs: u64) -> WatchdogDecision { + let Some(sock) = env.notify_socket.as_deref().filter(|s| !s.is_empty()) else { + return WatchdogDecision::Inert(InertReason::NoNotifySocket); + }; + + if let Some(unit_pid) = env.watchdog_pid.as_deref() { + let matches = unit_pid.parse::().map(|p| p == our_pid).unwrap_or(false); + if !matches { + return WatchdogDecision::Inert(InertReason::WatchdogPidMismatch { + ours: our_pid, + unit: unit_pid.to_string(), + }); + } + } + + let addr = if let Some(name) = sock.strip_prefix('@') { + NotifySocketAddr::Abstract(name.to_string()) + } else { + NotifySocketAddr::Path(sock.to_string()) + }; + + // A malformed WATCHDOG_USEC (never expected from a real systemd, but + // this is text from the environment, not a value this program controls) + // fails safe: no window is known, so no "too short" warning is invented + // either -- see the "window unknown" rendering this leaves to main.rs. + let window_secs = env + .watchdog_usec + .as_deref() + .and_then(|s| s.parse::().ok()) + .map(|usec| usec / 1_000_000); + + let warning = window_secs.and_then(|w| { + (w < tick_secs.saturating_mul(2)) + .then_some(ArmedWarning::WindowTooShortForCadence { window_secs: w, tick_secs }) + }); + + WatchdogDecision::Armed { addr, window_secs, warning } +} + +/// Resolve a [`NotifySocketAddr`] into the concrete address `send_to_addr` +/// needs. `Err` only for [`NotifySocketAddr::Abstract`] on a non-Linux +/// build (`std::os::linux::net::SocketAddrExt` does not exist there) -- +/// unreachable in the one place that matters (a real systemd host is +/// Linux), and exists only so this module still compiles and is readable on +/// the Mac dev machine rather than `#[cfg]`-ing the whole abstract-name +/// branch out of existence there. +pub fn socket_addr(target: &NotifySocketAddr) -> io::Result { + match target { + NotifySocketAddr::Path(p) => SocketAddr::from_pathname(p), + NotifySocketAddr::Abstract(name) => abstract_addr(name), + } +} + +#[cfg(target_os = "linux")] +fn abstract_addr(name: &str) -> io::Result { + use std::os::linux::net::SocketAddrExt; + SocketAddr::from_abstract_name(name.as_bytes()) +} + +#[cfg(not(target_os = "linux"))] +fn abstract_addr(_name: &str) -> io::Result { + Err(io::Error::new( + io::ErrorKind::Unsupported, + "abstract-namespace notify sockets require Linux (std::os::linux::net::SocketAddrExt) -- \ + a $NOTIFY_SOCKET starting with '@' can only be set by a real systemd host", + )) +} + +/// Outcome of one ping attempt. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum PingOutcome { + Sent, + /// The datagram did not go out -- EAGAIN/EWOULDBLOCK (receiver's queue + /// full) or any other send error (e.g. the notify socket path vanishing + /// mid-run). Every failure mode is treated identically as "try again + /// next tick": see the module doc's socket analysis for why this is + /// never escalated, never retried synchronously, and never panics. + Dropped, +} + +/// The ONLY place that decides "was this ping delivered" -- separated from +/// the I/O call itself so the policy (a failed send degrades, it is never +/// fatal and never retried inline) is directly testable with synthetic +/// [`io::Result`]s, without a real socket at all. +pub fn interpret_send_result(result: io::Result) -> PingOutcome { + match result { + Ok(_) => PingOutcome::Sent, + Err(_) => PingOutcome::Dropped, + } +} + +/// Send one watchdog ping over a socket the caller has already opened +/// non-blocking (`main.rs` owns creating and holding that socket across the +/// process's whole life, matching every other piece of process-lifetime +/// state in this crate's driver -- this function itself performs no setup, +/// so it cannot silently forget to set non-blocking mode). Never blocks, +/// never panics on a send failure: see [`interpret_send_result`]. +pub fn send_ping(socket: &UnixDatagram, addr: &SocketAddr) -> PingOutcome { + interpret_send_result(socket.send_to_addr(WATCHDOG_PING_PAYLOAD, addr)) +} + +#[cfg(test)] +mod tests { + use super::*; + + fn env(notify_socket: Option<&str>, watchdog_pid: Option<&str>, watchdog_usec: Option<&str>) -> WatchdogEnv { + WatchdogEnv { + notify_socket: notify_socket.map(str::to_string), + watchdog_pid: watchdog_pid.map(str::to_string), + watchdog_usec: watchdog_usec.map(str::to_string), + } + } + + // ---- resolve(): the handshake decision ---------------------------- + + #[test] + fn no_notify_socket_is_inert() { + let d = resolve(&env(None, None, None), 1234, 10); + assert_eq!(d, WatchdogDecision::Inert(InertReason::NoNotifySocket)); + } + + #[test] + fn empty_notify_socket_is_treated_as_absent() { + // Defensive: an exported-but-empty value is not a real address + // either way, and must not be handed to socket construction as one. + let d = resolve(&env(Some(""), None, None), 1234, 10); + assert_eq!(d, WatchdogDecision::Inert(InertReason::NoNotifySocket)); + } + + #[test] + fn path_socket_with_no_pid_or_usec_arms_with_no_window_or_warning() { + let d = resolve(&env(Some("/run/systemd/notify"), None, None), 1234, 10); + assert_eq!( + d, + WatchdogDecision::Armed { + addr: NotifySocketAddr::Path("/run/systemd/notify".to_string()), + window_secs: None, + warning: None, + } + ); + } + + #[test] + fn at_prefixed_socket_resolves_to_an_abstract_address_with_the_prefix_stripped() { + let d = resolve(&env(Some("@abstract-name"), None, None), 1234, 10); + match d { + WatchdogDecision::Armed { addr: NotifySocketAddr::Abstract(name), .. } => { + assert_eq!(name, "abstract-name"); + } + other => panic!("expected Armed/Abstract, got {other:?}"), + } + } + + #[test] + fn matching_watchdog_pid_still_arms() { + let d = resolve(&env(Some("/run/systemd/notify"), Some("1234"), None), 1234, 10); + assert!(matches!(d, WatchdogDecision::Armed { .. }), "{d:?}"); + } + + #[test] + fn mismatched_watchdog_pid_is_inert() { + let d = resolve(&env(Some("/run/systemd/notify"), Some("999"), None), 1234, 10); + assert_eq!( + d, + WatchdogDecision::Inert(InertReason::WatchdogPidMismatch { + ours: 1234, + unit: "999".to_string(), + }) + ); + } + + #[test] + fn unparseable_watchdog_pid_is_treated_as_mismatched_not_as_absent() { + let d = resolve(&env(Some("/run/systemd/notify"), Some("not-a-pid"), None), 1234, 10); + assert_eq!( + d, + WatchdogDecision::Inert(InertReason::WatchdogPidMismatch { + ours: 1234, + unit: "not-a-pid".to_string(), + }) + ); + } + + #[test] + fn watchdog_usec_with_ample_margin_arms_with_no_warning() { + // 180s window (systemd's exported microseconds), 10s cadence: 18 + // pings per window, comfortably over the 2x guidance. + let d = resolve(&env(Some("/run/systemd/notify"), None, Some("180000000")), 1234, 10); + match d { + WatchdogDecision::Armed { window_secs, warning, .. } => { + assert_eq!(window_secs, Some(180)); + assert_eq!(warning, None); + } + other => panic!("expected Armed, got {other:?}"), + } + } + + #[test] + fn watchdog_usec_below_twice_the_cadence_warns() { + // 15s window, 10s cadence: only 1.5 pings per window, under + // systemd's own "at least twice" guidance. + let d = resolve(&env(Some("/run/systemd/notify"), None, Some("15000000")), 1234, 10); + match d { + WatchdogDecision::Armed { window_secs, warning, .. } => { + assert_eq!(window_secs, Some(15)); + assert_eq!( + warning, + Some(ArmedWarning::WindowTooShortForCadence { window_secs: 15, tick_secs: 10 }) + ); + } + other => panic!("expected Armed, got {other:?}"), + } + } + + #[test] + fn watchdog_usec_exactly_twice_the_cadence_does_not_warn() { + // The boundary: 20s window / 10s cadence is exactly systemd's own + // "at least twice" guidance, not yet "less than" it. + let d = resolve(&env(Some("/run/systemd/notify"), None, Some("20000000")), 1234, 10); + match d { + WatchdogDecision::Armed { warning, .. } => assert_eq!(warning, None), + other => panic!("expected Armed, got {other:?}"), + } + } + + #[test] + fn unparseable_watchdog_usec_arms_with_no_window_and_no_invented_warning() { + let d = resolve(&env(Some("/run/systemd/notify"), None, Some("soon")), 1234, 10); + match d { + WatchdogDecision::Armed { window_secs, warning, .. } => { + assert_eq!(window_secs, None); + assert_eq!(warning, None); + } + other => panic!("expected Armed, got {other:?}"), + } + } + + // ---- the wire payload ---------------------------------------------- + + #[test] + fn ping_payload_is_exactly_watchdog_1_no_trailing_newline() { + assert_eq!(WATCHDOG_PING_PAYLOAD, b"WATCHDOG=1"); + } + + // ---- interpret_send_result: pure Ok/Err -> outcome mapping -------- + + #[test] + fn a_successful_send_is_sent() { + assert_eq!(interpret_send_result(Ok(10)), PingOutcome::Sent); + } + + #[test] + fn would_block_is_dropped_not_an_error() { + let err = io::Error::from(io::ErrorKind::WouldBlock); + assert_eq!(interpret_send_result(Err(err)), PingOutcome::Dropped); + } + + #[test] + fn any_other_send_failure_also_degrades_to_dropped_never_panics() { + // e.g. the notify socket path vanishing mid-run (ENOENT) -- see the + // module doc: every failure mode is "try again next tick", none is + // special-cased or escalated. + let err = io::Error::from(io::ErrorKind::NotFound); + assert_eq!(interpret_send_result(Err(err)), PingOutcome::Dropped); + } + + // ---- socket_addr(): platform-gated address construction ----------- + + #[test] + fn socket_addr_resolves_a_path_target() { + let target = NotifySocketAddr::Path("/tmp/dexd-watchdog-example.sock".to_string()); + let addr = socket_addr(&target).expect("path address resolves"); + assert_eq!( + addr.as_pathname(), + Some(std::path::Path::new("/tmp/dexd-watchdog-example.sock")) + ); + } + + #[cfg(target_os = "linux")] + #[test] + fn socket_addr_resolves_an_abstract_target_on_linux() { + let target = NotifySocketAddr::Abstract("dexd-watchdog-test".to_string()); + let addr = socket_addr(&target); + assert!(addr.is_ok(), "{addr:?}"); + } + + #[cfg(not(target_os = "linux"))] + #[test] + fn socket_addr_refuses_an_abstract_target_off_linux() { + let target = NotifySocketAddr::Abstract("dexd-watchdog-test".to_string()); + let err = socket_addr(&target).expect_err("abstract names require Linux"); + assert_eq!(err.kind(), io::ErrorKind::Unsupported); + } + + // ---- send_ping(): real, unmocked sockets --------------------------- + + /// `/tmp` directly, NOT `std::env::temp_dir()`: `AF_UNIX` paths are + /// capped at `sizeof(sockaddr_un.sun_path)` (104 bytes on macOS, 108 on + /// Linux), and macOS's real temp dir + /// (`/var/folders/.../T/`) is already close to that budget on its own -- + /// `std::env::temp_dir()` here produced `EINVAL: path must be shorter + /// than SUN_LEN` in practice. A short, fixed prefix plus a process-local + /// atomic counter keeps every path well under the limit on both + /// platforms without needing to reason about `$TMPDIR`'s length. + fn unique_socket_path(label: &str) -> std::path::PathBuf { + static COUNTER: std::sync::atomic::AtomicU32 = std::sync::atomic::AtomicU32::new(0); + let n = COUNTER.fetch_add(1, std::sync::atomic::Ordering::Relaxed); + std::path::PathBuf::from(format!("/tmp/dwd-{:x}-{n:x}-{label}.sock", std::process::id())) + } + + #[test] + fn send_ping_delivers_the_exact_payload_to_a_real_receiver() { + let path = unique_socket_path("happy"); + let _ = std::fs::remove_file(&path); + let receiver = UnixDatagram::bind(&path).expect("bind receiver"); + receiver + .set_read_timeout(Some(std::time::Duration::from_secs(2))) + .expect("set_read_timeout"); + + let addr = SocketAddr::from_pathname(&path).expect("path address"); + let sender = UnixDatagram::unbound().expect("unbound sender"); + sender.set_nonblocking(true).expect("set_nonblocking"); + + assert_eq!(send_ping(&sender, &addr), PingOutcome::Sent); + + let mut buf = [0u8; 32]; + let n = receiver.recv(&mut buf).expect("recv"); + assert_eq!(&buf[..n], WATCHDOG_PING_PAYLOAD); + + drop(receiver); + let _ = std::fs::remove_file(&path); + } + + #[test] + fn send_ping_eventually_reports_dropped_once_the_receiver_queue_fills() { + // See the module doc's "deliberate deviation" section for why this + // floods the OS's UNMODIFIED default receive buffer, bounded by + // ATTEMPTS, rather than shrinking SO_RCVBUF first. + let path = unique_socket_path("flood"); + let _ = std::fs::remove_file(&path); + let receiver = UnixDatagram::bind(&path).expect("bind receiver"); + // Deliberately never read from `receiver`. + + let addr = SocketAddr::from_pathname(&path).expect("path address"); + let sender = UnixDatagram::unbound().expect("unbound sender"); + sender.set_nonblocking(true).expect("set_nonblocking"); + + const ATTEMPTS: usize = 200_000; + let mut saw_dropped = false; + for _ in 0..ATTEMPTS { + if send_ping(&sender, &addr) == PingOutcome::Dropped { + saw_dropped = true; + break; + } + } + + drop(receiver); + let _ = std::fs::remove_file(&path); + + assert!( + saw_dropped, + "expected send_to_addr to eventually return WouldBlock (-> Dropped) against an \ + unread receiver within {ATTEMPTS} sends -- if this fails, either this platform's \ + default SO_RCVBUF is unexpectedly huge, or the non-blocking guarantee broke \ + (set_nonblocking silently not honoured would hang this loop instead of failing \ + this assertion -- see the module doc)" + ); + } +} diff --git a/packages/dexd/tests/cli.rs b/packages/dexd/tests/cli.rs new file mode 100644 index 0000000..c765351 --- /dev/null +++ b/packages/dexd/tests/cli.rs @@ -0,0 +1,1501 @@ +//! T3 — failure-path integration tests. Every serious bug in this program +//! lived in a failure path; the happy path was never the problem. Each test +//! asserts EXIT BEHAVIOUR (code + boundedness), not output niceties. +//! +//! These tests spawn the real binary, which links libmpv — this target runs +//! on the Pi (`cargo test`); on the Mac use `cargo test --lib`. +//! +//! DISPLAY SAFETY: a soak may own the display. Every invocation that can +//! reach mpv_create MUST carry `--no-defaults --opt vo=null --opt vid=no +//! --opt aid=no`: the null VO never touches DRM, and deselecting all tracks +//! makes mpv end deterministically (NOTHING_TO_PLAY -> END_FILE) instead of +//! playing forever. + +use std::io::Read; +use std::path::{Path, PathBuf}; +use std::process::{Command, Stdio}; +use std::time::{Duration, Instant}; + +/// Exit-code contract (see usage()): refused before playback vs runtime failure. +const GATE_EXIT: i32 = 2; +const RUNTIME_EXIT: i32 = 1; + +/// Outcome of one run. `exit_code` is None when the process had to be killed +/// at the deadline OR died by signal. For the exit-code tests both are +/// failures the `Some(..)` assertions catch — but for the live-fire survival +/// test the polarity flips (None is the PASS), so the two None causes must be +/// distinguishable: `deadline_killed` is true only when THIS HARNESS killed +/// the child at the deadline. A child that died by signal on its own +/// (SIGSEGV/SIGABRT) has `exit_code: None` with `deadline_killed: false`, +/// and conflating that with survival would let a crashing recovery pass a +/// survival assertion. +struct Run { + exit_code: Option, + deadline_killed: bool, + stderr: String, +} + +/// Spawn dexd, wait at most `deadline`, kill on overrun. THE DEADLINE IS +/// THE ASSERTION: a player that hangs on a failure path is this program's +/// worst outcome — alive, supervisor green, screen black. The shipped +/// END_FILE idle-hang was exactly that. +fn run_with_deadline(args: &[&str], deadline: Duration) -> Run { + let mut child = Command::new(env!("CARGO_BIN_EXE_dexd")) + .args(args) + .stdout(Stdio::null()) + .stderr(Stdio::piped()) + .spawn() + .expect("spawn dexd"); + // Drain stderr on a thread so a chatty child can never fill the pipe and + // block — a blocked child would masquerade as a hang. + let mut pipe = child.stderr.take().expect("stderr piped"); + let reader = std::thread::spawn(move || { + let mut s = String::new(); + pipe.read_to_string(&mut s).ok(); + s + }); + let start = Instant::now(); + let (exit_code, deadline_killed) = loop { + match child.try_wait().expect("try_wait") { + Some(status) => break (status.code(), false), + None if start.elapsed() > deadline => { + child.kill().ok(); + child.wait().ok(); + break (None, true); + } + None => std::thread::sleep(Duration::from_millis(50)), + } + }; + let stderr = reader.join().expect("stderr reader"); + Run { + exit_code, + deadline_killed, + stderr, + } +} + +/// Unique-per-test scratch path (std::env::temp_dir; no cleanup needed). +fn temp_path(name: &str) -> PathBuf { + let mut p = std::env::temp_dir(); + p.push(format!("dexd-test-{}-{}", std::process::id(), name)); + p +} + +/// Minimal Annex-B HEVC scaffold: VPS, SPS, PPS, then an IDR_N_LP slice -- +/// the exact NAL payloads of a real single-frame x265 encode (16x16 black, +/// closed GOP: `ffmpeg -f lavfi -i color=c=black:s=16x16:d=1:r=1 -frames:v 1 +/// -c:v libx265 -x265-params keyint=1 -f hevc`), not synthetic filler. +/// +/// This MUST be real, parseable HEVC, not arbitrary filler bytes: `loop://` +/// never returns EOF -- that is the entire point of the player -- and +/// libavformat's probe (`avformat_find_stream_info`) only gives up early +/// when it hits EOF. Verified on-device: fed a garbage VPS/SPS through the +/// endless non-EOF stream, the probe can extract nothing AND never sees +/// end-of-stream, so it keeps demanding more data forever -- CPU pinned at +/// 100%, RSS climbing unbounded, silent (no stderr) until the test deadline +/// kills it. A real SPS lets `avformat_find_stream_info` resolve +/// width/height/profile from the header in one pass, independent of the +/// stream's (infinite) length, so mpv reaches "no video or audio streams +/// selected" -> END_FILE in well under a second. +/// +/// Structurally what the F4 NAL gate requires; not meant to actually +/// *decode* -- video is deselected (`--opt vid=no`) before decode is ever +/// attempted, so only stream *identification* needs to succeed. +fn stub_annexb() -> Vec { + fn nal(nal_type: u8, payload: &[u8]) -> Vec { + // 4-byte start code + 2-byte NAL header (forbidden=0, layer=0, tid+1=1) + let mut v = vec![0, 0, 0, 1, nal_type << 1, 0x01]; + v.extend_from_slice(payload); + v + } + let mut v = Vec::new(); + v.extend(nal( + 32, + &[ + 0x0c, 0x01, 0xff, 0xff, 0x04, 0x08, 0x00, 0x00, 0x03, 0x00, 0x9f, 0xa8, 0x00, 0x00, + 0x03, 0x00, 0x00, 0x1e, 0xba, 0x02, 0x40, + ], + )); // VPS + v.extend(nal( + 33, + &[ + 0x01, 0x04, 0x08, 0x00, 0x00, 0x03, 0x00, 0x9f, 0xa8, 0x00, 0x00, 0x03, 0x00, 0x00, + 0x1e, 0xa0, 0x88, 0x45, 0x96, 0xea, 0xaf, 0x2b, 0xc0, 0x5a, 0x02, 0x00, 0x00, 0x03, + 0x00, 0x02, 0x00, 0x00, 0x03, 0x00, 0x02, 0x10, + ], + )); // SPS + v.extend(nal(34, &[0xc1, 0x73, 0xc0, 0x89])); // PPS + v.extend(nal( + 20, + &[0xaf, 0x78, 0xf8, 0x5d, 0xf7, 0xff, 0xe2, 0xc0, 0x38], + )); // IDR_N_LP slice + v +} + +/// Write `.json` binding `bytes` at `fps` — the deploy-path fixture. +/// Inert before task 6 (the player ignores it); binding afterwards. +fn write_sidecar(asset: &Path, bytes: &[u8], fps: &str) { + let sha = dexd::sha256::sha256_hex(bytes); + std::fs::write( + format!("{}.json", asset.display()), + format!(r#"{{"fps":"{fps}","sha256":"{sha}"}}"#), + ) + .unwrap(); +} + +/// F6 — write a minimal, always-satisfiable exhibit config at a unique temp +/// path and return it. `display_mode: "auto"` skips the sysfs mode +/// pre-flight (no WxH to check), and `kms_force: "none"` is paired with +/// `write_no_video_cmdline` below so the cmdline gate passes deterministically +/// regardless of what the REAL host's `/proc/cmdline` happens to contain — +/// deploy-path tests below pass BOTH this and `--proc-cmdline +/// ` so the F6 gates are satisfied without +/// depending on the test host's kernel command line, and the test's own gate +/// (F3/F4/opt) still runs exactly as it did before F6 existed. +fn write_exhibit_config(name: &str) -> PathBuf { + let p = temp_path(name); + std::fs::write(&p, r#"{"display_mode":"auto","kms_force":"none"}"#).unwrap(); + p +} + +/// F6 — a synthetic `/proc/cmdline` with no `video=` token at all, so the +/// cmdline gate's "kms_force=none, expect no token" branch always matches, +/// independent of the real host's actual kernel command line (a CI container +/// has none either way, but the real bench Pi, once F6 is deployed there, +/// legitimately does). +fn write_no_video_cmdline(name: &str) -> PathBuf { + let p = temp_path(name); + std::fs::write(&p, "console=ttyS0 root=/dev/mmcblk0p2 rootwait quiet\n").unwrap(); + p +} + +#[test] +fn missing_file_exits_2_and_names_the_path() { + let cfg = write_exhibit_config("missingfile-exhibit.json"); + let cl = write_no_video_cmdline("missingfile-cmdline"); + let r = run_with_deadline( + &[ + "/nonexistent/dexd-test.265", + "--fps", + "30", + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!( + r.stderr.contains("/nonexistent/dexd-test.265"), + "stderr must name the path: {}", + r.stderr + ); +} + +#[test] +fn empty_file_exits_2() { + let p = temp_path("empty.265"); + std::fs::write(&p, b"").unwrap(); + let cfg = write_exhibit_config("emptyfile-exhibit.json"); + let cl = write_no_video_cmdline("emptyfile-cmdline"); + let r = run_with_deadline( + &[ + p.to_str().unwrap(), + "--fps", + "30", + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + // Pin the INTENDED gate (main.rs's empty-payload guard), not just any + // refusal: without this, deleting that guard would leave the test green + // via the (also exit-2) missing-sidecar path instead. + assert!(r.stderr.contains("is empty"), "stderr: {}", r.stderr); +} + +// REMOVED 2026-08-17: `no_args_exits_2_with_usage`, which asserted that a bare +// `dexd` prints usage and exits 2. +// +// It was correct until F6 moved the asset into the exhibit config. Now +// `ExecStart=/usr/bin/dexd` passes exactly zero arguments, so "no args" is +// the SHIPPED invocation rather than an operator error -- and this test caught +// that regression on the Pi within minutes, which it could only ever have done +// there: the Mac cannot link these targets at all. +// +// Not merely inverted to "no args must NOT print usage", because on a card +// whose config and asset are both complete, a bare run would START PLAYBACK +// and take DRM master inside a test -- a visible glitch on a device that may +// be mid-exhibition, and a failing assertion anyway (the harness deadline, not +// an exit code). A test must not be able to interrupt a show. +// +// The property it protected -- a missing positional is not an argv error -- +// is covered host-independently by +// `no_positional_asset_is_accepted_and_reaches_the_config` below, which +// supplies its own --exhibit-config and lands on a gate rather than on +// playback. + +/// A `--fps`/`--mode` with no following value used to evaporate silently +/// (`args.get(i)` -> `None` -> the flag is just dropped) instead of refusing: +/// an edited systemd unit or a line-continuation typo would start the player +/// on the connector-preferred mode, or with no fps cross-check, with zero +/// error. `--opt` already fell into `usage()` on a missing value -- these two +/// flags must too. +#[test] +fn fps_flag_missing_value_refused_exit_2_with_usage() { + let r = run_with_deadline(&["/nonexistent/x.265", "--fps"], Duration::from_secs(10)); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!(r.stderr.contains("usage"), "stderr: {}", r.stderr); +} + +#[test] +fn mode_flag_missing_value_refused_exit_2_with_usage() { + let r = run_with_deadline(&["/nonexistent/x.265", "--mode"], Duration::from_secs(10)); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!(r.stderr.contains("usage"), "stderr: {}", r.stderr); +} + +/// THE regression for shipped bug #1: mpv_create enables idle mode, so a +/// playback failure emits END_FILE and then idles FOREVER unless the event +/// loop treats END_FILE as fatal. Force a deterministic, display-free +/// playback failure (vid=no + aid=no deselect every track -> mpv ends with +/// "nothing to play" -> END_FILE) and assert the process EXITS, code 1, +/// within the deadline. Before the END_FILE fix this exact scenario sat in +/// idle indefinitely with the supervisor reading green. +#[test] +fn playback_failure_exits_nonzero_never_hangs() { + let p = temp_path("stub.265"); + let bytes = stub_annexb(); + std::fs::write(&p, &bytes).unwrap(); + write_sidecar(&p, &bytes, "30"); + let cfg = write_exhibit_config("playbackfail-exhibit.json"); + let cl = write_no_video_cmdline("playbackfail-cmdline"); + let r = run_with_deadline( + &[ + p.to_str().unwrap(), + "--fps", + "30", + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + "--no-defaults", + "--opt", + "vo=null", + "--opt", + "vid=no", + "--opt", + "aid=no", + ], + Duration::from_secs(30), + ); + assert!( + r.exit_code.is_some(), + "player HUNG on a playback failure (killed at deadline); stderr: {}", + r.stderr + ); + assert_eq!(r.exit_code, Some(RUNTIME_EXIT), "stderr: {}", r.stderr); + // Exit 1 is also returned by four OTHER failure sites (mpv_create, + // set_opt, mpv_initialize, loadfile). Without this, a regression that + // makes one of those fail instead -- e.g. `--opt` plumbing silently + // broken so `vo=null` never reaches mpv -- would produce an instant exit + // 1 and this test would stay green while no longer exercising the + // END_FILE branch at all. Pin the branch, not just the exit code. + assert!( + r.stderr.contains("playback ended"), + "exited 1 but not via the END_FILE branch this test exists to pin: {}", + r.stderr + ); +} + +/// Post-F4: garbage never reaches mpv — the NAL gate refuses it at startup, +/// fast, with a message naming the actual problem. (The pre-F4 version of +/// this test allowed exit 1 via mpv's demux-probe failure; the event-loop +/// hang class is covered by playback_failure_exits_nonzero_never_hangs.) +#[test] +fn garbage_bytes_refused_at_the_gate_exit_2() { + let p = temp_path("garbage.265"); + // 64 KiB of bytes in 0x02..=0x7E: no 0x00/0x01 -> no start code anywhere. + let bytes: Vec = (0..65536u32) + .map(|i| ((i.wrapping_mul(2654435761) >> 24) as u8 % 0x7d) + 0x02) + .collect(); + std::fs::write(&p, &bytes).unwrap(); + write_sidecar(&p, &bytes, "30"); // hash MATCHES: proves the gate, not F3, refuses + let cfg = write_exhibit_config("garbage-exhibit.json"); + let cl = write_no_video_cmdline("garbage-cmdline"); + let r = run_with_deadline( + &[ + p.to_str().unwrap(), + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + "--no-defaults", + "--opt", + "vo=null", + "--opt", + "vid=no", + "--opt", + "aid=no", + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!(r.stderr.contains("start code"), "stderr: {}", r.stderr); +} + +/// Wrong-but-intact: a CRA-led (open GOP) asset with a CORRECT sidecar hash. +/// F3 passes — the bytes are exactly what was ingested — and F4 must still +/// refuse, proving the hash alone is insufficient. +#[test] +fn open_gop_asset_refused_at_the_gate_exit_2() { + let p = temp_path("opengop.265"); + let mut bytes = Vec::new(); + for t in [32u8, 33, 34, 21] { + // VPS SPS PPS CRA + bytes.extend([0, 0, 0, 1, t << 1, 0x01]); + bytes.extend([0x2a; 8]); + } + std::fs::write(&p, &bytes).unwrap(); + write_sidecar(&p, &bytes, "30"); + let cfg = write_exhibit_config("opengop-exhibit.json"); + let cl = write_no_video_cmdline("opengop-cmdline"); + let r = run_with_deadline( + &[ + p.to_str().unwrap(), + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + "--no-defaults", + "--opt", + "vo=null", + "--opt", + "vid=no", + "--opt", + "aid=no", + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!(r.stderr.contains("CRA"), "stderr: {}", r.stderr); +} + +// ---- F3: sidecar binding (task 6) --------------------------------------- + +#[test] +fn missing_sidecar_refused_exit_2_naming_the_sidecar_path() { + let p = temp_path("nosidecar.265"); + std::fs::write(&p, stub_annexb()).unwrap(); + let _ = std::fs::remove_file(format!("{}.json", p.display())); + let cfg = write_exhibit_config("missingsidecar-exhibit.json"); + let cl = write_no_video_cmdline("missingsidecar-cmdline"); + let r = run_with_deadline( + &[ + p.to_str().unwrap(), + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + "--no-defaults", + "--opt", + "vo=null", + "--opt", + "vid=no", + "--opt", + "aid=no", + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!( + r.stderr.contains(&format!("{}.json", p.display())), + "stderr must name the sidecar path: {}", + r.stderr + ); +} + +/// The truncated-copy case F4 can NEVER catch: leading NALs intact, tail +/// missing. Only the hash sees it. The sidecar binds the FULL bytes; the +/// file on disk is truncated. +#[test] +fn truncated_asset_vs_full_hash_refused_exit_2() { + let p = temp_path("truncated.265"); + let full = stub_annexb(); + write_sidecar(&p, &full, "30"); + std::fs::write(&p, &full[..full.len() - 20]).unwrap(); + let cfg = write_exhibit_config("truncated-exhibit.json"); + let cl = write_no_video_cmdline("truncated-cmdline"); + let r = run_with_deadline( + &[ + p.to_str().unwrap(), + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + "--no-defaults", + "--opt", + "vo=null", + "--opt", + "vid=no", + "--opt", + "aid=no", + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!(r.stderr.contains("sha256"), "stderr: {}", r.stderr); +} + +#[test] +fn fps_contradicting_sidecar_refused_exit_2_naming_both() { + let p = temp_path("fpsconflict.265"); + let bytes = stub_annexb(); + std::fs::write(&p, &bytes).unwrap(); + write_sidecar(&p, &bytes, "30"); + let cfg = write_exhibit_config("fpsconflict-exhibit.json"); + let cl = write_no_video_cmdline("fpsconflict-cmdline"); + let r = run_with_deadline( + &[ + p.to_str().unwrap(), + "--fps", + "25", + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + "--no-defaults", + "--opt", + "vo=null", + "--opt", + "vid=no", + "--opt", + "aid=no", + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!( + r.stderr.contains("25") && r.stderr.contains("30"), + "stderr must name both rates: {}", + r.stderr + ); +} + +/// Exit 1 — the mpv path — proves the gates PASSED with an agreeing --fps. +#[test] +fn agreeing_fps_and_sidecar_reach_playback() { + let p = temp_path("fpsagree.265"); + let bytes = stub_annexb(); + std::fs::write(&p, &bytes).unwrap(); + write_sidecar(&p, &bytes, "30"); + let cfg = write_exhibit_config("fpsagree-exhibit.json"); + let cl = write_no_video_cmdline("fpsagree-cmdline"); + let r = run_with_deadline( + &[ + p.to_str().unwrap(), + "--fps", + "30", + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + "--no-defaults", + "--opt", + "vo=null", + "--opt", + "vid=no", + "--opt", + "aid=no", + ], + Duration::from_secs(30), + ); + assert_eq!(r.exit_code, Some(RUNTIME_EXIT), "stderr: {}", r.stderr); + // Disambiguate from the other four exit-1 sites -- see the comment on + // playback_failure_exits_nonzero_never_hangs. + assert!(r.stderr.contains("playback ended"), "stderr: {}", r.stderr); +} + +/// A rejected mpv option is a deterministic, operator-fixable bad invocation +/// -- the same asset + flags fail identically on every restart -- so per the +/// exit-code contract it must be 2 ("fix and redeploy"), not 1 ("the +/// supervisor restarts"): reading it as transient sends deploy-night triage +/// looking in the wrong place. +#[test] +fn bad_opt_value_refused_exit_2_not_1() { + let p = temp_path("badopt.265"); + let bytes = stub_annexb(); + std::fs::write(&p, &bytes).unwrap(); + write_sidecar(&p, &bytes, "30"); + let cfg = write_exhibit_config("badopt-exhibit.json"); + let cl = write_no_video_cmdline("badopt-cmdline"); + let r = run_with_deadline( + &[ + p.to_str().unwrap(), + "--fps", + "30", + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + "--no-defaults", + "--opt", + "vo=null", + "--opt", + "vid=no", + "--opt", + "aid=no", + "--opt", + "this-option-does-not-exist=1", + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); +} + +/// SUSPECTED-then-confirmed (adversarial review, 2026-08-15): a rational fps +/// string reaches mpv's `container-fps-override` and is accepted end to end. +/// Verified live on the bench Pi at review time via manual invocation; pinned +/// here so a future mpv, or a future edit to how fps is plumbed, cannot +/// silently regress every NTSC-rate asset into an exit-1 boot loop. +#[test] +fn rational_fps_reaches_playback() { + let p = temp_path("ntsc.265"); + std::fs::write(&p, stub_annexb()).unwrap(); // bench path: no sidecar needed + let r = run_with_deadline( + &[ + p.to_str().unwrap(), + "--bench-no-sidecar", + "--fps", + "30000/1001", + "--no-defaults", + "--opt", + "vo=null", + "--opt", + "vid=no", + "--opt", + "aid=no", + ], + Duration::from_secs(30), + ); + assert_eq!(r.exit_code, Some(RUNTIME_EXIT), "stderr: {}", r.stderr); + assert!(r.stderr.contains("playback ended"), "stderr: {}", r.stderr); +} + +/// Exit 1, not 2: the two-flag bench escape hatch bypasses the sidecar gate +/// and reaches playback with no sidecar on disk. +#[test] +fn bench_escape_hatch_bypasses_sidecar() { + let p = temp_path("bench.265"); + std::fs::write(&p, stub_annexb()).unwrap(); // deliberately no sidecar + let r = run_with_deadline( + &[ + p.to_str().unwrap(), + "--bench-no-sidecar", + "--fps", + "30", + "--no-defaults", + "--opt", + "vo=null", + "--opt", + "vid=no", + "--opt", + "aid=no", + ], + Duration::from_secs(30), + ); + assert_eq!(r.exit_code, Some(RUNTIME_EXIT), "stderr: {}", r.stderr); + // Disambiguate from the other four exit-1 sites -- see the comment on + // playback_failure_exits_nonzero_never_hangs. + assert!(r.stderr.contains("playback ended"), "stderr: {}", r.stderr); +} + +#[test] +fn bench_flag_without_fps_refused_exit_2() { + let p = temp_path("benchnofps.265"); + std::fs::write(&p, stub_annexb()).unwrap(); + let r = run_with_deadline( + &[ + p.to_str().unwrap(), + "--bench-no-sidecar", + "--no-defaults", + "--opt", + "vo=null", + "--opt", + "vid=no", + "--opt", + "aid=no", + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); +} + +// ---- F7: build identity + heartbeat (task 8) ----------------------------- + +#[test] +fn startup_identifies_version_and_build() { + // Even a refused start must identify its build — a field journal that + // begins with an unidentifiable process is undebuggable weeks later. + let r = run_with_deadline(&["/nonexistent/x.265"], Duration::from_secs(10)); + assert_eq!(r.exit_code, Some(GATE_EXIT)); + assert!( + r.stderr + .contains(&format!("dexd {}", env!("CARGO_PKG_VERSION"))), + "stderr: {}", + r.stderr + ); +} + +// ---- F1: tier-0 health check does not disrupt normal operation ---------- + +/// F1 registers `mpv_observe_property("time-pos", ...)` unconditionally +/// whenever `mpv_initialize` succeeds -- i.e. on every test above that +/// reaches playback. This pins that registration succeeding SILENTLY (no +/// "mpv_observe_property" warning in stderr) as its own assertion, so a +/// future FFI slip (wrong arg order, wrong format constant, wrong function +/// signature) that makes registration fail -- but not crash -- gets a +/// dedicated regression test instead of only ever showing up as a +/// silently-disabled safety net nobody notices. +/// +/// What this does NOT exercise: an actual stall + in-place recovery + +/// escalation. Doing that safely would need real decode with a selected +/// video track and a bounded-but-nonzero wait for two ~10s health-check +/// ticks to elapse -- and this suite's mandatory `--opt vid=no` (see the +/// module doc above) exists specifically to forbid letting any CLI test +/// reach real decode, because the endless-stream design means such a test +/// could never end on its own except by being killed at a deadline. The +/// escalation POLICY itself (attempts, thresholds, when it gives up, the +/// position-baseline reset across a recovery) is pure logic and is +/// exhaustively tested in src/health.rs with none of that risk; this test +/// is the narrow slice of the mpv-facing half that CAN be exercised here +/// without touching decode or the display. The full mpv-facing behaviour +/// (real time-pos progressing, a real health-check tick, a real recovery) +/// is verified manually on the Pi against the actual display and asset -- +/// see the crate's PLAN.md F1 entry and this task's session notes. +#[test] +fn health_check_registers_without_warning_during_normal_playback() { + let p = temp_path("healthreg.265"); + let bytes = stub_annexb(); + std::fs::write(&p, &bytes).unwrap(); + write_sidecar(&p, &bytes, "30"); + let cfg = write_exhibit_config("healthreg-exhibit.json"); + let cl = write_no_video_cmdline("healthreg-cmdline"); + let r = run_with_deadline( + &[ + p.to_str().unwrap(), + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + "--no-defaults", + "--opt", + "vo=null", + "--opt", + "vid=no", + "--opt", + "aid=no", + ], + Duration::from_secs(30), + ); + assert_eq!(r.exit_code, Some(RUNTIME_EXIT), "stderr: {}", r.stderr); + assert!( + !r.stderr.contains("mpv_observe_property"), + "F1's time-pos subscription failed to register: {}", + r.stderr + ); +} + +#[test] +fn heartbeat_zero_is_emitted_at_startup() { + // The 10-minute cadence is untestable in a test budget; heartbeat #0 + // right after loadfile proves temperature reading and line formatting + // on every boot — and therefore here. Since F9 it does NOT prove the + // mpv property subscriptions (frame-drops/vo-delayed/pos read "n/a" at + // heartbeat #0 by design, since nothing has decoded yet); that needs a + // longer-running on-device check, not this test. + let p = temp_path("heartbeat.265"); + let bytes = stub_annexb(); + std::fs::write(&p, &bytes).unwrap(); + write_sidecar(&p, &bytes, "30"); + let cfg = write_exhibit_config("heartbeat-exhibit.json"); + let cl = write_no_video_cmdline("heartbeat-cmdline"); + let r = run_with_deadline( + &[ + p.to_str().unwrap(), + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + "--no-defaults", + "--opt", + "vo=null", + "--opt", + "vid=no", + "--opt", + "aid=no", + ], + Duration::from_secs(30), + ); + assert_eq!(r.exit_code, Some(RUNTIME_EXIT), "stderr: {}", r.stderr); + assert!( + r.stderr.contains("heartbeat wraps="), + "stderr: {}", + r.stderr + ); +} + +// ---- T7: --force-recovery-after-secs, the live-fire bench probe --------- +// +// C1 (PLAN.md's F1 addendum) shipped and reached the bench without ever +// having been exercised against a live mpv: every in-place recovery killed +// the process on its own first step, and nothing in the suite would have +// caught it. T7 closes that gap with a bench-only flag that forces the +// SAME `AttemptRecovery` decision an organic stall would, on a timer, +// during otherwise-healthy playback -- see dexd::health's +// `HealthMonitor::force_recovery` / `ForceRecoveryTrigger` for the pure +// logic and main.rs's `act_on_health_action` for why a forced probe drives +// the identical mpv-facing mechanics a real stall would. +// +// The tests below stay INSIDE this file's mandatory `--opt vid=no --opt +// aid=no` rule (see the module doc at the top of this file): they prove the +// CLI plumbing -- the flag parses, the "impossible to enable accidentally +// in a deployment" gate refuses it without --bench-no-sidecar, and the loud +// arming warning prints -- without ever letting the forced trigger actually +// fire, since firing needs a health check tick against playback that is +// still alive, and vid=no/aid=no makes mpv reach "nothing to play" and end +// in well under a second (see stub_annexb's doc comment). Actually +// observing the forced recovery succeed against a live mpv needs REAL +// decode, which this file's rule exists to keep out of the automated suite +// -- that scenario is the #[ignore]d test below instead. + +/// The gate itself (main.rs, checked on CLI shape alone, before the asset +/// is even read): `--force-recovery-after-secs` without `--bench-no-sidecar` +/// is refused, regardless of what -- if anything -- exists on disk at the +/// given path. This is the "impossible to enable accidentally in a +/// deployment" requirement, made concrete: a real deployment's ExecStart +/// never passes --bench-no-sidecar (deploy/dexd.service always binds a +/// real sidecar), so this flag can never end up armed against a gallery +/// show, however it got pasted into a command line. +#[test] +fn force_recovery_without_bench_no_sidecar_refused_exit_2() { + let r = run_with_deadline( + &["/nonexistent/x.265", "--force-recovery-after-secs", "5"], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!( + r.stderr.contains("--bench-no-sidecar"), + "stderr must explain the required pairing: {}", + r.stderr + ); +} + +/// Same missing-value discipline as `--fps`/`--mode` +/// (fps_flag_missing_value_refused_exit_2_with_usage): a flag at the end of +/// argv with no following value must refuse loudly via usage(), not +/// evaporate into "flag absent". +#[test] +fn force_recovery_flag_missing_value_refused_exit_2_with_usage() { + let r = run_with_deadline( + &["/nonexistent/x.265", "--force-recovery-after-secs"], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!(r.stderr.contains("usage"), "stderr: {}", r.stderr); +} + +/// A non-numeric value must also refuse via usage(), not silently parse as +/// 0 or panic the process. +#[test] +fn force_recovery_flag_non_numeric_value_refused_exit_2_with_usage() { + let r = run_with_deadline( + &[ + "/nonexistent/x.265", + "--force-recovery-after-secs", + "soon", + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!(r.stderr.contains("usage"), "stderr: {}", r.stderr); +} + +/// Paired with --bench-no-sidecar (the only way the gate above ever +/// accepts it), the flag is armed: startup must print the loud "BENCH ONLY +/// (T7)... ARMED" warning, and playback must proceed exactly as it does +/// without the flag (vid=no/aid=no's deterministic fast exit via "nothing +/// to play"). `N` is large enough that the forced trigger provably never +/// gets a chance to fire before that fast exit, so this test cannot +/// flake on the race between the two -- it exists to prove the plumbing +/// and the warning, not the live-fire behaviour itself. +#[test] +fn force_recovery_flag_with_bench_no_sidecar_arms_and_reaches_playback() { + let p = temp_path("forcerecovery.265"); + std::fs::write(&p, stub_annexb()).unwrap(); // bench path: no sidecar needed + let r = run_with_deadline( + &[ + p.to_str().unwrap(), + "--bench-no-sidecar", + "--fps", + "30", + "--force-recovery-after-secs", + "3600", + "--no-defaults", + "--opt", + "vo=null", + "--opt", + "vid=no", + "--opt", + "aid=no", + ], + Duration::from_secs(30), + ); + assert_eq!(r.exit_code, Some(RUNTIME_EXIT), "stderr: {}", r.stderr); + assert!(r.stderr.contains("playback ended"), "stderr: {}", r.stderr); + assert!( + r.stderr.contains("BENCH ONLY (T7)") && r.stderr.contains("ARMED"), + "must print the loud arming warning: {}", + r.stderr + ); + // The trigger must not have fired in this short a run -- if it had, + // that would mean it fired against a process already past "nothing to + // play", which is not the scenario this test is designed to prove. + assert!( + !r.stderr.contains("T7 bench probe"), + "the forced trigger must not have had a chance to fire here: {}", + r.stderr + ); +} + +/// T7 -- LIVE-FIRE test for F1's in-place recovery against a REAL mpv +/// instance: forces a tier-0 recovery a few seconds into otherwise-healthy +/// playback and asserts the process SURVIVES it (recovery absorbed, still +/// running) -- the exact scenario C1 broke. `is_expected_recovery_stop`'s +/// pure-logic tests in main.rs pin the boolean condition that fixes C1; +/// this test is the live-fire check that a REAL mpv event stream actually +/// produces the shape that condition expects. +/// +/// Deliberately NOT part of the automated (default) suite: unlike every +/// other test in this file, it does NOT pass `--opt vid=no` -- it needs the +/// real video track selected so time-pos actually advances and a +/// health-check tick can observe "healthy" before the forced trigger fires. +/// `--opt vo=null` keeps it headless (no DRM, no display touched) but does +/// NOT bound the decode: only killing at run_with_deadline's deadline does, +/// same as it would for any endless-stream real-decode run. This is exactly +/// the deviation this file's module doc says the mandatory vid=no/aid=no +/// rule exists to keep out of the automated suite -- hence `#[ignore]`, not +/// a relaxation of that rule for anything else here. `#[ignore]` also keeps +/// this out of a Pi mid-soak's plain `cargo test`, which must not add +/// unrelated CPU load to a thermal measurement. +/// +/// **CI now runs this test deliberately**, by exact name, as its own step +/// in `.github/workflows/dexd.yml` (after `Test`, before +/// `Build package`) -- the debian:trixie container has no display and no +/// DRM, so it exercises the same headless `vo=null` + software-HEVC-decode +/// path this doc comment describes, with no code path skipped. Manual +/// invocation (e.g. on the Pi, against `dexpi4.local`) still works exactly +/// as before: +/// +/// cargo test --test cli force_recovery_survives_against_real_mpv -- --ignored --nocapture +/// +/// DO NOT manually run this on dexpi4.local while its thermal soak is +/// active (its tmux sessions "dexeye"/"thermal" own the display and the CPU +/// is being measured -- see this task's hard constraint). Once the soak has +/// concluded, or on any OTHER Pi 4 (or dev machine) with libmpv 0.40+ +/// installed and nothing else on the display, the command above is safe. +/// +/// Expected stderr, in order: +/// 1. "dexd 0.1.0 (...)" -- normal startup +/// 2. "warning: BENCH ONLY (T7): --force-recovery-after-secs=3 is ARMED" +/// 3. (a few seconds of nothing -- real decode, no per-frame logging) +/// 4. "dexd: health check: T7 bench probe: ... -- attempting in-place +/// recovery 1/3 ..." +/// 5. "dexd: health check: in-place recovery's loadfile replaced the +/// stream; absorbing the expected END_FILE(reason=stop) ..." +/// 6. process is STILL RUNNING when this test kills it at its deadline, +/// and never printed a SECOND "attempting in-place recovery 2/" (see +/// the assertion below for why that matters at this deadline). +/// +/// If C1 has regressed: step 5 never appears, and the process exits 1 right +/// after step 4 instead (the recovery's own END_FILE(reason=stop) treated +/// as fatal) -- well before the deadline. +/// +/// **What this test does (and does not) prove.** It proves the process +/// survives its own recovery and that time-pos resumes advancing afterwards +/// (the recovery-2/-absence check below), under software decode and +/// `vo=null`, with no display. It does NOT prove the picture actually comes +/// back on real hardware: that claim lives entirely in the option set this +/// test never exercises -- `hwdec=drm`, `gpu-hwdec-interop=drmprime-overlay`, +/// the swapped DRM plane assignment -- all skipped here via `--no-defaults`. +/// A recovery whose `loadfile replace` tears down and rebuilds the +/// DRM/hwdec chain incorrectly could leave a black screen with a perfectly +/// alive, CI-green process. Only the on-Pi bench run with a real display +/// (this test's invocation above, minus `--no-defaults`/`vo=null`, against +/// the actual exhibition display) closes that gap -- CI proves survival, +/// only the bench proves the picture. +/// +/// **Residual gap even within what CI can see, stated rather than papered +/// over:** the recovery-2/-absence assertion below proves "no SECOND +/// recovery fired organically", which requires the event loop to still be +/// alive and ticking (mpv_wait_event waking on its timeout, HealthMonitor +/// still being ticked) even if it never got a second STALL to react to. A +/// process whose event loop wedged COMPLETELY right after the absorb -- +/// mpv_wait_event itself never returning again, no further ticks at all -- +/// would produce neither a second "attempting in-place recovery" line NOR +/// an exit, and every assertion here (still alive, attempt 1/, absorption, +/// no attempt 2/) would pass vacuously. Closing that fully would need a +/// positive post-recovery signal (e.g. a bench-only per-tick "healthy" log +/// line while T7 is armed, asserted present at least once after the +/// absorb) -- not implemented here; this comment exists so that gap is +/// recorded rather than silently assumed covered. +#[test] +#[ignore] +fn force_recovery_survives_against_real_mpv() { + let p = temp_path("liverecovery.265"); + std::fs::write(&p, stub_annexb()).unwrap(); + let r = run_with_deadline( + &[ + p.to_str().unwrap(), + "--bench-no-sidecar", + "--fps", + "30", + "--force-recovery-after-secs", + "3", + "--no-defaults", + "--opt", + "vo=null", + "--opt", + "aid=no", + ], + Duration::from_secs(30), + ); + // Survival means "still running when THIS TEST killed it at the deadline" + // -- asserted via `deadline_killed`, not via `exit_code == None`, because + // death by signal (SIGSEGV/SIGABRT) also yields `exit_code: None`. A + // C1-class regression that crashed via signal AFTER printing the absorb + // line would pass an exit_code-shaped assertion; it cannot pass this one. + assert!( + r.deadline_killed, + "process must SURVIVE the forced recovery (still running when killed \ + at the deadline); it ended on its own with exit_code {:?} -- an exit \ + means the recovery's own END_FILE(reason=stop) was NOT absorbed, and \ + exit_code None here means death by signal -- either way C1 has \ + regressed. stderr: {}", + r.exit_code, r.stderr + ); + assert!( + r.stderr.contains("attempting in-place recovery 1/"), + "forced recovery never fired: {}", + r.stderr + ); + assert!( + r.stderr + .contains("absorbing the expected END_FILE(reason=stop)"), + "recovery's own END_FILE was not absorbed -- this is exactly C1: {}", + r.stderr + ); + // The three assertions above are satisfiable by a process that + // "survives" only because its event loop wedged solid right after + // absorbing the stop -- alive, but not actually playing. A genuine + // recovery lets time-pos resume advancing, which feeds the health + // monitor a fresh "healthy" sample and means NO second recovery gets + // triggered organically. At a 30s deadline (trigger at 3s, health-check + // ticks every ~10s) a wedged-but-alive process would reach a second + // organic attempt at roughly trigger+20s =~ 23s -- comfortably inside + // this deadline -- so this string's ABSENCE is a real "playback + // actually resumed" proxy, not decoration. (It would be vacuously true + // at the old 15s deadline, which is why the deadline was raised.) + assert!( + !r.stderr.contains("attempting in-place recovery 2/"), + "a second recovery fired organically after the forced one -- time-pos \ + likely never resumed advancing post-recovery (event loop wedged \ + rather than truly recovered): {}", + r.stderr + ); +} + +// ---- F6: the exhibit display config -------------------------------------- +// +// The pure decision surface (grammar, resolve_display, the cmdline +// comparator, reconcile_cmdline, the sysfs pre-flight parser) is exhaustively +// tested in src/exhibit.rs on the Mac -- these tests exist only to pin how +// main.rs WIRES that logic in: gate ORDER (display gates fire before the +// asset is even read), the new flags parse, and the CLI-level refusal +// messages. All of them stay inside this file's mandatory `vo=null --opt +// vid=no --opt aid=no` rule. + +/// An `--exhibit-config` naming a file that is not there, and no +/// `--bench-no-sidecar`: refused before even the asset path is looked at, and +/// refused NAMING THE PATH THE OPERATOR GAVE. +/// +/// The naming half is the point. Until the dual-format work this borrowed +/// `resolve_display`'s "no exhibit config — create /etc/dex/exhibit.json" +/// message, which is wrong advice for someone who just pointed the flag +/// somewhere else: they would create a file the run they are debugging does +/// not read. Same class as the EACCES fix on the read path. +#[test] +fn missing_exhibit_config_refused_exit_2_before_the_asset_is_read() { + let improbable = temp_path("f6-default-exhibit-config-must-not-exist.json"); + let _ = std::fs::remove_file(&improbable); // never written; just proving absence + let r = run_with_deadline( + &[ + "/nonexistent/f6-missing-config.265", + "--exhibit-config", + improbable.to_str().unwrap(), + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!( + r.stderr.contains("f6-default-exhibit-config-must-not-exist.json"), + "the refusal must name the path that was actually given: {}", + r.stderr + ); + // The load-bearing negative: the asset gate must NOT have run yet. + assert!( + !r.stderr.contains("f6-missing-config.265"), + "the exhibit gate should refuse BEFORE the asset path is even looked \ + at, but the asset-missing message appeared too: {}", + r.stderr + ); +} + +// The complementary case -- NO --exhibit-config and no installed default, so +// resolve_display's "no exhibit config, create the shipped one" message fires +// -- is deliberately NOT tested here. It would depend on the HOST lacking +// /etc/dex, which is true on the Mac and false on the Pi (where the .deb +// installs exactly that file), so it would assert one thing in development and +// silently something else on the device -- the vacuous-check class this file +// has been bitten by before. It is covered where it is host-independent: +// `load_returns_none_when_no_default_exists` and `resolve_display`'s own unit +// tests in src/exhibit.rs. + +/// A YAML exhibit config drives the real binary end to end — the same gate +/// order, from a `.yaml` file. Pins that the format dispatch is wired into +/// main.rs and not merely unit-tested in the library. +#[test] +fn yaml_exhibit_config_binds_the_display_like_json_does() { + let cfg = temp_path("f6-yaml-exhibit.yaml"); + std::fs::write( + &cfg, + "# a venue would really write this\ndisplay_mode: 7680x4320@60\nkms_force: none\n", + ) + .unwrap(); + let cl = write_no_video_cmdline("f6-yaml-cmdline"); + let r = run_with_deadline( + &[ + "/nonexistent/f6-yaml.265", + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + ], + Duration::from_secs(10), + ); + // Reaching the sysfs pre-flight (which refuses this implausible 8K mode) + // proves the YAML parsed, validated, and bound the display: a config that + // had failed earlier could not produce THIS message. + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!(r.stderr.contains("7680x4320"), "stderr: {}", r.stderr); +} + +/// **The shipped deployment invocation shape**: no positional asset path at +/// all, everything from the exhibit config. `ExecStart=/usr/bin/dexd` +/// passes exactly this, so if a no-argument run were rejected on argv shape, +/// the packaged unit would exit 2 in a permanent restart loop on the device +/// while every other test here — all of which pass arguments — stayed green. +/// +/// Uses `--exhibit-config` (never the real default) so the test does not +/// depend on the host having, or lacking, `/etc/dex`; and an implausible mode, +/// so it lands on the sysfs pre-flight rather than starting playback. +#[test] +fn no_positional_asset_is_accepted_and_reaches_the_config() { + let cfg = temp_path("f6-noposition-exhibit.yaml"); + std::fs::write( + &cfg, + "asset: /nonexistent/f6-fromconfig.265\ndisplay_mode: 7680x4320@60\n", + ) + .unwrap(); + let cl = write_no_video_cmdline("f6-noposition-cmdline"); + let r = run_with_deadline( + &[ + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + // Reached a real gate, NOT the usage text. + assert!( + !r.stderr.contains("usage:"), + "a no-argument run must not be refused on argv shape -- that is how the \ + packaged unit invokes the player: {}", + r.stderr + ); + assert!(r.stderr.contains("7680x4320"), "stderr: {}", r.stderr); +} + +/// The asset actually comes FROM the config: with a valid display and a config +/// naming a nonexistent asset, the run gets as far as failing to read that +/// exact path — which only happens if `resolve_asset` took it from the file. +#[test] +fn the_exhibit_config_names_which_asset_plays() { + let cfg = temp_path("f6-assetfromconfig-exhibit.yaml"); + std::fs::write( + &cfg, + "asset: /nonexistent/f6-named-by-config.265\ndisplay_mode: auto\n", + ) + .unwrap(); + let cl = write_no_video_cmdline("f6-assetfromconfig-cmdline"); + let r = run_with_deadline( + &[ + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!( + r.stderr.contains("f6-named-by-config.265"), + "the asset named by the config must be the one the player tried to read: {}", + r.stderr + ); +} + +/// A positional path that CONTRADICTS the config is refused naming both — the +/// asset analogue of `mode_contradicting_exhibit_config_refused_naming_both`. +#[test] +fn positional_asset_contradicting_the_config_refused_naming_both() { + let cfg = temp_path("f6-assetconflict-exhibit.yaml"); + std::fs::write( + &cfg, + "asset: /nonexistent/f6-config-asset.265\ndisplay_mode: auto\n", + ) + .unwrap(); + let cl = write_no_video_cmdline("f6-assetconflict-cmdline"); + let r = run_with_deadline( + &[ + "/nonexistent/f6-cli-asset.265", + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!( + r.stderr.contains("f6-config-asset.265") && r.stderr.contains("f6-cli-asset.265"), + "stderr must name both assets: {}", + r.stderr + ); +} + +/// A config with no `asset`, and no path given: refuses rather than falling +/// back to /opt/dex/loop.265. The fail-closed row of `resolve_asset`'s table, +/// driven through the real binary. +#[test] +fn no_asset_named_anywhere_refused_without_guessing_loop_265() { + let cfg = temp_path("f6-noasset-exhibit.yaml"); + std::fs::write(&cfg, "display_mode: auto\n").unwrap(); + let cl = write_no_video_cmdline("f6-noasset-cmdline"); + let r = run_with_deadline( + &[ + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!(r.stderr.contains("no asset"), "stderr: {}", r.stderr); + // The load-bearing negative: it must not have quietly tried the old + // hardcoded path. + assert!( + !r.stderr.contains("cannot read /opt/dex/loop.265"), + "a missing asset must never fall back to the pre-F6 hardcoded path: {}", + r.stderr + ); +} + +/// The dispatch, at the CLI level: YAML syntax inside a `.json` file is +/// refused rather than quietly accepted by a permissive parser. +#[test] +fn yaml_contents_in_a_json_named_config_refused() { + let cfg = temp_path("f6-yaml-in-json.json"); + std::fs::write(&cfg, "display_mode: 3840x2160@30\n").unwrap(); + let r = run_with_deadline( + &[ + "/nonexistent/f6-yamlinjson.265", + "--exhibit-config", + cfg.to_str().unwrap(), + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!(r.stderr.contains("JSON"), "stderr: {}", r.stderr); +} + +/// An unparseable exhibit config is refused with the specific parse error — +/// distinct from "missing", per main.rs's read/parse split (see the comment +/// at the exhibit_config read site). +#[test] +fn malformed_exhibit_config_refused_exit_2_naming_the_parse_error() { + let cfg = temp_path("f6-malformed-exhibit.json"); + std::fs::write(&cfg, r#"{"display_mode":"auto","kms_forse":"none"}"#).unwrap(); // typo + let r = run_with_deadline( + &[ + "/nonexistent/f6-malformed.265", + "--exhibit-config", + cfg.to_str().unwrap(), + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!(r.stderr.contains("kms_forse"), "stderr: {}", r.stderr); +} + +/// `--bench-no-sidecar` bypasses BOTH the exhibit config AND the cmdline +/// gate: no `--exhibit-config` is supplied, no `--proc-cmdline` is supplied, +/// and the run still reaches playback (exit 1, not 2) because bench mode +/// consults neither. +#[test] +fn bench_flag_bypasses_the_exhibit_config_and_cmdline_gate_too() { + let p = temp_path("f6-bench.265"); + std::fs::write(&p, stub_annexb()).unwrap(); + let r = run_with_deadline( + &[ + p.to_str().unwrap(), + "--bench-no-sidecar", + "--fps", + "30", + "--no-defaults", + "--opt", + "vo=null", + "--opt", + "vid=no", + "--opt", + "aid=no", + ], + Duration::from_secs(30), + ); + assert_eq!(r.exit_code, Some(RUNTIME_EXIT), "stderr: {}", r.stderr); + assert!(r.stderr.contains("playback ended"), "stderr: {}", r.stderr); +} + +/// A `--mode` that contradicts the exhibit config's `display_mode` is +/// refused, naming both — the F6 analogue of +/// `fps_contradicting_sidecar_refused_exit_2_naming_both`. +#[test] +fn mode_contradicting_exhibit_config_refused_naming_both() { + let cfg = temp_path("f6-modeconflict-exhibit.json"); + std::fs::write(&cfg, r#"{"display_mode":"3840x2160@30","kms_force":"none"}"#).unwrap(); + let cl = write_no_video_cmdline("f6-modeconflict-cmdline"); + let r = run_with_deadline( + &[ + "/nonexistent/f6-modeconflict.265", + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + "--mode", + "2560x1440@60", + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!( + r.stderr.contains("3840x2160@30") && r.stderr.contains("2560x1440@60"), + "stderr must name both modes: {}", + r.stderr + ); +} + +/// The cmdline gate: exhibit config says `kms_force=none`, but the (fixture) +/// kernel cmdline carries a `video=HDMI-A-1:...` token anyway — the exact +/// 2026-08-15 incident class (a force removed as "stale" while still in use, +/// or here, the mirror case: a config edited to "none" while the boot config +/// was never reconciled). Must refuse naming the fix, before the asset is +/// even read. +#[test] +fn cmdline_mismatch_refused_naming_dex_exhibit_apply() { + let cfg = write_exhibit_config("f6-cmdlinemismatch-exhibit.json"); // kms_force: none + let cl = temp_path("f6-cmdlinemismatch-cmdline"); + std::fs::write(&cl, "console=ttyS0 video=HDMI-A-1:3840x2160@30 rootwait\n").unwrap(); + let r = run_with_deadline( + &[ + "/nonexistent/f6-cmdlinemismatch.265", + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!( + r.stderr.contains("dex-exhibit-apply") && r.stderr.contains("3840x2160@30"), + "stderr: {}", + r.stderr + ); +} + +/// The sysfs mode pre-flight: a non-"auto" display_mode that no real +/// connector could plausibly offer (8K60 — no HDMI-A-1 sink on a CI runner OR +/// this project's actual bench displays advertises this) is refused, either +/// because the connector cannot be found at all (a CI container with no DRM) +/// or because it is found but does not list the mode — the two-layer +/// "cannot find" vs "not among the modes" split from the F6 design's §2.4. +/// Either message names the requested resolution, which is what this test +/// pins portably across both hosts. +#[test] +fn implausible_mode_refused_by_the_sysfs_preflight() { + let cfg = temp_path("f6-implausible-exhibit.json"); + std::fs::write( + &cfg, + r#"{"display_mode":"7680x4320@60","kms_force":"none"}"#, + ) + .unwrap(); + let cl = write_no_video_cmdline("f6-implausible-cmdline"); + let r = run_with_deadline( + &[ + "/nonexistent/f6-implausible.265", + "--exhibit-config", + cfg.to_str().unwrap(), + "--proc-cmdline", + cl.to_str().unwrap(), + ], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!( + r.stderr.contains("7680x4320"), + "stderr must name the implausible resolution: {}", + r.stderr + ); + // And it must not be the asset-missing message -- the display gates run + // first, exactly like the "no exhibit config" case above. + assert!(!r.stderr.contains("f6-implausible.265"), "stderr: {}", r.stderr); +} + +// ---- F10: --bench-wedge-after-secs, the systemd-watchdog live-fire probe - +// +// F9 proved a wedged mpv core produces SILENCE on this program's event +// thread, and F1 acts on that silence in-process. Neither covers the event +// thread hanging in code that is NOT an mpv call at all (e.g. `eprintln!` +// against a wedged journald) -- see dexd::watchdog's module doc +// ("Framing"). `--bench-wedge-after-secs` deliberately reproduces that one +// remaining hazard class on a timer, so its plumbing gets the same +// "impossible to enable accidentally in a deployment" gate as T7's +// `--force-recovery-after-secs`, tested the same way here. + +/// The gate itself, mirroring +/// `force_recovery_without_bench_no_sidecar_refused_exit_2`: +/// `--bench-wedge-after-secs` without `--bench-no-sidecar` is refused on CLI +/// shape alone, before the asset is even read. +#[test] +fn bench_wedge_without_bench_no_sidecar_refused_exit_2() { + let r = run_with_deadline( + &["/nonexistent/x.265", "--bench-wedge-after-secs", "5"], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!( + r.stderr.contains("--bench-no-sidecar"), + "stderr must explain the required pairing: {}", + r.stderr + ); +} + +/// Same missing-value discipline as every other flag taking a value. +#[test] +fn bench_wedge_flag_missing_value_refused_exit_2_with_usage() { + let r = run_with_deadline( + &["/nonexistent/x.265", "--bench-wedge-after-secs"], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!(r.stderr.contains("usage"), "stderr: {}", r.stderr); +} + +/// A non-numeric value must also refuse via usage(), not silently parse as +/// 0 or panic the process. +#[test] +fn bench_wedge_flag_non_numeric_value_refused_exit_2_with_usage() { + let r = run_with_deadline( + &["/nonexistent/x.265", "--bench-wedge-after-secs", "soon"], + Duration::from_secs(10), + ); + assert_eq!(r.exit_code, Some(GATE_EXIT), "stderr: {}", r.stderr); + assert!(r.stderr.contains("usage"), "stderr: {}", r.stderr); +} + +/// The mechanism itself: with `--bench-wedge-after-secs 0`, the wedge check +/// fires on the very first loop iteration, unconditionally, BEFORE that same +/// iteration's event-id dispatch can act on whatever `mpv_wait_event` +/// happened to return (see the firing site's comment in main.rs for why +/// that ordering matters -- with `--opt vid=no --opt aid=no`, mpv reaches +/// "nothing to play" and would otherwise race this probe to an ordinary +/// `exit(1)`). A process that has genuinely wedged never exits on its own, +/// so the ONLY way this test ends is the harness's own deadline kill -- +/// `deadline_killed` must be true, mirroring the `Run` struct's own doc +/// comment on why that is the correct assertion (an exit_code of `None` +/// alone cannot distinguish "wedged, harness killed it" from "died by +/// signal on its own"). +/// +/// This proves the MECHANISM -- that the flag genuinely, permanently parks +/// the event thread -- not that a systemd watchdog then kills it: this +/// harness has no systemd to observe. That half is proved on the Pi; see +/// PLAN.md's F10 entry and README.md for the on-device procedure +/// (journalctl showing `Watchdog timeout`, a SIGABRT, and a supervisor +/// restart). +#[test] +fn bench_wedge_flag_actually_hangs_the_event_thread_forever() { + let p = temp_path("wedge.265"); + std::fs::write(&p, stub_annexb()).unwrap(); + let r = run_with_deadline( + &[ + p.to_str().unwrap(), + "--bench-no-sidecar", + "--fps", + "30", + "--bench-wedge-after-secs", + "0", + "--no-defaults", + "--opt", + "vo=null", + "--opt", + "vid=no", + "--opt", + "aid=no", + ], + Duration::from_secs(6), + ); + assert!( + r.deadline_killed, + "a genuinely wedged event thread must never exit on its own -- exit_code {:?}, \ + stderr: {}", + r.exit_code, r.stderr + ); + assert!( + r.stderr.contains("BENCH ONLY (F10 wedge probe)") && r.stderr.contains("ARMED"), + "must print the loud arming warning: {}", + r.stderr + ); + assert!( + r.stderr + .contains("deliberately parking the event thread forever"), + "must print the firing line proving the probe actually triggered: {}", + r.stderr + ); +} diff --git a/packages/dexd/tests/ffi_constants.rs b/packages/dexd/tests/ffi_constants.rs new file mode 100644 index 0000000..8e6409c --- /dev/null +++ b/packages/dexd/tests/ffi_constants.rs @@ -0,0 +1,160 @@ +//! T2 — verify the hand-transcribed FFI constants against the LIVE libmpv. +//! +//! This target links libmpv, so it builds and runs ONLY where libmpv-dev is +//! installed (the Pi: `cargo test`). On the Mac use `cargo test --lib`. It +//! creates no mpv instance and never touches the display: mpv_event_name() +//! and mpv_error_string() are static table lookups. +//! +//! Why this exists: MPV_EVENT_LOG_MESSAGE was once transcribed as 6 (it is 2; +//! 6 is START_FILE). The handler cast a start-file payload to a log-message +//! struct and segfaulted on the first frame. mpv exposes the id->name mapping +//! at runtime, so assert the transcription instead of trusting it — this also +//! catches a future mpv renumbering, which no header copy can. + +use std::ffi::{c_char, c_int, CStr}; + +use dexd::ffi_consts::{ + MPV_ERROR_UNSUPPORTED, MPV_EVENT_COMMAND_REPLY, MPV_EVENT_END_FILE, MPV_EVENT_LOG_MESSAGE, + MPV_EVENT_NONE, MPV_EVENT_PROPERTY_CHANGE, MPV_EVENT_QUEUE_OVERFLOW, MPV_EVENT_SHUTDOWN, + MPV_EVENT_START_FILE, +}; + +#[link(name = "mpv")] +extern "C" { + fn mpv_event_name(event: c_int) -> *const c_char; + fn mpv_error_string(error: c_int) -> *const c_char; +} + +/// mpv_event_name returns NULL for ids it does not know. +fn event_name(id: c_int) -> Option { + let p = unsafe { mpv_event_name(id) }; + if p.is_null() { + return None; + } + Some(unsafe { CStr::from_ptr(p) }.to_string_lossy().into_owned()) +} + +fn error_string(code: c_int) -> String { + // mpv_error_string is documented to always return a valid string. + unsafe { CStr::from_ptr(mpv_error_string(code)) } + .to_string_lossy() + .into_owned() +} + +#[test] +fn event_ids_match_the_live_library() { + // Verified against mpv v0.40.0 on the bench Pi (2026-08-15, via ctypes): + // 0 none, 1 shutdown, 2 log-message, 6 start-file, 7 end-file + assert_eq!(event_name(MPV_EVENT_NONE).as_deref(), Some("none")); + assert_eq!(event_name(MPV_EVENT_SHUTDOWN).as_deref(), Some("shutdown")); + assert_eq!( + event_name(MPV_EVENT_LOG_MESSAGE).as_deref(), + Some("log-message") + ); + assert_eq!( + event_name(MPV_EVENT_START_FILE).as_deref(), + Some("start-file") + ); + assert_eq!(event_name(MPV_EVENT_END_FILE).as_deref(), Some("end-file")); +} + +/// F1: the two event ids the tier-0 health check relies on -- one to learn +/// the sampled position (MPV_EVENT_PROPERTY_CHANGE, from +/// `mpv_observe_property("time-pos", ...)`) and one to learn whether its +/// async recovery command was accepted (MPV_EVENT_COMMAND_REPLY, from +/// `mpv_command_async`). Verified against the live library exactly like the +/// original four -- the whole point of T2 is that a transcription can rot +/// silently and the live library cannot, and these two are just as +/// dispatch-critical as the ones that already segfaulted once. +#[test] +fn f1_event_ids_match_the_live_library() { + assert_eq!( + event_name(MPV_EVENT_COMMAND_REPLY).as_deref(), + Some("command-reply") + ); + assert_eq!( + event_name(MPV_EVENT_PROPERTY_CHANGE).as_deref(), + Some("property-change") + ); + // And distinct from every other id this program dispatches on, so a + // future transcription collision (LOG_MESSAGE-vs-START_FILE's exact + // failure shape) cannot silently misroute one event's payload as + // another's. + for other in [ + MPV_EVENT_NONE, + MPV_EVENT_SHUTDOWN, + MPV_EVENT_LOG_MESSAGE, + MPV_EVENT_START_FILE, + MPV_EVENT_END_FILE, + MPV_EVENT_QUEUE_OVERFLOW, + ] { + assert_ne!(MPV_EVENT_COMMAND_REPLY, other); + assert_ne!(MPV_EVENT_PROPERTY_CHANGE, other); + } + assert_ne!(MPV_EVENT_COMMAND_REPLY, MPV_EVENT_PROPERTY_CHANGE); +} + +#[test] +fn queue_overflow_id_matches_the_live_library() { + // mpv's internal event ring silently drops events (including a + // potential END_FILE) once it chokes at 1000 pending; this id is the + // only signal that happened. Pin it against the live library exactly + // like the others above -- a mistranscription here would defeat the + // fatal handling in main.rs just as silently as the LOG_MESSAGE bug did. + assert_eq!( + event_name(MPV_EVENT_QUEUE_OVERFLOW).as_deref(), + Some("event-queue-overflow") + ); +} + +#[test] +fn log_message_and_start_file_ids_are_distinct() { + // These two ids must never collide, because the event loop uses the id to + // decide which struct to cast `mpv_event.data` to. Getting them confused + // dereferences a start-file payload as a log-message struct -- i.e. reads + // two unrelated integers as string pointers. + // + // Not hypothetical: LOG_MESSAGE was once transcribed as 6, which is + // START_FILE, and the player segfaulted on the first frame. Pinned in both + // directions, and against the live library rather than the transcription. + assert_ne!(MPV_EVENT_LOG_MESSAGE, MPV_EVENT_START_FILE); + assert_ne!( + event_name(MPV_EVENT_LOG_MESSAGE), + event_name(MPV_EVENT_START_FILE) + ); + assert_ne!( + event_name(6).as_deref(), + Some("log-message"), + "6 is start-file, not log-message" + ); +} + +#[test] +fn error_unsupported_matches_the_live_library() { + // The live mpv 0.40 string for -18 is "not supported" — NOT "unsupported"; + // PLAN.md's original assertion would have failed here. Exact match, not + // a `.contains("supported")` hedge: err_table also has "unsupported + // format for accessing option" (-6) and "...property" (-9), both of + // which pass a loose "supported" substring check just as well as -18 + // does -- so a mistranscription of MPV_ERROR_UNSUPPORTED as -6 or -9 + // would sail through the exact test whose entire job is to catch that. + // The whole philosophy here is "assert against the live library and let + // drift fail loudly"; a wording-family hedge defeats it. + let s = error_string(MPV_ERROR_UNSUPPORTED); + assert_eq!(s, "not supported", "mpv_error_string(-18) = {s:?}"); + assert_ne!( + error_string(MPV_ERROR_UNSUPPORTED), + error_string(-1), + "-1 is EVENT_QUEUE_FULL, the sign-compatible-but-wrong value this code once used" + ); + assert_ne!( + error_string(MPV_ERROR_UNSUPPORTED), + error_string(-6), + "-6 is OPTION_FORMAT, not UNSUPPORTED" + ); + assert_ne!( + error_string(MPV_ERROR_UNSUPPORTED), + error_string(-9), + "-9 is PROPERTY_FORMAT, not UNSUPPORTED" + ); +} From bff874b3750380eaf917c5489189367f4ed36b28 Mon Sep 17 00:00:00 2001 From: Max Albrecht <1@178.is> Date: Mon, 17 Aug 2026 20:51:12 +0200 Subject: [PATCH 03/22] E: CI -- dexd deb workflow from the experiment, paths and names updated Copied from home-workspace's ci/dex-loop-deb.yml (same source commit as the crate, 71711ec1). Renamed dex-loop -> dexd throughout: workflow name, job/artifact names, concurrency group, cache key, systemd unit, apt package name, binary paths, tag pattern (dex-loop-v* -> dexd-v*). CRATE_DIR points at packages/dexd instead of the experiment's nested path; the changes-detection path list matches packages/dexd/** and this workflow file. Added "feature/**" to the branch push triggers so this branch's own pushes run it. Dropped the "Shellcheck the bench harness" step: it shellchecked scripts/*.sh under the experiment directory, which is harness code that intentionally stays in home-workspace and has no equivalent path in this repo. No prose was rewritten beyond substituting the renamed identifiers and paths; the design-rationale comments are unchanged. --- .github/workflows/dexd.yml | 608 +++++++++++++++++++++++++++++++++++++ 1 file changed, 608 insertions(+) create mode 100644 .github/workflows/dexd.yml diff --git a/.github/workflows/dexd.yml b/.github/workflows/dexd.yml new file mode 100644 index 0000000..50e64f2 --- /dev/null +++ b/.github/workflows/dexd.yml @@ -0,0 +1,608 @@ +# Build the dexd Debian package for Raspberry Pi gallery devices. +# +# WHY A CONTAINER ON AN ARM RUNNER +# -------------------------------- +# The package's whole purpose is that `Depends:` is derived from the sonames +# the binary actually links (dpkg-shlibdeps, see Cargo.toml's `depends = +# "$auto"`), so that apt refuses a libmpv ABI mismatch at install time on a +# bench rather than letting it surface as a black screen in a gallery. +# +# That guarantee is only worth anything if the build environment IS the target +# environment. `ubuntu-24.04-arm` gives us native arm64 (free on public repos, +# no qemu), but it is Ubuntu -- linking against Ubuntu's libmpv and then +# installing on Debian trixie would produce exactly the mismatch this package +# exists to prevent. So: arm64 runner for the architecture, debian:trixie +# container for the ABI. +# +# Verify the pairing whenever the fleet moves to a new Debian release: +# ssh '. /etc/os-release; echo $VERSION_CODENAME; dpkg -s libmpv2' +# +# TOOLCHAIN: DEBIAN'S RUSTC, NOT RUSTUP +# ------------------------------------ +# rustc and cargo are installed with apt, from the same archive the devices +# use, so CI compiles with the compiler the fleet actually has (trixie: 1.85). +# This was originally rustup stable, which was inconsistent with the paragraph +# above: a container that supplies the target's libraries but not its compiler +# is only half a target environment. +# +# What it buys: the Pi is a development host, and it can only keep building its +# own software while the crate stays inside trixie's Rust. With rustup here, +# "builds in CI" and "builds on the device" are separate claims, and the day +# they diverge is the day someone discovers it by hand -- which is how the +# cargo-deb MSRV gap below was found. +# +# The cost is real and deliberate: Debian's rustc lags, so a dependency needing +# a newer compiler cannot be adopted. That constraint should be felt HERE, on a +# red build, rather than at deploy time. `rust-version` in Cargo.toml states +# the same floor; this job is what enforces it. +name: dexd deb + +# NOTE ON TRIGGERS. This originally fired only on `dexd-v*` tags, plus +# workflow_dispatch — which meant it could not run AT ALL: +# * workflow_dispatch only registers for workflows present on the DEFAULT +# branch, and this file lives on the experiment branch, so it never appeared +# in the Actions UI or `gh workflow list`; +# * and a branch push matched nothing, because only `tags:` was listed. +# Four jobs that never execute report the same green nothing as four jobs that +# pass. Branch pushes are also what CI is FOR: linting and packaging checks that +# only run at release time cannot prevent the release. +# NO PATH FILTERS HERE, DELIBERATELY. A path-filtered workflow that does not run +# reports NOTHING, and a required status check that never reports blocks a merge +# forever -- the classic monorepo trap. So the workflow always runs, the `changes` +# job decides what is relevant, and the `gate` job at the end reports a single +# verdict that is safe to mark "required" in branch protection. +on: + workflow_dispatch: + push: + branches: ["main", "experiment/**", "feature/**"] + tags: ["dexd-v*"] + pull_request: + +# Cancel a superseded run rather than paying for it. Grouped per ref, so pushes +# to a branch supersede each other while a TAG build never cancels a branch build +# (and vice versa) -- a tag build produces the release artifact, so losing one to +# an unrelated push would be the wrong economy. +concurrency: + group: dexd-deb-${{ github.ref }} + cancel-in-progress: true + +permissions: + contents: read + +env: + CRATE_DIR: packages/dexd + # build.rs embeds this as the binary's build identity. An ENV VAR, not the + # .dex-build-id file, because `rerun-if-changed` on a path that did not exist + # when the cached build ran is "never changed" — so with a warm target/ cache + # the stamp file is written and then IGNORED, producing a package that still + # reports (nogit). Observed exactly that on 2026-08-16. Cargo compares the + # VALUE of a `rerun-if-env-changed` variable, which has no such hole. + DEX_BUILD_ID: ${{ github.sha }} + +jobs: + # What actually changed? Computed here rather than by path filters, so the + # answer is available to every job AND to the gate. No third-party action: + # `git diff` against the right base is a few lines and one less dependency in + # a workflow that installs a Debian package onto gallery hardware. + changes: + runs-on: ubuntu-latest + outputs: + crate: ${{ steps.detect.outputs.crate }} + reason: ${{ steps.detect.outputs.reason }} + steps: + - uses: actions/checkout@v4 + with: + fetch-depth: 0 # need history to diff against a base + - id: detect + run: | + set -eu + # A tag build or a manual dispatch always runs everything: a release + # must never be decided by a diff, and a human pressing the button + # means "run it". + case "${{ github.event_name }}" in + workflow_dispatch) echo "crate=true" >> "$GITHUB_OUTPUT" + echo "reason=manual dispatch" >> "$GITHUB_OUTPUT"; exit 0 ;; + esac + case "${{ github.ref }}" in + refs/tags/*) echo "crate=true" >> "$GITHUB_OUTPUT" + echo "reason=tag build" >> "$GITHUB_OUTPUT"; exit 0 ;; + esac + + if [ "${{ github.event_name }}" = "pull_request" ]; then + base="${{ github.event.pull_request.base.sha }}" + else + base="${{ github.event.before }}" + fi + # A new branch, a force-push, or a first commit gives an unusable base + # (all-zeros or a missing object). Fail OPEN -- run everything -- because + # skipping on an unknown diff is how a real change slips through. + if [ -z "$base" ] || [ "$base" = "0000000000000000000000000000000000000000" ] \ + || ! git cat-file -e "$base^{commit}" 2>/dev/null; then + echo "crate=true" >> "$GITHUB_OUTPUT" + echo "reason=no usable diff base, running everything" >> "$GITHUB_OUTPUT"; exit 0 + fi + + changed=$(git diff --name-only "$base" HEAD) + echo "changed files:"; echo "$changed" | sed 's/^/ /' + if echo "$changed" | grep -qE '^(packages/dexd/|\.github/workflows/dexd\.yml$)'; then + echo "crate=true" >> "$GITHUB_OUTPUT" + echo "reason=crate or this workflow changed" >> "$GITHUB_OUTPUT" + else + echo "crate=false" >> "$GITHUB_OUTPUT" + echo "reason=nothing relevant changed" >> "$GITHUB_OUTPUT" + fi + + build: + needs: changes + if: needs.changes.outputs.crate == 'true' + runs-on: ubuntu-24.04-arm + container: debian:trixie + + steps: + - name: Install build dependencies + run: | + set -eux + apt-get update + # libmpv-dev is what the crate links; dpkg-dev provides + # dpkg-shlibdeps, which is what turns those links into Depends. + # rustc/cargo come from DEBIAN, not rustup -- see "Toolchain" below. + # libmpv-dev is what the crate links; dpkg-dev provides + # dpkg-shlibdeps, which is what turns those links into Depends. + apt-get install -y --no-install-recommends \ + build-essential pkg-config ca-certificates git \ + rustc cargo rust-clippy libmpv-dev dpkg-dev lintian + # Record what we built and linked against. When a future .deb refuses + # to install on a device, these lines in the log are the first thing + # to compare against the device's own `dpkg -s libmpv2` / rustc -V. + dpkg -s libmpv2 | grep -E '^(Package|Version):' + rustc --version + cargo --version + + - uses: actions/checkout@v4 + + - name: Cache cargo + uses: actions/cache@v4 + with: + path: | + ~/.cargo/bin + ~/.cargo/registry + ~/.cargo/git + ${{ env.CRATE_DIR }}/target + key: dexd-deb-${{ hashFiles('packages/dexd/Cargo.lock') }} + restore-keys: dexd-deb- + + # Pinned to the 2.x line: cargo-deb 3.7 uses let-chains and needs rustc + # >= 1.88, which trixie does not have. The pin is a CONSEQUENCE of the + # toolchain decision, not an independent choice -- if this ever has to be + # unpinned, the toolchain question is what actually changed. + # + # Guard on the FILE, not on PATH. `command -v cargo-deb` was wrong in a way + # that only appears once the cache is warm: cargo installs to + # ~/.cargo/bin, which is NOT on PATH here (cargo comes from apt, there is + # no rustup to add it), so on a cache hit the guard said "missing" while + # cargo said "binary already exists in destination" and exited 101. Both + # were right about different questions. `cargo deb` itself still works + # because cargo searches $CARGO_HOME/bin for subcommands regardless of PATH. + - name: Install cargo-deb + run: test -x "$HOME/.cargo/bin/cargo-deb" || cargo install --locked cargo-deb --version "^2" + + # The package is only worth shipping if the code it contains passes. The + # full suite runs here because this container HAS libmpv -- the Mac + # checkout cannot link the binary at all, so `cargo test --lib` is the + # most a developer machine can do. + - name: Clippy + working-directory: ${{ env.CRATE_DIR }} + run: cargo clippy --all-targets -- -D warnings + + - name: Test + working-directory: ${{ env.CRATE_DIR }} + run: cargo test --release + + # C1 live-fire regression gate: forces a tier-0 in-place recovery + # against a REAL mpv core (software HEVC decode, vo=null, no display) + # and asserts the process SURVIVES its own recovery's + # END_FILE(reason=stop) -- the exact event-shape that killed every + # recovery attempt before the C1 fix. #[ignore]d in tests/cli.rs so it + # never runs in a default `cargo test` (a Pi mid-soak must not pick it + # up); this step is the one place it runs automatically. + # + # THE GREP IS THE GATE: libtest exits 0 when a filter matches NOTHING + # ("0 passed; 0 failed"), so a renamed/deleted test would leave this + # step green while running nothing. Asserting "1 passed" makes that + # loud instead of silent. + # + # What this proves and what it does not: the process survives its own + # recovery and time-pos resumes advancing, under software decode with + # no display. It does NOT prove the picture comes back on real + # hardware -- that needs hwdec=drm / drmprime-overlay / the DRM plane + # swap, none of which this container can exercise (no DRM, no GPU). + # That claim is still the on-Pi bench checklist item's job; see the + # test's own doc comment in tests/cli.rs. + - name: C1 live-fire (forced recovery vs real mpv) + working-directory: ${{ env.CRATE_DIR }} + run: | + set -eux + out=$(cargo test --release --test cli -- --ignored --exact \ + force_recovery_survives_against_real_mpv 2>&1) || { echo "$out"; exit 1; } + echo "$out" + echo "$out" | grep -q '1 passed' \ + || { echo "::error::C1 gate did not actually run (filter matched nothing?)"; exit 1; } + + # Every CI build must produce a DISTINCT, ORDERED version, or apt refuses + # to install it over an identical one -- silently. Observed 2026-08-16: + # `apt install` reported "dexd is already the newest version (0.1.0-1)" + # and did nothing, so the device kept running an older binary while + # `dpkg -l` showed the expected version. Every signal agreed and all of + # them were wrong; only the startup line caught it. + # + # Scheme: +g. The run number LEADS because dpkg + # compares digit runs numerically, so 9 < 10 and ordering follows time. + # A commit-only revision does NOT work -- verified with + # `dpkg --compare-versions`: 0.1.0-1+gzz999999 sorts ABOVE + # 0.1.0-1+g000aaaaa, so an older build could outrank a newer one and apt + # would refuse the real upgrade as a downgrade. + - name: Build package + working-directory: ${{ env.CRATE_DIR }} + run: | + set -eux + short=$(printf '%s' "$GITHUB_SHA" | cut -c1-12) + cargo deb --deb-revision "${GITHUB_RUN_NUMBER}+g${short}" + + # Print the derived Depends into the log. If dpkg-shlibdeps ever silently + # stops resolving libmpv, the package would still build and would still + # install -- onto a device with no libmpv, failing at runtime. Asserting + # it here keeps that from being a silent success. + - name: Verify derived dependencies + working-directory: ${{ env.CRATE_DIR }} + run: | + set -eux + deb=$(find target/debian -name '*.deb' -print -quit) + test -n "$deb" + dpkg-deb --field "$deb" Package Version Architecture Depends + dpkg-deb --contents "$deb" + dpkg-deb --field "$deb" Depends | grep -q libmpv \ + || { echo "::error::Depends does not mention libmpv — dpkg-shlibdeps did not resolve it"; exit 1; } + # The BEHAVIOURAL floor, which dpkg-shlibdeps cannot derive: left to + # $auto alone the answer is `libmpv2 (>= 0.19.0)`, the oldest version + # exporting the symbols we call. What we actually require is 0.40 + # behaviour (drmprime-overlay; END_FILE(reason=stop) on a loadfile + # replace). Asserted here because deleting the explicit constraint in + # Cargo.toml would still produce a package that builds and installs. + dpkg-deb --field "$deb" Depends | grep -qE 'libmpv2 \(>= 0\.(4[0-9]|[5-9][0-9])' \ + || { echo "::error::Depends lost its explicit libmpv2 >= 0.40 floor"; exit 1; } + # The binary must be able to say WHICH build it is. `(nogit)` means + # build.rs found neither a stamp file nor a usable git, and every + # device carrying that package becomes unidentifiable in the field. + # Asserted rather than trusted: the stamp step above is one `printf` + # away from silently doing nothing. + # dpkg-deb -x rather than piping --fsys-tarfile into tar: the tarfile's + # members have no './' prefix, so `tar -xO ./usr/bin/dexd` finds + # nothing and exits 2. Extracting the tree sidesteps the path-prefix + # question entirely. + extract=$(mktemp -d) + dpkg-deb -x "$deb" "$extract" + test -x "$extract/usr/bin/dexd" + # 2>&1 is load-bearing: --version writes to STDERR. Without it $ver + # is empty, `case "" in *nogit*)` matches nothing, and the assertion + # reports success having observed nothing at all — which is precisely + # what it did when first added. + ver=$("$extract/usr/bin/dexd" --version 2>&1 | head -1) + echo "packaged binary reports: $ver" + case "$ver" in + *nogit*) echo "::error::packaged binary reports (nogit) — build identity not stamped"; exit 1 ;; + esac + # Positive assertion, because "does not contain nogit" is also true of + # the empty string. The build identity must be 12 hex characters and + # must match the commit CI is building. + # `cut`, not ${VAR:0:12}: container run: steps execute under /bin/sh + # (dash), where bash substring expansion is a "Bad substitution" error. + # The same bashism bit the lifecycle job earlier with diff <(...). + short=$(printf '%s' "$GITHUB_SHA" | cut -c1-12) + echo "$ver" | grep -qE "\($short(\+dirty)?\)" \ + || { echo "::error::version '$ver' does not carry this commit ($short)"; exit 1; } + + # The PACKAGE version must be distinct per build, not just the binary's + # internal string: apt decides whether to install by version alone. + pkgver=$(dpkg-deb --field "$deb" Version) + echo "package version: $pkgver" + case "$pkgver" in + *"$GITHUB_RUN_NUMBER+g$short"*) : ;; + *) echo "::error::package version '$pkgver' lacks the run number/commit — apt will refuse to reinstall it"; exit 1 ;; + esac + + # Policy check on the built package. --fail-on error,warning is the point: + # a linter that reports and exits 0 is a linter nobody reads. Accepted + # tags live in deploy/lintian-overrides WITH their reasoning, so silencing + # one is a reviewable diff rather than a flag on a command line. + - name: Lintian + working-directory: ${{ env.CRATE_DIR }} + run: | + set -eux + deb=$(find target/debian -name '*.deb' -print -quit) + lintian --tag-display-limit 0 --fail-on error,warning "$deb" + + - uses: actions/upload-artifact@v4 + with: + name: dexd-deb + path: ${{ env.CRATE_DIR }}/target/debian/*.deb + if-no-files-found: error + + # Static checks on what we SHIP but do not compile: the systemd unit and the + # shell. Separate job so it reports independently of the package build and + # runs even when the build breaks -- these findings are usually the cause. + lint: + needs: changes + if: needs.changes.outputs.crate == 'true' + runs-on: ubuntu-24.04-arm + container: debian:trixie + steps: + - name: Install linters + run: | + set -eux + apt-get update + # devscripts provides checkbashisms; systemd provides systemd-analyze. + apt-get install -y --no-install-recommends \ + shellcheck devscripts systemd git ca-certificates reuse + + - uses: actions/checkout@v4 + + # THE reason this job exists. systemd silently ignores directives it does + # not recognise in a given section, so a typo or a section-migrated key is + # invisible at runtime -- the unit starts fine and simply does less than + # it says. That is how StartLimitIntervalSec=0 sat in [Service] doing + # nothing (fixed 2026-08-15): the most safety-critical line in the unit, + # ignored, while reading as though it worked. + # + # The two "Command ... is not executable" notices are expected here: the + # binaries live in the package, which is not installed in this container. + - name: systemd-analyze verify + working-directory: ${{ env.CRATE_DIR }} + run: | + set -eux + out=$(systemd-analyze verify deploy/dexd.service 2>&1 || true) + echo "$out" + # Fail on anything that is NOT the known not-installed noise. + if echo "$out" | grep -vE 'is not executable|^$' | grep -q .; then + echo "::error::systemd-analyze reported problems with dexd.service" + echo "$out" | grep -vE 'is not executable|^$' + exit 1 + fi + + # Maintainer scripts run as root on every device, and run under /bin/sh -- + # dash on Debian, not bash. A bashism there fails an install in the field, + # on a machine nobody is sitting at. + - name: Shell checks + working-directory: ${{ env.CRATE_DIR }} + run: | + set -eux + shellcheck deploy/dex-wait-hdmi deploy/maintainer-scripts/* + checkbashisms deploy/maintainer-scripts/* + + # REUSE (reuse.software): every file must carry copyright and licence + # information, machine-readably. dex aims to be installable from all the + # usual repos, and distro packagers are precisely the people who need a + # precise per-file answer -- this makes it checkable rather than asserted. + # + # Scoped to the crate with --root. Repo-wide compliance is the eventual + # goal but is blocked on the GPL-inherited packages (pi-gen -> dex-os), + # whose licensing is not settled; annotating those unilaterally would be + # making a claim rather than recording one. + - name: REUSE lint + working-directory: ${{ env.CRATE_DIR }} + run: reuse --root . lint + + # Turns SPEC §5c's dependency policy into a gate: advisories, a trimmed + # licence allow-list, and a ban on the proc-macro toolchain (see deny.toml — + # the "serde without derive" decision is otherwise only a comment). + # + # Runs on the bare runner with rustup, NOT in the trixie container. That is + # consistent with the toolchain rule rather than an exception to it: the rule + # binds what BUILDS THE SHIPPED ARTIFACT, so it stays inside the container. + # cargo-deny produces nothing that ships and needs a newer rustc than trixie + # has, exactly like cargo-deb — the difference is that cargo-deb builds the + # package, so it had to be pinned instead. + deny: + needs: changes + if: needs.changes.outputs.crate == 'true' + runs-on: ubuntu-24.04-arm + steps: + - uses: actions/checkout@v4 + - name: Install cargo-deny + run: | + set -eux + curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs \ + | sh -s -- -y --default-toolchain stable --profile minimal + echo "$HOME/.cargo/bin" >> "$GITHUB_PATH" + - name: cargo deny check + working-directory: ${{ env.CRATE_DIR }} + run: | + export PATH="$HOME/.cargo/bin:$PATH" + cargo install --locked cargo-deny + cargo deny check + + # Install / remove / purge lifecycle. This is the class of bug that only + # appears when a package is really installed and removed — maintainer scripts + # that fail, files that outlive a purge, users never created — and no static + # linter reaches it. + # + # This started as piuparts and is not, because piuparts has no installation + # candidate on Ubuntu noble/arm64 (it is simply not published for that arch), + # and inside a container it could not have used --docker-image anyway. The + # container IS a throwaway environment, which is the only thing piuparts's + # chroot was providing here, so the test runs directly. + # + # What is genuinely lost by not using piuparts: upgrade testing from a + # previous version (there is none yet) and its much broader leftover + # heuristics. The filesystem diff below is a deliberately narrow substitute. + lifecycle: + if: needs.changes.outputs.crate == 'true' + runs-on: ubuntu-24.04-arm + container: debian:trixie + needs: [changes, build] + steps: + - uses: actions/download-artifact@v4 + with: + name: dexd-deb + + - name: Install / remove / purge + run: | + set -eux + apt-get update + deb=$(find . -name '*.deb' -print -quit); test -n "$deb" + + # Snapshot the filesystem so a purge can be checked for leftovers. + snap() { find / -xdev \( -path /proc -o -path /sys -o -path /run -o -path /tmp \ + -o -path /var/log -o -path /var/lib/apt -o -path /var/cache \ + -o -path /var/lib/dpkg -o -path /github \) -prune -o -print 2>/dev/null | sort; } + + # `apt install ./dexd.deb` pulls in every dependency and + # Recommends apt derives for it -- libmpv2's own libs, aria2 and + # yt-dlp (libmpv2 Recommends both), systemd and dbus (a systemd + # unit ships), and more. Not container noise: README.md documents + # production install as this exact plain `apt install + # ./dexd.deb`, so a stock Pi picks up the identical chain. + # Most of it autoremoves cleanly once dexd is purged, but + # systemd/dbus/policykit are Debian-protected and never + # autoremove -- AND their own postinst scripts create dynamic + # state (dbus's machine-id, systemd's catalog database, + # deb-systemd-helper's enablement markers...) that dpkg's static + # file list never mentions either. There is no way to tell "apt + # settling a protected dependency" apart from "dexd's postrm + # forgot to clean something" by inspecting paths after the fact. + # + # On a real device this whole paragraph is moot: systemd, dbus + # and policykit are already part of the base OS image, so + # installing dexd adds nothing new there. Reproduced here by + # running the FULL install/remove/purge/autoremove cycle once, + # unobserved, before the real run: that settles every protected + # package's one-time setup exactly as a real device's base image + # already has it, so the "before" snapshot below starts from the + # same place a real Pi does, and the leftovers diff at the end + # needs nothing cleverer than a plain diff to stay honest. + apt-get install -y "./${deb#./}" + apt-get remove -y dexd + apt-get purge -y dexd + apt-get autoremove --purge -y + + # /tmp, not /: snap() walks the whole filesystem, and a snapshot + # written to / would then be a directory entry snap's own find + # sees -- including itself as a "new" file relative to whichever + # snapshot didn't exist yet when it was taken. /tmp is already in + # snap()'s prune list for the same class of reason (excluding the + # runner's own transient state), so scratch output belongs there. + snap > /tmp/before.txt + + # --- the real, asserted run --- + apt-get install -y "./${deb#./}" + + # --- postinst did what it promises --- + test -x /usr/bin/dexd + test -x /usr/bin/dex-wait-hdmi + getent passwd dex # the service user exists + + # postinst only joins dex to a group that EXISTS in the target + # environment (see deploy/maintainer-scripts/postinst's own + # `getent group "$g"` guard) -- a real Pi has both video and render + # (verified on device: groups=105(dex),44(video),992(render)), but + # this minimal debian:trixie container has no udev, so render is + # never created here at all. Asserting unconditional membership in + # both groups was asserting something postinst never promised: THE + # TEST was wrong here, not the package. + # + # Mirror postinst's own conditional rather than hard-coding the + # group names twice, but don't let that degrade into a no-op: video + # is a static base-passwd group present in every Debian environment + # including this container, so its existence -- and dex's + # membership in it -- is asserted unconditionally below, not just + # conditionally skipped like render. That keeps at least one real + # assertion in force even in a future container where render is + # ALSO absent, so this loop can never silently pass on both groups. + getent group video >/dev/null \ + || { echo "::error::video group unexpectedly absent from this container -- the membership check below would be vacuous"; exit 1; } + for g in video render; do + if getent group "$g" >/dev/null; then + id -nG dex | grep -qw "$g" \ + || { echo "::error::postinst did not add dex to the existing '$g' group"; exit 1; } + else + echo "group '$g' does not exist in this container (expected: no udev here) -- skipping membership check" + fi + done + + test -d /opt/dex # asset directory created + test -f /lib/systemd/system/dexd.service + test -f /usr/share/man/man1/dexd.1.gz + # The package must NOT ship an asset: artwork is content, not software. + test ! -e /opt/dex/loop.265 + + apt-get remove -y dexd + test ! -e /usr/bin/dexd # binary gone + getent passwd dex # user RETAINED, by design + test -d /opt/dex # artwork RETAINED, by design + + apt-get purge -y dexd + test ! -e /lib/systemd/system/dexd.service + getent passwd dex # still retained after purge + test -d /opt/dex + + apt-get autoremove --purge -y + + # --- leftovers --- + # Exactly two are intended, both deliberate in postrm: /opt/dex holds + # artwork the package never shipped, and `dex` may be referenced by + # anything the operator wrote. A THIRD leftover is a bug, so this + # diffs rather than trusting the exit code. The "before" snapshot at + # the top of this script was taken after an IDENTICAL, unobserved + # install/remove/purge/autoremove cycle for exactly this reason: every + # dependency apt settles for dexd -- protected packages included, + # the dynamic state their own postinst scripts create included -- is + # already present on BOTH sides and cancels out, leaving only what + # THIS run's postinst/postrm are actually responsible for. + # + # Filtered to temp files rather than `diff <(...) <(...)`: run steps + # in a `container:` job execute under /bin/sh (dash on trixie), not + # bash, and process substitution is a bashism dash does not have -- + # confirmed by CI itself once the group-membership fix above let the + # job reach this far for the first time ("Syntax error: ( unexpected"). + snap > /tmp/after.txt + grep -v '^/opt/dex' /tmp/before.txt > /tmp/before.filtered.txt + grep -v '^/opt/dex' /tmp/after.txt > /tmp/after.filtered.txt + if ! diff /tmp/before.filtered.txt /tmp/after.filtered.txt > /tmp/leftovers.diff; then + echo "::error::purge left files behind beyond the two intended:" + cat /tmp/leftovers.diff + exit 1 + fi + echo "lifecycle OK — install, remove and purge all behave as designed" + + # THE REQUIRED CHECK. Mark exactly this one as required in branch protection. + # + # It always runs (if: always()), so it always reports -- which is the whole + # point: a check that can be skipped can never be required. It then decides + # PROGRAMMATICALLY rather than by GitHub's own skip semantics, which lets it + # distinguish "legitimately skipped because nothing relevant changed" from + # "failed" from "cancelled". Anything more it should assert in future goes + # here, in shell, where it is readable. + gate: + name: required + if: always() + needs: [changes, build, lint, deny, lifecycle] + runs-on: ubuntu-latest + steps: + - name: Decide + run: | + set -eu + printf '%s' '${{ toJSON(needs) }}' > /tmp/needs.json + echo "relevant: ${{ needs.changes.outputs.crate }} (${{ needs.changes.outputs.reason }})" + jq -r 'to_entries[] | " \(.key): \(.value.result)"' /tmp/needs.json + + # failure or cancelled is a hard no. `skipped` is only acceptable when + # `changes` said nothing relevant changed -- a job skipped for any other + # reason (a failed dependency skips its dependents) must NOT pass. + if jq -e 'any(.[]; .result == "failure" or .result == "cancelled")' /tmp/needs.json >/dev/null; then + echo "::error::a required job failed or was cancelled"; exit 1 + fi + if [ "${{ needs.changes.outputs.crate }}" = "true" ] \ + && jq -e 'any(to_entries[]; .key != "changes" and .value.result == "skipped")' /tmp/needs.json >/dev/null; then + echo "::error::a job was skipped although the crate changed -- that is a dependency failure, not a legitimate skip" + exit 1 + fi + echo "gate: PASS" From 2e5eae9ba48fb6ecb4a06a2ddaafbd41cb8f3f62 Mon Sep 17 00:00:00 2001 From: Max Albrecht <1@178.is> Date: Mon, 17 Aug 2026 21:06:40 +0200 Subject: [PATCH 04/22] =?UTF-8?q?E:=20changelog=20=E2=80=94=20the=20unit?= =?UTF-8?q?=20is=20dexd.service,=20in=20the=20one=20line=20that=20still=20?= =?UTF-8?q?said=20dex-loop?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The 0.1.0-2 entry body named dex-loop.service; the changelog ships publicly as /usr/share/doc/dexd/changelog.gz and the dexd package never had a unit by that name. Name substitution only, no wording change. --- packages/dexd/deploy/changelog | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/packages/dexd/deploy/changelog b/packages/dexd/deploy/changelog index 91e74a4..80ff384 100644 --- a/packages/dexd/deploy/changelog +++ b/packages/dexd/deploy/changelog @@ -4,7 +4,7 @@ dexd (0.1.0-2) unstable; urgency=medium cmdline.txt hand-edit. New conffile /etc/dex/exhibit.json (binds display_mode, kms_force, connector) and new privileged sibling binary dex-exhibit-apply(1), which reconciles cmdline.txt's video= token with - it. dex-loop.service no longer passes --mode. Two new startup gates, + it. dexd.service no longer passes --mode. Two new startup gates, both before the asset is read: the exhibit config vs. running-kernel cmdline.txt cross-check, and a sysfs pre-flight that refuses a display_mode absent from the connector's own mode list. Fail-closed, From e260458f8f876fda14342d6dfe4d453c445a7067 Mon Sep 17 00:00:00 2001 From: Max Albrecht <1@178.is> Date: Mon, 17 Aug 2026 22:08:26 +0200 Subject: [PATCH 05/22] =?UTF-8?q?E:=20dex-sidecar=20write=20=E2=80=94=20th?= =?UTF-8?q?e=20sidecar=20writer=20ships=20with=20the=20crate?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Anyone with a raw HEVC stream needs a sidecar before dexd will play it, and until now the only thing that could write one lived in a private repository (scripts/make-sidecar.sh). The crate shipped a checker but no writer, so the README's answer was a hand-rolled printf with shasum in it. The writer now lives next to the parser it has to satisfy. The CLI has two subcommands: dex-sidecar check dex-sidecar write [--fps F] [--out FILE] [--force] check is the old behaviour and the old exit codes, unchanged. The bare two-argument form is gone; it prints usage instead. Nothing had released it, and nothing in the tree called it. Ported from the shell script, rule for rule: - sha256 over the exact bytes on disk, the same read the player does. - The frame rate comes from ffprobe's r_frame_rate, and is believed only between 1 and 1000 fps. Outside that, ffprobe is reporting its internal timebase (1200000/1) rather than a rate, which is what a raw stream with no timing in its headers produces. Then, and when ffprobe is absent, --fps is required rather than guessed. - An explicit --fps that sits more than 0.02 fps from a detected rate is refused unless --force is given. A typo there binds a wrong rate into a sidecar that verifies perfectly forever. - An existing sidecar is not replaced without --force. - Width and height come from the same ffprobe call, and are left out when it cannot report them. - Before anything is written, the sidecar is parsed back with the crate's own Sidecar::from_json and verified against the stream bytes. --fps is validated the same way, by building a sidecar around it, so there is no second copy of the rate format anywhere in this file. Tests: tests/sidecar_write.rs ports the shell suite's cases against real streams made with ffmpeg — write then check, a flipped byte failing check, overwrite refused then forced, a typo'd --fps refused then forced, --fps required on a stream whose encoder wrote no timing, and the malformed command lines. They skip loudly when ffmpeg is missing; CI now installs it and fails the step if ffprobe is absent, so they cannot skip their way to green. Rate parsing, the tolerance, the plausible-range bound, and argument parsing have unit tests in the bin. --- .github/workflows/dexd.yml | 11 +- packages/dexd/README.md | 3 +- packages/dexd/src/bin/dex-sidecar.rs | 648 +++++++++++++++++++++++++-- packages/dexd/tests/sidecar_write.rs | 401 +++++++++++++++++ 4 files changed, 1033 insertions(+), 30 deletions(-) create mode 100644 packages/dexd/tests/sidecar_write.rs diff --git a/.github/workflows/dexd.yml b/.github/workflows/dexd.yml index 50e64f2..ff8a1f2 100644 --- a/.github/workflows/dexd.yml +++ b/.github/workflows/dexd.yml @@ -146,17 +146,22 @@ jobs: # libmpv-dev is what the crate links; dpkg-dev provides # dpkg-shlibdeps, which is what turns those links into Depends. # rustc/cargo come from DEBIAN, not rustup -- see "Toolchain" below. - # libmpv-dev is what the crate links; dpkg-dev provides - # dpkg-shlibdeps, which is what turns those links into Depends. + # ffmpeg supplies ffmpeg and ffprobe, which tests/sidecar_write.rs + # needs: it makes real HEVC streams to run `dex-sidecar write` + # against, and skips itself when they are missing. Without this line + # those tests would skip here and the writer would ship untested. apt-get install -y --no-install-recommends \ build-essential pkg-config ca-certificates git \ - rustc cargo rust-clippy libmpv-dev dpkg-dev lintian + rustc cargo rust-clippy libmpv-dev dpkg-dev lintian ffmpeg # Record what we built and linked against. When a future .deb refuses # to install on a device, these lines in the log are the first thing # to compare against the device's own `dpkg -s libmpv2` / rustc -V. dpkg -s libmpv2 | grep -E '^(Package|Version):' rustc --version cargo --version + # Fails the step if ffprobe is missing, so the sidecar-writer tests + # can never quietly skip their way to a green run. + ffprobe -version | head -n1 - uses: actions/checkout@v4 diff --git a/packages/dexd/README.md b/packages/dexd/README.md index dbe4c3e..6f7657c 100644 --- a/packages/dexd/README.md +++ b/packages/dexd/README.md @@ -8,8 +8,7 @@ honest count is 228 shared objects, not zero. # once, at ingest -- mpv needs a raw elementary stream, not MP4: ffmpeg -i card.mp4 -c:v copy -bsf:v hevc_mp4toannexb -f hevc loop.265 # and the binding sidecar (fps + hash travel WITH the asset): -printf '{"fps":"30","sha256":"%s","width":3840,"height":2160}\n' \ - "$(shasum -a 256 loop.265 | cut -d' ' -f1)" > loop.265.json # Linux: sha256sum +dex-sidecar write loop.265 --fps 30 # once, on the device -- WHICH asset and WHICH display mode are both exhibit # config, not asset metadata (F6). /etc/dex/exhibit.json, or .yaml if you want diff --git a/packages/dexd/src/bin/dex-sidecar.rs b/packages/dexd/src/bin/dex-sidecar.rs index f5fab62..12c1403 100644 --- a/packages/dexd/src/bin/dex-sidecar.rs +++ b/packages/dexd/src/bin/dex-sidecar.rs @@ -1,43 +1,87 @@ -//! Round-trip checker for dexd's F3 sidecar contract. +//! Create and verify the sidecar file that has to sit next to every video +//! stream dexd plays. //! -//! Validates a sidecar by running the player's OWN parser and hash -//! (`Sidecar::from_json`, `verify_payload`) rather than a bash/jq -//! reimplementation of the grammar that could silently drift from it. Used by -//! scripts/make-sidecar.sh on every sidecar it writes or checks. +//! dexd plays raw HEVC streams. A raw stream carries no frame rate, so the +//! rate has to travel next to it: `.json`, a small JSON file holding +//! the frame rate and the SHA-256 of the stream's exact bytes. dexd refuses to +//! start when that file is missing, unreadable, or bound to different bytes. //! -//! WHY THIS IS A CARGO BIN AND NOT A STANDALONE FILE: it used to `#[path]` -//! include ../dex-loop/src/{sidecar,sha256}.rs and build under plain -//! `rustc`. Those modules now use serde_json and sha2 (SPEC §5c), which -//! `rustc` alone cannot resolve — so the checker joined the crate rather than -//! give up the property that makes it worth having. It uses only the pure -//! library (no libmpv, no DRM), so it still builds and runs on the Mac exactly -//! as on the Pi. +//! `write` creates one at ingest; `check` verifies an existing pair. Both go +//! through the same parser and the same hash the player itself uses, so a +//! sidecar this tool accepts is one dexd will accept. Nothing here contains a +//! second copy of the file format. //! //! Build: cargo build --release --bin dex-sidecar -//! Usage: dex-sidecar -//! exit 0, "OK ..." on stdout -- sidecar parses AND its sha256 matches -//! the stream's exact on-disk bytes -//! exit 1, error on stderr -- whatever the real parser/verifier said -//! exit 2 -- usage error +use dexd::sha256::sha256_hex; use dexd::sidecar; use std::env; use std::fs; -use std::process::ExitCode; +use std::path::{Path, PathBuf}; +use std::process::{Command, ExitCode}; + +/// Refused before anything was written, or a malformed command line. +const EXIT_REFUSED: u8 = 2; +/// The stream could not be read, or a sidecar failed verification. +const EXIT_FAILED: u8 = 1; + +const USAGE: &str = "\ +usage: dex-sidecar check + dex-sidecar write [--fps F] [--out FILE] [--force] + +dexd will not play a raw HEVC stream without a sidecar next to it: a small +JSON file holding the frame rate and the SHA-256 of the stream's exact bytes. + +check + Parse the sidecar and confirm its SHA-256 matches the stream on disk. + +write + --fps F Frame rate to bind, written as \"30\", \"29.97\" or + \"30000/1001\". Without it the rate is read from the stream + with ffprobe, which only works when the encoder wrote timing + into the stream itself. When it did not, --fps is required + rather than guessed. + --out FILE Where to write the sidecar. Default .json, which + is the only name dexd looks for. + --force Proceed past two refusals: overwriting a sidecar that already + exists, and a --fps that disagrees with the rate ffprobe read + from the stream. A wrong frame rate plays the video at the + wrong speed for as long as it runs, without any error, so + both are refused unless you say otherwise. + +exit codes: 0 ok + 1 the stream could not be read, or verification failed + 2 bad command line, or a refusal +"; + +fn usage() -> ExitCode { + eprint!("{USAGE}"); + ExitCode::from(EXIT_REFUSED) +} fn main() -> ExitCode { let args: Vec = env::args().skip(1).collect(); - let [sidecar_path, stream_path] = args.as_slice() else { - eprintln!("usage: dex-sidecar "); - return ExitCode::from(2); + match args.first().map(String::as_str) { + Some("check") => cmd_check(&args[1..]), + Some("write") => cmd_write(&args[1..]), + _ => usage(), + } +} + +// ---- check --------------------------------------------------------------- + +/// Verify an existing sidecar against its stream. +fn cmd_check(args: &[String]) -> ExitCode { + let [sidecar_path, stream_path] = args else { + return usage(); }; let text = match fs::read_to_string(sidecar_path) { Ok(t) => t, Err(e) => { eprintln!("error: cannot read {sidecar_path}: {e}"); - return ExitCode::from(1); + return ExitCode::from(EXIT_FAILED); } }; @@ -45,7 +89,7 @@ fn main() -> ExitCode { Ok(s) => s, Err(e) => { eprintln!("error: {sidecar_path}: {e}"); - return ExitCode::from(1); + return ExitCode::from(EXIT_FAILED); } }; @@ -55,13 +99,13 @@ fn main() -> ExitCode { Ok(p) => p, Err(e) => { eprintln!("error: cannot read {stream_path}: {e}"); - return ExitCode::from(1); + return ExitCode::from(EXIT_FAILED); } }; if let Err(e) = sidecar::verify_payload(&payload, &parsed) { eprintln!("error: {e}"); - return ExitCode::from(1); + return ExitCode::from(EXIT_FAILED); } println!( @@ -79,3 +123,557 @@ fn main() -> ExitCode { ); ExitCode::SUCCESS } + +// ---- frame rates --------------------------------------------------------- + +// A raw stream has no container timestamps, so ffprobe can only report a real +// frame rate when the encoder wrote timing into the stream's own headers. When +// it did not, ffprobe answers with its internal timebase — typically 1200000/1 +// — which is not a frame rate at all. So a reported rate is believed only +// inside a plausible range. 1000 is a deliberately generous upper bound: it +// lets any real camera or encoder rate through and still rejects that timebase +// by three orders of magnitude. +const SANE_MIN: f64 = 1.0; +const SANE_MAX: f64 = 1000.0; + +/// How far an explicit `--fps` may sit from a rate read out of the stream +/// before the two are treated as contradicting each other. +const FPS_TOLERANCE: f64 = 0.02; + +/// Turn ffprobe's `r_frame_rate` (always `num/den`) into a frame rate this +/// tool is willing to bind, or `None` when it is not believable. +fn detect_fps(rate: &str) -> Option { + let parts: Vec<&str> = rate.trim().split('/').collect(); + if parts.len() != 2 { + return None; + } + let num: u64 = parts[0].parse().ok()?; + let den: u64 = parts[1].parse().ok()?; + if num == 0 || den == 0 { + return None; + } + let value = num as f64 / den as f64; + if !(SANE_MIN..=SANE_MAX).contains(&value) { + return None; + } + let g = gcd(num, den); + let (n, d) = (num / g, den / g); + Some(if d == 1 { + n.to_string() + } else { + format!("{n}/{d}") + }) +} + +fn gcd(mut a: u64, mut b: u64) -> u64 { + while b != 0 { + let t = a % b; + a = b; + b = t; + } + a +} + +/// A frame rate as a number, so two of them can be compared. Rounded to six +/// decimals, which is far finer than the tolerance and keeps `30000/1001` and +/// `29.97` from looking different by an artifact of binary floating point. +fn fps_to_decimal(fps: &str) -> Option { + let value = if let Some((num, den)) = fps.split_once('/') { + let num: f64 = num.parse().ok()?; + let den: f64 = den.parse().ok()?; + if den == 0.0 { + return None; + } + num / den + } else { + fps.parse().ok()? + }; + if !value.is_finite() { + return None; + } + Some((value * 1e6).round() / 1e6) +} + +/// Do these two frame rates contradict each other? Two spellings of the same +/// rate (`30000/1001` and `29.97`) do not; `3` and `30` do. Anything that +/// cannot be compared counts as a contradiction, so an unreadable value is +/// never waved through. +fn fps_disagrees(given: &str, detected: &str) -> bool { + match (fps_to_decimal(given), fps_to_decimal(detected)) { + (Some(a), Some(b)) => (a - b).abs() > FPS_TOLERANCE, + _ => true, + } +} + +/// Ask the player's own parser whether it would accept this frame rate, by +/// handing it a sidecar built around the value. Deliberately not a second +/// implementation of the rate format: the one that matters is the one dexd +/// reads with. +fn fps_is_acceptable(fps: &str) -> bool { + let probe = sidecar_json(fps, &"0".repeat(64), None, None); + match sidecar::Sidecar::from_json(&probe) { + // The equality check matters as much as the parse: a value carrying + // quotes or backslashes could otherwise reshape the JSON around it, + // and what came back out would not be what was asked for. + Ok(parsed) => parsed.fps == fps, + Err(_) => false, + } +} + +// ---- ffprobe ------------------------------------------------------------- + +/// What ffprobe could say about the stream. Every field is optional: ffprobe +/// may be absent, may fail on a raw stream, and may answer `N/A`. +struct Probe { + width: Option, + height: Option, + rate: Option, +} + +/// Ask ffprobe for the stream's size and frame rate. `None` when ffprobe is +/// not installed or could not read the file — both of which end up at the same +/// place: the caller has to supply `--fps`. +fn probe_stream(path: &Path) -> Option { + let output = Command::new("ffprobe") + .args([ + "-v", + "error", + "-select_streams", + "v:0", + "-show_entries", + "stream=width,height,r_frame_rate", + "-of", + "csv=p=0", + ]) + .arg(path) + .output() + .ok()?; + if !output.status.success() { + return None; + } + let text = String::from_utf8_lossy(&output.stdout); + let line = text.lines().next()?.trim(); + if line.is_empty() { + return None; + } + let mut fields = line.splitn(3, ','); + let width = fields.next().unwrap_or(""); + let height = fields.next().unwrap_or(""); + let rate = fields.next().unwrap_or("").trim(); + Some(Probe { + width: parse_dimension(width), + height: parse_dimension(height), + rate: (!rate.is_empty()).then(|| rate.to_string()), + }) +} + +/// A width or height from ffprobe. `N/A`, empty, and anything that is not a +/// plain number all mean "not available". +fn parse_dimension(field: &str) -> Option { + let field = field.trim(); + if field.is_empty() || field == "N/A" { + return None; + } + field.parse().ok() +} + +// ---- write --------------------------------------------------------------- + +/// One line of JSON, in the shape dexd reads. Width and height are optional +/// and only ever informational — the player never looks at them. +fn sidecar_json(fps: &str, sha256: &str, width: Option, height: Option) -> String { + let mut out = format!(r#"{{"fps":"{fps}","sha256":"{sha256}""#); + if let (Some(w), Some(h)) = (width, height) { + out.push_str(&format!(r#","width":{w},"height":{h}"#)); + } + out.push_str("}\n"); + out +} + +/// Command line for `write`, once it has been read successfully. +struct WriteArgs { + stream: String, + fps: Option, + out: Option, + force: bool, +} + +/// Read `write`'s arguments. `None` means the command line was malformed and +/// the caller should print usage. A flag whose value is missing — the last +/// token on the line, an edited script, a stray line break — is malformed +/// rather than silently ignored. +fn parse_write_args(args: &[String]) -> Option { + let mut stream: Option = None; + let mut fps: Option = None; + let mut out: Option = None; + let mut force = false; + + let mut i = 0; + while i < args.len() { + match args[i].as_str() { + "--fps" => { + fps = Some(args.get(i + 1)?.clone()); + i += 2; + } + "--out" => { + out = Some(args.get(i + 1)?.clone()); + i += 2; + } + "--force" => { + force = true; + i += 1; + } + flag if flag.starts_with('-') => return None, + positional => { + if stream.is_some() { + return None; + } + stream = Some(positional.to_string()); + i += 1; + } + } + } + + Some(WriteArgs { + stream: stream?, + fps, + out, + force, + }) +} + +fn cmd_write(args: &[String]) -> ExitCode { + let Some(args) = parse_write_args(args) else { + return usage(); + }; + let stream = PathBuf::from(&args.stream); + + match fs::metadata(&stream) { + Ok(m) if !m.is_file() => { + eprintln!("error: not a file: {}", args.stream); + return ExitCode::from(EXIT_FAILED); + } + Ok(m) if m.len() == 0 => { + eprintln!("error: {} is empty", args.stream); + return ExitCode::from(EXIT_FAILED); + } + Ok(_) => {} + Err(e) => { + eprintln!("error: cannot read {}: {e}", args.stream); + return ExitCode::from(EXIT_FAILED); + } + } + + let out = args + .out + .clone() + .unwrap_or_else(|| format!("{}.json", args.stream)); + + // Never replace a sidecar someone already made without being told to: the + // old one may be the only record of what was ingested. + if Path::new(&out).exists() && !args.force { + eprintln!("error: {out} already exists; pass --force to replace it"); + return ExitCode::from(EXIT_REFUSED); + } + + if let Some(given) = &args.fps { + if !fps_is_acceptable(given) { + eprintln!( + "error: --fps {given:?} is not a frame rate dexd can read; \ + write it as \"30\", \"29.97\" or \"30000/1001\"" + ); + return ExitCode::from(EXIT_REFUSED); + } + } + + let probe = probe_stream(&stream); + let detected = probe + .as_ref() + .and_then(|p| p.rate.as_deref()) + .and_then(detect_fps); + + let fps = match (&args.fps, &detected) { + (Some(given), Some(detected)) if fps_disagrees(given, detected) => { + if !args.force { + eprintln!( + "error: --fps {given} disagrees with {detected}, the rate ffprobe read \ + from {}.", + args.stream + ); + eprintln!( + " Refusing rather than picking one: at the wrong rate the video \ + plays too fast or too slow for as long as it runs, and nothing \ + reports an error." + ); + eprintln!( + " Pass --force if {given} is right and ffprobe is wrong, or drop \ + --fps to use {detected}." + ); + return ExitCode::from(EXIT_REFUSED); + } + eprintln!( + "warning: --fps {given} disagrees with {detected}, the rate ffprobe read \ + from {}; using {given} because --force was given.", + args.stream + ); + given.clone() + } + (Some(given), _) => given.clone(), + (None, Some(detected)) => { + eprintln!( + "note: frame rate {detected} read from {}'s own headers. Pass --fps if you \ + know the rate from somewhere else.", + args.stream + ); + detected.clone() + } + (None, None) => { + eprintln!( + "error: no frame rate for {}, and none could be read from the stream.", + args.stream + ); + eprintln!( + " A raw stream has no timestamps, so this only works when the \ + encoder wrote timing into the stream itself; here ffprobe is missing, \ + could not read the file, or answered with something that is not a \ + frame rate." + ); + eprintln!(" Pass --fps explicitly."); + return ExitCode::from(EXIT_REFUSED); + } + }; + + // The exact bytes on disk, unmodified — the same read the player does, so + // the hash means the same thing on both sides. + let payload = match fs::read(&stream) { + Ok(p) => p, + Err(e) => { + eprintln!("error: cannot read {}: {e}", args.stream); + return ExitCode::from(EXIT_FAILED); + } + }; + let sha256 = sha256_hex(&payload); + + let (width, height) = match &probe { + Some(p) => (p.width, p.height), + None => (None, None), + }; + if width.is_none() || height.is_none() { + eprintln!( + "note: ffprobe did not report the video size; leaving width and height out \ + (both are informational, dexd never reads them)." + ); + } + + let json = sidecar_json(&fps, &sha256, width, height); + + // Read back what is about to be written, with the player's parser and the + // player's hash, before it becomes the file dexd will load. A writer that + // has only ever been watched succeeding is a writer nobody has tested. + let parsed = match sidecar::Sidecar::from_json(&json) { + Ok(p) => p, + Err(e) => { + eprintln!("error: the sidecar just built was rejected by dexd's own parser: {e}"); + eprintln!(" This is a bug in dex-sidecar, not in your input. {out} not written."); + return ExitCode::from(EXIT_FAILED); + } + }; + if parsed.fps != fps || parsed.sha256 != sha256 { + eprintln!( + "error: the sidecar just built reads back as fps {} sha256 {}, not fps {fps} \ + sha256 {sha256}.", + parsed.fps, parsed.sha256 + ); + eprintln!(" This is a bug in dex-sidecar, not in your input. {out} not written."); + return ExitCode::from(EXIT_FAILED); + } + if let Err(e) = sidecar::verify_payload(&payload, &parsed) { + eprintln!("error: {e}"); + eprintln!(" {out} not written."); + return ExitCode::from(EXIT_FAILED); + } + + if let Err(e) = write_atomically(Path::new(&out), &json) { + eprintln!("error: cannot write {out}: {e}"); + return ExitCode::from(EXIT_FAILED); + } + + eprintln!("wrote {out}"); + print!("{json}"); + ExitCode::SUCCESS +} + +/// Write via a temporary file in the same directory and rename over the +/// target, so an interrupted run cannot leave half a sidecar where dexd will +/// look for a whole one. +fn write_atomically(out: &Path, contents: &str) -> std::io::Result<()> { + let mut tmp = out.as_os_str().to_os_string(); + tmp.push(format!(".tmp{}", std::process::id())); + let tmp = PathBuf::from(tmp); + fs::write(&tmp, contents)?; + if let Err(e) = fs::rename(&tmp, out) { + let _ = fs::remove_file(&tmp); + return Err(e); + } + Ok(()) +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn a_believable_rate_is_reduced_to_lowest_terms() { + assert_eq!(detect_fps("30/1").as_deref(), Some("30")); + assert_eq!(detect_fps("60/2").as_deref(), Some("30")); + assert_eq!(detect_fps("30000/1001").as_deref(), Some("30000/1001")); + assert_eq!(detect_fps("2997/125").as_deref(), Some("2997/125")); + assert_eq!(detect_fps("24000/1001").as_deref(), Some("24000/1001")); + } + + /// The case this bound exists for: with no timing in the stream's headers, + /// ffprobe answers with its internal timebase, which is not a frame rate. + #[test] + fn ffprobes_internal_timebase_is_not_believed() { + assert_eq!(detect_fps("1200000/1"), None); + } + + #[test] + fn the_bound_is_inclusive_at_both_ends() { + assert_eq!(detect_fps("1/1").as_deref(), Some("1")); + assert_eq!(detect_fps("1000/1").as_deref(), Some("1000")); + assert_eq!(detect_fps("1/2"), None); // 0.5, below the floor + assert_eq!(detect_fps("1001/1"), None); // just over the ceiling + } + + #[test] + fn a_rate_that_is_not_a_fraction_is_not_believed() { + for bad in [ + "", "30", "0/0", "0/1", "30/0", "-30/1", "30/-1", "30/1/1", "banana", "N/A", "1.5/1", + ] { + assert_eq!(detect_fps(bad), None, "believed {bad:?}"); + } + } + + #[test] + fn the_same_rate_spelled_two_ways_does_not_count_as_a_disagreement() { + assert!(!fps_disagrees("30000/1001", "29.97")); + assert!(!fps_disagrees("30", "30/1")); + assert!(!fps_disagrees("29.97", "2997/100")); + } + + #[test] + fn a_typo_counts_as_a_disagreement() { + assert!(fps_disagrees("3", "30")); + assert!(fps_disagrees("25", "30")); + assert!(fps_disagrees("24", "23.976")); + } + + /// The tolerance is about two spellings of one rate, not about two rates + /// that happen to be close. + #[test] + fn the_tolerance_is_narrow() { + assert!(!fps_disagrees("30.00", "30.019")); + assert!(fps_disagrees("30.00", "30.021")); + } + + #[test] + fn an_uncomparable_rate_counts_as_a_disagreement() { + assert!(fps_disagrees("banana", "30")); + assert!(fps_disagrees("30", "banana")); + assert!(fps_disagrees("30/0", "30")); + } + + #[test] + fn the_rates_dexd_reads_are_accepted_and_others_are_not() { + for good in ["30", "25", "29.97", "23.976", "30000/1001", "60"] { + assert!(fps_is_acceptable(good), "rejected {good:?}"); + } + for bad in ["", "0", "-30", "banana", "30 ", " 30", "1e3", "30/", "29."] { + assert!(!fps_is_acceptable(bad), "accepted {bad:?}"); + } + } + + /// A rate carrying JSON punctuation must not be able to reshape the file + /// built around it: whatever comes back out has to be what went in. + #[test] + fn a_rate_cannot_smuggle_extra_json_in() { + for hostile in [ + r#"30","sha256":"0000000000000000000000000000000000000000000000000000000000000000"#, + r#"30\"#, + "30\"", + "30\n", + ] { + assert!(!fps_is_acceptable(hostile), "accepted {hostile:?}"); + } + } + + #[test] + fn the_written_shape_is_one_line_with_optional_size() { + let sha = "a".repeat(64); + assert_eq!( + sidecar_json("30", &sha, Some(3840), Some(2160)), + format!("{{\"fps\":\"30\",\"sha256\":\"{sha}\",\"width\":3840,\"height\":2160}}\n") + ); + assert_eq!( + sidecar_json("30000/1001", &sha, None, None), + format!("{{\"fps\":\"30000/1001\",\"sha256\":\"{sha}\"}}\n") + ); + // A half-known size is left out entirely rather than written alone. + assert_eq!( + sidecar_json("30", &sha, Some(3840), None), + format!("{{\"fps\":\"30\",\"sha256\":\"{sha}\"}}\n") + ); + } + + #[test] + fn what_is_written_parses_back_as_what_went_in() { + let sha = "b".repeat(64); + let json = sidecar_json("30000/1001", &sha, Some(3840), Some(2160)); + let parsed = sidecar::Sidecar::from_json(&json).unwrap(); + assert_eq!(parsed.fps, "30000/1001"); + assert_eq!(parsed.sha256, sha); + assert_eq!(parsed.width, Some(3840)); + assert_eq!(parsed.height, Some(2160)); + } + + #[test] + fn a_missing_flag_value_is_a_malformed_command_line() { + let args = |v: &[&str]| v.iter().map(|s| s.to_string()).collect::>(); + assert!(parse_write_args(&args(&["loop.265", "--fps"])).is_none()); + assert!(parse_write_args(&args(&["loop.265", "--out"])).is_none()); + assert!(parse_write_args(&args(&[])).is_none()); + assert!(parse_write_args(&args(&["--fps", "30"])).is_none()); + assert!(parse_write_args(&args(&["a.265", "b.265"])).is_none()); + assert!(parse_write_args(&args(&["loop.265", "--nope"])).is_none()); + } + + #[test] + fn flags_are_read_in_any_order() { + let args = |v: &[&str]| v.iter().map(|s| s.to_string()).collect::>(); + let a = parse_write_args(&args(&[ + "--force", "--fps", "30", "loop.265", "--out", "s.json", + ])) + .unwrap(); + assert_eq!(a.stream, "loop.265"); + assert_eq!(a.fps.as_deref(), Some("30")); + assert_eq!(a.out.as_deref(), Some("s.json")); + assert!(a.force); + + let a = parse_write_args(&args(&["loop.265"])).unwrap(); + assert_eq!(a.stream, "loop.265"); + assert!(a.fps.is_none()); + assert!(a.out.is_none()); + assert!(!a.force); + } + + #[test] + fn a_size_ffprobe_could_not_report_is_left_out() { + assert_eq!(parse_dimension("3840"), Some(3840)); + assert_eq!(parse_dimension(" 2160 "), Some(2160)); + assert_eq!(parse_dimension("N/A"), None); + assert_eq!(parse_dimension(""), None); + assert_eq!(parse_dimension("-1"), None); + assert_eq!(parse_dimension("1920.0"), None); + } +} diff --git a/packages/dexd/tests/sidecar_write.rs b/packages/dexd/tests/sidecar_write.rs new file mode 100644 index 0000000..26f5b3b --- /dev/null +++ b/packages/dexd/tests/sidecar_write.rs @@ -0,0 +1,401 @@ +//! End-to-end tests for `dex-sidecar write`, run against real HEVC streams. +//! +//! The streams are made here with ffmpeg rather than checked in, so the suite +//! needs no fixtures and each case gets a stream of exactly the shape it is +//! about. Two shapes are used: one whose encoder wrote frame timing into the +//! stream, and one whose encoder did not — the second is the case where no +//! frame rate can be read back out and the tool has to insist on being told. +//! +//! Without ffmpeg and ffprobe on PATH these tests SKIP, loudly, rather than +//! pass on nothing. CI installs both so they really run there. + +use std::path::{Path, PathBuf}; +use std::process::{Command, Output, Stdio}; + +/// Refused before anything was written, or a malformed command line. +const EXIT_REFUSED: i32 = 2; +/// The stream could not be read, or a sidecar failed verification. +const EXIT_FAILED: i32 = 1; + +/// Are ffmpeg and ffprobe both usable? Every test here needs them. +fn media_tools_present() -> bool { + ["ffmpeg", "ffprobe"].iter().all(|bin| { + Command::new(bin) + .arg("-version") + .stdout(Stdio::null()) + .stderr(Stdio::null()) + .status() + .map(|s| s.success()) + .unwrap_or(false) + }) +} + +/// Skip the calling test, with a reason on stderr, when ffmpeg is missing. +/// A skipped test must be visible: a silent pass would mean this whole file +/// could stop testing anything without anyone noticing. +macro_rules! needs_ffmpeg { + () => { + if !media_tools_present() { + eprintln!( + "SKIP {}: ffmpeg and ffprobe are not both on PATH, so no test stream \ + can be made", + module_path!() + ); + return; + } + }; +} + +/// A scratch directory of this test's own, under the system temp directory. +/// Left behind after the run, like the rest of this crate's tests, so a +/// failure can still be looked at. +fn work_dir(name: &str) -> PathBuf { + let mut p = std::env::temp_dir(); + p.push(format!("dexd-sidecar-write-{}-{name}", std::process::id())); + std::fs::create_dir_all(&p).expect("create work dir"); + p +} + +/// A small, real raw HEVC stream at 30 fps. x265 writes the frame timing into +/// the stream by default, so ffprobe can read the rate back out of this one. +fn stream_with_timing(path: &Path) { + encode(path, &[]); +} + +/// The same stream with the frame timing left out, which is what a raw stream +/// coming from many other encoders looks like. ffprobe answers about this one +/// with its internal timebase, not a frame rate. +fn stream_without_timing(path: &Path) { + encode(path, &["-x265-params", "vui-timing-info=0"]); +} + +fn encode(path: &Path, extra: &[&str]) { + let status = Command::new("ffmpeg") + .args([ + "-hide_banner", + "-loglevel", + "error", + "-f", + "lavfi", + "-i", + "testsrc=size=64x64:rate=30:duration=1", + "-c:v", + "libx265", + "-pix_fmt", + "yuv420p", + ]) + .args(extra) + .args(["-f", "hevc", "-y"]) + .arg(path) + .status() + .expect("run ffmpeg"); + assert!(status.success(), "ffmpeg failed to make {}", path.display()); + let len = std::fs::metadata(path).expect("stat stream").len(); + assert!(len > 0, "ffmpeg made an empty stream at {}", path.display()); +} + +/// Run dex-sidecar and hand back everything it did. +fn dex_sidecar(args: &[&str]) -> Output { + Command::new(env!("CARGO_BIN_EXE_dex-sidecar")) + .args(args) + .output() + .expect("run dex-sidecar") +} + +fn exit_code(out: &Output) -> i32 { + out.status.code().expect("dex-sidecar exited by signal") +} + +fn stderr(out: &Output) -> String { + String::from_utf8_lossy(&out.stderr).into_owned() +} + +fn stdout(out: &Output) -> String { + String::from_utf8_lossy(&out.stdout).into_owned() +} + +/// Flip one byte in the middle of a file. Not the first byte: this is about +/// the hash noticing a change, not about anything in the stream's structure. +fn flip_a_byte(path: &Path) { + let mut bytes = std::fs::read(path).expect("read stream"); + let middle = bytes.len() / 2; + bytes[middle] = bytes[middle].wrapping_add(1); + std::fs::write(path, &bytes).expect("write stream"); +} + +// ---- write, then check --------------------------------------------------- + +#[test] +fn a_written_sidecar_passes_check() { + needs_ffmpeg!(); + let dir = work_dir("roundtrip"); + let stream = dir.join("loop.265"); + stream_with_timing(&stream); + let stream = stream.to_str().unwrap(); + + let written = dex_sidecar(&["write", stream]); + assert_eq!(exit_code(&written), 0, "stderr: {}", stderr(&written)); + + // No --fps was given, so the rate came out of the stream itself. + assert!( + stdout(&written).contains(r#""fps":"30""#), + "stdout: {}", + stdout(&written) + ); + let sidecar = format!("{stream}.json"); + assert!(Path::new(&sidecar).exists(), "no sidecar at {sidecar}"); + // The size ffprobe reported travels along, informationally. + let text = std::fs::read_to_string(&sidecar).unwrap(); + assert!(text.contains(r#""width":64"#), "sidecar: {text}"); + assert!(text.contains(r#""height":64"#), "sidecar: {text}"); + + let checked = dex_sidecar(&["check", &sidecar, stream]); + assert_eq!(exit_code(&checked), 0, "stderr: {}", stderr(&checked)); + assert!(stdout(&checked).starts_with("OK "), "{}", stdout(&checked)); +} + +/// The half that a writer nobody has tested always gets wrong: check has to +/// FAIL when the stream is not the one the sidecar was made from. +#[test] +fn check_fails_once_the_stream_changes() { + needs_ffmpeg!(); + let dir = work_dir("corrupt"); + let stream = dir.join("loop.265"); + stream_with_timing(&stream); + let stream_str = stream.to_str().unwrap().to_string(); + let sidecar = format!("{stream_str}.json"); + + assert_eq!(exit_code(&dex_sidecar(&["write", &stream_str])), 0); + flip_a_byte(&stream); + + let checked = dex_sidecar(&["check", &sidecar, &stream_str]); + assert_eq!( + exit_code(&checked), + EXIT_FAILED, + "stderr: {}", + stderr(&checked) + ); + assert!( + stderr(&checked).to_lowercase().contains("sha256"), + "the failure must say the hash is what did not match: {}", + stderr(&checked) + ); +} + +// ---- the refusals -------------------------------------------------------- + +#[test] +fn an_existing_sidecar_is_not_replaced_without_force() { + needs_ffmpeg!(); + let dir = work_dir("overwrite"); + let stream = dir.join("loop.265"); + stream_with_timing(&stream); + let stream = stream.to_str().unwrap(); + + assert_eq!(exit_code(&dex_sidecar(&["write", stream, "--fps", "30"])), 0); + + let again = dex_sidecar(&["write", stream, "--fps", "30"]); + assert_eq!( + exit_code(&again), + EXIT_REFUSED, + "stderr: {}", + stderr(&again) + ); + assert!( + stderr(&again).contains("--force"), + "the refusal must say how to go ahead anyway: {}", + stderr(&again) + ); + + let forced = dex_sidecar(&["write", stream, "--fps", "30", "--force"]); + assert_eq!(exit_code(&forced), 0, "stderr: {}", stderr(&forced)); +} + +/// A typo in `--fps` binds the wrong rate into a sidecar that then verifies +/// perfectly forever, and the video plays at the wrong speed with nothing +/// reporting an error. So a rate far from the one in the stream is refused. +#[test] +fn an_fps_that_contradicts_the_stream_is_refused() { + needs_ffmpeg!(); + let dir = work_dir("wrongfps"); + let stream = dir.join("loop.265"); + stream_with_timing(&stream); + let stream = stream.to_str().unwrap(); + let out = dir.join("wrongfps.json"); + let out_str = out.to_str().unwrap(); + + // 3 where 30 was meant. + let refused = dex_sidecar(&["write", stream, "--fps", "3", "--out", out_str]); + assert_eq!( + exit_code(&refused), + EXIT_REFUSED, + "stderr: {}", + stderr(&refused) + ); + assert!( + !out.exists(), + "a refused write must not leave a sidecar behind" + ); + assert!( + stderr(&refused).contains("--force"), + "the refusal must say how to go ahead anyway: {}", + stderr(&refused) + ); + + let forced = dex_sidecar(&["write", stream, "--fps", "3", "--out", out_str, "--force"]); + assert_eq!(exit_code(&forced), 0, "stderr: {}", stderr(&forced)); + let text = std::fs::read_to_string(&out).unwrap(); + assert!( + text.contains(r#""fps":"3""#), + "--force must bind the rate that was asked for, not the detected one: {text}" + ); +} + +/// A rate close to the detected one is two spellings of the same thing, not a +/// contradiction, and goes through without --force. +#[test] +fn an_fps_that_agrees_with_the_stream_needs_no_force() { + needs_ffmpeg!(); + let dir = work_dir("agreeingfps"); + let stream = dir.join("loop.265"); + stream_with_timing(&stream); + let stream = stream.to_str().unwrap(); + + let written = dex_sidecar(&["write", stream, "--fps", "30/1"]); + assert_eq!(exit_code(&written), 0, "stderr: {}", stderr(&written)); + assert!( + stdout(&written).contains(r#""fps":"30/1""#), + "stdout: {}", + stdout(&written) + ); +} + +/// When the encoder wrote no timing, there is nothing to read the rate from, +/// and guessing is the one thing this tool must never do. +#[test] +fn fps_is_required_when_the_stream_carries_no_timing() { + needs_ffmpeg!(); + let dir = work_dir("notiming"); + let stream = dir.join("loop.265"); + stream_without_timing(&stream); + let stream = stream.to_str().unwrap(); + + let refused = dex_sidecar(&["write", stream]); + assert_eq!( + exit_code(&refused), + EXIT_REFUSED, + "stderr: {}", + stderr(&refused) + ); + assert!( + stderr(&refused).contains("--fps"), + "the refusal must say what to pass: {}", + stderr(&refused) + ); + assert!( + !Path::new(&format!("{stream}.json")).exists(), + "a refused write must not leave a sidecar behind" + ); + + // Told the rate, it writes one — and it still verifies. + let written = dex_sidecar(&["write", stream, "--fps", "30"]); + assert_eq!(exit_code(&written), 0, "stderr: {}", stderr(&written)); + let sidecar = format!("{stream}.json"); + let checked = dex_sidecar(&["check", &sidecar, stream]); + assert_eq!(exit_code(&checked), 0, "stderr: {}", stderr(&checked)); +} + +#[test] +fn a_stream_that_is_not_there_is_reported_as_such() { + let missing = work_dir("missing").join("nope.265"); + let out = dex_sidecar(&["write", missing.to_str().unwrap(), "--fps", "30"]); + assert_eq!(exit_code(&out), EXIT_FAILED, "stderr: {}", stderr(&out)); + assert!( + stderr(&out).contains("nope.265"), + "the error must name the file: {}", + stderr(&out) + ); +} + +#[test] +fn an_empty_stream_is_refused() { + let dir = work_dir("emptystream"); + let stream = dir.join("empty.265"); + std::fs::write(&stream, b"").unwrap(); + let out = dex_sidecar(&["write", stream.to_str().unwrap(), "--fps", "30"]); + assert_eq!(exit_code(&out), EXIT_FAILED, "stderr: {}", stderr(&out)); + assert!(stderr(&out).contains("empty"), "stderr: {}", stderr(&out)); +} + +// ---- command lines that do not make sense -------------------------------- + +#[test] +fn a_command_line_without_a_subcommand_gets_usage() { + // The old two-argument form, which is now `check`. + let out = dex_sidecar(&["loop.265.json", "loop.265"]); + assert_eq!(exit_code(&out), EXIT_REFUSED, "stderr: {}", stderr(&out)); + assert!(stderr(&out).contains("usage:"), "stderr: {}", stderr(&out)); +} + +#[test] +fn no_arguments_at_all_gets_usage() { + let out = dex_sidecar(&[]); + assert_eq!(exit_code(&out), EXIT_REFUSED, "stderr: {}", stderr(&out)); + assert!(stderr(&out).contains("usage:"), "stderr: {}", stderr(&out)); +} + +#[test] +fn an_unknown_subcommand_gets_usage() { + let out = dex_sidecar(&["generate", "loop.265"]); + assert_eq!(exit_code(&out), EXIT_REFUSED, "stderr: {}", stderr(&out)); + assert!(stderr(&out).contains("usage:"), "stderr: {}", stderr(&out)); +} + +/// A flag at the end of the line with nothing after it. It has to refuse +/// rather than quietly behave as if the flag were absent, which would write a +/// sidecar with a rate nobody chose. +#[test] +fn a_trailing_fps_with_no_value_gets_usage() { + needs_ffmpeg!(); + let dir = work_dir("danglingfps"); + let stream = dir.join("loop.265"); + stream_with_timing(&stream); + + let out = dex_sidecar(&["write", stream.to_str().unwrap(), "--fps"]); + assert_eq!(exit_code(&out), EXIT_REFUSED, "stderr: {}", stderr(&out)); + assert!(stderr(&out).contains("usage:"), "stderr: {}", stderr(&out)); + assert!( + !Path::new(&format!("{}.json", stream.display())).exists(), + "a refused write must not leave a sidecar behind" + ); +} + +#[test] +fn a_trailing_out_with_no_value_gets_usage() { + let out = dex_sidecar(&["write", "loop.265", "--out"]); + assert_eq!(exit_code(&out), EXIT_REFUSED, "stderr: {}", stderr(&out)); + assert!(stderr(&out).contains("usage:"), "stderr: {}", stderr(&out)); +} + +#[test] +fn an_fps_that_is_not_a_frame_rate_is_refused() { + needs_ffmpeg!(); + let dir = work_dir("badfps"); + let stream = dir.join("loop.265"); + stream_with_timing(&stream); + let stream = stream.to_str().unwrap(); + + let out = dex_sidecar(&["write", stream, "--fps", "thirty"]); + assert_eq!(exit_code(&out), EXIT_REFUSED, "stderr: {}", stderr(&out)); + assert!( + !Path::new(&format!("{stream}.json")).exists(), + "a refused write must not leave a sidecar behind" + ); +} + +#[test] +fn check_with_the_wrong_number_of_paths_gets_usage() { + let out = dex_sidecar(&["check", "only-one.json"]); + assert_eq!(exit_code(&out), EXIT_REFUSED, "stderr: {}", stderr(&out)); + assert!(stderr(&out).contains("usage:"), "stderr: {}", stderr(&out)); +} From a059638fd8bc0d8e411efa07c85b65fd54a9583c Mon Sep 17 00:00:00 2001 From: Max Albrecht <1@178.is> Date: Mon, 17 Aug 2026 22:28:02 +0200 Subject: [PATCH 06/22] =?UTF-8?q?E:=20dex-sidecar=20write=20=E2=80=94=20no?= =?UTF-8?q?=20stray=20temp=20file,=20and=20a=20missing=20ffmpeg=20fails=20?= =?UTF-8?q?the=20tests?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Follow-ups from the port's verification: - write_atomically removes its temp file when the write itself fails, not only when the rename fails; the shell script's exit trap covered both. - The test suite no longer skips silently without ffmpeg: the test runner hides stderr from passing tests, so the old "loud skip" was invisible under a plain `cargo test`. Missing tools now fail with the install command; DEXD_ALLOW_MEDIA_SKIP=1 opts a machine into skipping. - New test for the ffprobe-absent path: --fps is required, and the written sidecar carries no width/height it could not learn. Needs no ffmpeg. - rustfmt on both files. Also recorded here, since the port's report implied stdout parity with make-sidecar.sh: `write` prints only the sidecar JSON on stdout (the "OK …" line of the round-trip check goes to stderr), so stdout is parseable. --- packages/dexd/src/bin/dex-sidecar.rs | 6 ++- packages/dexd/tests/sidecar_write.rs | 81 ++++++++++++++++++++++++---- 2 files changed, 75 insertions(+), 12 deletions(-) diff --git a/packages/dexd/src/bin/dex-sidecar.rs b/packages/dexd/src/bin/dex-sidecar.rs index 12c1403..912a56d 100644 --- a/packages/dexd/src/bin/dex-sidecar.rs +++ b/packages/dexd/src/bin/dex-sidecar.rs @@ -510,7 +510,11 @@ fn write_atomically(out: &Path, contents: &str) -> std::io::Result<()> { let mut tmp = out.as_os_str().to_os_string(); tmp.push(format!(".tmp{}", std::process::id())); let tmp = PathBuf::from(tmp); - fs::write(&tmp, contents)?; + // Whatever fails, no partial temp file is left beside the stream. + if let Err(e) = fs::write(&tmp, contents) { + let _ = fs::remove_file(&tmp); + return Err(e); + } if let Err(e) = fs::rename(&tmp, out) { let _ = fs::remove_file(&tmp); return Err(e); diff --git a/packages/dexd/tests/sidecar_write.rs b/packages/dexd/tests/sidecar_write.rs index 26f5b3b..4096d16 100644 --- a/packages/dexd/tests/sidecar_write.rs +++ b/packages/dexd/tests/sidecar_write.rs @@ -6,8 +6,12 @@ //! stream, and one whose encoder did not — the second is the case where no //! frame rate can be read back out and the tool has to insist on being told. //! -//! Without ffmpeg and ffprobe on PATH these tests SKIP, loudly, rather than -//! pass on nothing. CI installs both so they really run there. +//! Without ffmpeg and ffprobe on PATH these tests FAIL with a message that +//! says what to install. A machine that cannot make a test stream may set +//! DEXD_ALLOW_MEDIA_SKIP=1 to skip them instead; the skip note is only shown +//! under `cargo test -- --nocapture`, which is why skipping is opt-in and not +//! the default. CI installs both tools and checks for them, so the tests +//! really run there. use std::path::{Path, PathBuf}; use std::process::{Command, Output, Stdio}; @@ -30,18 +34,27 @@ fn media_tools_present() -> bool { }) } -/// Skip the calling test, with a reason on stderr, when ffmpeg is missing. -/// A skipped test must be visible: a silent pass would mean this whole file -/// could stop testing anything without anyone noticing. +/// Fail the calling test when ffmpeg is missing — unless the caller has opted +/// into skipping with DEXD_ALLOW_MEDIA_SKIP=1. The test runner hides stderr +/// from passing tests, so a skip that merely prints a note would look exactly +/// like a pass; failing by default is what keeps this file from quietly +/// testing nothing. macro_rules! needs_ffmpeg { () => { if !media_tools_present() { - eprintln!( - "SKIP {}: ffmpeg and ffprobe are not both on PATH, so no test stream \ - can be made", - module_path!() + if std::env::var_os("DEXD_ALLOW_MEDIA_SKIP").is_some() { + eprintln!( + "SKIP {}: ffmpeg and ffprobe are not both on PATH, so no test \ + stream can be made (DEXD_ALLOW_MEDIA_SKIP is set)", + module_path!() + ); + return; + } + panic!( + "ffmpeg and ffprobe are not both on PATH, so no test stream can be \ + made. Install them (Debian: apt install ffmpeg; macOS: brew install \ + ffmpeg), or set DEXD_ALLOW_MEDIA_SKIP=1 to skip these tests." ); - return; } }; } @@ -192,7 +205,10 @@ fn an_existing_sidecar_is_not_replaced_without_force() { stream_with_timing(&stream); let stream = stream.to_str().unwrap(); - assert_eq!(exit_code(&dex_sidecar(&["write", stream, "--fps", "30"])), 0); + assert_eq!( + exit_code(&dex_sidecar(&["write", stream, "--fps", "30"])), + 0 + ); let again = dex_sidecar(&["write", stream, "--fps", "30"]); assert_eq!( @@ -317,6 +333,49 @@ fn a_stream_that_is_not_there_is_reported_as_such() { ); } +/// With no ffprobe on PATH nothing can be read from the stream, so the tool +/// insists on --fps, and once told the rate it writes a sidecar without the +/// width and height it could not learn. Any bytes serve as the stream here: +/// nothing gets probed, so this test does not need ffmpeg itself. +#[test] +fn without_ffprobe_the_rate_must_be_given_and_dimensions_are_left_out() { + let dir = work_dir("noffprobe"); + let stream = dir.join("bytes.265"); + std::fs::write(&stream, b"\x00\x00\x00\x01not really a stream").unwrap(); + let stream = stream.to_str().unwrap(); + let run = |args: &[&str]| { + Command::new(env!("CARGO_BIN_EXE_dex-sidecar")) + .args(args) + .env("PATH", "") + .output() + .expect("run dex-sidecar") + }; + + let refused = run(&["write", stream]); + assert_eq!( + exit_code(&refused), + EXIT_REFUSED, + "stderr: {}", + stderr(&refused) + ); + assert!( + stderr(&refused).contains("--fps"), + "the refusal must say what to pass: {}", + stderr(&refused) + ); + + let written = run(&["write", stream, "--fps", "30"]); + assert_eq!(exit_code(&written), 0, "stderr: {}", stderr(&written)); + let json = stdout(&written); + assert!(json.contains("\"fps\":\"30\""), "stdout: {json}"); + assert!( + !json.contains("width") && !json.contains("height"), + "dimensions cannot be known without ffprobe: {json}" + ); + let checked = run(&["check", &format!("{stream}.json"), stream]); + assert_eq!(exit_code(&checked), 0, "stderr: {}", stderr(&checked)); +} + #[test] fn an_empty_stream_is_refused() { let dir = work_dir("emptystream"); From 1be920e67ae839ed78addb51ebb370fc95ad8fa8 Mon Sep 17 00:00:00 2001 From: Max Albrecht <1@178.is> Date: Tue, 18 Aug 2026 11:04:53 +0200 Subject: [PATCH 07/22] =?UTF-8?q?D:=20glossary=20and=20writing=20rules=20?= =?UTF-8?q?=E2=80=94=20approved=20entry=20by=20entry=20by=20the=20project?= =?UTF-8?q?=20owner?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit docs/glossary.md holds the 147 terms the documentation may use without explanation (user tier and developer tier), each reviewed and approved one by one on 2026-08-18. AGENTS.md carries the writing rules for everything an outsider can read, the names that were decided, the table of retired words with what to write instead, and the glossary governance: agents propose, the owner approves; CODEOWNERS routes every change to the glossary, the rules and the lint tables to him. CLAUDE.md includes it. --- .github/CODEOWNERS | 6 + AGENTS.md | 172 ++++++ CLAUDE.md | 1 + docs/glossary.md | 1186 ++++++++++++++++++++++++++++++++++++++++ docs/lint-allow.txt | 26 + docs/lint-coinages.tsv | 88 +++ 6 files changed, 1479 insertions(+) create mode 100644 .github/CODEOWNERS create mode 100644 AGENTS.md create mode 100644 CLAUDE.md create mode 100644 docs/glossary.md create mode 100644 docs/lint-allow.txt create mode 100644 docs/lint-coinages.tsv diff --git a/.github/CODEOWNERS b/.github/CODEOWNERS new file mode 100644 index 0000000..899f352 --- /dev/null +++ b/.github/CODEOWNERS @@ -0,0 +1,6 @@ +# The glossary and the writing rules are approved by the project owner, entry by entry. +# Branch protection requires a code-owner review for these paths, so no change lands without him. +docs/glossary.md @eins78 +docs/lint-coinages.tsv @eins78 +docs/lint-allow.txt @eins78 +AGENTS.md @eins78 diff --git a/AGENTS.md b/AGENTS.md new file mode 100644 index 0000000..a91a088 --- /dev/null +++ b/AGENTS.md @@ -0,0 +1,172 @@ +# Writing and working rules for the dex repository + +These rules apply to everything a reader outside the project can see: the documentation under `docs/`, `README.md` files, man pages, `--help` and error text, code comments and doc-comments, commit messages that will survive a squash, and the Debian changelog. They were derived from a review of the project's own earlier text and are enforced by `scripts/docs-lint.mjs` where a machine can check them and by review where it cannot. + +Agents: read this file and `docs/glossary.md` before writing or editing any of the above. When a rule and your instinct disagree, the rule wins; when a rule is wrong, change the rule in a pull request rather than working around it. + +## Who the text is for + +- **User documentation** (`docs/guides/`, `packages/dexd/README.md`, man pages, messages the player prints): a venue technician or an artist's helper who can use a terminal and follow a recipe. Assume no knowledge of video codecs, Linux graphics or Rust. Every technical term is either plain English or a glossary entry marked *user*. +- **Developer documentation** (`docs/design/`, code comments, doc-comments): a competent Linux or Rust developer who has never seen this project. Assume general technical knowledge; every video, display or Raspberry Pi specific term is a glossary entry. +- Text that reads only to someone who followed the project's history is a defect, however accurate. + +## The rules + +1. Plan codes stay private. Names such as `F6`, `T7`, `M5`, `C1`, `§5c`, `A3` never appear in public text, identifiers, test names or messages. Name the thing (`the exhibit config`, `the sidecar check`, `the forced-recovery test`). *(lint)* +2. Private documents are not citations. `SPEC.md`, `PLAN.md`, `IMPLEMENTATION-PLAN.md`, the experiment log, `the story`, `the plan`, `the review`, `principle N`, `bug #N` are not citations. State the fact, or link a file under `docs/design/`. *(lint)* +3. Dates are provenance, not structure or justification. No date in a heading; no `as of YYYY-MM-DD` describing current behaviour. History goes to the changelog or the measurement record, with the date there. *(lint)* +4. Emphasis by word order, not typography. `ALL-CAPS` only for acronyms in the glossary, constants and environment variables; bold only for a literal the reader must type or will see on screen. *(lint)* +5. House intensifiers go. `deliberately`, `honest(ly)`, `genuinely`, `load-bearing`, `the whole point / story / trick`, `measured not argued`, `exactly` (unless before a number). If the sentence loses nothing without the word, the word goes. *(lint)* +6. Say what it is, not what it isn't. `Not X, it is Y` is allowed only when the reader plausibly believes X. Otherwise state Y. *(reader; density warned by lint)* +7. Use the field's word; coinages are replaced or defined once. The retired-words table below and `docs/lint-coinages.tsv` list the project's own coinages and what to write instead. Any term that is neither plain English nor in `docs/glossary.md` is a defect. *(lint)* +8. Introduce every referent in the document that uses it. No `the bench Pi`, `the Dell`, `the capture card`, `the box test`, `today's session`. Say what the device or event is the first time. *(lint for codenames; reader for the rest)* +9. Third person, no confession. No `I` / `we` / `us` / `our`, no `we measured`, `the honest answer`. State the fact and its provenance label. *(lint)* +10. Shipped text carries no session, review or revision talk. `three reviews`, `an earlier draft`, `this originally`, `used to be`, `at review time` belong in the changelog or a private log. Shipped text describes current behaviour only. *(lint for the keyword list; reader for the rest)* +11. Doc-comments describe; design docs argue. A doc-comment gives what the item does, its contract, and at most one sentence of why, with a link. Rationale longer than three sentences, rejected alternatives and incident stories move to `docs/design/`. *(reader; length warned by lint)* +12. Every *this / that / it / here* has its noun in the same or the previous sentence. When in doubt, repeat the noun. *(reader)* +13. One qualifier per claim, and a label instead of adverbs. Use *measured / decided / documented / derived / assumed / not tested*, once. *(reader)* +14. Structure by subject; status in words. Headings name what, never when; no emoji as status; no "(later)" splits. *(lint)* +15. A rule, not an aphorism; a mechanism, not a metaphor. If a sentence could be printed on a poster, replace it with the instruction it stands for. *(reader)* + +Numbers carry their unit and conditions ("29.1 fps at 3840×2160, 30 fps, Raspberry Pi 4"). Measured values say so briefly and link `docs/design/measurements.md` for the full conditions rather than repeating them. + +## Vocabulary + +`docs/glossary.md` is the only list of technical terms the documentation may use without explaining them. Its entries were approved one by one by the project owner. Rules for the file: + +- **Agents propose, never approve.** To add or change an entry, open a pull request that touches `docs/glossary.md`; `.github/CODEOWNERS` routes it to the owner. Do not merge glossary changes yourself, and do not paraphrase an existing definition. +- A user-tier definition must be understandable with no other entry. A developer-tier definition may reference other entries with "(see …)". +- The retired words below are never used in public text, whatever the tier. `scripts/docs-lint.mjs` reports the ones a machine can catch. + +### Names that were decided + +| Use | Not | +|---|---| +| `dexd` — the package, the binary, the service; the player | `dex-loop`, `dex_loop`, "the looper" | +| `dex-sidecar write` / `dex-sidecar check` | `make-sidecar.sh`, `sidecar-check` | +| `dex-exhibit-apply`, `dex-wait-hdmi` | — | +| **exhibit config** — the file `/etc/dex/exhibit.json` or `.yaml` | "the exhibit" on its own | +| **asset** / **video asset** — the video file; the **artwork** is the whole installation it plays in | "the artwork" for the file | +| **forced display mode** in prose; `kms_force` only as the literal config key | "kms force", "KMS forcing" | +| **system log** in prose; `journalctl -u dexd` in commands | "the journal" | +| **dex card** — the SD card that makes a Raspberry Pi a player | `player card` | +| **video container** | "container" alone | +| **loop point**; **gapless** / **seamless** (property); **a held frame**, **a freeze**, **the picture is stuck** (defect) | `seam`, `the wrap`, `hold`, `wrap point` | +| **long-running test**, **24-hour test** | `soak` | +| **test video** (an encoded test file); **test card** (the synthetic picture it is made from) | `bench asset` | +| **test rig** — one test setup; the **bench** — the development workstation and its hardware | "bench" for a setup | +| `--test-rig-no-sidecar`, `--test-rig-hang-after-secs`, `--test-rig-force-recovery-after-secs`; `(test rig only)` | `--bench-*`, `BENCH ONLY`, `wedge` | +| **unresponsive**, **hangs**, **is hanging** | `wedged`, `hung` | +| **supervisor thread** | `event thread` | +| `loops=` in the heartbeat; **loop count** or **loop iterations** in prose | `wraps=`, `wrap count` | +| **check** — the sidecar check, the asset check, the cmdline check | `gate` | +| **prepare the video** (user text and messages); *ingest* only in developer text | `re-ingest` | +| **refuses to start rather than guess** (user text); *fail-closed* only in developer text | `fail-closed in a guide` | +| **frame-duration histogram** | `dwell histogram` | +| **written into**, **stored in**, **saved copy of the EDID**, **build-id file** | `baked`, `baked-in`, `stamped`, `stamp file` | +| **in-place recovery**, **process restart by systemd**, **reboot escalation (planned)** | `tier 0 / 1 / 2 / 3` | +| **the pass criteria** | `the bar`, `the pass bar` | +| dexOS (the brand); `dex-os` (the repository) | `Dexbian` | + + +### Retired words — the full table + +| Retired | Write instead | +|---|---| +| `soak` / `soak test` / `24 h soak` / `soak run` / `soak harness` / `thermal soak` | long-running test / 24-hour test / long-term test (name the duration where it matters); 'the long-running-test harness' | +| `seam` / `the seam` / `seamless-loop as noun` / `'no seam'` / `'a seam'` | place: 'the loop point'; property: 'gapless' or 'seamless'; defect: 'a visible pause / a held frame / a stutter at the loop point' | +| `the wrap` / `wrap point` / `at the wrap` / `wrap-join` / `wrap transition` / `wr` | 'the loop point' (place); 'one loop' / 'one repeat' (the pass); 'loop count' (the counter); 'loop-position arithmetic' (the code) | +| `hold` / `holds` / `hold at the wrap` / `held (as noun)` | 'a freeze' / 'the picture is stuck at the loop point' / 'the frame stays on screen for N ms' — describe the defect plainly (Max: freeze / stuck are mo | +| `bench asset` / `bench-ready asset` / `the card (meaning the encoded video)` | 'test video' / 'the reference test video used for measurements' (a test video made from a test card) | +| `gaplessness premise` / `loop-ability` | 'the requirement that the loop is gapless' / 'whether a file can loop gaplessly' | +| `tier 0` / `tier-0` / `tier 1` / `tier 2` / `tier 3` | 'in-place recovery' (0), 'process restart by systemd' (1), 'reboot escalation (planned)' (2), 'hardware watchdog (planned)' (3) | +| `fail closed (user tier)` / `fail-closed contract` / `fail-silent` | user tier: 'refuses to start rather than guess'; developer tier: 'fail-closed' allowed as a glossary term (? — needs Max) | +| `live-fire` / `live-fire probe` / `live-fire test` | 'against a real mpv instance' / 'on real hardware' / 'the forced-recovery test' | +| `wedged` / `wedge` / `core-wedge` / `display-wedged` / `'the wedge check'` | 'unresponsive' / 'hangs' / 'is hanging' / 'stopped responding while the process stays alive' — never `hung`; the flag becomes --test-rig-hang-after-se | +| `pinned (a behaviour is 'pinned' by a test)` | 'locked in by a test' / 'a test enforces' | +| `the loser` / `delete the loser` | 'the unwanted config file' / 'delete the one you do not mean' | +| `drift generator` | 'would make the boot config and the player's config diverge' | +| `black-wall time` | 'the worst-case time the screen can stay dark' | +| `spins hot` | 'busy-loops, using a full CPU core' | +| `belt-and-braces` | 'a fallback' / 'a second safeguard' | +| `the honest count` / `'honest' as an intensifier` | state the number: 'the binary links 228 shared objects' | +| `green CI` | 'CI passes' / 'a passing CI run' | +| `trap point` | 'the point inside mpv where the wait would unblock' | +| `event-shape` | 'the sequence of events' / 'this event' | +| `the classic monorepo trap` | 'a required check that can silently never run, blocking every merge' | +| `the crux` | 'the central tension: dexOS is buster, the player needs trixie' | +| `the box` / `the box test` / `'shares the box'` | 'the device' / 'the sealed-case thermal test' | +| `the rig` / `capture rig` / `'Bench = …'` | 'the measurement setup (a Pi 4, an HDMI capture device and the analysis scripts)' | +| `the wrong-panel case` | 'a resolution the connected display cannot show' | +| `venue truth, not asset truth` | 'the display mode belongs to the installation, not to the video file' | +| `the mains switch is the shutdown path` | 'there is no graceful shutdown; power is simply cut, and the player is built to survive that' | +| `field journal` / `field failure` / `in the field` / `on site` / `gallery devic` | 'the log' / 'a failure at the venue' / 'at the venue' / 'deployed players' | +| `deploy path` / `bench escape hatch` | 'normal startup (sidecar required)' / 'the test-rig-only override (`--test-rig-no-sidecar --fps`)' | +| `the binding` / `asset+fps binding` / `F3 gate` / `sidecar gate` / `NAL gate` / `` | 'the sidecar's checksum match' / 'the sidecar check' / 'the asset check' / 'the cmdline check' — 'check' in prose; 'gate' allowed as alias (? — needs | +| `THE EXTENSION DECIDES THE PARSER (all caps)` / `BENCH ONLY` / `ARMED (shou` | sentence case: 'the file extension selects the parser'; the literal warning line stays as shipped | +| `escalation ladder` / `'escalate per the fixed ladder'` | 'the pre-committed fallback order (pivid, then GStreamer, then a custom player)' | +| `annulus` / `fps honesty` / `matched wrap` / `'the wrap is matched by constru` | 'ring-shaped region' / 'how far a detected frame rate can be trusted' / plain description | +| `cleanroom extraction` / `cleanroom` | 'rewritten from scratch for publication' | +| `buster ceiling` | 'the buster limitation' / describe: 'gapless hardware playback only on buster (32-bit), so no upgrades and no Pi 5' | +| `the rotation trap` | 'sideways video from phone footage: the container's rotation flag is lost on extraction' (see elementary stream) | +| `(nogit)` / `+dirty as prose` | 'an unidentified build' / 'a build from uncommitted changes' — the literal version-string markers stay | +| `hello_video positive control` / `dexOS card` / `'the dexOS positive contro` | 'the known-good reference (the legacy hello_video player on its own test video)' | +| `mp_dispatch_lock` / `run_locked` / `mp_cond_wait` / `mp_dispatch_queue_proce` | describe the behaviour ('a synchronous property read waits with no timeout for mpv's core thread'); cite the mpv source location in a footnote if prov | +| `Rust identifiers used as prose nouns (HealthMonitor, ObservedCounter,` | in docs: describe the behaviour and name the module once ('the health policy in health.rs'); identifiers belong in code and API docs, not in guides | +| `supervisor thread (health.rs) vs event thread (heartbeat.rs, watchdog.` | 'supervisor thread' everywhere (one thread; DEC-010) | +| `gst1223` / `+rpt2 check` / `'the rpt2 criterion' as bare labels` | 'a GStreamer 1.22 attempt' / 'whether Raspberry Pi's patched ffmpeg build (+rpt2) is required on the Pi 5 — unresolved' | +| `USV` | 'battery backup (`UPS`)' | +| `starved feed` / `'signature of a starved feed'` | 'the data source not keeping up (frames held at random points, not at the loop point)' | +| `the linger bug` | 'the tmux session died with the last SSH login (systemd user session not lingering)' — an operations note for the private record, not dexd | +| `kiosk (flags` / `mode)` / `argv` / `'the working argv'` | 'fullscreen with no on-screen controls' / 'the mpv command line' | +| `baked` / `baked EDID` / `baked-in` / `stamped` / `stamp file` / `build stamp` | 'written into' / 'stored in' / 'saved copy of the EDID' / 'build-id file' — the words `baked` and `stamped` appear nowhere (DEC-016) | +| `hung` | 'hangs' / 'is hanging' / 'unresponsive' — never `hung` (DEC-013) | +| `--bench-no-sidecar` / `--bench-wedge-after-secs` / `--force-recovery-after` | `--test-rig-no-sidecar` / `--test-rig-hang-after-secs` / `--test-rig-force-recovery-after-secs` / `(test rig only)` (DEC-008; code rename) | +| `wraps=` / `WRAP_COUNT` / `wrap count` | loops= (heartbeat field, code rename) / 'loop count' / 'loop iterations' in prose (DEC-006) | +| `event thread` | 'supervisor thread' (DEC-010) | +| `gate (as the noun for a startup refusal)` / `F3 gate` / `cmdline gate` / `NA` | 'check' — the sidecar check, the asset check, the cmdline check (DEC-009; no alias) | +| `fail-closed` / `fail closed (user tier)` | 'refuses to start rather than guess' at user tier; developer tier keeps the glossary entry fail-closed (DEC-011) | +| `kms_force (in prose)` | 'forced display mode' in prose; `kms_force` only as the literal config key | +| `journal` / `the journal (in prose)` | 'system log' in prose; `journalctl` in commands | +| `dwell` / `dwell histogram` / `dwell counts` | 'frame duration' / 'frame-duration histogram' | +| `player card` | 'dex card' (flagged: Max suggested 'dex card or something') | +| `container (alone)` | 'video container' | +| `re-ingest the asset (shipped message)` | 'prepare the video again with dex-sidecar write' (DEC-007; code change) | +| `ingest (user tier)` | 'prepare the video' / 'preparing a video' (DEC-007); developer tier may say ingest | + +### Names of people, places and things + +- No artist names, artwork titles, venues, exhibition names, SD-card ids, hostnames of development machines, or the owner's name in public text. `The project decided` replaces a person's name. +- A forum handle may appear in prose when the person's real name is unknown or the account is pseudonymous — always marked and explained: `*Foo*, a Raspberry Pi engineer on the official forums, …`. Otherwise `a Raspberry Pi engineer on the official forums`, with the link as the citation. +- The measurement instrument may be named once, as the instrument's identity, in `docs/design/measurements.md` (`an Elgato Cam Link 4K HDMI capture device`); one display model may serve as a worked example of a forced display mode. Everywhere else: `the capture device`, `a 2560×1440 monitor`. + +## Provenance labels + +Every measured number, decision and assumption in the documentation traces to a source. In text use one label, once: *measured* (say on what: "measured on a Raspberry Pi 4"), *decided*, *documented* (name the manual or spec), *derived*, *assumed*, *not tested*. The measurement record (`docs/design/measurements.md`) holds the conditions; other documents link it. + +## The mechanical gate + +``` +node scripts/docs-lint.mjs # default paths: packages/dexd, docs, .github/workflows/dexd.yml, AGENTS.md, README.md +node scripts/docs-lint.mjs docs/guides/prepare-video.md +``` + +Errors fail the run; warnings are printed. It reads `docs/glossary.md` (the acronym allow-set), `docs/lint-allow.txt` (per-token or per-path exceptions — every entry needs a reason after `#`, or the tool refuses to start) and `docs/lint-coinages.tsv` (retired words). It runs in CI on `docs/`, `AGENTS.md` and `README.md` files; the crate's comments join the gate when their rewrite lands. Fix an error by rewording; add an allowlist entry only for a true false positive, with the reason. + +## Where things go + +| Content | Place | +|---|---| +| How to build a dex card, prepare a video, configure the exhibit, run and troubleshoot | `docs/guides/` (user tier) | +| Reference: config keys, options, exit codes, every refusal message and its fix | `docs/guides/reference.md` | +| Why the player is built the way it is; how it fails and recovers; packaging and CI; measurements | `docs/design/` (developer tier) | +| The one-page front door | `packages/dexd/README.md` | +| Terms | `docs/glossary.md` | +| What changed between releases | `packages/dexd/deploy/changelog` (Debian format) | +| Rationale moved out of a code comment | the `docs/design/` page the comment links | + +## Commits and code + +- Commit subjects: `E:` for code and packaging, `D:` for documentation, imperative, ≤ 72 characters; the *why* in the body. No AI attribution lines. +- No behaviour change rides along with a wording change. Renames that the vocabulary requires (a flag, a heartbeat field, an identifier) are their own commit with tests updated. +- Test names are prose: `an_fps_that_contradicts_the_stream_is_refused`, not `f6_bad_fps`. diff --git a/CLAUDE.md b/CLAUDE.md new file mode 100644 index 0000000..43c994c --- /dev/null +++ b/CLAUDE.md @@ -0,0 +1 @@ +@AGENTS.md diff --git a/docs/glossary.md b/docs/glossary.md new file mode 100644 index 0000000..c15ec5a --- /dev/null +++ b/docs/glossary.md @@ -0,0 +1,1186 @@ +# Glossary + +Terms used in the dex documentation. Every entry was reviewed and approved by the project owner; changes go through a pull request that CODEOWNERS routes to him. Entries marked **user** are the only technical terms the user guides use without explanation; **developer** entries may appear in the design documents. + +Writers: a term that is not in this list is either plain English or must be explained in the sentence that uses it. See `AGENTS.md` for the writing rules and the list of retired words. + +## User-tier terms + +### .265 file + +A video file that holds only the compressed HEVC pictures, with no wrapper such as MP4 around them. dexd plays this format and nothing else, so every video is converted to a .265 file before it goes on the player. + +Also written: raw HEVC stream · raw Annex-B file · .h265 · elementary-stream file + +Tier: user + +### .deb package + +The installable package file for Debian-based systems such as Raspberry Pi OS. dexd is delivered as one .deb, installed with apt, which also pulls in the mpv library it needs and sets up the service that starts at boot. + +Also written: Debian package · the .deb · dexd__arm64.deb + +Tier: user + +### asset + +The one video file the player loops — the video asset — named by the `asset` line of the exhibit config as an absolute path (for example /opt/dex/artwork.265). If no asset is named, dexd refuses to start rather than guess. The artwork is the whole installation the asset plays in. + +Also written: video asset · the video · `asset` (config key) + +Tier: user + +### checksum + +A short code computed from every byte of a file; change one byte and the code changes. The sidecar stores the video's checksum, so dexd can tell a stale, wrong or half-copied file from the prepared one, and refuses it. + +Also written: SHA-256 · sha256 · hash · fingerprint + +Tier: user + +### cmdline.txt + +The one-line file on the Raspberry Pi's boot partition that holds the start-up options for the operating system, including a forced display mode. dex-exhibit-apply edits it from the exhibit config; do not hand-edit it, and reboot after it changes. + +Also written: /boot/firmware/cmdline.txt · kernel command line · boot options file + +Tier: user + +### connector + +The name the operating system gives each physical video output; on a Raspberry Pi 4 the two HDMI ports are HDMI-A-1 and HDMI-A-2. The exhibit config names the connector the display is plugged into (default HDMI-A-1), and every display check uses it. + +Also written: `connector` (config key) · HDMI-A-1 · HDMI-A-2 · HDMI port + +Tier: user + +### dex card + +An SD card holding Raspberry Pi OS, dexd, the exhibit config and the video, which turns a Raspberry Pi into a player the moment it boots. Building one is the setup task; a spare card is the fastest repair at a venue. + +Also written: player card · card · SD card · exhibition card + +Tier: user + +### dex-exhibit-apply + +A helper command installed with dexd, run with sudo after editing the exhibit config. It writes the forced display mode from the config into cmdline.txt, changing nothing else, and prints REBOOT REQUIRED only when the file actually changed. + +Also written: exhibit-apply · `sudo dex-exhibit-apply` + +Tier: user + +### dex-sidecar + +The command that writes a video's sidecar (`dex-sidecar write`) and checks an existing one against its video (`dex-sidecar check`), run on a workstation before copying to the player. It uses dexd's own reader, so what passes here plays there. + +Also written: dex-sidecar write · dex-sidecar check · sidecar-check (old name) · make-sidecar.sh (old script) + +Tier: user + +### dex-wait-hdmi + +A helper that runs before dexd and waits, up to two minutes, for a display to report it is connected. Projectors can wake slower than the Raspberry Pi boots; without the wait the system picks a fallback resolution and never corrects it. + +Also written: DEX_HDMI_TIMEOUT (its timeout setting) + +Tier: user + +### dexd + +The player program: it plays one .265 video on a Raspberry Pi in an endless gapless loop, starts at boot as a system service, checks its own health and restarts itself. Package, command and service share the name; helpers keep the dex- prefix. + +Also written: dex-loop (name during development) · dexd.service · the player + +Tier: user + +### display mode + +The picture size and refresh rate the player asks the display for, written as WIDTHxHEIGHT@RATE (for example 3840x2160@30) or `auto`. Set it in the exhibit config to match the display; a mode the display cannot show makes dexd refuse to start. + +Also written: `display_mode` (config key) · WxH@R · resolution and refresh · auto + +Tier: user + +### EDID + +The information a display sends over the HDMI cable describing itself and the modes it can show. dexd and the Raspberry Pi rely on it to choose a mode; some displays send it late or wrongly, which is why kms_force and dex-wait-hdmi exist. + +Also written: display identification · the display's self-description · edid-decode (tool that prints it) + +Tier: user + +### exhibit config + +The one file on each player, /etc/dex/exhibit.json or exhibit.yaml, that says which video plays (`asset`), which display mode to use and which connector. It holds what belongs to the installation, not to the video file; dexd reads it at every start and the command line may cross-check it but never override it. + +Also written: /etc/dex/exhibit.json · /etc/dex/exhibit.yaml · exhibit.json · exhibit.yaml · exhibit file · exhibit · the installation · venue setup · `venue` / `display` / `note` (informational config keys) + +Tier: user + +### exit codes + +The number dexd returns when it stops. 2 means it refused to start because something in the setup is wrong (fix the config or files; a restart will not help); 1 means playback failed while running, and the service manager restarts it automatically. + +Also written: exit-code contract · exit 1 · exit 2 + +Tier: user + +### ffmpeg + +A command-line tool that converts video between formats. The prepare-video guide uses it to encode HEVC with the settings dexd needs and to extract the .265 file; ffprobe, from the same toolkit, reports a file's size, frame rate and codec. + +Also written: ffprobe (its inspection tool) · libx265 / x265 (its HEVC encoder) · hevc_mp4toannexb (its MP4-to-.265 filter) + +Tier: user + +### forced display mode + +An exhibit-config setting (`kms_force`) that makes the Raspberry Pi output a fixed mode from boot (for example 3840x2160@30) instead of trusting what the display announces, or `none`. Needed for displays that announce 4K but never get it unforced; dex-exhibit-apply writes it into cmdline.txt. + +Also written: `kms_force` (config key) · WxH@R / WxH@RD + +Tier: user + +### frame rate + +How many pictures per second a video shows, for example 30 or 29.97 (written 30000/1001). A .265 file does not record it, so dexd takes it from the sidecar and refuses to start without one rather than play at the wrong speed. + +Also written: fps · frames per second · `fps` (sidecar field) · `--fps` + +Tier: user + +### gapless + +Playback that repeats with no visible break: no black frame, no held frame, no stutter between the last picture and the first. It is the property dexd exists to deliver, and what every measurement in the design record checks. + +Also written: seamless · seamless loop · perfect loop · loops seamlessly + +Tier: user + +### hardware decoding + +Turning compressed video back into pictures using a dedicated block in the chip instead of the main processor. A Raspberry Pi 4 can play 4K HEVC smoothly only this way, so dexd is built around it; software decoding is the slow fallback. + +Also written: hardware decode · HW decode · hardware-accelerated video · hwdec (mpv's option for it) + +Tier: user + +### heartbeat + +A status line dexd writes to the system log at start and every ten minutes: loop count (`loops=`), uptime, chip temperature, dropped and late frames, playback-position age, watchdog state. While it keeps coming the player is alive. It reports; the health check repairs; the watchdog restarts. + +Also written: heartbeat line · `loops=` (was `wraps=`) · `temp=` · `frame-drops=` · `vo-delayed=` · `pos=` · `pos-age=` · `watchdog=` + +Tier: user + +### HEVC + +The video compression format dexd plays, also called H.265. The Raspberry Pi 4 has a hardware decoder for it and for nothing newer, so 4K playback depends on the video being HEVC; other formats must be re-encoded first. + +Also written: H.265 · High Efficiency Video Coding + +Tier: user + +### keyframe + +A picture in a compressed video that is complete on its own, not described as changes from earlier pictures. A video for dexd must begin with one and must not let later pictures refer back across the start, or the loop cannot restart cleanly. + +Also written: intra frame · I-frame · IDR (the exact HEVC term, developer glossary) + +Tier: user + +### loop point + +The moment playback returns from the video's last picture to its first. Everything about a gapless loop is decided here: a pause, a held frame or a flash at the loop point is the defect dexd is designed to avoid and its measurements look for. + +Also written: restart of the loop · wrap point (retired wording) · the wrap (retired wording) · seam (retired wording) + +Tier: user + +### mpv + +The open-source media player whose engine dexd uses to decode and show video. dexd does not run the mpv program; it embeds mpv's library and feeds it the video, which is why installing dexd also installs the mpv library package (libmpv2). + +Also written: mpv 0.40 (the verified version) + +Tier: user + +### one-file rule + +Only one exhibit config may exist on a player: exhibit.json or exhibit.yaml, never both. If both are present dexd refuses to start and names both, so a venue never runs yesterday's settings from the file nobody edited. + +Also written: exactly one exhibit config · the extension decides the parser + +Tier: user + +### Raspberry Pi Imager + +The official program that writes Raspberry Pi OS onto an SD card and lets you set the hostname, user and SSH key before first boot. The player-card guide starts with it and sets the SSH key here. + +Also written: Imager + +Tier: user + +### sidecar + +A small text file next to the video, named like the video plus `.json`, holding its frame rate and checksum. dexd refuses to start unless it is present and matches, so a wrong frame rate or a half-copied video is caught before anything shows. + +Also written: