# JellyfinSyncPlus — Plan ## Problem Jellyfin's built-in SyncPlay drifts noticeably when watching with remote friends (tested with 2). Symptom: one participant is consistently ahead or behind the others, not just momentarily but persistently. ## Root cause Confirmed against known Jellyfin issues (jellyfin-web#6210, jellyfin#5557, jellyfin#7579): when group members end up on different playback paths (one Direct Plays, another transcodes -- due to codec/bitrate/device differences), the transcoding client's decode/render pipeline runs measurably behind, and pause/resume compounds the gap. SyncPlay's own resync mechanism (periodic hard seeks) is also fragile and doesn't fully compensate. Immediate mitigation (no code, do this first / always): check Dashboard -> Active Sessions during a session and confirm all clients show Direct Play, not Transcode. Force "Maximum/Original" quality on every client. Fully leave/rejoin the SyncPlay group rather than trusting in-place resync. ## Why build something custom Jellyfin's own SyncPlay is the only real option today -- nothing has superseded it as of mid-2026. `jellyfin-mpv-shim` is actively maintained (now under the official `jellyfin` GitHub org) but rides the same underlying SyncPlay protocol; it only helps indirectly by avoiding transcode on capable devices. Standalone `syncplay.pl` + Jellyfin direct-stream URLs is undocumented/unproven for this use case. So: build a plugin that replaces SyncPlay's weak points rather than switching tools. ## Architecture -- three pillars ### 1. Group quality negotiation (the easy piece) On group formation, resolve each member's device capability/bitrate profile (Jellyfin already does this per-client via its existing `GetPlaybackInfo` / `DeviceProfile` negotiation -- reuse that logic, don't reinvent it) and pick the most restrictive result across the whole group. Every member streams at that one common profile. ### 2. Shared transcode (the hard, novel piece) Problem with (1) alone: even at matched quality, N independently-running ffmpeg transcode jobs (one per client, Jellyfin's default) have their own segment boundaries and start offsets -- literally different encodes, which is its own source of drift. Jellyfin's `ITranscodeManager` ties one transcode job to one `PlaySessionId` by design; there's no built-in dedup for identical (item, profile) requests. Patching that is real core surgery -- **not** the v1 approach. **Pragmatic v1: host-relay.** One group member is the "host" -- their session transcodes/direct-plays completely normally, untouched. The plugin exposes a relay API route that proxies the host's already-generated HLS playlist/segments to the other group members, so everyone reads byte-identical media without touching Jellyfin's transcode internals at all. Gets ~90% of the benefit of true shared-session transcoding for a fraction of the engineering risk. ### 3. Active drift correction (still needed even with #1 + #2) Identical bytes doesn't mean identical render timing -- per-client network fetch latency and decode speed still differ. Needed on top: - **Virtual playhead as the single authority.** Server tracks `{position, wall-clock timestamp of that position, playing/paused, speed}` and computes `expected_position = last_position + (now - last_command_time)`. All clients (including the host) sync against this abstraction, never against each other -- avoids the host always looking "correct" to itself. - **NTP-style clock calibration per client on join** -- ping/pong to estimate RTT and clock offset. Needed both for scheduling (below) and for correctly interpreting heartbeat reports. - **Heartbeats + EMA-smoothed drift.** Each client reports its real player-engine position (not a fetch/buffer estimate) every 1-2s. Server diffs against the virtual playhead (RTT/2-adjusted) and smooths with an EMA filter rather than reacting to every raw sample. - **Two-tier correction:** - Small drift (~150-300ms): nudge local playback rate briefly (1.0 -> ~1.02-1.05x or 0.97x) until closed -- imperceptible, no visible jump. (Same trick `syncplay.pl` uses for local-file sync.) - Large drift (~1.5-2s+, typically post-stall): hard-seek that client back onto the playhead instead -- speed-nudging alone would take too long and be noticeable. - **Scheduled commands, not "do it now."** Play/pause/seek carry an absolute UTC execution timestamp; each client fires locally at exactly `T` using its calibrated offset, so message-delivery latency itself isn't a desync source. - **Late joiners / reconnects** both reduce to: query current virtual playhead, buffer ahead until healthy, join via a scheduled seek to the interpolated position. ## Scope confirmed with user - Disconnects, late joiners, seek/pause propagation: in scope for v1. - Quality-profile picker menu (AVC-only override etc.): explicitly out of scope for the *bot* project (MR-Discord) but not discussed here -- this project only forces a common profile for sync purposes, doesn't add a UI picker. ## Target environment (real, not hypothetical) - Jellyfin runs in the `entertainment` namespace of the user's k8s cluster, image `linuxserver/jellyfin:10.11.6` (pinned -- build the plugin against this exact server version's ABI). - Config volume: hostPath PVC at `/mnt/redundant/k8s/jellyfin/config` (`jellyfin-config-pvc`, `jellyfin-config` StorageClass). - Transcoding scratch: hostPath PVC at `/mnt/nvme/k8s/jellyfin/transcoding` (`jellyfin-transcoding-pvc`, currently on the `local-retain` StorageClass -- scheduled for consolidation into `nvme-retain` in the separate, paused k8s/entertainment hardening plan, not part of this project). See [entertainment/base/jellyfin/deployment.yaml](../entertainment/base/jellyfin/deployment.yaml) and [entertainment/base/jellyfin/config-claim.yaml](../entertainment/base/jellyfin/config-claim.yaml). - Deployed via ArgoCD, app name `entertainment`. **Does not auto-sync** -- every deploy needs a manual `kubectl patch application entertainment -n argocd --type merge -p '{"operation":{"sync":{"revision":"HEAD"}}}'` (same pattern used for MR-Discord releases). - Sibling project [MR-Discord](../MR-Discord) already has a working tag-push -> Gitea Actions -> build -> deploy pipeline against this same cluster; reuse that pattern for this project's releases once it's stable. ## Dev / test / iterate workflow Two different loops -- do not conflate them: 1. **Fast dev loop (default, use this for almost everything):** run a throwaway local Jellyfin container (`docker run linuxserver/jellyfin:10.11.6`, same tag as prod -- avoids ABI surprises) with the plugins folder bind-mounted. Loop: `dotnet build` -> copy DLL into the mounted folder -> restart the local container -> check logs -> test. Jellyfin has no plugin hot-reload, so a restart is unavoidable per change, but a local container restart is seconds, not a k8s rollout. 2. **Milestone testing against the real pod (deliberate, ask first):** copy the built plugin into the real hostPath config dir on the node, then `kubectl rollout restart deployment/jellyfin -n entertainment`. This briefly interrupts the live, shared Jellyfin instance -- **do not do this routinely; only for actual multi-friend sync test sessions, and confirm with the user first.** 3. **Release distribution (later, once stable):** package as a versioned zip + `meta.json`, host a plugin-repository manifest via Gitea (raw file URLs or release attachments both work, Jellyfin just needs HTTP(S) access to the manifest + zip), add the repo URL in Jellyfin's dashboard. Mirrors MR-Discord's tag -> Gitea Actions -> build pipeline, just producing a plugin zip instead of a Docker image. Slow iteration loop -- not for development, only for shipping tested versions. ## Known friction / risks (flagged going in, not discovered later) - **ABI/version pinning**: `meta.json`'s `targetAbi` must match the running server version. Already fine since the deployment is pinned to `10.11.6`, not floating -- build against that exact version's SDK. - **Reaches past the stable plugin API.** SyncPlay group internals and transcode-session state aren't part of Jellyfin's published plugin SDK surface. Realistically requires building against Jellyfin server source at the matching tag (`git` checkout/submodule of `jellyfin/jellyfin` @ `v10.11.6`) rather than just NuGet packages. This is closer to "semi-fork with a plugin wrapper" than a typical plugin -- sizing decisions should account for that. - Host-relay approach (pillar 2) means the "host" member is a single point of failure for the group's stream -- needs a defined handoff/reselection behavior if the host disconnects mid-session (not just late-joiners). ## Rough phase breakdown - **Phase 0 -- scaffold & dev loop.** `dotnet` plugin project scaffold against `Jellyfin.Controller` (or a source checkout of `jellyfin/jellyfin@v10.11.6` if the public SDK isn't sufficient), local docker dev Jellyfin with bind-mounted plugins folder, confirm a trivial "hello world" plugin loads and shows up in Dashboard -> Plugins. - **Phase 1 -- group quality negotiation.** Hook group-join, resolve each member's device profile, compute the common lowest profile, verify against real device/browser combos. - **Phase 2 -- host-relay shared transcode.** Custom API route serving the host session's HLS playlist/segments to other members; handle host disconnect/reselection. - **Phase 3 -- active drift correction.** Virtual playhead, clock calibration, heartbeat + EMA, two-tier nudge/reseek correction, scheduled absolute-time commands. - **Phase 4 -- disconnects, late joiners, seek/pause polish.** - **Phase 5 -- packaging & release pipeline.** Gitea-hosted plugin repository manifest, versioned zip releases, mirroring MR-Discord's pipeline. ## Open questions (revisit before/at the relevant phase, not now) - Exact client compatibility target for v1 -- web client only, or also `jellyfin-mpv-shim`/mobile apps? (Web client is simplest to control fully; others may need companion-side changes.) - Host reselection policy on disconnect -- pick another member automatically, or pause the group and prompt?