10 KiB
JellyfinSyncPlus — Plan
Problem
Jellyfin's built-in SyncPlay drifts noticeably when watching with remote friends (tested with 2). Symptom: one participant is consistently ahead or behind the others, not just momentarily but persistently.
Root cause
Confirmed against known Jellyfin issues (jellyfin-web#6210, jellyfin#5557, jellyfin#7579): when group members end up on different playback paths (one Direct Plays, another transcodes -- due to codec/bitrate/device differences), the transcoding client's decode/render pipeline runs measurably behind, and pause/resume compounds the gap. SyncPlay's own resync mechanism (periodic hard seeks) is also fragile and doesn't fully compensate.
Immediate mitigation (no code, do this first / always): check Dashboard -> Active Sessions during a session and confirm all clients show Direct Play, not Transcode. Force "Maximum/Original" quality on every client. Fully leave/rejoin the SyncPlay group rather than trusting in-place resync.
Why build something custom
Jellyfin's own SyncPlay is the only real option today -- nothing has superseded
it as of mid-2026. jellyfin-mpv-shim is actively maintained (now under the
official jellyfin GitHub org) but rides the same underlying SyncPlay protocol;
it only helps indirectly by avoiding transcode on capable devices. Standalone
syncplay.pl + Jellyfin direct-stream URLs is undocumented/unproven for this
use case. So: build a plugin that replaces SyncPlay's weak points rather than
switching tools.
Architecture -- three pillars
1. Group quality negotiation (the easy piece)
On group formation, resolve each member's device capability/bitrate profile
(Jellyfin already does this per-client via its existing GetPlaybackInfo /
DeviceProfile negotiation -- reuse that logic, don't reinvent it) and pick the
most restrictive result across the whole group. Every member streams at that
one common profile.
2. Shared transcode (the hard, novel piece)
Problem with (1) alone: even at matched quality, N independently-running ffmpeg transcode jobs (one per client, Jellyfin's default) have their own segment boundaries and start offsets -- literally different encodes, which is its own source of drift.
Jellyfin's ITranscodeManager ties one transcode job to one PlaySessionId by
design; there's no built-in dedup for identical (item, profile) requests.
Patching that is real core surgery -- not the v1 approach.
Pragmatic v1: host-relay. One group member is the "host" -- their session transcodes/direct-plays completely normally, untouched. The plugin exposes a relay API route that proxies the host's already-generated HLS playlist/segments to the other group members, so everyone reads byte-identical media without touching Jellyfin's transcode internals at all. Gets ~90% of the benefit of true shared-session transcoding for a fraction of the engineering risk.
3. Active drift correction (still needed even with #1 + #2)
Identical bytes doesn't mean identical render timing -- per-client network fetch latency and decode speed still differ. Needed on top:
- Virtual playhead as the single authority. Server tracks
{position, wall-clock timestamp of that position, playing/paused, speed}and computesexpected_position = last_position + (now - last_command_time). All clients (including the host) sync against this abstraction, never against each other -- avoids the host always looking "correct" to itself. - NTP-style clock calibration per client on join -- ping/pong to estimate RTT and clock offset. Needed both for scheduling (below) and for correctly interpreting heartbeat reports.
- Heartbeats + EMA-smoothed drift. Each client reports its real player-engine position (not a fetch/buffer estimate) every 1-2s. Server diffs against the virtual playhead (RTT/2-adjusted) and smooths with an EMA filter rather than reacting to every raw sample.
- Two-tier correction:
- Small drift (~150-300ms): nudge local playback rate briefly
(1.0 -> ~1.02-1.05x or 0.97x) until closed -- imperceptible, no visible
jump. (Same trick
syncplay.pluses for local-file sync.) - Large drift (~1.5-2s+, typically post-stall): hard-seek that client back onto the playhead instead -- speed-nudging alone would take too long and be noticeable.
- Small drift (~150-300ms): nudge local playback rate briefly
(1.0 -> ~1.02-1.05x or 0.97x) until closed -- imperceptible, no visible
jump. (Same trick
- Scheduled commands, not "do it now." Play/pause/seek carry an absolute
UTC execution timestamp; each client fires locally at exactly
Tusing its calibrated offset, so message-delivery latency itself isn't a desync source. - Late joiners / reconnects both reduce to: query current virtual playhead, buffer ahead until healthy, join via a scheduled seek to the interpolated position.
Scope confirmed with user
- Disconnects, late joiners, seek/pause propagation: in scope for v1.
- Quality-profile picker menu (AVC-only override etc.): explicitly out of scope for the bot project (MR-Discord) but not discussed here -- this project only forces a common profile for sync purposes, doesn't add a UI picker.
Target environment (real, not hypothetical)
- Jellyfin runs in the
entertainmentnamespace of the user's k8s cluster, imagelinuxserver/jellyfin:10.11.6(pinned -- build the plugin against this exact server version's ABI). - Config volume: hostPath PVC at
/mnt/redundant/k8s/jellyfin/config(jellyfin-config-pvc,jellyfin-configStorageClass). - Transcoding scratch: hostPath PVC at
/mnt/nvme/k8s/jellyfin/transcoding(jellyfin-transcoding-pvc, currently on thelocal-retainStorageClass -- scheduled for consolidation intonvme-retainin the separate, paused k8s/entertainment hardening plan, not part of this project). See entertainment/base/jellyfin/deployment.yaml and entertainment/base/jellyfin/config-claim.yaml. - Deployed via ArgoCD, app name
entertainment. Does not auto-sync -- every deploy needs a manualkubectl patch application entertainment -n argocd --type merge -p '{"operation":{"sync":{"revision":"HEAD"}}}'(same pattern used for MR-Discord releases). - Sibling project MR-Discord already has a working tag-push -> Gitea Actions -> build -> deploy pipeline against this same cluster; reuse that pattern for this project's releases once it's stable.
Dev / test / iterate workflow
Two different loops -- do not conflate them:
- Fast dev loop (default, use this for almost everything): run a
throwaway local Jellyfin container (
docker run linuxserver/jellyfin:10.11.6, same tag as prod -- avoids ABI surprises) with the plugins folder bind-mounted. Loop:dotnet build-> copy DLL into the mounted folder -> restart the local container -> check logs -> test. Jellyfin has no plugin hot-reload, so a restart is unavoidable per change, but a local container restart is seconds, not a k8s rollout. - Milestone testing against the real pod (deliberate, ask first): copy
the built plugin into the real hostPath config dir on the node, then
kubectl rollout restart deployment/jellyfin -n entertainment. This briefly interrupts the live, shared Jellyfin instance -- do not do this routinely; only for actual multi-friend sync test sessions, and confirm with the user first. - Release distribution (later, once stable): package as a versioned zip
meta.json, host a plugin-repository manifest via Gitea (raw file URLs or release attachments both work, Jellyfin just needs HTTP(S) access to the manifest + zip), add the repo URL in Jellyfin's dashboard. Mirrors MR-Discord's tag -> Gitea Actions -> build pipeline, just producing a plugin zip instead of a Docker image. Slow iteration loop -- not for development, only for shipping tested versions.
Known friction / risks (flagged going in, not discovered later)
- ABI/version pinning:
meta.json'stargetAbimust match the running server version. Already fine since the deployment is pinned to10.11.6, not floating -- build against that exact version's SDK. - Reaches past the stable plugin API. SyncPlay group internals and
transcode-session state aren't part of Jellyfin's published plugin SDK
surface. Realistically requires building against Jellyfin server source at
the matching tag (
gitcheckout/submodule ofjellyfin/jellyfin@v10.11.6) rather than just NuGet packages. This is closer to "semi-fork with a plugin wrapper" than a typical plugin -- sizing decisions should account for that. - Host-relay approach (pillar 2) means the "host" member is a single point of failure for the group's stream -- needs a defined handoff/reselection behavior if the host disconnects mid-session (not just late-joiners).
Rough phase breakdown
- Phase 0 -- scaffold & dev loop.
dotnetplugin project scaffold againstJellyfin.Controller(or a source checkout ofjellyfin/jellyfin@v10.11.6if the public SDK isn't sufficient), local docker dev Jellyfin with bind-mounted plugins folder, confirm a trivial "hello world" plugin loads and shows up in Dashboard -> Plugins. - Phase 1 -- group quality negotiation. Hook group-join, resolve each member's device profile, compute the common lowest profile, verify against real device/browser combos.
- Phase 2 -- host-relay shared transcode. Custom API route serving the host session's HLS playlist/segments to other members; handle host disconnect/reselection.
- Phase 3 -- active drift correction. Virtual playhead, clock calibration, heartbeat + EMA, two-tier nudge/reseek correction, scheduled absolute-time commands.
- Phase 4 -- disconnects, late joiners, seek/pause polish.
- Phase 5 -- packaging & release pipeline. Gitea-hosted plugin repository manifest, versioned zip releases, mirroring MR-Discord's pipeline.
Open questions (revisit before/at the relevant phase, not now)
- Exact client compatibility target for v1 -- web client only, or also
jellyfin-mpv-shim/mobile apps? (Web client is simplest to control fully; others may need companion-side changes.) - Host reselection policy on disconnect -- pick another member automatically, or pause the group and prompt?