Skip to content

feat: recoverable daemon, drm-state lifecycle, and native libhybris/CGO build - #72

Open
silentone12725 wants to merge 4 commits into
WorldObservationLog:mainfrom
silentone12725:main
Open

silentone12725 wants to merge 4 commits into
WorldObservationLog:mainfrom
silentone12725:main

Conversation

@silentone12725

Copy link
Copy Markdown

Summary

This PR turns the wrapper from a process that exits on Apple playback lease errors into a long-running, self-recovering daemon. It also adds lifecycle reporting, a drift check between the two wrapper implementations, and a host-native build that can run without a container.

It contains four commits:

  1. daemon: replace exit(1) callbacks with recoverable state machine
  2. Add DRM state tracking and wrapper drift check
  3. Add host-native DRM wrapper build via libhybris, in-process CGO bridge and deploy script
  4. native: locate libCoreLSKD via /proc/self/maps for the R1 hook

1. Recoverable daemon

Both SVPlaybackLeaseManager callbacks (endLeaseCb, pbErrCb) used to call exit(1), so any lease termination killed the wrapper and needed an external supervisor to restart it.

State machine. main.cpp tracks Running, Scheduled, Refreshing and Failed. get_recovery_state() is exported as extern "C" (0=Running, 1=Scheduled, 2=Refreshing, 3=Failed), and is_recovery_active() is a simple gate for request handlers.

Non-blocking callbacks. The callbacks only enqueue a recovery event and return. No FairPlay or network call runs in the lease-manager callback context, which avoids reentrancy and deadlock risk.

Recovery worker. A persistent thread owns context reacquisition. It coalesces bursts of queued events into one refresh, calls refresh_decrypt_ctx(), and checks the result with is_preshare_ctx_ready(). The failure counter resets only once a usable preshareCtx exists. After a failed refresh the worker schedules its own retry, so recovery continues without another Apple event. Backoff is 1s → 2s → 5s → 10s → 30s, clamped at 30s rather than giving up.

Thread safety. preshareCtx is protected by a mutex shared by the decrypt path and the recovery worker. The lock is released before long-running FairPlay or network work.

Request gating. While recovery is active, the decrypt endpoint returns EOF and the M3U8 endpoint returns an empty response. The decrypt, M3U8 and account servers stay up throughout. External supervision is then only needed for real process failures (crash, OOM).

lease error → endLeaseCb / pbErrCb → enqueue event → return
                                          │
                                   recovery worker
                          coalesce → refresh ctx → ready? ─ yes → RUNNING
                                                      └─ no → backoff → retry

2. Lifecycle state and drift check

State file. The wrapper writes its lifecycle to <base-dir>/drm-state, so a supervisor or UI can observe it without inferring state from process existence. States are STARTING, LOGIN, WAITING_2FA, INITIALIZING_FAIRPLAY, RUNNING, RECOVERY, FAILED and STOPPED.

Rootless wrapper. wrapper-rootless.c follows the same recovery model instead of exiting on lease termination.

Drift check. scripts/check-drift.py compares the privileged and rootless wrapper sources so they don't diverge as recovery and lifecycle behaviour evolves. It runs in the x86_64 GitHub Actions workflow. Dockerfile.build provides a reproducible build environment.


3. Host-native build and in-process library

This adds a third execution mode alongside Docker and rootless: loading the Android libraries into a native glibc process through libhybris. It avoids container or proot overhead.

Binaries. build-native.sh builds drm-native and libdrm-native.so with host gcc/g++. hybris_stubs.c, hybris_ctor.c, hybris_types.h and import.h hold the host-side glue. A Bionic std::function ABI mismatch is fixed (__f_-first layout).

Deploy script. build-and-deploy.sh builds and copies drm-native, libdrm-native.so, libhybris-core.so and hybris-linker/q.so into a target directory and sets the rpath to $ORIGIN. It is configured entirely through environment variables:

Variable Purpose
DEPLOY_DIR Target directory (required)
DEPLOY_DIR_EXTRA Optional second target
HYBRIS_INC libhybris include directory (required)
HYBRIS_LIB, LINKER_SO Override paths to libhybris-core.so and q.so
NATIVE_BIN, NATIVE_SO Override paths to the built artefacts
DOBBY_SRC, DOBBY_BUILD Dobby source and build directories

In-process library (drm_lib). drm_lib.h/drm_lib.c expose drm_lib_init, drm_lib_decrypt and drm_lib_shutdown, with auth (2FA) and state callbacks. A Go (CGO) or C/C++ host can link the DRM engine directly instead of using socket IPC. Library mode builds with -DDRM_LIB_BUILD, which compiles out main(). Recovery state is available through get_recovery_state() and drm_lib_is_recovery_active().

Fixes included:

  • GUID is copied into a null-terminated buffer (Data::bytes() is not null-terminated).
  • get_music_user_token and get_dev_token use heap buffers, and the request body is no longer freed before run().
  • Android library load failures return an error instead of calling exit().
  • drm_lib_decrypt uses the key context, not its slot.
  • r1_capture_cb has the current Dobby callback signature, (void *address, DobbyRegisterContext *ctx).
  • FHinstance, preshareCtx and getKdContext are non-static again, since drm_lib.c links against them.

R1 hook in native mode (commit 4). libCoreLSKD.so is now located through /proc/self/maps instead of dl_iterate_phdr. In native mode libhybris' own linker loads the Android libraries, so glibc's dl_iterate_phdr never sees them: the lookup failed with libCoreLSKD not loaded and the key service returned only contentKey. With the /proc/self/maps lookup the hook installs and the key service returns the full template (ctx, state, rcx/rax/rdx/r9/rbp). The now-unused _find_lib_cb is removed.


Behaviour changes to review

  • --mv-port now defaults to 50020 (the itun decrypt listener is mv-port + 10000 = 60020). It previously shared 40020 with --key-port. Every listener sets SO_REUSEPORT, so both binds could succeed and traffic could reach either service. Explicit --mv-port is unaffected. Scripts relying on the old default need updating. Docker users need -p 50020:50020 -p 60020:60020; the README example is updated.
  • build-native.sh links Dobby. The R1 key-server hook is no longer gated by MyRelease, so native builds need dobby.h and a built libdobby.a.

Build prerequisites (native mode)

  • Dobby, tested at commit e9fe7fb. On newer GCC, configure with -DCMAKE_C_FLAGS="-include sys/time.h". At that commit, Logger::Shared() in external/logging/logging/logging.h needs inline, or linking fails with multiple definitions. This is a Dobby-side patch, not part of this PR.
  • A built libhybris-core.so and the linker plugin q.so, plus the libhybris headers (HYBRIS_INC).
  • At runtime: an Android rootfs (system/lib64) and a base directory, passed through HYBRIS_LD_LIBRARY_PATH, HYBRIS_ANDROID_LIB64 and --base-dir.

Rebase notes

  • Rebased onto current origin/main; conflicts were resolved in cmdline.*, main.c, wrapper.ggo and README.md.
  • --key-port/-K (from origin/main) and --mv-port/-G (from this PR) are both kept.
  • cmdline.c/cmdline.h are hand-merged in gengetopt 2.23 style. Regenerating with 2.23.1 differs only in version-related output.
  • The follow-up commits to the host-native build were squashed into commit 3.

Testing

  • drm-native and libdrm-native.so build against Dobby and libhybris, deploy with DEPLOY_DIR, and resolve libhybris-core.so through the $ORIGIN runpath.
  • drm-native --help and --version run (wrapper 1.2.0); -K defaults to 40020 and -G to 50020.
  • libdrm-native.so exports drm_lib_init, drm_lib_decrypt, drm_lib_shutdown, get_recovery_state and drm_lib_is_recovery_active.
  • scripts/check-drift.py passes on this revision.
  • Daemon starts against a real rootfs and logged-in data dir, reaches RUNNING in drm-state, starts the recovery thread, and all six ports (10020, 20020, 30020, 40020, 50020, 60020) listen.
  • Account endpoint (30020) returns the storefront and tokens.
  • Key service (40020) returns the full R1 template on a real track in native mode.
  • Decrypt (10020), M3U8 (20020) and MV (50020/60020) servers with real playback. (not yet run)
  • Recovery path exercised with a forced lease error: RECOVERY → RUNNING, process stays alive, and gated requests return EOF/empty responses. (not yet run)

CI for this fork PR is waiting for maintainer approval ("Approve and run").

Docs

The README covers the three execution modes, the CGO API, the recovery state machine, both port options and the extra Docker ports. The Protocol column for the 50020 and 60020 rows is a placeholder.

Previously, both SVPlaybackLeaseManager callbacks (endLeaseCb, pbErrCb)
called exit(1) unconditionally, making every Apple lease termination a
fatal process death. wrapper-rootless was behaving like a short-lived
utility rather than a persistent background service.

This commit introduces a proper recovery lifecycle:

Recovery state machine (main.cpp)
- RecoveryState enum: Running / Scheduled / Refreshing / Failed
- get_recovery_state() exported as extern C int for status endpoints
  and Electron IPC (0=Running, 1=Scheduled, 2=Refreshing, 3=Failed)
- is_recovery_active() derived from state, used to gate HTTP requests

Non-blocking callbacks
- endLeaseCb / pbErrCb now push a code onto a queue and return
  immediately — no library calls from within the lease manager thread
- Eliminates reentrancy and deadlock risk from callback context

Dedicated recovery worker thread
- Drains the entire queue on each wake (coalescing): a burst of
  3084+3084+PLAYBACK_ERR produces one refresh cycle, not three
- Exponential backoff: 1s -> 2s -> 5s -> 10s -> 30s (clamped, never stops)
- Calls refresh_decrypt_ctx() as the sole owner of reacquisition logic
- Verifies success via is_preshare_ctx_ready() — resets consec_fails
  only when preshareCtx is non-null after the refresh
- Self-schedules retry on failure (kRetryInternal) so the daemon keeps
  retrying even when Apple sends no further lease-end event (e.g.
  transient network failure during refresh)

Thread safety (main.c)
- g_ctx_mutex (PTHREAD_MUTEX_INITIALIZER) protects all preshareCtx
  reads and writes against concurrent access from the decrypt thread
  and the recovery worker
- Lock is released before FairPlay network calls to avoid blocking
  decryption during reacquisition
- is_preshare_ctx_ready() reads preshareCtx under g_ctx_mutex

Client-visible state gating (main.c)
- handle() and handle_m3u8() check is_recovery_active() per request
- Decrypt server: returns immediately (EOF) during recovery
- M3U8 server: writes empty line, continues loop — no hanging requests
- Prevents FairPlay key-delivery calls from racing the recovery worker

HTTP servers (decrypt, m3u8, account) remain alive across all lease
events. The Electron supervisor continues to provide the outer safety
net for true process deaths (segfault, OOM, etc.).
@silentone12725
silentone12725 marked this pull request as ready for review October 1, 2026 09:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant