Skip to content

feat: workspace lifecycle, retention, and honest resume for script steps - #597

Open
hertznsk wants to merge 12 commits into
microsoft:mainfrom
hertznsk:workspace-lifecycle
Open

hertznsk wants to merge 12 commits into
microsoft:mainfrom
hertznsk:workspace-lifecycle

Conversation

@hertznsk

@hertznsk hertznsk commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds the workspace lifecycle layer for script-step execution: a retained-workspace
contract (WorkspaceIdentity + verify-only attach_run), checkpoint lifecycle
fields (workspace, resume_contract, interrupted_step), an honest resume
digest gate, docker backend attach/marker-staging/retention with a
fail-closed disposition table, restart policy with attempt tracking, local
claim fencing for the retained-volume lease, and full test coverage including a
serialized real-daemon lane.

Behavior for workflows without a workspace block is byte-identical to before
(checkpoint JSON, CLI argv/labels, events, checkpoint list columns).

Details

  • workflow.workspace YAML block (mode + persistence durable/on-failure)
    and restart.mode (rerun|fail) schema with dual enforcement at compile
    and validate time.
  • Checkpoints record workspace identities, the resume contract (workflow /
    environment / manifest / bundle digests), and the interrupted step with its
    attempt id; protocol key names are exempt from secret scrubbing while values
    are still scrubbed.
  • conductor resume verifies the resume contract before touching the retained
    volume and refuses with a named, actionable error on mismatch; --environment
    is honored on resume.
  • Docker backend: marker-protocol v2 staging (separate marker publish carrying
    the bundle digest after the tree copy), attach verification is read-only,
    retained volumes survive failures and are removed only after a successful
    resumed run; attached leases never restage.
  • Local-machine claim fencing (advisory lock + owner token, Windows byte-range
    locking) serializes retained-volume ownership per run.
  • Restart policy re-runs or refuses interrupted steps; attempt_id threads to
    docker labels only under a retention policy.
  • Docs: design document, workflow syntax, CLI reference, runnable example with
    pinned image digest, changelog fragment.

Notes

… backend

Implement the retained-workspace lifecycle for the docker runner:

- marker protocol v2: the .conductor-staged marker (carrying the bundle digest) is published by a separate docker cp after the tree copy; staged detection requires labels match plus a marker probe (scratch cp-out) plus digest comparison, never labels alone

- verify-only attach_run with keyword-only expect_staged: volume inspect plus label verification (managed, run_id, resource, incarnation); never creates, copies, or removes; digest mismatch is always fatal

- fail-closed command path for attached leases: per-command label re-verification and per-command marker re-probe (each command probes independently, even under concurrency); post-create label re-check detects daemon auto-created volumes and removes the container

- eager retained-volume creation in prepare_run with the io.conductor.retention label (durable/on-failure); finalize_run retain keeps the volume while always reclaiming containers; CommandSpec.attempt_id lands on the io.conductor.attempt_id label

- LocalRunnerBackend accepts the shared expect_staged staging policy but refuses attach (retained workspaces are docker-only in this delivery)

- stateful fake docker CLI (disk-backed per-volume store persisting across backend instances, label-less auto-create simulation, serialized store access via fcntl/msvcrt) and lifecycle test coverage incl. tamper and concurrency regressions
Wire the execution session for retained workspace lifecycles:

- prepare_leases/ensure_backends map RunSpec.workspace_persistence per backend capability (retained_workspace backends keep the policy, others get None; the same mapping is applied uniformly before attach)

- resume identity routing: prepare_leases accepts resume_identities and per-backend expect_staged, attaching stored WorkspaceIdentities instead of preparing fresh workspaces; identities for unknown backends are rejected; workspace_identities() exposes only retained-capable backends

- executed-backends session history with union-restore semantics (a fresh process cannot erase checkpoint history); finalize_leases forwards the retain disposition to every lease

- new stdlib-only workspace_claim module: per-run_id advisory lock (fcntl.flock / bounded msvcrt byte-range locking), pid liveness with ESRCH/EPERM distinction, O_EXCL claim creation, token-checked ABA-proof release, lock file never deleted

- all test backend doubles gain the keyword-only retain parameter
…e it

Pin the run-global workspace policy into the audit manifest and enforce it consistently at compile and validate time:

- WorkspacePolicy (mode/persistence) on ResolvedRunManifest with explicit None-pop in every dump/digest path, keeping pre-policy manifests byte-stable; manifest_semantic_digest covers workflow/environment/profiles/secret-references/workspace-policy and excludes producer metadata (conductor_version, audit)

- compile_run_manifest rejects reserved modes (isolated, restart reuse) and retained persistence on backends without retained_workspace capability, naming the follow-up delivery; statically resolvable sub-workflows declaring their own workspace block are rejected (inherited from the root)

- validator: _validate_workspace_policy mirrors the docker-profiles three-tier behavior — bare validate warns that the capability cross-check needs --environment, --environment runs the full check; the execution-resolution report shows the policy only when present

- the backend-capability registry import is function-local to avoid the config-engine-bundle import cycle; a child workflow that fails to parse keeps its legacy reach-time error semantics
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant