feat(upgrades): add cluster release upgrade orchestration - #909
Draft
hemanthnakkina wants to merge 20 commits into
Draft
hemanthnakkina wants to merge 20 commits into
hemanthnakkina wants to merge 20 commits into
Conversation
…dination Introduces an advisory lock in clusterd that serializes upgrade-related state mutations. The lock carries a monotonically increasing fencing token: every state write compare-and-swaps on the token, so a stale holder whose TTL expired cannot silently corrupt state after another process acquires the lock. Go side (sunbeam-microcluster): - upgrade_lock table (single row, seeded at schema apply) - AcquireUpgradeLock / RefreshUpgradeLock / ReleaseUpgradeLock / VerifyToken - HTTP endpoints: POST/PUT/DELETE /1.0/upgrade/lock, GET/PUT /1.0/upgrade/state, GET /1.0/upgrade/active - Stale-token writes return TokenMismatchError (HTTP 409) - Held-lock acquires return LockHeldError (HTTP 409) - Tests: monotonic token, stale-token rejection on verify/refresh, token preservation across release Python side (sunbeam-python): - UpgradeLockHeldException / UpgradeTokenMismatchException - HTTP 409 error translation in BaseService._request - ClusterService methods: acquire/refresh/release_upgrade_lock, get/update_upgrade_state, is_upgrade_active - AcquireUpgradeLockResponse pydantic model - Tests: acquire returns token, 409 surfaces as typed exceptions, CAS-guarded state write sends token + state
Defines the persisted upgrade_state JSON blob as pydantic models matching the section 6.1 state structure, with two design resolutions applied: - metadata_build_id: typed field on Hop, sourced from snap revision at preflight. Lets the engine detect a mid-hop snap refresh and validate engine compatibility against the persisted metadata version. - active_hop is a reference (hop_history_index), not a duplicate of the hop's live state. hop_history[index] is the single canonical record - one source of truth, no dual-write drift on SIGKILL between two writes. The model includes idempotency helpers (is_step_complete, mark_step_complete) for the coordinator's resume logic. Every write to this blob is serialized through the fencing token (commit 26cc78e). 21 tests: example round-trip, metadata_build_id required + preserved, one-source-of-truth (active_hop is just an index), idempotency helpers, empty state safety, fresh hop construction for hop creation.
RELEASE_TRACKS replaces the hardcoded release string fallback in versions.py. Maps each supported OpenStack release (2024.1, 2025.1, 2026.1) to its infrastructure channel mappings. The upgrade engine uses this to validate SLURP hop validity and detect which release a cluster is running at. - RELEASE_TRACKS: dict keyed by release with channel mappings - SLURP_HOPS: set of valid (from, to) release pairs - detect_deployed_release(): reads charm channels, returns release key - detect_snap_release(): reads the snap's deployment.version config - is_valid_hop(): checks if a release pair is a valid upgrade path - DEFAULT_RELEASE replaces the hardcoded "2026.1" fallback Error code catalog defines a closed set of error codes for the last_error.code field on persisted state. Every phase handler that sets last_error uses a code from this catalog. The status command surfaces code + message to the operator. - UpgradeErrorCode enum: 22 codes following <COMPONENT>_<FAILURE> convention - ERROR_MESSAGES: human-readable, actionable message per code - get_error_message(): resolves code to message, falls back to code 29 tests: release track structure, SLURP hop validation, deployed release detection, error code uniqueness, message coverage, actionable length.
Collaborator
Author
|
This is still a PoC code and hence kept as draft |
Defines the orchestration metadata that drives the release-upgrade engine. The metadata tells the engine what to do (group ordering, which actions to run, step sequences, timeouts) — not how to do it (action implementations live in charms, engine steps in the coordinator). This is distinct from the deployment manifest (charm channels and config). The upgrade metadata carries orchestration: which groups exist, their order, which actions to run on which apps, and per-phase step sequences. Adding a new release is a new manifests/<release>/upgrade.yml file — zero Python code changes. Schema: - HopMetadata: top-level (from, to, groups, compatibility, dataplane, storage, finalize, prerequisites) - ControlPlaneGroup: name, apps, ready_timeout_sec, pre/post_actions - ActionSpec: action name, target apps, scope (leader vs all-units) - FinalizeStep: type=action (juju action on apps) or type=engine (built-in handler) - DataplaneConfig / StorageConfig: principal, auxiliary, steps, timeouts Ships manifests/2026.1/upgrade.yml for the 2025.1->2026.1 hop: 9 control-plane groups (identity-core through optional-features), 9-step dataplane sequence, 5-step storage sequence, 6 finalize steps including rpc-cache-refresh on all units of nova-k8s, openstack-hypervisor, cinder-k8s, cinder-volume. 21 tests: schema validation, loader against the shipped YAML, group structure, finalize step types, round-trip serialization.
The coordinator is the central integration point for the upgrade engine. It ties together the advisory lock (fencing token), the typed state model, the orchestration metadata, the error code catalog, and the release tracks table into a single lifecycle: acquire lock, load state, load metadata, dispatch to phase handlers, persist state, release lock. The coordinator is generic — it knows the lifecycle pattern but nothing about specific releases, charms, or actions. Those live in the metadata and the phase handlers (W3-W6, to be implemented). PhaseHandler protocol defines the interface for phase handlers: run() is called with the coordinator, metadata, and state; it returns a PhaseResult. Each phase (preflight, control-plane, dataplane, storage, finalize) will implement this protocol. State machine: hop/phase transitions are validated against transition tables. Invalid transitions raise TransitionError. Terminal states (completed, abandoned) have no outgoing transitions. Resume: load_state reads persisted state from clusterd. If a hop is in_progress, the coordinator finds the current phase and step. Completed steps are skipped; steps with status in_progress are treated as failed and re-run from scratch. Abandon: marks the hop abandoned, releases the lock. Restore artifacts are retained for manual recovery. 33 tests: lock acquire/release/refresh, state load/persist, hop creation, resume (no hop, completed, dataplane, control-plane), phase execution (success, failure with error code, exception), transitions (valid, invalid, terminal), abandon, full lifecycle.
UpgradeLogger writes one JSON object per line to /var/snap/openstack/common/logs/upgrade.log. Every state mutation in the coordinator emits a line: lock acquire/release/refresh, phase started/completed/failed, hop abandoned. Lines are flushed immediately so a SIGKILL does not lose log lines that were already written. Three log entry types: - log_state_change: phase/group/node/step transitions with status and optional error_code + error_message (error catalog) - log_lock_event: advisory lock lifecycle (acquired, refreshed, released, stale) - log_command: CLI command invocations for audit trail The log is append-only — resume appends to the same file. The status command shows current state; upgrade.log shows the history. gather-logs will include this file in the log bundle. Coordinator integration: acquire_lock, refresh_lock, release_lock, run_phase (start/complete/fail), and abandon all emit log lines via the logger. The logger is injected via the coordinator constructor and defaults to UpgradeLogger() with the standard log path. 12 tests: JSON line format, error field inclusion/omission, metadata, append behavior, lock events, command logging, timestamp ISO format, one-line-per-entry validation.
Adds the health-check framework for the upgrade preflight phase. Checks subclass sunbeam.core.checks.Check and carry an exit_code (1 = operational, 2 = invalid hop/metadata). Four checks in order: - SnapVersionCheck: snap release matches target (exit 2) - HopMetadataCheck: validates hop is supported, metadata loads and from/to match the requested hop (exit 2) - ClusterHealthCheck: all apps in both control-plane and machines models are healthy, with tolerated-blocked-message set (exit 1) - MySQLQuorumCheck: runs get-cluster-status action on mysql-k8s leader to verify quorum (exit 1) CheckContext derives client, JujuHelper, and model names from the Deployment object (provider-aware). The runner run_upgrade_preflight_checks short-circuits on first failure. 23 tests, all CI green.
Adds CapacityCheck to the preflight sequence. For each openstack-hypervisor unit, runs the running-guests juju action. A node is free if the action returns an empty list. Fails if free node count or percentage is below the policy threshold. CapacityPolicy (min_free_percentage=25, min_free_nodes=1 by default) is stored in clusterd under the upgrade_capacity_policy config key. The --capacity-policy-override flag (wired via build_preflight_checks capacity_override param) skips the check. build_preflight_checks now returns 5 checks in order: SnapVersion, HopMetadata, ClusterHealth, Capacity, MySQLQuorum. 16 new tests, all CI green.
Adds GuardedGroup (subclass of CatchGroup) to sunbeam/utils.py. Before invoking any subcommand in GUARDED_COMMANDS, checks is_upgrade_active() on clusterd. If an upgrade is in progress, raises ClickException with a clear message directing the operator to check status or abandon. Applied to cluster, enable, disable, configure, and storage command groups. Read-only commands (list, show, status) are not in the guarded set and pass through. The guard checks both ctx.info_name (for top-level groups like enable/disable) and ctx.invoked_subcommand (for nested groups like cluster refresh/join/bootstrap). 10 tests, all CI green.
Adds the 'sunbeam cluster upgrade' command group with the 'abandon' subcommand. Abandon marks the active hop as abandoned (terminal state), releases the upgrade lock, and prints recovery guidance pointing at 'sunbeam restore' for operator-driven recovery. Confirmation prompt by default showing the hop's from->to release pair. --yes flag skips the prompt for automation. Fails with a clear message if no active hop exists or if the upgrade lock is held by another process. The upgrade command group is registered under 'cluster' in both local and MaaS providers. It does not use GuardedGroup — abandon must be runnable during an active upgrade. 8 tests, all CI green.
Adds create_hop_after_preflight() which creates the active upgrade hop atomically after all preflight checks pass. The hop is created with status pending — the first mutating command transitions it to in_progress. Steps: acquire lock, load state, verify no active hop exists, get snap revision as metadata_build_id, create hop via coordinator, copy orchestration metadata to clusterd config key upgrade_metadata (so all nodes can read it), release lock. Releases the lock on any failure. Raises if the lock is held by another process or if an active hop already exists. 8 tests, all CI green.
Adds 'sunbeam cluster upgrade preflight' subcommand. Runs all preflight checks (snap version, hop/metadata, cluster health, capacity, MySQL quorum) and creates the active hop if all pass. --from auto-detected from deployed charm channels via detect_deployed_release() if omitted. --to auto-detected from detect_snap_release(). --capacity-policy-override skips the capacity check. On success: prints backup note, creates active hop with status pending, prints next step (control-plane --auto). On failure: exits with the check's exit code (1 = operational, 2 = invalid hop/metadata). 8 tests, all CI green.
… apply Adds ControlPlaneHandler implementing the PhaseHandler protocol. Reads upgrade groups from orchestration metadata, upgrades each group via scoped terraform apply (update_partial_tfvars_and_apply_tf), waits for convergence via wait_until_desired_status, and persists per-group state transitions. Groups are upgraded in metadata-defined order. Completed groups are skipped on resume. Failed groups are re-attempted on retry (terraform apply is idempotent). On terraform failure: CONTROL_PLANE_APPLY_FAILED. On convergence timeout: CONTROL_PLANE_CONVERGENCE_TIMEOUT. The handler lazily loads tfhelper, manifest, and jhelper from the Deployment object. 10 tests, all CI green.
…-plane groups Adds run_pre_actions() and run_post_actions() to dispatch juju actions declared in group metadata. Pre-actions run before terraform apply with a short propagation delay. Post-actions run after convergence (or after failure, as cleanup). Actions support leader-only and all-units scope. On pre-action failure, post-actions still run as cleanup before returning the error. On post-action failure, the group is marked failed with CONTROL_PLANE_ACTION_FAILED. Wired into ControlPlaneHandler._upgrade_group flow: pre-actions -> terraform apply -> convergence -> post-actions. 15 new tests, all CI green.
Adds 'sunbeam cluster upgrade control-plane' subcommand with six flags: --auto, --group, --application, --status, --retry-group, --dry-run. --auto upgrades all remaining groups in order. --group and --application upgrade a single group or app. --application is rejected on failed groups (must use --retry-group). --status shows per-group state. --retry-group resets a failed/blocked group and re-runs it. --dry-run runs terraform plan per group and shows the actual changes that would be applied — not just a metadata listing. Adds terraform_plan() and update_partial_tfvars_and_plan_tf() to TerraformHelper for this purpose. Also adds run_group() and run_application() methods to ControlPlaneHandler for single-group and single-app execution. plan_group() runs terraform plan for a single group. 16 tests, all CI green.
Manifest Software charms that are no longer in the snap's default software config (e.g. removed between releases) are logged as warnings and stripped from the merged manifest rather than causing a hard error. Carved out of the consolidated testing-fixes commit 371dfbe5.
…handler A snap refresh can bump the bundled juju terraform provider version, leaving the .terraform directory stale. Without init, terraform apply fails with 'unavailable provider registry.terraform.io/juju/juju'. Every other apply path in the codebase runs TerraformInitStep before apply; ControlPlaneHandler was the only one that skipped it. Add self.tfhelper.init() (which passes -upgrade, re-resolving providers from the snap filesystem mirror) before each terraform op: plan_group, run_application, and _upgrade_group.
…ip undeployed charms
hemanthnakkina
force-pushed
the
feat/clusterupgrades
branch
from
August 28, 2026 09:54
d2846f3 to
29511ff
Compare
specs updated in docs section for clsuter upgrades. Assisted-by: z.ai/glm-5.2 Signed-off-by: Hemanth Nakkina <hemanth.nakkina@canonical.com>
…completion The control-plane handler never transitioned phase status or hop.phase, leaving pending forever. W6.3's status command needs these to show correct phase-level progress. Sets: - hop.phase = 'control_plane' on first group/app start - phase status = IN_PROGRESS on first group/app start - phase status = COMPLETED when all groups completed (checked via all_completed guard, not just last-group-is-done) Also adds all_completed guard to run_group() and run_application() completion paths — phase is only marked COMPLETED if every group reports COMPLETED, preventing premature phase completion when a group was retried and failed.
hemanthnakkina
force-pushed
the
feat/clusterupgrades
branch
from
August 30, 2026 06:16
cf6bcd4 to
7b58fcb
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds end-to-end orchestration for upgrading a Sunbeam cluster across
OpenStack releases (e.g. 2025.1 -> 2026.1), built around a hop-based
state machine persisted in the cluster database.
Core engine
upgrades/state.py) with hop lifecycle(pending -> active -> completed/failed/abandoned).
RELEASE_TRACKStable and catalogued error codes(
upgrades/errors.py) for deterministic failure reporting.(
upgrades/metadata.py,manifests/2026.1/upgrade.yml) describingper-release control-plane groups, charm ordering, and pre/post
actions.
upgrades/coordinator.py) driving hopsand group execution, with a structured upgrade logger
(
upgrades/observability.py).Safety
(
database/upgrade_lock.go, exposed via new/upgradeAPI andclusterdclient) so only one upgrade runs at a time and staleholders cannot mutate state.
upgrades/preflight/) including acapacity policy check to ensure compute nodes can be drained safely.
sunbeam cluster upgrade abandonto abort an in-flight hop.Control-plane upgrades
upgrades/control_plane/groups.py)performing scoped terraform plan/apply per group, with pre/post
upgrade actions (
control_plane/actions.py).plandry-run.sunbeam cluster upgrade control-planeCLI with--dry-runplan.sunbeam cluster upgrade preflightCLI to run checks and create theactive hop.
This change is need to support major upgrades for Sunbeam OpenStack.
Assisted-By: z.ai/glm-5.2
QA steps
Links
Jira card: OPEN-4687