Skip to content

feat(upgrades): add cluster release upgrade orchestration - #909

Draft
hemanthnakkina wants to merge 20 commits into
canonical:mainfrom
hemanthnakkina:feat/clusterupgrades
Draft

hemanthnakkina wants to merge 20 commits into
canonical:mainfrom
hemanthnakkina:feat/clusterupgrades

Conversation

@hemanthnakkina

@hemanthnakkina hemanthnakkina commented Aug 28, 2026 •

Copy link
Copy Markdown
Collaborator

Adds end-to-end orchestration for upgrading a Sunbeam cluster across
OpenStack releases (e.g. 2025.1 -> 2026.1), built around a hop-based
state machine persisted in the cluster database.

Core engine

  • Typed upgrade state model (upgrades/state.py) with hop lifecycle
    (pending -> active -> completed/failed/abandoned).
  • RELEASE_TRACKS table and catalogued error codes
    (upgrades/errors.py) for deterministic failure reporting.
  • Declarative upgrade metadata schema and loader
    (upgrades/metadata.py, manifests/2026.1/upgrade.yml) describing
    per-release control-plane groups, charm ordering, and pre/post
    actions.
  • Release upgrade coordinator (upgrades/coordinator.py) driving hops
    and group execution, with a structured upgrade logger
    (upgrades/observability.py).

Safety

  • Advisory lock with fencing token in sunbeam-microcluster
    (database/upgrade_lock.go, exposed via new /upgrade API and
    clusterd client) so only one upgrade runs at a time and stale
    holders cannot mutate state.
  • Mutating CLI commands are guarded while an upgrade hop is active.
  • Preflight health-check framework (upgrades/preflight/) including a
    capacity policy check to ensure compute nodes can be drained safely.
  • sunbeam cluster upgrade abandon to abort an in-flight hop.

Control-plane upgrades

  • Control-plane group handler (upgrades/control_plane/groups.py)
    performing scoped terraform plan/apply per group, with pre/post
    upgrade actions (control_plane/actions.py).
  • Terraform helper supports targeted applies and plan dry-run.
  • sunbeam cluster upgrade control-plane CLI with --dry-run plan.
  • sunbeam cluster upgrade preflight CLI to run checks and create the
    active hop.

This change is need to support major upgrades for Sunbeam OpenStack.

Assisted-By: z.ai/glm-5.2

QA steps

  1. Deploy sunbeam epoxy 2025.1
  2. Install the snap built from these changes 2026.1/edge/upgrade
  3. sunbeam cluster upgrade preflight
  4. sunbeam cluster upgrade control-plane --dry-run
  5. sunbeam cluster upgrade control-plane --group identity-core
  6. sunbeam cluster upgrade control-plane --group image
  7. sunbeam cluster upgrade control-plane --group placement
  8. sunbeam cluster upgrade control-plane --group block-storage-api
  9. sunbeam cluster upgrade control-plane --group network-api
  10. sunbeam cluster upgrade control-plane --group compute-control
  11. sunbeam cluster upgrade control-plane --group dashboard
  12. sunbeam cluster upgrade control-plane --group optional-features

Links

Jira card: OPEN-4687

…dination

Introduces an advisory lock in clusterd that serializes upgrade-related
state mutations. The lock carries a monotonically increasing fencing
token: every state write compare-and-swaps on the token, so a stale
holder whose TTL expired cannot silently corrupt state after another
process acquires the lock.

Go side (sunbeam-microcluster):
- upgrade_lock table (single row, seeded at schema apply)
- AcquireUpgradeLock / RefreshUpgradeLock / ReleaseUpgradeLock / VerifyToken
- HTTP endpoints: POST/PUT/DELETE /1.0/upgrade/lock,
  GET/PUT /1.0/upgrade/state, GET /1.0/upgrade/active
- Stale-token writes return TokenMismatchError (HTTP 409)
- Held-lock acquires return LockHeldError (HTTP 409)
- Tests: monotonic token, stale-token rejection on verify/refresh,
  token preservation across release

Python side (sunbeam-python):
- UpgradeLockHeldException / UpgradeTokenMismatchException
- HTTP 409 error translation in BaseService._request
- ClusterService methods: acquire/refresh/release_upgrade_lock,
  get/update_upgrade_state, is_upgrade_active
- AcquireUpgradeLockResponse pydantic model
- Tests: acquire returns token, 409 surfaces as typed exceptions,
  CAS-guarded state write sends token + state
Defines the persisted upgrade_state JSON blob as pydantic models matching
the section 6.1 state structure, with two design resolutions applied:

- metadata_build_id: typed field on Hop, sourced from snap revision at
  preflight. Lets the engine detect a mid-hop snap refresh and validate
  engine compatibility against the persisted metadata version.

- active_hop is a reference (hop_history_index), not a duplicate of the
  hop's live state. hop_history[index] is the single canonical record -
  one source of truth, no dual-write drift on SIGKILL between two writes.

The model includes idempotency helpers (is_step_complete, mark_step_complete)
for the coordinator's resume logic. Every write to this blob is serialized
through the fencing token (commit 26cc78e).

21 tests: example round-trip, metadata_build_id required + preserved,
one-source-of-truth (active_hop is just an index), idempotency helpers,
empty state safety, fresh hop construction for hop creation.
RELEASE_TRACKS replaces the hardcoded release string fallback in
versions.py. Maps each supported OpenStack release (2024.1, 2025.1,
2026.1) to its infrastructure channel mappings. The upgrade engine
uses this to validate SLURP hop validity and detect which release a
cluster is running at.

- RELEASE_TRACKS: dict keyed by release with channel mappings
- SLURP_HOPS: set of valid (from, to) release pairs
- detect_deployed_release(): reads charm channels, returns release key
- detect_snap_release(): reads the snap's deployment.version config
- is_valid_hop(): checks if a release pair is a valid upgrade path
- DEFAULT_RELEASE replaces the hardcoded "2026.1" fallback

Error code catalog  defines a closed set of error codes for the
last_error.code field on persisted state. Every phase handler that sets
last_error uses a code from this catalog. The status command surfaces
code + message to the operator.

- UpgradeErrorCode enum: 22 codes following <COMPONENT>_<FAILURE> convention
- ERROR_MESSAGES: human-readable, actionable message per code
- get_error_message(): resolves code to message, falls back to code

29 tests: release track structure, SLURP hop validation, deployed release
detection, error code uniqueness, message coverage, actionable length.
@hemanthnakkina

Copy link
Copy Markdown
Collaborator Author

This is still a PoC code and hence kept as draft

Defines the orchestration metadata that drives the release-upgrade engine.
The metadata tells the engine what to do (group ordering, which actions to
run, step sequences, timeouts) — not how to do it (action implementations
live in charms, engine steps in the coordinator).

This is distinct from the deployment manifest (charm channels and config).
The upgrade metadata carries orchestration: which groups exist, their
order, which actions to run on which apps, and per-phase step sequences.

Adding a new release is a new manifests/<release>/upgrade.yml file —
zero Python code changes.

Schema:
- HopMetadata: top-level (from, to, groups, compatibility, dataplane,
  storage, finalize, prerequisites)
- ControlPlaneGroup: name, apps, ready_timeout_sec, pre/post_actions
- ActionSpec: action name, target apps, scope (leader vs all-units)
- FinalizeStep: type=action (juju action on apps) or type=engine
  (built-in handler)
- DataplaneConfig / StorageConfig: principal, auxiliary, steps, timeouts

Ships manifests/2026.1/upgrade.yml for the 2025.1->2026.1 hop: 9
control-plane groups (identity-core through optional-features), 9-step
dataplane sequence, 5-step storage sequence, 6 finalize steps including
rpc-cache-refresh on all units of nova-k8s, openstack-hypervisor,
cinder-k8s, cinder-volume.

21 tests: schema validation, loader against the shipped YAML, group
structure, finalize step types, round-trip serialization.
The coordinator is the central integration point for the upgrade engine.
It ties together the advisory lock (fencing token), the typed state model,
the orchestration metadata, the error code catalog, and the release tracks
table into a single lifecycle: acquire lock, load state, load metadata,
dispatch to phase handlers, persist state, release lock.

The coordinator is generic — it knows the lifecycle pattern but nothing
about specific releases, charms, or actions. Those live in the metadata
and the phase handlers (W3-W6, to be implemented).

PhaseHandler protocol defines the interface for phase handlers: run() is
called with the coordinator, metadata, and state; it returns a PhaseResult.
Each phase (preflight, control-plane, dataplane, storage, finalize) will
implement this protocol.

State machine: hop/phase transitions are validated against transition
tables. Invalid transitions raise TransitionError. Terminal states
(completed, abandoned) have no outgoing transitions.

Resume: load_state reads persisted state from clusterd. If a hop is
in_progress, the coordinator finds the current phase and step. Completed
steps are skipped; steps with status in_progress are treated as failed
and re-run from scratch.

Abandon: marks the hop abandoned, releases the lock. Restore artifacts
are retained for manual recovery.

33 tests: lock acquire/release/refresh, state load/persist, hop creation,
resume (no hop, completed, dataplane, control-plane), phase execution
(success, failure with error code, exception), transitions (valid, invalid,
terminal), abandon, full lifecycle.
UpgradeLogger writes one JSON object per line to
/var/snap/openstack/common/logs/upgrade.log. Every state mutation in the
coordinator emits a line: lock acquire/release/refresh, phase
started/completed/failed, hop abandoned. Lines are flushed immediately
so a SIGKILL does not lose log lines that were already written.

Three log entry types:
- log_state_change: phase/group/node/step transitions with status and
  optional error_code + error_message (error catalog)
- log_lock_event: advisory lock lifecycle (acquired, refreshed,
  released, stale)
- log_command: CLI command invocations for audit trail

The log is append-only — resume appends to the same file. The status
command  shows current state; upgrade.log shows the history.
gather-logs  will include this file in the log bundle.

Coordinator integration: acquire_lock, refresh_lock, release_lock,
run_phase (start/complete/fail), and abandon all emit log lines via
the logger. The logger is injected via the coordinator constructor and
defaults to UpgradeLogger() with the standard log path.

12 tests: JSON line format, error field inclusion/omission, metadata,
append behavior, lock events, command logging, timestamp ISO format,
one-line-per-entry validation.
Adds the health-check framework for the upgrade preflight phase.
Checks subclass sunbeam.core.checks.Check and carry an exit_code
(1 = operational, 2 = invalid hop/metadata).

Four checks in order:
- SnapVersionCheck: snap release matches target (exit 2)
- HopMetadataCheck: validates hop is supported, metadata loads and
  from/to match the requested hop (exit 2)
- ClusterHealthCheck: all apps in both control-plane and machines
  models are healthy, with tolerated-blocked-message set (exit 1)
- MySQLQuorumCheck: runs get-cluster-status action on mysql-k8s
  leader to verify quorum (exit 1)

CheckContext derives client, JujuHelper, and model names from the
Deployment object (provider-aware). The runner
run_upgrade_preflight_checks short-circuits on first failure.

23 tests, all CI green.
Adds CapacityCheck to the preflight sequence. For each
openstack-hypervisor unit, runs the running-guests juju action.
A node is free if the action returns an empty list. Fails if free
node count or percentage is below the policy threshold.

CapacityPolicy (min_free_percentage=25, min_free_nodes=1 by default)
is stored in clusterd under the upgrade_capacity_policy config key.
The --capacity-policy-override flag (wired via build_preflight_checks
capacity_override param) skips the check.

build_preflight_checks now returns 5 checks in order:
SnapVersion, HopMetadata, ClusterHealth, Capacity, MySQLQuorum.

16 new tests, all CI green.
Adds GuardedGroup (subclass of CatchGroup) to sunbeam/utils.py.
Before invoking any subcommand in GUARDED_COMMANDS, checks
is_upgrade_active() on clusterd. If an upgrade is in progress,
raises ClickException with a clear message directing the operator
to check status or abandon.

Applied to cluster, enable, disable, configure, and storage
command groups. Read-only commands (list, show, status) are not
in the guarded set and pass through.

The guard checks both ctx.info_name (for top-level groups like
enable/disable) and ctx.invoked_subcommand (for nested groups
like cluster refresh/join/bootstrap).

10 tests, all CI green.
Adds the 'sunbeam cluster upgrade' command group with the 'abandon'
subcommand. Abandon marks the active hop as abandoned (terminal
state), releases the upgrade lock, and prints recovery guidance
pointing at 'sunbeam restore' for operator-driven recovery.

Confirmation prompt by default showing the hop's from->to release
pair. --yes flag skips the prompt for automation.

Fails with a clear message if no active hop exists or if the
upgrade lock is held by another process.

The upgrade command group is registered under 'cluster' in both
local and MaaS providers. It does not use GuardedGroup — abandon
must be runnable during an active upgrade.

8 tests, all CI green.
Adds create_hop_after_preflight() which creates the active upgrade
hop atomically after all preflight checks pass. The hop is created
with status pending — the first mutating command transitions it to
in_progress.

Steps: acquire lock, load state, verify no active hop exists, get
snap revision as metadata_build_id, create hop via coordinator,
copy orchestration metadata to clusterd config key upgrade_metadata
(so all nodes can read it), release lock.

Releases the lock on any failure. Raises if the lock is held by
another process or if an active hop already exists.

8 tests, all CI green.
Adds 'sunbeam cluster upgrade preflight' subcommand. Runs all
preflight checks (snap version, hop/metadata, cluster health,
capacity, MySQL quorum) and creates the active hop if all pass.

--from auto-detected from deployed charm channels via
detect_deployed_release() if omitted. --to auto-detected from
detect_snap_release(). --capacity-policy-override skips the
capacity check.

On success: prints backup note, creates active hop with status
pending, prints next step (control-plane --auto).
On failure: exits with the check's exit code (1 = operational,
2 = invalid hop/metadata).

8 tests, all CI green.
… apply

Adds ControlPlaneHandler implementing the PhaseHandler protocol.
Reads upgrade groups from orchestration metadata, upgrades each group
via scoped terraform apply (update_partial_tfvars_and_apply_tf), waits
for convergence via wait_until_desired_status, and persists per-group
state transitions.

Groups are upgraded in metadata-defined order. Completed groups are
skipped on resume. Failed groups are re-attempted on retry (terraform
apply is idempotent). On terraform failure: CONTROL_PLANE_APPLY_FAILED.
On convergence timeout: CONTROL_PLANE_CONVERGENCE_TIMEOUT.

The handler lazily loads tfhelper, manifest, and jhelper from the
Deployment object.

10 tests, all CI green.
…-plane groups

Adds run_pre_actions() and run_post_actions() to dispatch juju
actions declared in group metadata. Pre-actions run before terraform
apply with a short propagation delay. Post-actions run after
convergence (or after failure, as cleanup).

Actions support leader-only and all-units scope. On pre-action
failure, post-actions still run as cleanup before returning the
error. On post-action failure, the group is marked failed with
CONTROL_PLANE_ACTION_FAILED.

Wired into ControlPlaneHandler._upgrade_group flow:
pre-actions -> terraform apply -> convergence -> post-actions.

15 new tests, all CI green.
Adds 'sunbeam cluster upgrade control-plane' subcommand with six
flags: --auto, --group, --application, --status, --retry-group,
--dry-run.

--auto upgrades all remaining groups in order. --group and
--application upgrade a single group or app. --application is
rejected on failed groups (must use --retry-group). --status shows
per-group state. --retry-group resets a failed/blocked group and
re-runs it.

--dry-run runs terraform plan per group and shows the actual changes
that would be applied — not just a metadata listing. Adds
terraform_plan() and update_partial_tfvars_and_plan_tf() to
TerraformHelper for this purpose.

Also adds run_group() and run_application() methods to
ControlPlaneHandler for single-group and single-app execution.
plan_group() runs terraform plan for a single group.

16 tests, all CI green.
Manifest Software charms that are no longer in the snap's default
software config (e.g. removed between releases) are logged as warnings
and stripped from the merged manifest rather than causing a hard error.

Carved out of the consolidated testing-fixes commit 371dfbe5.
…handler

A snap refresh can bump the bundled juju terraform provider version,
leaving the .terraform directory stale. Without init, terraform apply
fails with 'unavailable provider registry.terraform.io/juju/juju'.

Every other apply path in the codebase runs TerraformInitStep before
apply; ControlPlaneHandler was the only one that skipped it. Add
self.tfhelper.init() (which passes -upgrade, re-resolving providers
from the snap filesystem mirror) before each terraform op:
plan_group, run_application, and _upgrade_group.
specs updated in docs section for clsuter upgrades.

Assisted-by: z.ai/glm-5.2
Signed-off-by: Hemanth Nakkina <hemanth.nakkina@canonical.com>
…completion

The control-plane handler never transitioned phase status or hop.phase,
leaving pending forever. W6.3's status command needs these to show
correct phase-level progress.

Sets:
- hop.phase = 'control_plane' on first group/app start
- phase status = IN_PROGRESS on first group/app start
- phase status = COMPLETED when all groups completed (checked via
  all_completed guard, not just last-group-is-done)

Also adds all_completed guard to run_group() and run_application()
completion paths — phase is only marked COMPLETED if every group
reports COMPLETED, preventing premature phase completion when a
group was retried and failed.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant