Skip to content

astro deployment create panics with nil pointer dereference when the app-config lookup fails for a newly-registered cluster #2258

Description

@leestro

Summary

astro deployment create --cluster-id <id> panics with a nil pointer dereference (SIGSEGV) instead of returning a clean error, when the GetAppConfig lookup for that cluster fails.

Steps to reproduce

  1. Register a new cluster against a Houston-backed APC control plane (e.g. via the registerCluster GraphQL mutation) and confirm it reaches ACTIVE status.
  2. Run:
    astro deployment create --label test --executor celery --cluster-id <newly-registered-cluster-id>
    
  3. Observe the panic (reproduced twice on 1.43.1):
    panic: runtime error: invalid memory address or nil pointer dereference
    [signal SIGSEGV: segmentation violation code=0x2 addr=0x39 pc=...]
    
    goroutine 1 [running]:
    github.com/astronomer/astro-cli/software/deployment.Create(...)
    	.../software/deployment/deployment.go:100 +0x23c
    github.com/astronomer/astro-cli/cmd/software.deploymentCreate(...)
    	.../cmd/software/deployment.go:481 +0x698
    github.com/astronomer/astro-cli/cmd/software.newDeploymentCreateCmd.func1(...)
    	.../cmd/software/deployment.go:161 +0x20
    

Expected vs actual

Expected: either the deployment is created, or astro deployment create returns a clean, actionable error (e.g. "could not fetch app config for cluster X").
Actual: the CLI panics with a raw Go stack trace.

Root cause

  • cmd/software/deployment.go:489: appConfig, _ = houston.Call(houstonClient.GetAppConfig)(houston.GetAppConfigRequest{ClusterID: clusterID, WorkspaceUUID: ws}) discards the error from GetAppConfig. When this call fails for the given cluster ID, appConfig is left nil.
  • software/deployment/deployment.go:103: if appConfig.Flags.ManualNamespaceNames { dereferences appConfig unconditionally, with no nil check — even though the caller, two lines below its own call site (cmd/software/deployment.go:505, if appConfig != nil { ... }), explicitly treats appConfig as nilable. That mismatch between the two call sites' assumptions is what turns an app-config lookup failure into a segfault instead of a clean error.
  • Confirmed still present on current main as of 2026-09-16: diffed v1.43.1..HEAD on both files above — the only changes in range are an added --mode flag and the unrelated Adopt/Unadopt deployment commands. The missing nil-guard is unchanged, so upgrading to the latest release would not fix this.

Environment

  • astro-cli 1.43.1 (bug confirmed unfixed on current main too)
  • Astro Private Cloud (Software) context, local APC 2.1.0 k3d install
  • Reproduced against a data-plane cluster registered seconds earlier via registerCluster

Suggested fix

Either:

  • Return the error from GetAppConfig in cmd/software/deployment.go (appConfig, err = ...) and handle a non-nil err instead of discarding it, or
  • Add the same if appConfig != nil guard inside software/deployment.Create before accessing appConfig.Flags.*, matching the caller's own defensive pattern two lines away.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions