Skip to content

fix(platform): one contract on aws and gcp keeps two states, and GCP fails closed on sovereignty - #676

Merged
fas89 merged 13 commits into
mainfrom
feat/gcp-platform-safety
Sep 28, 2026
Merged

fas89 merged 13 commits into
mainfrom
feat/gcp-platform-safety

Conversation

@fas89

@fas89 fas89 commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Platform safety for deploying one contract to aws and gcp through --env overlays (PR 1 of 3: state, stage 8, sovereignty, Command Center run reports, pipeline collisions). The embedded-SQL-over-BigQuery work and the GCP governance emitters are the other two PRs; this one keeps its edits in iac/providers/gcp.py, providers/gcp/provider.py and the policy engine local to sovereignty.

The gaps, with the evidence

An offline verification of the three demo products on 0.16.5 (no real cloud) found:

# Gap Measured
F1 aws and gcp applies of one contract share one remote state resolve_state_target with FLUID_STATE_BACKEND=s3://fluid-demo-lab-state-… gave both fluid/bronze.customer_subscriptions/terraform.tfstate. The gcp plan would read the aws resources as orphans to destroy, and --allow-data-loss would destroy them. Re-measured on origin/main in this PR: SAME.
F2 stage 8 crashes on every gcp build fluid policy-apply … --mode enforce exit 1, TypeError: info() got multiple values for argument 'message' (cli/policy_apply.py:136), on all three products. Re-measured on origin/main with the demo's bronze gcp bindings.json: rc 1, same TypeError.
F5 sovereignty fails open on GCP A gcp overlay with no region: fluid validate rc 0, plan --check-sovereignty printed PASS, and the dataset landed in US. us-central1 was refused by validate only; fluid generate iac emitted it with rc 0 (no GCP provider hook). A region given as location.location, which the GCP emitter reads, was never checked. The BigQuery load defaulted its job location to US.
F10 nothing reports a Jenkins run to the Command Center get_reporter (cli/bootstrap.py:81) has no callers; a scratch CC test recorded 0 pipeline_executions rows.
F11a the aws and gcp DAGs of one product collide Both dag_id=<product>__<build> in the same schedule/<product>/ directory, which schedule-sync --delete-scope product mirrors with delete.
F11b the generated stage 6 never checks sovereignty No --check-sovereignty in any rendered CI system; a strict violation first failed at stage 7.
F11c --env gcp with no gcp overlay quietly runs the local base silver: validate rc 0 and plan rc 0 with only a warning.

Design

F1: provider-keyed state, moved with OpenTofu's own migration

  • iac/backend.py: every per-contract default key gains the provider, fluid/<id>/<provider>/terraform.tfstate (GCS prefix fluid/<id>/<provider>). That covers the FLUID_STATE_BACKEND bucket-only default and the packaging-contract default. The shared legacy key fluid/terraform.tfstate and any explicit key or prefix are unchanged. parse_backend without provider is byte-for-byte what it was. legacy_default_backend names where the previous release kept the state.
  • iac/state_migration.py (new): after the workdir's own tofu init, tofu state pull reads the new key. When it holds state, nothing else happens. When it holds none, a provider-less scratch module is initialised on the old key and pulled. The state is attributed by its resources' own provider["registry…/<ns>/<type>"] addresses, checked against each IaC plugin's required_providers:
    • all resources this provider's → the old state is moved;
    • all resources exactly one other provider's (the gcp apply finding the aws state) → left alone;
    • anything else (two clouds, an unknown provider, only shared utilities) → refused with state_migration_ambiguous, naming both keys.
  • The copy is OpenTofu's. The scratch directory's .terraform/ recorded the old backend. Its module is rewritten to the new one, and tofu init -force-copy (which implies -migrate-state) copies the state. This is the backend-reconfiguration path of terraform init -migrate-state: it locks both states where the backend supports locking, and it leaves the source untouched.
    • -force-copy overwrites a non-empty destination, so the new key is checked twice: before deciding, and again just before the copy. A different state appearing between the two checks is state_migration_raced.
    • The copy is verified by reading it back. OpenTofu 1.12 gives a copy into an empty destination a fresh lineage and serial 1 (measured), so the check is on the resources.
    • The old object is never deleted.
  • A kept workdir whose .terraform/ recorded the old key is re-initialised with -reconfigure (a plain init stops at "Backend configuration changed"). A wiped CI workspace needs nothing.
  • fluid diff / verify --state-drift and fluid apply --dry-run are read-only. While the move is pending they plan against the old key, so a stage 5 drift gate, or a plan-only stage 7, before the first real upgraded apply still sees the real state and writes nothing. Only a real apply moves it.
  • The probe installs nothing. OpenTofu's init installs the providers the state names (a {"terraform": {}} module beside a hashicorp/null state installed the latest hashicorp/null). The probe runs on every apply whose new key is empty, so it inits with -plugin-dir on an empty directory: the init stops at its provider step after recording the old backend, and the state is pulled from exactly that recorded backend (checked with records_backend, never assumed). If a future OpenTofu did not record it first, the probe falls back to a plain init. The one-time copy installs what the old state names at the IaC plugin's own pins (~> 5.0, not the latest).
  • A shared key holding the other cloud is refused. A remote key that does not name the provider (the shared fluid/terraform.tfstate that a bucket-only --state-backend gives a contract without packaging, or an explicit key reused for both clouds) is read once; if it holds another cloud's resources, the apply refuses with state_shared_with_another_provider instead of planning their destruction.

F2: the log helpers take logger and message positionally only

info / warn / error(logger, message, /, **payload) (PEP 570). A payload key that names an envelope key (time, level, name, message) is kept as extra_<key>, so the GCP policy result's own sentence survives and the event keeps its name. This fixes it for every provider that returns a message, and for every other **dict call site (contract_tests, the apply engine).

F5: sovereignty fails closed, and GCP gets a provider hook

  • Policy engine (policy/sovereignty.py):
    • New check 0: a binding on a region-placed platform (aws, gcp, azure) with a sovereignty block and no region is a finding at severity_for(mode). Strict refuses, advisory warns, audit logs. local bindings, and platforms whose region is not on the binding, keep the old behaviour.
    • binding_region reads location.location for GCP, as the emitter does.
    • US / EU (the BigQuery and Cloud Storage multi-regions) resolve to their jurisdictions.
    • Strict fails closed on an unplaceable location. On an aws / gcp / azure binding, a jurisdiction the table cannot resolve is refused under strict, unless the operator named the region in allowedRegions. It was a warning, so me-central2, northamerica-south1 and the asia multi-region passed validate and generate iac on an EU-only contract. Advisory and audit, and other platforms, keep the warning.
    • The GCP table is complete. The nine Google regions the vendored dgl/cloud-regions csv lacks (europe-north2, europe-west10, europe-west12, me-central1/2, me-west1, africa-south1, northamerica-south1, asia-southeast3), and the dual-regions within one jurisdiction (EUR4 → EU, NAM4 → US, ASIA1 → JP), are mapped from Google's own location pages, in source (the csv is ODbL and stays pristine). EUR5/EUR7/EUR8 and the ASIA multi-region span jurisdictions and stay unresolved, so strict refuses them.
    • The per-region checks were refactored into check_placements(sovereignty, [(where, region)]), so a provider hook runs the same rules on its own placements.
  • GCP hook (providers/gcp/util/sovereignty.py):
    • GcpIacPlugin.emit checks every emitted location / region. This is the chokepoint of both fluid apply and fluid generate iac, and it runs whether or not a native provider can be built.
    • GcpProvider.plan checks every planned one, raising the typed SovereigntyViolationError through ProviderError, the way AwsProvider.plan does. native_actions re-raises it.
    • A region the provider filled in by default is therefore checked where it is used: the planner's US dataset default, and the provider region a Cloud Scheduler job or staging bucket inherits (the --region default europe-west3, the SDK's us-central1).
    • With no sovereignty block, the US default is unchanged. Under advisory it is emitted and warned about, so it is never silent.
    • Pub/Sub keeps its region. A gcp pubsub_topic binding's region becomes message_storage_policy.allowed_persistence_regions on the google_pubsub_topic (hashicorp/google), and the hook checks it. It was dropped, so the topic passed validate under allowedRegions: [europe-west1] and could store messages anywhere.
  • BigQuery load: a binding with no region gives location: None. The job runs where the table is (table.location), never at a guessed US.

F10: fluid apply reports each run to the Command Center

  • The API. The CC's executions API is app/api/v1/executions.py, with ExecutionCreate / ExecutionUpdate. It takes X-API-Key or a bearer token, and scopes the run by X-Organization-Id. CommandCenterReporter (async queue, circuit breaker, SSRF host gate) already speaks it, so it is wired in, not replaced. observability/config.py gains headers, for a bearer credential and the organization.
  • Authentication. The run authenticates the way fluid publish does: the fluid-command-center catalog configuration (FLUID_CC_ENDPOINT, FLUID_API_KEY or FLUID_BEARER_TOKEN). The organization comes from FLUID_CC_ORG_ID, organization_id, or the product's fluid.config.yaml slug, which the demo lab uses, all through the publisher's own resolve_organization_id. The reporter's FLUID_COMMAND_CENTER_URL / _API_KEY are the fallback.
  • What is sent. POST /api/v1/executions when the engine knows the product (status running, provider, environment, runner: jenkins|cli, cli_version, and in metadata: product_id, contract_version, fluid_version, contract_hash, platform, state, mode). Then PATCH /api/v1/executions/{id} at the end: success / failed, and a result with the planned and applied counts, the resource addresses, dry_run, exit_code, started_at, finished_at and duration_seconds, plus the typed error_event on failure.
  • Best effort. An outage, a 401 or a 500 costs a warning line and never the exit code.
  • Never a secret. The credential is only ever a header. tofu's output is never sent (its text can carry attribute values); a failure is its event name.
  • Early failures. The engine identifies the run (product, contract version, hash, platform, env) as soon as the contract and provider are known, so a refusal before the run is registered (a sovereignty refusal in the emitter) still carries its product. A run refused even earlier (a declared env with no overlay, a provider mismatch caught while the apply picks its engine, plan binding) is described from the base contract: id, name and versions, no hash (the Command Center keys versions by the compiled contract's hash).
  • Refusals. With no credential configured (the built-in default endpoint has none), nothing is sent and nothing is said. Without a settled organization, the run is not sent (the CC would store it untagged, invisible to every read). The SSRF gate runs before the organization lookup sends the credential. FLUID_COMMAND_CENTER_ENABLED=false turns it off.

F11

  • (a) DAG ids and directories. dag_id_for(product, build, env) gives <product>__<env>__<build>, and schedule_scope_for gives schedule/<product>__<env>/. No env keeps both as they were.
  • (b) Stage 6. --check-sovereignty is in the shared command table (GitHub Actions, GitLab, Azure DevOps, Bitbucket, CircleCI), the 11-stage spec (Tekton) and Jenkins' stage 6. A contract with no sovereignty block prints NOT CHECKED and passes.
  • (c) Missing overlays. forge-cli did not know expected-environments; it now reads it as a declaration.
    • --env X with no overlay is overlay_declared_but_missing when the workspace's expected-environments declares X for the product (keyed by the contract's directory name, then its id, as fluid-demo-env writes it).
    • The contract's own environments block does not refuse: it is schema-valid, forge-cli applies nothing from it, and refusing on it broke contracts that validated on 0.16.5. The warning names it instead.
    • dev (the base by convention) and an env the base is already bound to (local for a local base) stay the base.
    • The refusal applies to validate, plan, apply, bundle and generate.
    • Otherwise the warning stays, and now names what the base binds to: "it binds to local, not to 'prod'".

Prior art (borrow-before-build receipts)

Searched:

  • opentofu init -migrate-state -force-copy backend configuration change copy state non-interactive: the OpenTofu init docs. -force-copy implies -migrate-state and answers yes to every prompt.
  • terragrunt remote_state key change migrate existing state terraform init -migrate-state: Terragrunt backend migrate (renamed units); the HashiCorp help article on init -migrate-state.
  • structlog event key collision "got multiple values for argument" positional-only: Can't set the "logger" key in "get_logger(initial_values=...)" hynek/structlog#295, the same collision on logger.
  • python logging reserved LogRecord attributes "Attempt to overwrite" message extra key collision: logging: rename colliding extra keys instead of raising KeyError WebbPulse/webbpulse-python#129 renames to extra_<key>; Drop event dict keys colliding with LogRecord attributes in render_to_log_* hynek/structlog#842 drops the key instead.
  • terraform google_bigquery_dataset location default "US": BigQuery defaults an unplaced dataset to the US multi-region.
  • checkov OR tfsec OR conftest policy require resource location region allowed list: region allow-lists are policy-as-code (conftest over plan JSON). Here the contract's own engine is that policy.
  • openlineage python client http transport timeout errors do not fail the job: the client logs emit failures and never fails the job (continue_on_failure).
  • airflow dag_id unique across deployments naming convention environment suffix: dag_id must be unique per Airflow, and renaming one orphans its history.
  • dbt "does not have a target named": dbt refuses a requested target its profile does not declare.
  • (review round) BigQuery dataset locations list regions multi-regions dual-regions EUR4 NAM4 ASIA1: Google's location pages.
  • (review round) terraform google_pubsub_topic message_storage_policy allowed_persistence_regions: the hashicorp/google argument (message_storage_policy { allowed_persistence_regions = [...] }).
  • (review round) terraform plan -lock=false read-only state backend dry run migrate-state: nothing plans across a pending backend move; plan against the old backend, move on apply.
  • (review round) opentofu init installs providers required by state plugin cache TF_PLUGIN_CACHE_DIR: init installs state-required providers; Conflict when TF_PLUGIN_CACHE_DIR is set to an local provider search directory opentofu/opentofu#3721 (a plugin cache that is also a search directory conflicts), which is why the apply's own .terraform/providers is never the probe's plugin directory.

Fetched:

Reuse strategy:

  • Depend on: OpenTofu's backend migration, invoked rather than re-implemented; nothing reads or edits a state document's content.
  • Reuse in repo: CommandCenterReporter, FluidCommandCenterProvider.resolve_organization_id, the SSRF gate, and severity_for.
  • Adapt the pattern: the extra_<key> rename; the "an explicitly requested target must exist" rule from dbt; the provider's own message_storage_policy for Pub/Sub placement; dbt-style "plan reads, apply writes" for the pending move.
  • Diverge intentionally: Terragrunt's backend migrate needs an operator-named source and destination; here both are derived, and the attribution rule replaces the operator.

Each borrowed pattern is credited in a code comment.

Schema changes

None. No contract field was added (0.7.5 GA and the 0.7.6 preview are untouched). expected-environments is a key of fluid.workspace.yaml, read if present; the contract's environments block already exists in the schema.

Behaviour changes (release notes)

  • State key. The default remote state key for a bucket-only FLUID_STATE_BACKEND, or a packaging contract, is now fluid/<id>/<provider>/terraform.tfstate (GCS prefix fluid/<id>/<provider>). The first apply after the upgrade moves an existing state there with tofu init -migrate-state, prints state move: moved N resource(s) from … to …, and leaves the old object in place. A state it cannot attribute to one provider is refused (state_migration_ambiguous). The shared legacy key and explicit keys are unchanged.
  • Sovereignty. Under strict sovereignty, an aws / gcp / azure binding with no region now fails fluid validate, plan --check-sovereignty, generate iac and apply, and so does one whose region's jurisdiction cannot be resolved (unless the region is in allowedRegions). Advisory warns and audit logs. The demo's committed bronze gcp overlay already sets region: europe-west1 (fluid-demo-env d225844) and is unaffected; the silver and gold gcp overlays (D1) must set a region when they are added. A GCP location.location is checked. US / EU multi-regions resolve to US / EU, and the GCP regions and dual-regions the vendored table lacked now resolve.
  • Pub/Sub. A gcp pubsub_topic binding with a region now emits message_storage_policy.allowed_persistence_regions. An existing topic plans an in-place update.
  • Dry-run. fluid apply --dry-run never moves state; the first real apply does.
  • Shared state key. An apply whose remote key names no provider and holds another cloud's resources is refused (state_shared_with_another_provider).
  • GCP placement. fluid apply and fluid generate iac on gcp refuse a placement outside the sovereignty policy, including a region a planned resource inherits from the provider default.
  • BigQuery load. A load for a binding with no region runs in the table's own location.
  • Command Center reports. fluid apply reports each run to the Command Center wherever fluid publish is configured. Off with FLUID_COMMAND_CENTER_ENABLED=false. A Command Center on a private address needs FLUID_COMMAND_CENTER_HOST_ALLOWLIST (loopback is always allowed).
  • Stage 8. fluid policy-apply on gcp no longer crashes. The structured log event carries the provider's message as extra_message.
  • DAGs. An env-bound DAG's id is <product>__<env>__<build>, in schedule/<product>__<env>/. Upgrade step: a DAG an earlier release synced for an env stays at the scheduler under <product>/ with its old id, because schedule-sync only mirrors directories the source has. Delete that directory once, or both DAGs run.
  • Stage 6. Every generated stage 6 runs fluid plan --check-sovereignty.
  • Missing overlays. --env X with no overlay is an error when the workspace's expected-environments declares X for the product. The contract's environments block only warns.

Tests (all new ones fail on origin/main, recorded)

Each file was run against an origin/main worktree (b5040ef) with the file copied in.

Test file On the branch On origin/main
tests/iac/test_iac_state_key_migration_moto.py (real tofu 1.12 against moto S3, AWS APIs and the S3 backend) 7 passed ImportError: state_migration
tests/iac/test_state_key_per_provider.py 21 passed ImportError: state_migration
tests/cli/test_policy_apply_provider_message.py 6 passed 6 failed (the TypeError)
tests/providers/test_gcp_sovereignty_fail_closed.py 20 passed ImportError: binding_region
tests/cli/test_apply_reports_to_command_center.py (loopback stub CC server) 9 passed 7 failed, 2 passed. The two that pass assert that nothing is sent, which is also true on main.
tests/cli/test_one_contract_two_clouds_pipeline.py 36 passed 33 failed, 3 passed

What the moto file pins:

  • The old release's state moves, and a plan after it shows +0 ~0 -0, for a kept and for a wiped workdir.
  • The old object is untouched after the move.
  • The second apply moves nothing.
  • The data-loss gate still refuses a destroy after the move.
  • The aws state at the old key is left alone by the gcp apply.
  • A state of two clouds is refused, with nothing written.
  • A new key that holds state is never overwritten.
  • fluid apply --dry-run leaves the bucket's key set unchanged and plans +0 ~0 -0 on the old key; the real apply after it moves the state.
  • The drift pass (check_state_drift, what fluid diff and verify --state-drift run) reads the old key while the move is pending: status checked, every managed resource, nothing written. (In the first round this line claimed a pin the file did not have: the test called reconcile_state_key(migrate=False) directly. A mutant pointing diff at the new key passed the suite; this test now fails it.)

For the three import-error files, behaviour was measured on origin/main with a probe:

  • aws and gcp keys: SAME.
  • The policy engine, for a no-region binding and for location.location=us-central1: valid, 0 findings.
  • The emitter, no-region strict: dataset US.
  • The emitter, us-central1 strict: emitted.
  • fluid validate, no region: rc 0.
  • fluid generate iac us-central1: rc 0.
  • The load target, no region: US.

Updated for the intended changes:

  • tests/iac/test_state_backend_env_default.py and tests/cli/test_diff_state_drift.py: the keys gain /aws/, and the move is stubbed there.
  • tests/forge/test_artifact_fanout_schedule.py: the env DAG path.
  • tests/cli/test_generate_schedule_fluid_apply.py: JENKINS_URL and ProgramData join NOT_PASSED, with their reasons.

CI: a Stage 1 step in iac-tests.yml runs the moto migration file with a coverage assert, as the state-drift step does. .secrets.baseline changes by one line number only, because that step shifts an existing entry.

Review round (tests fail before the fix, pass after)

Each file below was run on the first-round head (6eaa58d) and on origin/main (8d82828) with the file copied in.

Test Branch 6eaa58d origin/main
tests/iac/test_state_migration_safety.py (fault-injected raced / unverified / probe paths, the probe offline with real tofu, the shared-key guard, dry-run asks no move) 20 passed 12 failed, 5 passed, 3 skipped collection error (no state_migration)
tests/providers/test_gcp_locations_and_pubsub_placement.py 72 passed 27 failed, 45 passed collection error
tests/policy/test_sovereignty_enforcement_mode.py (unknown region under strict) 16 passed 1 failed 1 failed
tests/cli/test_apply_reports_to_command_center.py (3 new early-failure tests) 12 passed 3 failed 10 failed
tests/cli/test_one_contract_two_clouds_pipeline.py (the environments block warns, validate --env prod rc 0) 37 passed 2 failed 33 failed
tests/iac/test_iac_state_key_migration_moto.py (real tofu 1.12 on moto) 9 passed (4 min 40 s) the new dry-run test fails: the key set gained fluid/<id>/aws/terraform.tfstate (reproduced) ImportError

Mutants, each run against these tests (all killed):

  • the pre-copy re-check removed;
  • the post-copy verification disabled;
  • the probe reading an unrecorded backend;
  • the probe without -plugin-dir;
  • the move unpinned;
  • the dry-run migrating;
  • the shared-key guard removed, or blind;
  • strict warning on an unknown jurisdiction;
  • the GCP gap-fill removed;
  • the GCP hook not marking its placements;
  • the Pub/Sub policy dropped, or unread;
  • the engine's early identify removed;
  • the base-contract fallback removed;
  • the environments block refusing again.

fluid diff pointed at the new key (read = target) and the dry-run migrating were checked against the moto file: see the Gates section.

One design correction found while testing: the probe first used the apply's own .terraform/providers as its plugin directory. Under a plugin cache those entries link into the cache, and the probe's init broke the workdir ("no package for hashicorp/aws 5.100.0 cached in .terraform/providers", measured on moto; the class of opentofu/opentofu#3721). The probe now uses an empty directory.

Not changed, with the evidence:

  • _adopt_existing imports into state during a dry-run. Brownfield tofu import of a declared resource missing from state runs before the dry-run return, on main and on this branch. That is a state write in a plan-only run, but it predates this PR and the key set does not change. It is not touched here.
  • An interpolated region skips the emit-time check. A region written as an OpenTofu interpolation ('${"us-central1"}') is skipped by the GCP emit hook (references are not places). fluid validate and plan --check-sovereignty still refuse it under strict.
  • Event-trigger topics carry no storage policy. Topics from execution.trigger (ps.ensure_topic from schedule.py) carry no region, so they get no storage policy.

Proofs

  • Real tofu against moto: the migration and a plan with no changes after it, 7 of 7.
  • tofu validate on the emitted bronze gcp module (europe-west1, hashicorp/google 6.x): Success! The configuration is valid.
  • Demo bronze contract (fluid-demo-env copy), gcp overlay:
    • no region: fluid validate --env gcp rc 1, generate iac rc 1, plan --check-sovereignty rc 1;
    • us-central1: generate iac rc 1;
    • europe-west1: rc 0;
    • --env local and --env aws: rc 0, unchanged.
  • Demo silver with expected-environments: [local, aws, gcp] and no gcp overlay: validate and plan rc 1 (overlay_declared_but_missing); --env local and --env aws rc 0.
  • Stage 8 on the demo's bronze gcp bindings.json: this branch rc 0, origin/main rc 1 (TypeError).
  • Every CI system and complexity, rendered: each fluid plan command carries --check-sovereignty. Basic Tekton renders no plan command.

Not proven (no emulator can)

  • No real cloud was called.
    • The migration was run against moto's S3 only. A GCS backend migration, S3 with DynamoDB or lockfile locking, and IAM denials on a real bucket are not exercised.
    • A real GCP apply and a real Command Center (the stub implements only the four routes used) are not exercised.
  • Concurrency. The S3 backend forge-cli emits has no locking. A concurrent aws-* and gcp-* first run of the same product cannot corrupt each other any more (different keys). Two runs of the same provider racing during the move are narrowed by the second check, but not excluded.
  • US / EU jurisdictions come from BigQuery's and Cloud Storage's documented multi-region definitions, not from a dataset.
  • Legacy DAGs. The one-time cleanup of an env-bound DAG synced by an earlier release to <product>/ is documented, not automated.
  • The probe's recorded-backend read. That OpenTofu records the backend before its provider step is measured on 1.12 (local and S3 backends), not documented. It is checked on every read, and a plain init is the fallback.
  • Plan-only IAM. A dry-run no longer writes state objects, but a real plan-only role with read-only state access was not exercised (moto does not enforce IAM).
  • GCP locations. The table is Google's lists as read on 2026-09-28; a region Google adds later resolves Unknown, which strict refuses until it is named in allowedRegions.
  • Report scope. The CC report covers the OpenTofu engine fully; a native (local) apply is reported with status and timings only. The CC-side "one product, a deployment per cloud" view is a separate CC change.

Gates

On cc9784c, python 3.12, black 24.10.0:

  • Lint: ruff check fluid_build/ tests/ clean. black --check clean (1776 files). python scripts/mypy_strict.py no issues. lint-imports 6 kept, 0 broken.
  • Hygiene: add_license_headers.py added nothing (0 added, 1809 skipped). detect-secrets-hook --baseline .secrets.baseline on the changed files: rc 0. The merge of origin/main moved one baseline line number (the iac-tests.yml AWS_SECRET_ACCESS_KEY: test entry, 222 → 231).
  • Gate tests: pytest -p no:randomly -n 4 with an empty HOME and a temp cwd, on tests/build_runners tests/providers tests/cli tests/iac tests/forge tests/policy tests/test_sovereignty.py (the moto file deselected and run on its own): 5723 passed, 0 failed, 401 skipped, 1 xfailed, 1 xpassed. The moto file: 9 passed (real tofu 1.12, 4 min 40 s, sequential as CI runs it).
  • Whole suite: not re-run in this round; the first round's run is above (19308 passed, 8 environmental failures).

Security review

The security-review skill loaded against the session's working directory, a different repository, in both rounds. So its procedure was run by hand on this branch's diff. Outcome: no high- or medium-confidence findings.

First round:

  • The credential. It is sent only as a header, to the endpoint fluid publish already uses. The reporter's SSRF host gate now runs before the organization lookup sends it. Header values are never logged, and CommandCenterConfig.__repr__ prints header names only.
  • The payload. Ids, versions, hash, state location (whose bucket grammar refuses credentials), resource addresses, counts and timings, and the typed event name of a failure. No tofu output, no environment values.
  • Subprocess calls. Every tofu call is an argv list. Backend blocks are built by the validated parse_backend. Scratch directories come from tempfile. A provider name must be one [a-z][a-z0-9_-]* key segment.
  • YAML. The workspace file goes through the repository's safe YAML loader.
  • Hardening. A failed tofu state pull is reported from stderr only, because its stdout is the state document. _logging serialization is unchanged.

This round:

  • -plugin-dir. An argv element naming a tempfile directory.
  • Scratch modules. Written with json.dumps from parse_backend blocks and the plugin's static pins.
  • Reading state. The shared-key guard reads the state in memory; only provider names and counts leave it, and its error names the bucket and key only.
  • records_backend. Reads only the scratch directory's own .terraform/.
  • CC fallback. The base-contract fallback reads the contract named on the command line through the safe loader and sends the same identity facts, minus the hash.
  • Sovereignty. Only stricter.
  • Considered, below the threshold: the interpolated-region skip listed under "Not changed" (same trust domain: the contract author writes the policy).

…fails closed on sovereignty

Platform safety for deploying one contract to two clouds through --env
overlays, from the offline verification of the demo products on 0.16.5.

State (F1). The default remote state key names the provider:
fluid/<id>/<provider>/terraform.tfstate (GCS prefix fluid/<id>/<provider>),
for FLUID_STATE_BACKEND bucket-only specs and packaging contracts. The aws
and the gcp apply of one contract shared fluid/<id>/terraform.tfstate, so
each plan read the other cloud's resources as orphans to destroy. The
first apply after the upgrade moves the old key's state with OpenTofu's
own `tofu init -force-copy` (implies -migrate-state), only when the new key
holds no state and the old one holds this provider's resources; another
provider's state is left alone, a state of two clouds is refused, the old
object is never deleted, and the copy is verified by reading it back.
fluid diff / verify --state-drift read the old key while the move is
pending. The shared legacy key and explicit keys are unchanged.

Stage 8 (F2). info()/warn()/error() take logger and message positionally
only, and a payload key named like an envelope key is kept as extra_<key>:
the GCP policy result's "message" raised a TypeError on every gcp build.

Sovereignty (F5). Under enforcementMode strict a binding on aws, gcp or
azure that names no region is refused (advisory warns, audit logs); a GCP
region given as location.location is checked; BigQuery's US and EU
multi-regions resolve to their jurisdictions. The GCP IaC plugin checks
every emitted location and the GCP provider every planned one (the way
AwsProvider.plan does), so fluid apply and fluid generate iac refuse an
out-of-jurisdiction region, including a provider-default region a resource
inherits. The BigQuery load runs where the table is, never a guessed US.

Command Center (F10). fluid apply registers each run at
POST /api/v1/executions and closes it at PATCH, through the existing
CommandCenterReporter, with the credential and organization fluid publish
uses: product id, contract version and hash, environment, provider, mode,
state location, resource addresses, change counts, timings and status.
Best effort: an outage or refusal never changes the exit code, and no
secret, header or tofu output is sent. FLUID_COMMAND_CENTER_ENABLED=false
turns it off.

Small items (F11). An env's Airflow DAG id is <product>__<env>__<build>
and its schedule directory <product>__<env>; every generated stage 6 runs
plan --check-sovereignty; --env X with no overlay is an error when the
contract's environments block or the workspace's expected-environments
declares X for the product (dev and the base's own platform stay the base),
and the warning otherwise names what the base binds to.

Borrowed: OpenTofu backend migration (meta_backend_migrate.go), Terragrunt
backend migrate, WebbPulse/webbpulse-python#129 (rename colliding log keys),
the OpenLineage client's log-don't-fail posture, dbt's unknown-target error.
@github-actions github-actions Bot added provider Changes or requests related to providers ci Continuous integration and automation changes security Security-related changes or reports cli Changes to the CLI surface or implementation tests Test coverage or test infrastructure changes needs-docs Pull request needs a linked docs update or justification labels Sep 28, 2026
@github-actions

Copy link
Copy Markdown

📄 Documentation Reminder

This PR appears to be missing a documentation reference. Our docs live in a separate repo.

Please update the PR description with one of:

  • Link a docs PR — check the "Docs PR linked" box and paste the URL
  • Mark as no docs needed — check "No docs needed" with a justification
  • Acknowledge docs TODO — check "Docs TODO" and create the docs PR before merge

See the Contributing Guide for details.

Comment thread tests/cli/test_one_contract_two_clouds_pipeline.py Fixed
# Conflicts:
#	.secrets.baseline
…ils closed on an unplaceable location, and a Pub/Sub topic keeps its region

Review round on the one-contract-two-clouds platform safety change.

State. `fluid apply --dry-run` (the generated Jenkins default APPLY_MODE)
ran the state move and wrote the new key; measured with real tofu against
moto, the key set grew by fluid/<id>/aws/terraform.tfstate. A dry-run now
asks the move not to run and plans against the old key while it is pending,
the read-only path fluid diff already takes; the first real apply moves it.
The move's probe no longer installs anything: OpenTofu's init installs the
providers the state names (the latest hashicorp/aws for an aws state), and
the probe ran on every apply whose new key was empty. It now inits with
-plugin-dir on an empty directory and pulls from the backend that init is
shown to have recorded, falling back to a plain init. The one-time copy
installs what the old state names at the plugin's own pins. Using the
apply's .terraform/providers as the plugin directory was tried and broke
the workdir under a plugin cache, so it is not used. A remote key that does
not name the provider (the shared fluid/terraform.tfstate a bucket-only
--state-backend gives a contract without packaging) is refused when it
holds another cloud's resources (state_shared_with_another_provider).

Sovereignty. Under strict, a cloud-region binding (aws/gcp/azure) whose
jurisdiction cannot be resolved is refused unless the region is named in
allowedRegions; it was a warning, so me-central2, northamerica-south1 and
the asia multi-region passed validate and generate iac on an EU-only
contract. The nine GCP regions the vendored table lacks, and the dual-regions
within one jurisdiction (EUR4, NAM4, ASIA1), are mapped from Google's own
location lists, in source, not in the ODbL csv. A gcp pubsub_topic
binding's region becomes message_storage_policy.allowed_persistence_regions
(hashicorp/google), which the GCP hook checks; it was dropped.

Command Center. A run refused before the engine registered it now carries
its product: the engine identifies it as soon as the contract and provider
are known, and a run refused earlier is described from the base contract.

Overlays. Only the workspace's expected-environments refuses --env X with
no overlay; the contract's own environments block warns again, as on
0.16.5, since forge-cli applies nothing from it.

Tests: the raced and unverified branches of the move under fault
injection, the probe installing nothing (real tofu, no registry reachable),
a dry-run leaving the bucket's key set unchanged and the drift pass reading
the old key (real tofu against moto), and each of the fixes above.
…ed at their dataset's multi-region, and plan --check-sovereignty checks what apply emits

Cloud KMS names BigQuery's EU and US multi-regions europe and us, and Data
Catalog names them eu and us. The GCP sovereignty hook checked those ids as
places, so with the governance branch's CMEK and policy tags an EU-only strict
contract with an EU dataset was refused at apply (Region 'europe' not in
allowed regions list, jurisdiction Unknown), after validate --strict and
plan --check-sovereignty passed. resource_placements now places both at the
multi-region, spelled as the emitted dataset spells it.

fluid plan --check-sovereignty ran only the policy engine over the bindings
for gcp (the provider had no hook), so a placement the planner or the
emitter derived passed stage 6 and was refused at stage 7. GcpProvider now
has validate_sovereignty: the engine's contract checks plus the planned
actions' and the emitted resources' placements, through the same engine, and
it returns the error-severity findings. GcpIacPlugin.emit takes
enforce_sovereignty=False for it. A hook that gives no verdict is now named
as such in the fallback line, not as a missing hook.
…' clouds, not by OpenTofu's built-in provider

A gcp product whose table's partitions expire keeps a terraform_data beside
the table (the trigger that replaces it when its partitioning changes), under
provider["terraform.io/builtin/terraform"]. state_sources counted
builtin/terraform as a provider no plugin emits, so classify called every such
legacy state ambiguous and refused the move, and with it every apply, dry-run
and diff of the product (measured on the integration with the governance
branch). The built-in provider ships inside OpenTofu and names no cloud; it is
now left out, and a state of nothing but built-in resources is still not
guessed.
…at each build did, and why the run failed

fluid apply --mode amend-and-build reported only the infrastructure: the
PATCH carried planned and applied changes and resources in phase apply, so a
silver run that loaded 28 rows said nothing of it, and a bronze run whose load
failed was closed failed with no error event and no message. The builds
return an exit code, not an exception, and run_builds_from_args never touched
the run report.

run_builds_from_args now records each build on the open report: its id, its
status (succeeded, failed, skipped) and, from the run record it wrote, the run
id, the table facets.bigquery_load names, the landed destinations and the
rows. The phase is build, and a failed build phase names its reason:
build_failed:<build id>, build_not_found:<build id> or builds_all_skipped.
Nothing is read when no apply report is open.
…tires the DAG it replaced

An env's DAGs moved from <product>/ (dag id <product>__<build>) to
<product>__<env>/ (<product>__<env>__<build>), and --delete-scope product
mirrors only the new directory. On the demo lab the first sync after the
upgrade left bronze.customer_subscriptions/ingest_subscriptions_dag.py beside
bronze.customer_subscriptions__aws/, both FLUID_ENV_NAME 'aws' on the same
cron, and the lab's Airflow unpauses a DAG as it parses it: two fluid apply
--env aws of one product against one state at the same minute.

Under --delete-scope product the file and git+ssh transports, whose
destination is readable here, now retire the old directory's DAGs rendered for
the same product and env with the old id, read from each file's syntax tree
(never imported): an rsync --delete from an empty directory filtered to
exactly those files, after the new directory synced. An env-less DAG, another
env's, and any file that is not a rendered DAG stay. The other transports
print the one step left to do, and the report records each superseded scope
and whether its old DAGs were retired.
…ld_runners records builds without importing the CLI

The previous commit read the run report through a function-local import of
fluid_build.cli._apply_cc_report from build_runners, the reverse edge
tests/observability/test_import_hygiene.py forbids. The context variable
now lives in observability/apply_run.py, which build_runners already
depends on; _apply_cc_report sets and reads it there.
fas89 added a commit that referenced this pull request Sep 28, 2026
…p-parity

The fix round on #676: the GCP sovereignty hook places a KMS key ring and a
Data Catalog taxonomy at their dataset's multi-region, and plan
--check-sovereignty runs the placements apply runs; the state-key migration
ignores OpenTofu's built-in provider; a build-augmented apply reports its
builds to the Command Center; schedule-sync retires the DAG an env's
directory replaced. Also main 0.16.6, merged into the branch.

Merged cleanly. build_runners/base.py and iac/providers/gcp.py auto-merged
with #673's and #674's changes to the same files.
fas89 added a commit that referenced this pull request Sep 28, 2026
fas89 added a commit that referenced this pull request Sep 28, 2026
@fas89
fas89 merged commit 0e6d195 into main Sep 28, 2026
40 checks passed
fas89 added a commit that referenced this pull request Sep 28, 2026
…#676's load target now says

#676 made bigquery_load_target name no location for a binding that declares
none, so the load job runs where the table itself is. The embedded-SQL target
here filled the IaC's default US back in; it now passes the undeclared
location through to the load and keeps US only for the sovereignty checks,
which reason about the dataset the IaC creates. A test pins that the load
names no location. Also: a staged file already gone is Path.unlink's
missing_ok, not an empty except.
fas89 added a commit that referenced this pull request Sep 28, 2026
…rify counts and checks masking on BigQuery (#673)

* feat(build): an embedded-SQL build on DuckDB reads BigQuery upstreams and lands in BigQuery; fluid verify counts the table and checks its masked columns

On 0.16.5 a silver or gold product could not build on GCP: a consumes[]
entry whose upstream is a gcp bigquery_table was UnreadableBindingError, and
with parameters.inputs the result was written to a local file named after
the overlay's gs:// staging path while the table stayed empty and the build
reported success.

Reads: the upstream table is named by the IaC's own rule
(_bigquery_load.bigquery_load_target) and read with python-bigquery's Arrow
pages into a staged Parquet file the view reads; a BigQuery TIMESTAMP reads
as the UTC wall clock (the DuckDB BigQuery extension's documented mapping),
as the same SQL reads it on local and aws. The staged copy is removed after
the build.

Landing: a first expose bound to a bigquery_table is staged as Parquet and
loaded by the acquisition runner's own load_file (WRITE_TRUNCATE,
CREATE_NEVER, the table's schema). A gs:// or other non-S3 URI, a GCS
bucket, an Azure/Snowflake/Databricks binding, a second BigQuery output and
a BigQuery landing without inline SQL are refused before the SQL runs.

Load: TIMESTAMP columns held as naive Parquet timestamps are sent
UTC-adjusted (python-bigquery's TIMESTAMP -> timestamp(us, UTC) mapping), so
BigQuery does not read them as DATETIME. With BIGQUERY_EMULATOR_HOST set the
client is anonymous, so no ADC token reaches an emulator; a job with no
outputRows is checked by COUNT(*) and never taken as zero or as success.

Verify: one GoogleSQL count query adds row_count (held to the acquisition
run's facets.bigquery_load, by the Athena verifier's rules) and masking
(COUNTIF ... NOT REGEXP_CONTAINS against a bound, anchored shape), each
CRITICAL under --strict. The num_rows None crash is fixed, and verify names
the table the way the load does.

pyarrow is declared in the gcp and dev extras; the heavy emulated lane
installs local and requires the new emulator chain test.

* fix(build): the BigQuery embedded-SQL path holds its reads and load to sovereignty, never loads into its own input, refuses outputs it cannot land, and records its load for verify

Review of the first push found seven defects; each was reproduced on the
branch before it was fixed.

- CI was red: three of the new unit tests needed google-cloud-bigquery or
  google-auth, which the .[dev,local] unit lanes do not install. The two
  resolve tests now use the fake BigQuery, and AnonymousCredentials goes
  through a seam (_bigquery_load._anonymous_credentials) the fake replaces.
  In a .[dev,local] venv the old file failed 3 tests, and the new one passes.
- Sovereignty failed open. An EU-only silver whose gcp overlay named no
  region read europe-west1 and loaded into US, the IaC default, with rc 0.
  Before anything is read, a contract with a sovereignty block now has every
  BigQuery binding the build reads or loads checked by fluid validate's rules
  (refuse_sovereignty_breach). The binding must name a region. The landing
  must be outside deniedRegions, inside allowedRegions and in the declared
  jurisdiction. With no crossBorderTransfer, every input must be in the
  landing's jurisdiction. enforcementMode sets the severity. BigQuery's EU
  and US multi-regions map to EU and US, as Google's in:eu-locations and
  in:us-locations value groups list them.
- A landing that resolved to one of its own inputs WRITE_TRUNCATEd another
  product's table, and bronze went from 5 rows to 3. This is now refused on
  what the two resolve to, the way Dagster's _validate_self_deps compares
  asset keys: BigQuery table ids are compared case-insensitively, and a
  project left to the client matches any project. An S3 landing inside a
  prefix an input reads is refused too.
- A further output bound to GCS, S3 or another cloud was planned and then
  silently not landed. It is now refused. A further local expose or output
  port is still not written, and the build now prints a warning saying so.
- verify looked up ".dataset.table" when the binding named no project. It
  now uses the client's project, as the load does, and with no project at
  all it is an error naming the table.
- silver and gold row_count was never compared with the load. An
  embedded-SQL BigQuery load now writes the acquisition runner's run record
  (facets.bigquery_load, full_refresh, rows_from: write), and the BigQuery
  verifier reads embedded-SQL runs. The emulator chain now holds silver and
  gold to their loads.
- The failed-load and short-load tests passed on main and under a mutant
  that breaks staging. They now assert the load was attempted, the error
  was printed, and the table was left unchanged.

* fix(build): an undeclared BigQuery region loads where the table is, as #676's load target now says

#676 made bigquery_load_target name no location for a binding that declares
none, so the load job runs where the table itself is. The embedded-SQL target
here filled the IaC's default US back in; it now passes the undeclared
location through to the load and keeps US only for the sovereignty checks,
which reason about the dataset the IaC creates. A test pins that the load
names no location. Also: a staged file already gone is Path.unlink's
missing_ok, not an empty except.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci Continuous integration and automation changes cli Changes to the CLI surface or implementation needs-docs Pull request needs a linked docs update or justification provider Changes or requests related to providers security Security-related changes or reports tests Test coverage or test infrastructure changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant