fix(platform): one contract on aws and gcp keeps two states, and GCP fails closed on sovereignty - #676
Merged
Conversation
…fails closed on sovereignty Platform safety for deploying one contract to two clouds through --env overlays, from the offline verification of the demo products on 0.16.5. State (F1). The default remote state key names the provider: fluid/<id>/<provider>/terraform.tfstate (GCS prefix fluid/<id>/<provider>), for FLUID_STATE_BACKEND bucket-only specs and packaging contracts. The aws and the gcp apply of one contract shared fluid/<id>/terraform.tfstate, so each plan read the other cloud's resources as orphans to destroy. The first apply after the upgrade moves the old key's state with OpenTofu's own `tofu init -force-copy` (implies -migrate-state), only when the new key holds no state and the old one holds this provider's resources; another provider's state is left alone, a state of two clouds is refused, the old object is never deleted, and the copy is verified by reading it back. fluid diff / verify --state-drift read the old key while the move is pending. The shared legacy key and explicit keys are unchanged. Stage 8 (F2). info()/warn()/error() take logger and message positionally only, and a payload key named like an envelope key is kept as extra_<key>: the GCP policy result's "message" raised a TypeError on every gcp build. Sovereignty (F5). Under enforcementMode strict a binding on aws, gcp or azure that names no region is refused (advisory warns, audit logs); a GCP region given as location.location is checked; BigQuery's US and EU multi-regions resolve to their jurisdictions. The GCP IaC plugin checks every emitted location and the GCP provider every planned one (the way AwsProvider.plan does), so fluid apply and fluid generate iac refuse an out-of-jurisdiction region, including a provider-default region a resource inherits. The BigQuery load runs where the table is, never a guessed US. Command Center (F10). fluid apply registers each run at POST /api/v1/executions and closes it at PATCH, through the existing CommandCenterReporter, with the credential and organization fluid publish uses: product id, contract version and hash, environment, provider, mode, state location, resource addresses, change counts, timings and status. Best effort: an outage or refusal never changes the exit code, and no secret, header or tofu output is sent. FLUID_COMMAND_CENTER_ENABLED=false turns it off. Small items (F11). An env's Airflow DAG id is <product>__<env>__<build> and its schedule directory <product>__<env>; every generated stage 6 runs plan --check-sovereignty; --env X with no overlay is an error when the contract's environments block or the workspace's expected-environments declares X for the product (dev and the base's own platform stay the base), and the warning otherwise names what the base binds to. Borrowed: OpenTofu backend migration (meta_backend_migrate.go), Terragrunt backend migrate, WebbPulse/webbpulse-python#129 (rename colliding log keys), the OpenLineage client's log-don't-fail posture, dbt's unknown-target error.
📄 Documentation ReminderThis PR appears to be missing a documentation reference. Our docs live in a separate repo. Please update the PR description with one of:
See the Contributing Guide for details. |
# Conflicts: # .secrets.baseline
…ils closed on an unplaceable location, and a Pub/Sub topic keeps its region Review round on the one-contract-two-clouds platform safety change. State. `fluid apply --dry-run` (the generated Jenkins default APPLY_MODE) ran the state move and wrote the new key; measured with real tofu against moto, the key set grew by fluid/<id>/aws/terraform.tfstate. A dry-run now asks the move not to run and plans against the old key while it is pending, the read-only path fluid diff already takes; the first real apply moves it. The move's probe no longer installs anything: OpenTofu's init installs the providers the state names (the latest hashicorp/aws for an aws state), and the probe ran on every apply whose new key was empty. It now inits with -plugin-dir on an empty directory and pulls from the backend that init is shown to have recorded, falling back to a plain init. The one-time copy installs what the old state names at the plugin's own pins. Using the apply's .terraform/providers as the plugin directory was tried and broke the workdir under a plugin cache, so it is not used. A remote key that does not name the provider (the shared fluid/terraform.tfstate a bucket-only --state-backend gives a contract without packaging) is refused when it holds another cloud's resources (state_shared_with_another_provider). Sovereignty. Under strict, a cloud-region binding (aws/gcp/azure) whose jurisdiction cannot be resolved is refused unless the region is named in allowedRegions; it was a warning, so me-central2, northamerica-south1 and the asia multi-region passed validate and generate iac on an EU-only contract. The nine GCP regions the vendored table lacks, and the dual-regions within one jurisdiction (EUR4, NAM4, ASIA1), are mapped from Google's own location lists, in source, not in the ODbL csv. A gcp pubsub_topic binding's region becomes message_storage_policy.allowed_persistence_regions (hashicorp/google), which the GCP hook checks; it was dropped. Command Center. A run refused before the engine registered it now carries its product: the engine identifies it as soon as the contract and provider are known, and a run refused earlier is described from the base contract. Overlays. Only the workspace's expected-environments refuses --env X with no overlay; the contract's own environments block warns again, as on 0.16.5, since forge-cli applies nothing from it. Tests: the raced and unverified branches of the move under fault injection, the probe installing nothing (real tofu, no registry reachable), a dry-run leaving the bucket's key set unchanged and the drift pass reading the old key (real tofu against moto), and each of the fixes above.
…ed at their dataset's multi-region, and plan --check-sovereignty checks what apply emits Cloud KMS names BigQuery's EU and US multi-regions europe and us, and Data Catalog names them eu and us. The GCP sovereignty hook checked those ids as places, so with the governance branch's CMEK and policy tags an EU-only strict contract with an EU dataset was refused at apply (Region 'europe' not in allowed regions list, jurisdiction Unknown), after validate --strict and plan --check-sovereignty passed. resource_placements now places both at the multi-region, spelled as the emitted dataset spells it. fluid plan --check-sovereignty ran only the policy engine over the bindings for gcp (the provider had no hook), so a placement the planner or the emitter derived passed stage 6 and was refused at stage 7. GcpProvider now has validate_sovereignty: the engine's contract checks plus the planned actions' and the emitted resources' placements, through the same engine, and it returns the error-severity findings. GcpIacPlugin.emit takes enforce_sovereignty=False for it. A hook that gives no verdict is now named as such in the fallback line, not as a missing hook.
…' clouds, not by OpenTofu's built-in provider A gcp product whose table's partitions expire keeps a terraform_data beside the table (the trigger that replaces it when its partitioning changes), under provider["terraform.io/builtin/terraform"]. state_sources counted builtin/terraform as a provider no plugin emits, so classify called every such legacy state ambiguous and refused the move, and with it every apply, dry-run and diff of the product (measured on the integration with the governance branch). The built-in provider ships inside OpenTofu and names no cloud; it is now left out, and a state of nothing but built-in resources is still not guessed.
…at each build did, and why the run failed fluid apply --mode amend-and-build reported only the infrastructure: the PATCH carried planned and applied changes and resources in phase apply, so a silver run that loaded 28 rows said nothing of it, and a bronze run whose load failed was closed failed with no error event and no message. The builds return an exit code, not an exception, and run_builds_from_args never touched the run report. run_builds_from_args now records each build on the open report: its id, its status (succeeded, failed, skipped) and, from the run record it wrote, the run id, the table facets.bigquery_load names, the landed destinations and the rows. The phase is build, and a failed build phase names its reason: build_failed:<build id>, build_not_found:<build id> or builds_all_skipped. Nothing is read when no apply report is open.
…tires the DAG it replaced An env's DAGs moved from <product>/ (dag id <product>__<build>) to <product>__<env>/ (<product>__<env>__<build>), and --delete-scope product mirrors only the new directory. On the demo lab the first sync after the upgrade left bronze.customer_subscriptions/ingest_subscriptions_dag.py beside bronze.customer_subscriptions__aws/, both FLUID_ENV_NAME 'aws' on the same cron, and the lab's Airflow unpauses a DAG as it parses it: two fluid apply --env aws of one product against one state at the same minute. Under --delete-scope product the file and git+ssh transports, whose destination is readable here, now retire the old directory's DAGs rendered for the same product and env with the old id, read from each file's syntax tree (never imported): an rsync --delete from an empty directory filtered to exactly those files, after the new directory synced. An env-less DAG, another env's, and any file that is not a rendered DAG stay. The other transports print the one step left to do, and the report records each superseded scope and whether its old DAGs were retired.
…out retires each env's old DAG
…ld_runners records builds without importing the CLI The previous commit read the run report through a function-local import of fluid_build.cli._apply_cc_report from build_runners, the reverse edge tests/observability/test_import_hygiene.py forbids. The context variable now lives in observability/apply_run.py, which build_runners already depends on; _apply_cc_report sets and reads it there.
fas89
added a commit
that referenced
this pull request
Sep 28, 2026
…p-parity The fix round on #676: the GCP sovereignty hook places a KMS key ring and a Data Catalog taxonomy at their dataset's multi-region, and plan --check-sovereignty runs the placements apply runs; the state-key migration ignores OpenTofu's built-in provider; a build-augmented apply reports its builds to the Command Center; schedule-sync retires the DAG an env's directory replaced. Also main 0.16.6, merged into the branch. Merged cleanly. build_runners/base.py and iac/providers/gcp.py auto-merged with #673's and #674's changes to the same files.
fas89
added a commit
that referenced
this pull request
Sep 28, 2026
fas89
added a commit
that referenced
this pull request
Sep 28, 2026
…verify client answers #673's row count
fas89
added a commit
that referenced
this pull request
Sep 28, 2026
…#676's load target now says #676 made bigquery_load_target name no location for a binding that declares none, so the load job runs where the table itself is. The embedded-SQL target here filled the IaC's default US back in; it now passes the undeclared location through to the load and keeps US only for the sovereignty checks, which reason about the dataset the IaC creates. A test pins that the load names no location. Also: a staged file already gone is Path.unlink's missing_ok, not an empty except.
fas89
added a commit
that referenced
this pull request
Sep 28, 2026
…rify counts and checks masking on BigQuery (#673) * feat(build): an embedded-SQL build on DuckDB reads BigQuery upstreams and lands in BigQuery; fluid verify counts the table and checks its masked columns On 0.16.5 a silver or gold product could not build on GCP: a consumes[] entry whose upstream is a gcp bigquery_table was UnreadableBindingError, and with parameters.inputs the result was written to a local file named after the overlay's gs:// staging path while the table stayed empty and the build reported success. Reads: the upstream table is named by the IaC's own rule (_bigquery_load.bigquery_load_target) and read with python-bigquery's Arrow pages into a staged Parquet file the view reads; a BigQuery TIMESTAMP reads as the UTC wall clock (the DuckDB BigQuery extension's documented mapping), as the same SQL reads it on local and aws. The staged copy is removed after the build. Landing: a first expose bound to a bigquery_table is staged as Parquet and loaded by the acquisition runner's own load_file (WRITE_TRUNCATE, CREATE_NEVER, the table's schema). A gs:// or other non-S3 URI, a GCS bucket, an Azure/Snowflake/Databricks binding, a second BigQuery output and a BigQuery landing without inline SQL are refused before the SQL runs. Load: TIMESTAMP columns held as naive Parquet timestamps are sent UTC-adjusted (python-bigquery's TIMESTAMP -> timestamp(us, UTC) mapping), so BigQuery does not read them as DATETIME. With BIGQUERY_EMULATOR_HOST set the client is anonymous, so no ADC token reaches an emulator; a job with no outputRows is checked by COUNT(*) and never taken as zero or as success. Verify: one GoogleSQL count query adds row_count (held to the acquisition run's facets.bigquery_load, by the Athena verifier's rules) and masking (COUNTIF ... NOT REGEXP_CONTAINS against a bound, anchored shape), each CRITICAL under --strict. The num_rows None crash is fixed, and verify names the table the way the load does. pyarrow is declared in the gcp and dev extras; the heavy emulated lane installs local and requires the new emulator chain test. * fix(build): the BigQuery embedded-SQL path holds its reads and load to sovereignty, never loads into its own input, refuses outputs it cannot land, and records its load for verify Review of the first push found seven defects; each was reproduced on the branch before it was fixed. - CI was red: three of the new unit tests needed google-cloud-bigquery or google-auth, which the .[dev,local] unit lanes do not install. The two resolve tests now use the fake BigQuery, and AnonymousCredentials goes through a seam (_bigquery_load._anonymous_credentials) the fake replaces. In a .[dev,local] venv the old file failed 3 tests, and the new one passes. - Sovereignty failed open. An EU-only silver whose gcp overlay named no region read europe-west1 and loaded into US, the IaC default, with rc 0. Before anything is read, a contract with a sovereignty block now has every BigQuery binding the build reads or loads checked by fluid validate's rules (refuse_sovereignty_breach). The binding must name a region. The landing must be outside deniedRegions, inside allowedRegions and in the declared jurisdiction. With no crossBorderTransfer, every input must be in the landing's jurisdiction. enforcementMode sets the severity. BigQuery's EU and US multi-regions map to EU and US, as Google's in:eu-locations and in:us-locations value groups list them. - A landing that resolved to one of its own inputs WRITE_TRUNCATEd another product's table, and bronze went from 5 rows to 3. This is now refused on what the two resolve to, the way Dagster's _validate_self_deps compares asset keys: BigQuery table ids are compared case-insensitively, and a project left to the client matches any project. An S3 landing inside a prefix an input reads is refused too. - A further output bound to GCS, S3 or another cloud was planned and then silently not landed. It is now refused. A further local expose or output port is still not written, and the build now prints a warning saying so. - verify looked up ".dataset.table" when the binding named no project. It now uses the client's project, as the load does, and with no project at all it is an error naming the table. - silver and gold row_count was never compared with the load. An embedded-SQL BigQuery load now writes the acquisition runner's run record (facets.bigquery_load, full_refresh, rows_from: write), and the BigQuery verifier reads embedded-SQL runs. The emulator chain now holds silver and gold to their loads. - The failed-load and short-load tests passed on main and under a mutant that breaks staging. They now assert the load was attempted, the error was printed, and the table was left unchanged. * fix(build): an undeclared BigQuery region loads where the table is, as #676's load target now says #676 made bigquery_load_target name no location for a binding that declares none, so the load job runs where the table itself is. The embedded-SQL target here filled the IaC's default US back in; it now passes the undeclared location through to the load and keeps US only for the sovereignty checks, which reason about the dataset the IaC creates. A test pins that the load names no location. Also: a staged file already gone is Path.unlink's missing_ok, not an empty except.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Platform safety for deploying one contract to aws and gcp through
--envoverlays (PR 1 of 3: state, stage 8, sovereignty, Command Center run reports, pipeline collisions). The embedded-SQL-over-BigQuery work and the GCP governance emitters are the other two PRs; this one keeps its edits iniac/providers/gcp.py,providers/gcp/provider.pyand the policy engine local to sovereignty.The gaps, with the evidence
An offline verification of the three demo products on 0.16.5 (no real cloud) found:
resolve_state_targetwithFLUID_STATE_BACKEND=s3://fluid-demo-lab-state-…gave bothfluid/bronze.customer_subscriptions/terraform.tfstate. The gcp plan would read the aws resources as orphans to destroy, and--allow-data-losswould destroy them. Re-measured on origin/main in this PR:SAME.fluid policy-apply … --mode enforceexit 1,TypeError: info() got multiple values for argument 'message'(cli/policy_apply.py:136), on all three products. Re-measured on origin/main with the demo's bronze gcpbindings.json: rc 1, same TypeError.fluid validaterc 0,plan --check-sovereigntyprinted PASS, and the dataset landed inUS.us-central1was refused by validate only;fluid generate iacemitted it with rc 0 (no GCP provider hook). A region given aslocation.location, which the GCP emitter reads, was never checked. The BigQuery load defaulted its job location toUS.get_reporter(cli/bootstrap.py:81) has no callers; a scratch CC test recorded 0pipeline_executionsrows.dag_id=<product>__<build>in the sameschedule/<product>/directory, whichschedule-sync --delete-scope productmirrors with delete.--check-sovereigntyin any rendered CI system; a strict violation first failed at stage 7.--env gcpwith no gcp overlay quietly runs the local baseDesign
F1: provider-keyed state, moved with OpenTofu's own migration
iac/backend.py: every per-contract default key gains the provider,fluid/<id>/<provider>/terraform.tfstate(GCS prefixfluid/<id>/<provider>). That covers theFLUID_STATE_BACKENDbucket-only default and the packaging-contract default. The shared legacy keyfluid/terraform.tfstateand any explicit key or prefix are unchanged.parse_backendwithoutprovideris byte-for-byte what it was.legacy_default_backendnames where the previous release kept the state.iac/state_migration.py(new): after the workdir's owntofu init,tofu state pullreads the new key. When it holds state, nothing else happens. When it holds none, a provider-less scratch module is initialised on the old key and pulled. The state is attributed by its resources' ownprovider["registry…/<ns>/<type>"]addresses, checked against each IaC plugin'srequired_providers:state_migration_ambiguous, naming both keys..terraform/recorded the old backend. Its module is rewritten to the new one, andtofu init -force-copy(which implies-migrate-state) copies the state. This is the backend-reconfiguration path ofterraform init -migrate-state: it locks both states where the backend supports locking, and it leaves the source untouched.-force-copyoverwrites a non-empty destination, so the new key is checked twice: before deciding, and again just before the copy. A different state appearing between the two checks isstate_migration_raced..terraform/recorded the old key is re-initialised with-reconfigure(a plain init stops at "Backend configuration changed"). A wiped CI workspace needs nothing.fluid diff/verify --state-driftandfluid apply --dry-runare read-only. While the move is pending they plan against the old key, so a stage 5 drift gate, or a plan-only stage 7, before the first real upgraded apply still sees the real state and writes nothing. Only a real apply moves it.{"terraform": {}}module beside ahashicorp/nullstate installed the latesthashicorp/null). The probe runs on every apply whose new key is empty, so it inits with-plugin-diron an empty directory: the init stops at its provider step after recording the old backend, and the state is pulled from exactly that recorded backend (checked withrecords_backend, never assumed). If a future OpenTofu did not record it first, the probe falls back to a plain init. The one-time copy installs what the old state names at the IaC plugin's own pins (~> 5.0, not the latest).fluid/terraform.tfstatethat a bucket-only--state-backendgives a contract without packaging, or an explicit key reused for both clouds) is read once; if it holds another cloud's resources, the apply refuses withstate_shared_with_another_providerinstead of planning their destruction.F2: the log helpers take
loggerandmessagepositionally onlyinfo/warn/error(logger, message, /, **payload)(PEP 570). A payload key that names an envelope key (time,level,name,message) is kept asextra_<key>, so the GCP policy result's own sentence survives and the event keeps its name. This fixes it for every provider that returns amessage, and for every other**dictcall site (contract_tests, the apply engine).F5: sovereignty fails closed, and GCP gets a provider hook
policy/sovereignty.py):aws,gcp,azure) with a sovereignty block and no region is a finding atseverity_for(mode). Strict refuses, advisory warns, audit logs.localbindings, and platforms whose region is not on the binding, keep the old behaviour.binding_regionreadslocation.locationfor GCP, as the emitter does.US/EU(the BigQuery and Cloud Storage multi-regions) resolve to their jurisdictions.allowedRegions. It was a warning, some-central2,northamerica-south1and theasiamulti-region passed validate andgenerate iacon an EU-only contract. Advisory and audit, and other platforms, keep the warning.check_placements(sovereignty, [(where, region)]), so a provider hook runs the same rules on its own placements.providers/gcp/util/sovereignty.py):GcpIacPlugin.emitchecks every emittedlocation/region. This is the chokepoint of bothfluid applyandfluid generate iac, and it runs whether or not a native provider can be built.GcpProvider.planchecks every planned one, raising the typedSovereigntyViolationErrorthroughProviderError, the wayAwsProvider.plandoes.native_actionsre-raises it.USdataset default, and the provider region a Cloud Scheduler job or staging bucket inherits (the--regiondefaulteurope-west3, the SDK'sus-central1).USdefault is unchanged. Under advisory it is emitted and warned about, so it is never silent.pubsub_topicbinding's region becomesmessage_storage_policy.allowed_persistence_regionson thegoogle_pubsub_topic(hashicorp/google), and the hook checks it. It was dropped, so the topic passed validate underallowedRegions: [europe-west1]and could store messages anywhere.location: None. The job runs where the table is (table.location), never at a guessedUS.F10:
fluid applyreports each run to the Command Centerapp/api/v1/executions.py, withExecutionCreate/ExecutionUpdate. It takesX-API-Keyor a bearer token, and scopes the run byX-Organization-Id.CommandCenterReporter(async queue, circuit breaker, SSRF host gate) already speaks it, so it is wired in, not replaced.observability/config.pygainsheaders, for a bearer credential and the organization.fluid publishdoes: thefluid-command-centercatalog configuration (FLUID_CC_ENDPOINT,FLUID_API_KEYorFLUID_BEARER_TOKEN). The organization comes fromFLUID_CC_ORG_ID,organization_id, or the product'sfluid.config.yamlslug, which the demo lab uses, all through the publisher's ownresolve_organization_id. The reporter'sFLUID_COMMAND_CENTER_URL/_API_KEYare the fallback.POST /api/v1/executionswhen the engine knows the product (statusrunning, provider, environment,runner: jenkins|cli,cli_version, and inmetadata:product_id,contract_version,fluid_version,contract_hash,platform,state,mode). ThenPATCH /api/v1/executions/{id}at the end:success/failed, and a result with the planned and applied counts, the resource addresses,dry_run,exit_code,started_at,finished_atandduration_seconds, plus the typederror_eventon failure.FLUID_COMMAND_CENTER_ENABLED=falseturns it off.F11
dag_id_for(product, build, env)gives<product>__<env>__<build>, andschedule_scope_forgivesschedule/<product>__<env>/. No env keeps both as they were.--check-sovereigntyis in the shared command table (GitHub Actions, GitLab, Azure DevOps, Bitbucket, CircleCI), the 11-stage spec (Tekton) and Jenkins' stage 6. A contract with no sovereignty block prints NOT CHECKED and passes.expected-environments; it now reads it as a declaration.--env Xwith no overlay isoverlay_declared_but_missingwhen the workspace'sexpected-environmentsdeclares X for the product (keyed by the contract's directory name, then its id, as fluid-demo-env writes it).environmentsblock does not refuse: it is schema-valid, forge-cli applies nothing from it, and refusing on it broke contracts that validated on 0.16.5. The warning names it instead.dev(the base by convention) and an env the base is already bound to (localfor a local base) stay the base.Prior art (borrow-before-build receipts)
Searched:
opentofu init -migrate-state -force-copy backend configuration change copy state non-interactive: the OpenTofu init docs.-force-copyimplies-migrate-stateand answers yes to every prompt.terragrunt remote_state key change migrate existing state terraform init -migrate-state: Terragruntbackend migrate(renamed units); the HashiCorp help article oninit -migrate-state.structlog event key collision "got multiple values for argument" positional-only: Can't set the "logger" key in "get_logger(initial_values=...)" hynek/structlog#295, the same collision onlogger.python logging reserved LogRecord attributes "Attempt to overwrite" message extra key collision: logging: rename colliding extra keys instead of raising KeyError WebbPulse/webbpulse-python#129 renames toextra_<key>; Drop event dict keys colliding with LogRecord attributes in render_to_log_* hynek/structlog#842 drops the key instead.terraform google_bigquery_dataset location default "US": BigQuery defaults an unplaced dataset to the US multi-region.checkov OR tfsec OR conftest policy require resource location region allowed list: region allow-lists are policy-as-code (conftest over plan JSON). Here the contract's own engine is that policy.openlineage python client http transport timeout errors do not fail the job: the client logs emit failures and never fails the job (continue_on_failure).airflow dag_id unique across deployments naming convention environment suffix:dag_idmust be unique per Airflow, and renaming one orphans its history.dbt "does not have a target named": dbt refuses a requested target its profile does not declare.BigQuery dataset locations list regions multi-regions dual-regions EUR4 NAM4 ASIA1: Google's location pages.terraform google_pubsub_topic message_storage_policy allowed_persistence_regions: the hashicorp/google argument (message_storage_policy { allowed_persistence_regions = [...] }).terraform plan -lock=false read-only state backend dry run migrate-state: nothing plans across a pending backend move; plan against the old backend, move on apply.opentofu init installs providers required by state plugin cache TF_PLUGIN_CACHE_DIR: init installs state-required providers; Conflict when TF_PLUGIN_CACHE_DIR is set to an local provider search directory opentofu/opentofu#3721 (a plugin cache that is also a search directory conflicts), which is why the apply's own.terraform/providersis never the probe's plugin directory.Fetched:
-migrate-state,-force-copyand-reconfiguresemantics.backendMigrateState_s_s,-force-copyoverwrites a non-empty destination; the source stays untouched; both states are locked.gcp.resourceLocations(no multi-region detail).-plugin-dirreads plugins only from that directory, "as if it had been configured as afilesystem_mirror".app/api/v1/executions.py,app/models/execution.pyandapp/api/dependencies.py(fluid_cc_backend 0d70235, read-only), for the payload and authentication contract.Reuse strategy:
CommandCenterReporter,FluidCommandCenterProvider.resolve_organization_id, the SSRF gate, andseverity_for.extra_<key>rename; the "an explicitly requested target must exist" rule from dbt; the provider's ownmessage_storage_policyfor Pub/Sub placement; dbt-style "plan reads, apply writes" for the pending move.backend migrateneeds an operator-named source and destination; here both are derived, and the attribution rule replaces the operator.Each borrowed pattern is credited in a code comment.
Schema changes
None. No contract field was added (0.7.5 GA and the 0.7.6 preview are untouched).
expected-environmentsis a key offluid.workspace.yaml, read if present; the contract'senvironmentsblock already exists in the schema.Behaviour changes (release notes)
FLUID_STATE_BACKEND, or a packaging contract, is nowfluid/<id>/<provider>/terraform.tfstate(GCS prefixfluid/<id>/<provider>). The first apply after the upgrade moves an existing state there withtofu init -migrate-state, printsstate move: moved N resource(s) from … to …, and leaves the old object in place. A state it cannot attribute to one provider is refused (state_migration_ambiguous). The shared legacy key and explicit keys are unchanged.fluid validate,plan --check-sovereignty,generate iacandapply, and so does one whose region's jurisdiction cannot be resolved (unless the region is inallowedRegions). Advisory warns and audit logs. The demo's committed bronze gcp overlay already setsregion: europe-west1(fluid-demo-envd225844) and is unaffected; the silver and gold gcp overlays (D1) must set a region when they are added. A GCPlocation.locationis checked.US/EUmulti-regions resolve to US / EU, and the GCP regions and dual-regions the vendored table lacked now resolve.pubsub_topicbinding with a region now emitsmessage_storage_policy.allowed_persistence_regions. An existing topic plans an in-place update.fluid apply --dry-runnever moves state; the first real apply does.state_shared_with_another_provider).fluid applyandfluid generate iacon gcp refuse a placement outside the sovereignty policy, including a region a planned resource inherits from the provider default.fluid applyreports each run to the Command Center whereverfluid publishis configured. Off withFLUID_COMMAND_CENTER_ENABLED=false. A Command Center on a private address needsFLUID_COMMAND_CENTER_HOST_ALLOWLIST(loopback is always allowed).fluid policy-applyon gcp no longer crashes. The structured log event carries the provider'smessageasextra_message.<product>__<env>__<build>, inschedule/<product>__<env>/. Upgrade step: a DAG an earlier release synced for an env stays at the scheduler under<product>/with its old id, becauseschedule-synconly mirrors directories the source has. Delete that directory once, or both DAGs run.fluid plan --check-sovereignty.--env Xwith no overlay is an error when the workspace'sexpected-environmentsdeclares X for the product. The contract'senvironmentsblock only warns.Tests (all new ones fail on origin/main, recorded)
Each file was run against an origin/main worktree (
b5040ef) with the file copied in.tests/iac/test_iac_state_key_migration_moto.py(real tofu 1.12 against moto S3, AWS APIs and the S3 backend)state_migrationtests/iac/test_state_key_per_provider.pystate_migrationtests/cli/test_policy_apply_provider_message.pytests/providers/test_gcp_sovereignty_fail_closed.pybinding_regiontests/cli/test_apply_reports_to_command_center.py(loopback stub CC server)tests/cli/test_one_contract_two_clouds_pipeline.pyWhat the moto file pins:
+0 ~0 -0, for a kept and for a wiped workdir.fluid apply --dry-runleaves the bucket's key set unchanged and plans+0 ~0 -0on the old key; the real apply after it moves the state.check_state_drift, whatfluid diffandverify --state-driftrun) reads the old key while the move is pending: statuschecked, every managed resource, nothing written. (In the first round this line claimed a pin the file did not have: the test calledreconcile_state_key(migrate=False)directly. A mutant pointing diff at the new key passed the suite; this test now fails it.)For the three import-error files, behaviour was measured on origin/main with a probe:
location.location=us-central1: valid, 0 findings.US.us-central1strict: emitted.fluid validate, no region: rc 0.fluid generate iac us-central1: rc 0.US.Updated for the intended changes:
tests/iac/test_state_backend_env_default.pyandtests/cli/test_diff_state_drift.py: the keys gain/aws/, and the move is stubbed there.tests/forge/test_artifact_fanout_schedule.py: the env DAG path.tests/cli/test_generate_schedule_fluid_apply.py:JENKINS_URLandProgramDatajoinNOT_PASSED, with their reasons.CI: a Stage 1 step in
iac-tests.ymlruns the moto migration file with a coverage assert, as the state-drift step does..secrets.baselinechanges by one line number only, because that step shifts an existing entry.Review round (tests fail before the fix, pass after)
Each file below was run on the first-round head (
6eaa58d) and on origin/main (8d82828) with the file copied in.6eaa58dtests/iac/test_state_migration_safety.py(fault-injected raced / unverified / probe paths, the probe offline with real tofu, the shared-key guard, dry-run asks no move)state_migration)tests/providers/test_gcp_locations_and_pubsub_placement.pytests/policy/test_sovereignty_enforcement_mode.py(unknown region under strict)tests/cli/test_apply_reports_to_command_center.py(3 new early-failure tests)tests/cli/test_one_contract_two_clouds_pipeline.py(theenvironmentsblock warns,validate --env prodrc 0)tests/iac/test_iac_state_key_migration_moto.py(real tofu 1.12 on moto)fluid/<id>/aws/terraform.tfstate(reproduced)Mutants, each run against these tests (all killed):
-plugin-dir;environmentsblock refusing again.fluid diffpointed at the new key (read = target) and the dry-run migrating were checked against the moto file: see the Gates section.One design correction found while testing: the probe first used the apply's own
.terraform/providersas its plugin directory. Under a plugin cache those entries link into the cache, and the probe's init broke the workdir ("no package for hashicorp/aws 5.100.0 cached in .terraform/providers", measured on moto; the class of opentofu/opentofu#3721). The probe now uses an empty directory.Not changed, with the evidence:
_adopt_existingimports into state during a dry-run. Brownfieldtofu importof a declared resource missing from state runs before the dry-run return, on main and on this branch. That is a state write in a plan-only run, but it predates this PR and the key set does not change. It is not touched here.'${"us-central1"}') is skipped by the GCP emit hook (references are not places).fluid validateandplan --check-sovereigntystill refuse it under strict.execution.trigger(ps.ensure_topicfromschedule.py) carry no region, so they get no storage policy.Proofs
tofu validateon the emitted bronze gcp module (europe-west1, hashicorp/google 6.x):Success! The configuration is valid.fluid validate --env gcprc 1,generate iacrc 1,plan --check-sovereigntyrc 1;us-central1:generate iacrc 1;europe-west1: rc 0;--env localand--env aws: rc 0, unchanged.expected-environments: [local, aws, gcp]and no gcp overlay: validate and plan rc 1 (overlay_declared_but_missing);--env localand--env awsrc 0.bindings.json: this branch rc 0, origin/main rc 1 (TypeError).fluid plancommand carries--check-sovereignty. Basic Tekton renders no plan command.Not proven (no emulator can)
US/EUjurisdictions come from BigQuery's and Cloud Storage's documented multi-region definitions, not from a dataset.<product>/is documented, not automated.allowedRegions.local) apply is reported with status and timings only. The CC-side "one product, a deployment per cloud" view is a separate CC change.Gates
On
cc9784c, python 3.12, black 24.10.0:ruff check fluid_build/ tests/clean.black --checkclean (1776 files).python scripts/mypy_strict.pyno issues.lint-imports6 kept, 0 broken.add_license_headers.pyadded nothing (0 added, 1809 skipped).detect-secrets-hook --baseline .secrets.baselineon the changed files: rc 0. The merge of origin/main moved one baseline line number (the iac-tests.ymlAWS_SECRET_ACCESS_KEY: testentry, 222 → 231).pytest -p no:randomly -n 4with an empty HOME and a temp cwd, ontests/build_runners tests/providers tests/cli tests/iac tests/forge tests/policy tests/test_sovereignty.py(the moto file deselected and run on its own): 5723 passed, 0 failed, 401 skipped, 1 xfailed, 1 xpassed. The moto file: 9 passed (real tofu 1.12, 4 min 40 s, sequential as CI runs it).Security review
The
security-reviewskill loaded against the session's working directory, a different repository, in both rounds. So its procedure was run by hand on this branch's diff. Outcome: no high- or medium-confidence findings.First round:
fluid publishalready uses. The reporter's SSRF host gate now runs before the organization lookup sends it. Header values are never logged, andCommandCenterConfig.__repr__prints header names only.parse_backend. Scratch directories come fromtempfile. A provider name must be one[a-z][a-z0-9_-]*key segment.tofu state pullis reported from stderr only, because its stdout is the state document._loggingserialization is unchanged.This round:
-plugin-dir. An argv element naming atempfiledirectory.json.dumpsfromparse_backendblocks and the plugin's static pins.records_backend. Reads only the scratch directory's own.terraform/.