Conversation
…ion, CMEK, column restrictions and mapped principals reach BigQuery, and fluid verify checks them The offline verification of the demo products on forge-cli 0.16.5 found a gcp binding's governance dropped without a warning: lifecycle.expire, binding.encryption.kms and policy.authz.columnRestrictions emitted nothing (the module was byte-identical with and without them), and the contract's accessPolicy principals, placeholders such as group:data-platform@northwind.example, became the dataset's AUTHORITATIVE access list, which real BigQuery refuses and which replaces every other entry on the dataset. Retention (F7): expire: true + retention is BigQuery partition expiration on daily partitions (time_partitioning type DAY, expiration_ms), on binding.location.partitionBy when it names one DATE/TIMESTAMP/DATETIME column, else by ingestion time. Never a table TTL. BigQuery cannot partition an existing table and the provider plans adding ingestion-time partitioning as an in-place update, so the table's lifecycle.replace_triggered_by names a terraform_data holding the partitioning's shape: adding it to a live table plans a replacement, which the data-loss gate refuses without --allow-data-loss. A new retention is an in-place change. Encryption (F6): kms: product on a gcp binding creates a key ring and crypto key per product and dataset in the dataset's location (EU/US -> europe/us), 90-day rotation, grants the BigQuery service agent cryptoKeyEncrypterDecrypter, and sets the dataset's default_encryption_configuration and the table's encryption_configuration. A Cloud KMS key name validates only on gcp; alias/ and arn: values are refused there, and a Cloud KMS name on aws. none = Google-managed. Key ring and key are import candidates, since GCP cannot delete them. Column restrictions (F8): the existing columnRestrictions field becomes a Data Catalog taxonomy per product and dataset with fine-grained access control, a policy tag per set of restricted columns sharing their readers (attached through the schema's policyTags), and categoryFineGrainedReader for exactly the allowed readers. On AWS it becomes each Lake Formation grant's excluded columns; the overlay's own excludedColumns keeps working and must agree when both are set. The LF emit now writes wildcard beside excluded_column_names, without which every excludedColumns grant failed tofu plan. Principals (D9): dataset grants are non-authoritative google_bigquery_dataset_iam_member resources. binding.principals (fluid-schema 0.7.6) maps each logical principal to the binding platform's identities; with it, an unmapped principal is refused, and without it a principal in a reserved TLD is refused as a placeholder. fluid verify: BigQuery gains retention (partition type, field, expiration, no table TTL), encryption (kmsKeyName of dataset and table) and columnRestrictions (tags present, fine-grained readers exactly the derived set, taxonomy enforcing) dimensions; S3+Glue gains a Lake Formation column check (no SELECT reaches a restricted column, IAM_ALLOWED_PRINCIPALS included). fluid validate reports every refusal at stage 2. Prior art: hashicorp/google provider docs and source (bigquery_dataset_iam, bigquery_table time_partitioning and encryption_configuration, data catalog taxonomy/policy tag), terraform-provider-aws lakeformation_permissions docs, dbt-bigquery configs (partition_expiration_days, kms_key_name, policy_tags, grants), cloud-foundation-fabric data-catalog-policy-tag, terraform-google-modules kms, OpenTofu terraform_data + replace_triggered_by, ODCS v3 roles and servers[].roles, BigQuery CMEK, column-level security and partition docs.
… and lands in BigQuery; fluid verify counts the table and checks its masked columns On 0.16.5 a silver or gold product could not build on GCP: a consumes[] entry whose upstream is a gcp bigquery_table was UnreadableBindingError, and with parameters.inputs the result was written to a local file named after the overlay's gs:// staging path while the table stayed empty and the build reported success. Reads: the upstream table is named by the IaC's own rule (_bigquery_load.bigquery_load_target) and read with python-bigquery's Arrow pages into a staged Parquet file the view reads; a BigQuery TIMESTAMP reads as the UTC wall clock (the DuckDB BigQuery extension's documented mapping), as the same SQL reads it on local and aws. The staged copy is removed after the build. Landing: a first expose bound to a bigquery_table is staged as Parquet and loaded by the acquisition runner's own load_file (WRITE_TRUNCATE, CREATE_NEVER, the table's schema). A gs:// or other non-S3 URI, a GCS bucket, an Azure/Snowflake/Databricks binding, a second BigQuery output and a BigQuery landing without inline SQL are refused before the SQL runs. Load: TIMESTAMP columns held as naive Parquet timestamps are sent UTC-adjusted (python-bigquery's TIMESTAMP -> timestamp(us, UTC) mapping), so BigQuery does not read them as DATETIME. With BIGQUERY_EMULATOR_HOST set the client is anonymous, so no ADC token reaches an emulator; a job with no outputRows is checked by COUNT(*) and never taken as zero or as success. Verify: one GoogleSQL count query adds row_count (held to the acquisition run's facets.bigquery_load, by the Athena verifier's rules) and masking (COUNTIF ... NOT REGEXP_CONTAINS against a bound, anchored shape), each CRITICAL under --strict. The num_rows None crash is fixed, and verify names the table the way the load does. pyarrow is declared in the gcp and dev extras; the heavy emulated lane installs local and requires the new emulator chain test.
…s import id, and a partial Lake Formation listing is an error, not a pass Two findings from the security review of the governance change. 1. The product key ring's import id is built from the dataset's region, so a region such as "europe-west1/keyRings/other/../.." would have made the apply adopt another key ring (and its key) into this product's state. The derived Cloud KMS location must now be a location id ([a-z0-9-]), or the binding is refused (encryption-kms-location). 2. lakeformation:ListPermissions returns only what the caller may see, so a verifying identity that is not a Lake Formation administrator got a partial listing in which a leaking grant was simply absent, and the check passed. The contract's own grants must be in the listing now; when one is not, the dimension is an error naming it. Also: docs say what non-authoritative dataset grants change (BigQuery's default entries for the project's basic roles stay; a grant made outside the contract is no longer removed), and iac-tests.yml Stage 1 runs the governance files, failing if any of them only skipped.
…fails closed on sovereignty Platform safety for deploying one contract to two clouds through --env overlays, from the offline verification of the demo products on 0.16.5. State (F1). The default remote state key names the provider: fluid/<id>/<provider>/terraform.tfstate (GCS prefix fluid/<id>/<provider>), for FLUID_STATE_BACKEND bucket-only specs and packaging contracts. The aws and the gcp apply of one contract shared fluid/<id>/terraform.tfstate, so each plan read the other cloud's resources as orphans to destroy. The first apply after the upgrade moves the old key's state with OpenTofu's own `tofu init -force-copy` (implies -migrate-state), only when the new key holds no state and the old one holds this provider's resources; another provider's state is left alone, a state of two clouds is refused, the old object is never deleted, and the copy is verified by reading it back. fluid diff / verify --state-drift read the old key while the move is pending. The shared legacy key and explicit keys are unchanged. Stage 8 (F2). info()/warn()/error() take logger and message positionally only, and a payload key named like an envelope key is kept as extra_<key>: the GCP policy result's "message" raised a TypeError on every gcp build. Sovereignty (F5). Under enforcementMode strict a binding on aws, gcp or azure that names no region is refused (advisory warns, audit logs); a GCP region given as location.location is checked; BigQuery's US and EU multi-regions resolve to their jurisdictions. The GCP IaC plugin checks every emitted location and the GCP provider every planned one (the way AwsProvider.plan does), so fluid apply and fluid generate iac refuse an out-of-jurisdiction region, including a provider-default region a resource inherits. The BigQuery load runs where the table is, never a guessed US. Command Center (F10). fluid apply registers each run at POST /api/v1/executions and closes it at PATCH, through the existing CommandCenterReporter, with the credential and organization fluid publish uses: product id, contract version and hash, environment, provider, mode, state location, resource addresses, change counts, timings and status. Best effort: an outage or refusal never changes the exit code, and no secret, header or tofu output is sent. FLUID_COMMAND_CENTER_ENABLED=false turns it off. Small items (F11). An env's Airflow DAG id is <product>__<env>__<build> and its schedule directory <product>__<env>; every generated stage 6 runs plan --check-sovereignty; --env X with no overlay is an error when the contract's environments block or the workspace's expected-environments declares X for the product (dev and the base's own platform stay the base), and the warning otherwise names what the base binds to. Borrowed: OpenTofu backend migration (meta_backend_migrate.go), Terragrunt backend migrate, WebbPulse/webbpulse-python#129 (rename colliding log keys), the OpenLineage client's log-don't-fail posture, dbt's unknown-target error.
…s as it will emit
fluid validate does not resolve {{ env.* }} (fluid apply does, before the
emitter runs), so an overlay mapping the pipeline to
serviceAccount:fluid-pipeline@{{ env.FLUID_DEMO_GCP_PROJECT }}.iam.gserviceaccount.com
was refused as not an IAM member, by the shape check and by the schema pattern.
Each template now stands in for one plain segment when the shape is checked, and
the schema's GCP member pattern allows it, as the Lake Formation principal pattern
already does.
The governance-parity lane in iac-tests.yml moved the baseline's known test-credential line from 217 to 233; detect-secrets-hook refuses a stale baseline (exit 3), as the Lint & Format secret scan did on this PR.
…o sovereignty, never loads into its own input, refuses outputs it cannot land, and records its load for verify Review of the first push found seven defects; each was reproduced on the branch before it was fixed. - CI was red: three of the new unit tests needed google-cloud-bigquery or google-auth, which the .[dev,local] unit lanes do not install. The two resolve tests now use the fake BigQuery, and AnonymousCredentials goes through a seam (_bigquery_load._anonymous_credentials) the fake replaces. In a .[dev,local] venv the old file failed 3 tests, and the new one passes. - Sovereignty failed open. An EU-only silver whose gcp overlay named no region read europe-west1 and loaded into US, the IaC default, with rc 0. Before anything is read, a contract with a sovereignty block now has every BigQuery binding the build reads or loads checked by fluid validate's rules (refuse_sovereignty_breach). The binding must name a region. The landing must be outside deniedRegions, inside allowedRegions and in the declared jurisdiction. With no crossBorderTransfer, every input must be in the landing's jurisdiction. enforcementMode sets the severity. BigQuery's EU and US multi-regions map to EU and US, as Google's in:eu-locations and in:us-locations value groups list them. - A landing that resolved to one of its own inputs WRITE_TRUNCATEd another product's table, and bronze went from 5 rows to 3. This is now refused on what the two resolve to, the way Dagster's _validate_self_deps compares asset keys: BigQuery table ids are compared case-insensitively, and a project left to the client matches any project. An S3 landing inside a prefix an input reads is refused too. - A further output bound to GCS, S3 or another cloud was planned and then silently not landed. It is now refused. A further local expose or output port is still not written, and the build now prints a warning saying so. - verify looked up ".dataset.table" when the binding named no project. It now uses the client's project, as the load does, and with no project at all it is an error naming the table. - silver and gold row_count was never compared with the load. An embedded-SQL BigQuery load now writes the acquisition runner's run record (facets.bigquery_load, full_refresh, rows_from: write), and the BigQuery verifier reads embedded-SQL runs. The emulator chain now holds silver and gold to their loads. - The failed-load and short-load tests passed on main and under a mutant that breaks staging. They now assert the load was attempted, the error was printed, and the table was left unchanged.
# Conflicts: # .secrets.baseline
…n AWS, restriction readers, IAM names, revocation, migration and one key per dataset - fluid validate --strict failed every AWS contract that carries accessPolicy for its gcp deployment (the three demo pipelines' stage 2). It now warns only for an aws binding with no governance.lakeFormation.grants, where the access intent really is unenforced. - validate dispatches on the platform first: an aws Iceberg binding was checked as GCP Iceberg storage and its retention, key and logical principals refused. - A column restriction's readers now include exposes[].policy.authz.readers, mapped through binding.principals. A deny for one group gave the policy tag no reader, locking the columns for everyone. A restriction with no reader at all is refused (column-restriction-no-readers), as AWS refuses one with no LF grant. - Dataset and policy-tag IAM member names carry a hash of the exact role and member: data.eng@, data-eng@ and data_eng@ shared one name and two grants were silently dropped. - The data-loss gate counts only data-bearing removals: revoking a dataset, table, bucket or policy-tag grant, a policy tag, a taxonomy or an LF permission no longer needs --allow-data-loss. It fails closed when the plan does not itemise every removal. - The first apply on a dataset 0.16.5 applied with an authoritative access list revokes the entries no member resource covers (a principal removed or remapped in the same change), once, through the module's access list; fluid diff shows it. The provider keeps access as Computed, so they stayed (provider issue 8165). - Tables of one dataset with different keys are refused (encryption-kms-mixed-dataset), and the dataset's default key no longer depends on expose order (provider issue 26193: an unkeyed table in a keyed dataset is replaced on every plan). - An unmapped GCP principal that is not an IAM member (group:data-platform, analysts, role:analyst) is refused as a placeholder. - A restriction's tags and labels are written into the policy tag description. - Retention on a partition column is documented as event-time age (a backfill of older rows is deleted at once), and logged. - Tests pin the fluid verify wiring and a TableWildcard LF grant to a denied principal (both mutants survived the suite before).
#675 fixed the Lake Formation excludedColumns wildcard on main; this branch's own copy of that fix is dropped in favour of main's. The secrets baseline follows iac-tests.yml's merged line (238). Wording that named 0.16.5 as the last release with the authoritative dataset access list now names 0.16.6.
…ils closed on an unplaceable location, and a Pub/Sub topic keeps its region Review round on the one-contract-two-clouds platform safety change. State. `fluid apply --dry-run` (the generated Jenkins default APPLY_MODE) ran the state move and wrote the new key; measured with real tofu against moto, the key set grew by fluid/<id>/aws/terraform.tfstate. A dry-run now asks the move not to run and plans against the old key while it is pending, the read-only path fluid diff already takes; the first real apply moves it. The move's probe no longer installs anything: OpenTofu's init installs the providers the state names (the latest hashicorp/aws for an aws state), and the probe ran on every apply whose new key was empty. It now inits with -plugin-dir on an empty directory and pulls from the backend that init is shown to have recorded, falling back to a plain init. The one-time copy installs what the old state names at the plugin's own pins. Using the apply's .terraform/providers as the plugin directory was tried and broke the workdir under a plugin cache, so it is not used. A remote key that does not name the provider (the shared fluid/terraform.tfstate a bucket-only --state-backend gives a contract without packaging) is refused when it holds another cloud's resources (state_shared_with_another_provider). Sovereignty. Under strict, a cloud-region binding (aws/gcp/azure) whose jurisdiction cannot be resolved is refused unless the region is named in allowedRegions; it was a warning, so me-central2, northamerica-south1 and the asia multi-region passed validate and generate iac on an EU-only contract. The nine GCP regions the vendored table lacks, and the dual-regions within one jurisdiction (EUR4, NAM4, ASIA1), are mapped from Google's own location lists, in source, not in the ODbL csv. A gcp pubsub_topic binding's region becomes message_storage_policy.allowed_persistence_regions (hashicorp/google), which the GCP hook checks; it was dropped. Command Center. A run refused before the engine registered it now carries its product: the engine identifies it as soon as the contract and provider are known, and a run refused earlier is described from the base contract. Overlays. Only the workspace's expected-environments refuses --env X with no overlay; the contract's own environments block warns again, as on 0.16.5, since forge-cli applies nothing from it. Tests: the raced and unverified branches of the move under fault injection, the probe installing nothing (real tofu, no registry reachable), a dry-run leaving the bucket's key set unchanged and the drift pass reading the old key (real tofu against moto), and each of the fixes above.
📄 Documentation ReminderThis PR appears to be missing a documentation reference. Our docs live in a separate repo. Please update the PR description with one of:
See the Contributing Guide for details. |
| contract = _with_readers("readers@corp-a.com") | ||
| contract["exposes"][0]["lifecycle"] = {"retention": "P30D", "expire": True} | ||
| _write(contract, tmp_path, fake.endpoint) | ||
| _summary, (data_changes, revoked) = _data_bearing(tmp_path, tofu_env) |
|
|
||
| import argparse | ||
| import logging | ||
| from pathlib import Path |
| for path in paths: | ||
| try: | ||
| Path(path).unlink() | ||
| except FileNotFoundError: |
…ed at their dataset's multi-region, and plan --check-sovereignty checks what apply emits Cloud KMS names BigQuery's EU and US multi-regions europe and us, and Data Catalog names them eu and us. The GCP sovereignty hook checked those ids as places, so with the governance branch's CMEK and policy tags an EU-only strict contract with an EU dataset was refused at apply (Region 'europe' not in allowed regions list, jurisdiction Unknown), after validate --strict and plan --check-sovereignty passed. resource_placements now places both at the multi-region, spelled as the emitted dataset spells it. fluid plan --check-sovereignty ran only the policy engine over the bindings for gcp (the provider had no hook), so a placement the planner or the emitter derived passed stage 6 and was refused at stage 7. GcpProvider now has validate_sovereignty: the engine's contract checks plus the planned actions' and the emitted resources' placements, through the same engine, and it returns the error-severity findings. GcpIacPlugin.emit takes enforce_sovereignty=False for it. A hook that gives no verdict is now named as such in the fallback line, not as a missing hook.
…' clouds, not by OpenTofu's built-in provider A gcp product whose table's partitions expire keeps a terraform_data beside the table (the trigger that replaces it when its partitioning changes), under provider["terraform.io/builtin/terraform"]. state_sources counted builtin/terraform as a provider no plugin emits, so classify called every such legacy state ambiguous and refused the move, and with it every apply, dry-run and diff of the product (measured on the integration with the governance branch). The built-in provider ships inside OpenTofu and names no cloud; it is now left out, and a state of nothing but built-in resources is still not guessed.
…at each build did, and why the run failed fluid apply --mode amend-and-build reported only the infrastructure: the PATCH carried planned and applied changes and resources in phase apply, so a silver run that loaded 28 rows said nothing of it, and a bronze run whose load failed was closed failed with no error event and no message. The builds return an exit code, not an exception, and run_builds_from_args never touched the run report. run_builds_from_args now records each build on the open report: its id, its status (succeeded, failed, skipped) and, from the run record it wrote, the run id, the table facets.bigquery_load names, the landed destinations and the rows. The phase is build, and a failed build phase names its reason: build_failed:<build id>, build_not_found:<build id> or builds_all_skipped. Nothing is read when no apply report is open.
…tires the DAG it replaced An env's DAGs moved from <product>/ (dag id <product>__<build>) to <product>__<env>/ (<product>__<env>__<build>), and --delete-scope product mirrors only the new directory. On the demo lab the first sync after the upgrade left bronze.customer_subscriptions/ingest_subscriptions_dag.py beside bronze.customer_subscriptions__aws/, both FLUID_ENV_NAME 'aws' on the same cron, and the lab's Airflow unpauses a DAG as it parses it: two fluid apply --env aws of one product against one state at the same minute. Under --delete-scope product the file and git+ssh transports, whose destination is readable here, now retire the old directory's DAGs rendered for the same product and env with the old id, read from each file's syntax tree (never imported): an rsync --delete from an empty directory filtered to exactly those files, after the new directory synced. An env-less DAG, another env's, and any file that is not a rendered DAG stay. The other transports print the one step left to do, and the report records each superseded scope and whether its old DAGs were retired.
…out retires each env's old DAG
…ith no column is refused, as a hand-written one is 0.16.6 (#675) refuses a grant whose excludedColumns name every column of the table: it would grant nothing to read, and emitted it is a column wildcard that excludes everything. The exclusions columnRestrictions derive bypassed that check: _check_lf_grant_columns read only the grant's own excludedColumns, while _emit_lakeformation wrote the derived ones. A deny on every column for the analyst, or an allow-list naming only the steward, emitted the analyst's grant as wildcard = true excluding customer_id, status and msisdn. lf_column_exclusions, which fluid validate and the emitter both run, now refuses a read grant the restrictions leave with no column, naming the grant and its principal, so stage 2 refuses it as stage 7 does. The grant check reads the derived exclusions too, so what is checked is what is written.
…e Data Catalog, and say so on the emulator governance_dimensions opened the Data Catalog session (ADC) before _column_access_dimension looked at the table's policy tags. On the BigQuery emulator, with no ADC, a table whose restricted msisdn carried no tag was reported as an error (DefaultCredentialsError) instead of the failure it is, and with ADC present verify would have asked datacatalog.googleapis.com about the emulator's table. The session is now opened only when a tagged column's readers must be read. A column with no tag fails from the table alone, and when Data Catalog cannot be reached what the table showed is still the result, with a note of what was not checked. Under BIGQUERY_EMULATOR_HOST, which no Data Catalog emulator answers, a table whose restricted columns are all tagged gets status unsupported (not verifiable there) rather than a call to the real API. Elsewhere, a catalog that cannot be read is still an error.
…ld_runners records builds without importing the CLI The previous commit read the run report through a function-local import of fluid_build.cli._apply_cc_report from build_runners, the reverse edge tests/observability/test_import_hygiene.py forbids. The context variable now lives in observability/apply_run.py, which build_runners already depends on; _apply_cc_report sets and reads it there.
…ccess through Lake Formation grants
The governance gate warns for an aws binding whose accessPolicy no Lake
Formation grant enforces, and fails `fluid validate --strict` on it; it
refuses an unmapped placeholder principal on GCP. Six shipped examples failed
it. Each is fixed the way the feature intends, none is exempted.
AWS (aws-glue-data-lake contract-database and contract-iceberg,
aws-iceberg-lakehouse, aws-medallion-lake, aws-s3-glue-athena): the binding
declares governance.lakeFormation, registerLocation plus one grant per
accessPolicy grant, as an IAM role in the applying account
(arn:aws:iam::{{ env.AWS_ACCOUNT_ID }}:role/...): read/select/query ->
SELECT + DESCRIBE, write/create -> INSERT, update/delete -> DELETE. Same-
account grantees get no bucket-policy statement (count = 0 at plan). The
two aws-glue-data-lake contracts move from 0.7.1 to 0.7.5, the first GA
schema with binding.governance. accessPolicy stays as the neutral intent.
bitcoin-price-api-declarative-part-b: binding.principals (0.7.6) maps the
four readers and the interns group to example.com identities, including
serviceAccount:looker@<<YOUR_PROJECT_HERE>>..., so the column restriction
now emits its policy tag, readable by the four readers.
Tests: the AWS example pins count the Lake Formation resources and
validate with --strict; part-b's readers are pinned, and its looker
placeholder is still refused once binding.principals is removed.
…t-b, which now emits part-b maps its looker reader through binding.principals, so it goes through the byte-identity path; the branch stays for any contract the emitter refuses.
# Conflicts: # CHANGELOG.md
…s a hand-written one does
…ion-gcp-parity
…gcp-parity # Conflicts: # .github/workflows/iac-tests.yml # .secrets.baseline # fluid_build/cli/_apply_opentofu_engine.py # fluid_build/cli/_diff_state.py # fluid_build/cli/verify.py
…verify client answers #673's row count
fas89
force-pushed
the
ci/integration-gcp-parity
branch
from
September 28, 2026 18:18
d050db8 to
852dde4
Compare
Collaborator
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Integration check for #676 + #673 + #674 on main at 0.16.7 (do not merge).
Rebuilt from main: #676 (fc9adbd), then #673 (8d1e62f), then #674 (5422bec), then the cross-PR glue the old integration head carried inside its merge commits, now as its own commit (852dde4):
_embedded_sql_io.py: feat(build): embedded-SQL builds on DuckDB read and land BigQuery; verify counts and checks masking on BigQuery #673's BigQuery target follows fix(platform): one contract on aws and gcp keeps two states, and GCP fails closed on sovereignty #676's rule that an undeclared region names no load location, and a test that the load then runs where the table is.tests/cli/test_verify_bigquery_governance.py: the fake client answers the row count feat(build): embedded-SQL builds on DuckDB read and land BigQuery; verify counts and checks masking on BigQuery #673 adds to verify.That glue lands with #673 and #674 when each is brought up to date after the one before it merges.