Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
31 commits
Select commit Hold shift + click to select a range
6a1c099
feat(iac): one contract is governed the same on GCP as on AWS: retent…
fas89 Sep 28, 2026
f91f7d8
feat(build): an embedded-SQL build on DuckDB reads BigQuery upstreams…
fas89 Sep 28, 2026
f2c87ff
fix(iac,verify): a region that is not a location never reaches a key'…
fas89 Sep 28, 2026
6eaa58d
fix(platform): one contract on aws and gcp keeps two states, and GCP …
fas89 Sep 28, 2026
93de43d
fix(iac): a mapped identity carrying an {{ env.* }} template validate…
fas89 Sep 28, 2026
35ba1a3
chore(secrets): the baseline follows iac-tests.yml's moved line
fas89 Sep 28, 2026
96a70db
fix(build): the BigQuery embedded-SQL path holds its reads and load t…
fas89 Sep 28, 2026
29f831d
Merge origin/main into feat/gcp-platform-safety
fas89 Sep 28, 2026
8f72348
fix(iac,validate): the review's governance defects: strict validate o…
fas89 Sep 28, 2026
7bcd695
Merge origin/main (0.16.6) into feat/gcp-governance-parity
fas89 Sep 28, 2026
cc9784c
fix(platform): a dry-run never moves state, strict GCP sovereignty fa…
fas89 Sep 28, 2026
fbc31fd
fix(sovereignty): a KMS key ring and a Data Catalog taxonomy are plac…
fas89 Sep 28, 2026
b321998
fix(iac): the state-key migration attributes a state by its resources…
fas89 Sep 28, 2026
093e209
fix(apply): the Command Center run of a build-augmented apply says wh…
fas89 Sep 28, 2026
6d456fc
fix(schedule-sync): the first sync after the env-scoped DAG layout re…
fas89 Sep 28, 2026
15ccfbc
Merge origin/main (0.16.6) into feat/gcp-platform-safety
fas89 Sep 28, 2026
b6af57e
docs(changelog): the first schedule-sync after the env-scoped DAG lay…
fas89 Sep 28, 2026
d68304a
fix(iac): a Lake Formation read grant the column restrictions leave w…
fas89 Sep 28, 2026
d9a0cbf
fix(verify): BigQuery column restrictions read the table's tags befor…
fas89 Sep 28, 2026
fc37787
refactor(apply): the current apply run lives in observability, so bui…
fas89 Sep 28, 2026
7aeafe4
fix(examples): the shipped examples pass the governance checks, AWS a…
fas89 Sep 28, 2026
1f9dddc
test(iac): the byte-identity pin's refusal branch no longer names par…
fas89 Sep 28, 2026
16705ff
Merge remote-tracking branch 'origin/main' into feat/gcp-platform-safety
fas89 Sep 28, 2026
fc9adbd
docs(schedule-sync): 0.16.7 is the last release with the old DAG layout
fas89 Sep 28, 2026
8d1e62f
Merge remote-tracking branch 'origin/main' into feat/gcp-embedded-sql…
fas89 Sep 28, 2026
dc5a00c
Merge remote-tracking branch 'origin/main' into feat/gcp-governance-p…
fas89 Sep 28, 2026
5422bec
test(iac): a restriction-derived column limit carries SELECT alone, a…
fas89 Sep 28, 2026
0c0ca8b
Merge #676 (feat/gcp-platform-safety, fc9adbd) into ci/integration-gc…
fas89 Sep 28, 2026
a4feb9f
Merge #673 (feat/gcp-embedded-sql-bigquery, 8d1e62f) into ci/integrat…
fas89 Sep 28, 2026
a082bd8
Merge #674 (feat/gcp-governance-parity, 5422bec) into ci/integration-…
fas89 Sep 28, 2026
852dde4
integration: #673 on #676's undeclared BigQuery location, and #674's …
fas89 Sep 28, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
25 changes: 25 additions & 0 deletions .github/workflows/iac-tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -137,6 +137,31 @@ jobs:
-q -p no:randomly --junitxml=state-drift.xml
.venv/bin/python scripts/ci/assert_lane_coverage.py state-drift.xml \
--require "state drift=tests/iac/test_iac_state_drift_moto.py"
# The default state key names the provider, and the first apply after the
# upgrade moves the old key's state with `tofu init -migrate-state`. Only a
# real apply, move and plan prove the plan after it changes nothing.
- name: Stage 1 — per-provider state key and its migration (tofu vs moto, creds-free)
run: |
.venv/bin/python -m pytest tests/iac/test_iac_state_key_migration_moto.py \
-q -p no:randomly --junitxml=state-key-migration.xml
.venv/bin/python scripts/ci/assert_lane_coverage.py state-key-migration.xml \
--require "state key migration=tests/iac/test_iac_state_key_migration_moto.py"
# Governance parity (retention, CMEK, column restrictions, mapped
# principals): tofu validate of each governed GCP shape; a real plan and
# apply with terraform-provider-google against an in-process stand-in for
# the BigQuery REST API (the goccy emulator crashes the provider on
# apply), which is where the data-loss gate decides; and the Lake
# Formation excluded columns planned against moto. Fails if a file only
# skipped.
- name: Stage 1 — governance parity (tofu vs BigQuery stand-in and moto, creds-free)
run: |
.venv/bin/python -m pytest tests/iac/test_iac_gcp_governance_plan.py \
tests/iac/test_iac_aws_column_restrictions.py tests/iac/test_iac_gcp_governance.py \
-q -p no:randomly --junitxml=governance-parity.xml
.venv/bin/python scripts/ci/assert_lane_coverage.py governance-parity.xml \
--require "GCP governance plan=tests/iac/test_iac_gcp_governance_plan.py" \
--require "AWS column restrictions=tests/iac/test_iac_aws_column_restrictions.py" \
--require "GCP governance validate=tests/iac/test_iac_gcp_governance.py"

# -------------------------------------------------------------------
# Stage 2 — docker emulators. PR + push.
Expand Down
6 changes: 5 additions & 1 deletion .github/workflows/integration-emulated-heavy.yml
Original file line number Diff line number Diff line change
Expand Up @@ -96,10 +96,13 @@ jobs:
with:
tofu_version: 1.12.0

# ``local`` brings DuckDB: the embedded-SQL chain on the bigquery-emulator
# runs its SQL there. Without it that test skips, and the lane-coverage
# assert below fails the job.
- name: Install with heavy-emulator extras
run: |
python -m pip install --upgrade pip
python -m pip install -e ".[dev,test-emulators-heavy]"
python -m pip install -e ".[dev,test-emulators-heavy,local]"

# Skip-clean: a secret cannot gate a job-level `if:`. With no token
# the LocalStack steps skip and the GCP half still runs — EXCEPT on
Expand Down Expand Up @@ -198,6 +201,7 @@ jobs:
--require "GCP emulator e2e=tests/iac/test_iac_gcp_emulator_e2e.py"
--require "GCP cross-project=tests/iac/test_iac_cross_project_gcp_emulator.py"
--require "BigQuery happy path=tests/providers/test_bigquery_emulated_happy_path.py"
--require "BigQuery embedded-SQL chain=tests/providers/test_bigquery_emulated_embedded_sql_chain.py"
)
if [ "${{ steps.gate.outputs.localstack }}" = "true" ]; then
requirements+=(--require "AWS LocalStack=tests/providers/test_aws_localstack_happy_path.py")
Expand Down
4 changes: 2 additions & 2 deletions .secrets.baseline

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

15 changes: 15 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,21 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

### Upgrade notes

- **The first `fluid schedule-sync` after upgrading must retire the old DAG
of each env.** An env's Airflow DAGs now live in `<product-id>__<env>/` with
dag id `<product>__<env>__<build>`; 0.16.7 and earlier wrote them to
`<product-id>/` as `<product>__<build>`. Left in place, the old DAG runs
beside the new one, and both apply the same product against the same state.
With `--delete-scope product` (the default) to a local path or a `git+ssh`
repository, the sync retires them itself: it deletes from `<product-id>/`
only the DAGs rendered for the same product and env under the old id, and
keeps any other file. For every other transport it prints the step to take:
delete those DAG files at the destination once. `--delete-scope destination`
removes the old directory with the rest of what the sync does not ship.
The report's `superseded_scopes` records which case applied.

## [0.16.7] — 2026-09-28

A Lake Formation grant that hides columns now applies on AWS, not only plans.
Expand Down
3 changes: 2 additions & 1 deletion HONESTLY_TESTED.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,7 @@ Markers declared in `pyproject.toml`.
| Stage 1 LF bucket-policy modes | `tests/iac/test_iac_lakeformation_bucket_policy.py`, `tests/iac/test_iac_lakeformation_bucket_policy_plan.py`, `tests/iac/test_iac_tofu_validate.py::test_lakeformation_bucket_policy_modes_pass_tofu_validate` | ✅ — the three `bucketPolicy` modes rendered (all-grantees byte-identical to the previous emit), `tofu validate` of each, and a real `tofu plan` against moto's STS proving the default keeps a cross-account grantee's statements and plans no bucket policy for a same-account one. |
| Stage 1 LF column grants | `tests/iac/test_iac_lakeformation_column_grants.py`, `tests/iac/test_iac_lakeformation_column_permissions.py`, `tests/iac/test_iac_lakeformation_column_grants_plan.py` | ✅ — the three grant shapes rendered (no columns: `table`; `columns`: `table_with_columns.column_names`, no wildcard; `excludedColumns`: `wildcard = true` + `excluded_column_names`; refused at emit: both together, a column name the expose's schema does not declare, an exclusion of every column, and a column limit on a binding with no `location.table`), and a real `tofu plan` against moto of all three. Plan, not validate: the excluded shape without the wildcard passes `tofu validate` and fails `tofu plan` with "Missing required argument", and the file pins that too. A column-limited grant is emitted and planned with `SELECT` only: Lake Formation takes nothing else on `table_with_columns` ("Permissions modification is invalid" on the demo lab's apply, 28 Sept 2026), a `DESCRIBE` beside it is dropped because the column-limited `SELECT` already lets the principal see the table and Lake Formation refuses `DESCRIBE` to a principal holding a partial `SELECT`, and `ALTER`, `DROP`, `DELETE`, `INSERT` or `ALL` beside a column limit is refused at emit, as is a second grant for the same principal on the same table (Lake Formation refuses the table-level permissions to it, a table-level `SELECT` would read the withheld columns, and the provider reads the pair back as one). `DESCRIBE`'s grant option beside a column limit is refused (no resource can carry it). `SELECT`'s grant option is kept beside `columns` and refused beside `excludedColumns`: the guide's pages disagree (the permissions reference: "you can't include the grant option if column filtering is applied"; the console page offers it for simple column-based access), and *Data filtering limitations* is the specific rule followed: "To grant SELECT with the grant option and column filtering, you must use an include list, not an exclude list." **Not proven here**: moto stores any permission on any resource and `tofu plan` checks none, so the permission rules rest on the AWS LF developer guide and that failed apply; the fixed grant has **not yet been applied live**, and `SELECT` with the grant option on an include list has **not been tried live** (a live `GrantPermissions` of it settles the conflict). |
| Stage 1 retention + encryption at rest | `tests/iac/test_iac_aws_retention_encryption.py`, `tests/iac/test_iac_aws_retention_encryption_moto.py`, `tests/iac/test_iac_tofu_validate.py::test_retention_and_kms_pass_tofu_validate`, `tests/cli/test_verify_athena.py` (storage section) | ✅ — `exposes[].lifecycle {retention, expire: true}` and `binding.encryption.kms` (fluid-schema 0.7.6) rendered (prefix-scoped rules, never a whole-bucket rule; key, alias, default SSE-KMS with a bucket key; key policy), `tofu validate` of each shape, and a real `tofu plan` / `apply` / `destroy` against moto: the key policy evaluated against the applying account, an object written with no encryption header landing SSE-KMS under the product key, `fluid verify`'s storage checks passing and then failing on a changed rule and an SSE-S3 object, and dropping the fields deleting the lifecycle configuration and scheduling the key's deletion 7 days out. An existing key is planned against moto too: named by alias, S3 is handed its key ARN; a key pending deletion and an RSA key fail the plan on the SSE configuration's preconditions; the AWS managed key named by its key ARN (which passes the name check) fails it with registerLocation (moto records every key as customer managed, so the test sets its record of alias/aws/s3 to `AWS`). A brownfield bucket carrying an operator's lifecycle rule and SSE-S3 is adopted through `fluid apply`'s own `_adopt_existing`, and the plan shows both configurations as in-place updates. `fluid verify` is covered with botocore stubs, including a key that is not `Enabled` and lifecycle rules filtered by tag, size or a narrower prefix that expire objects sooner. **Not proven here**: moto enforces neither key policies nor Lake Formation, so Athena reading an SSE-KMS location through Lake Formation-vended credentials (the service-linked role's key-policy statement) rests on the AWS documentation and has **not been run live**. |
| Stage 1 GCP governance parity (retention, CMEK, column restrictions, mapped principals) | `tests/iac/test_iac_gcp_governance.py`, `tests/iac/test_iac_gcp_governance_plan.py`, `tests/iac/test_iac_aws_column_restrictions.py`, `tests/cli/test_verify_bigquery_governance.py` | ✅ — rendered and `tofu validate`d: partition expiration (never a table TTL), a Cloud KMS key ring, key, service-agent grant and the dataset and table keys, a Data Catalog taxonomy, policy tags and fine-grained readers, and non-authoritative dataset members with the contract's logical principals mapped by `binding.principals` (placeholders and unmapped principals refused). A real `tofu plan` / `apply` with terraform-provider-google against an in-process stand-in for the BigQuery REST API (`tests/iac/_fake_bigquery.py`; the goccy emulator crashes the provider on apply): adding retention or a key to a live table plans its replacement, which the data-loss gate refuses without `--allow-data-loss`; a new retention is in place; the governed table plans clean; the dataset keeps its own access entries. The same stand-in shows a revoked reader is a member destroy the data-loss gate lets through, and that the first apply after the authoritative access list of 0.16.6 and earlier revokes a grant removed in the same change and then plans clean. AWS: the restriction becomes each Lake Formation grant's excluded columns, a real `tofu plan` against moto accepts them, and `fluid verify`'s Lake Formation check runs against moto's stored grants (and against a listed `TableWildcard` grant). `fluid verify` on BigQuery is covered with real `google.cloud.bigquery` Table/Dataset objects and a fake Data Catalog session. **Not proven here**: nothing ran against real BigQuery, Cloud KMS, Data Catalog or Lake Formation; no emulator enforces policy tags, keys or IAM, so a denied principal's refusal, the service agent's use of the key and a load into a tagged column rest on the providers' documentation. |
| Stage 2 cross-account (two-account LocalStack) | `tests/integration/test_iac_cross_account_localstack.py` | ✅ 10 tests, gated on `FLUID_IAC_LIVE_XACCT=1` + a second LocalStack Pro carrying `lakeformation`. Producer stack applies in account A (`000000000000`); every grant names a role in account B (`222222222222`) — LocalStack derives the account from the access-key id. Verified live: LF permission + LF prefix-registration + `aws_s3_bucket_policy` all land and read back; the shared-**pool** bucket is referenced via `data.aws_s3_bucket`, its `GetObject` grant is scoped to `location.path` and its `ListBucket` carries the `s3:prefix` condition. Under `ENFORCE_IAM=1` the emitted bucket policy is proven to be the **deciding** access control: with a deliberately broad identity policy on the account-B role, the granted prefix reads and the sibling tenant's prefix is denied — and widening only the bucket policy flips that denial (causal control). **Still NOT proven here**: LF cross-account *authorization* (LocalStack rejects `GrantPermissions` under IAM enforcement even for a registered `DataLakeAdmin` — `docs/upstream-issues/localstack-lakeformation-grant-auth.md`), and cross-account Glue catalog sharing (no AWS RAM in LocalStack; account B gets `EntityNotFoundException`, which is state isolation, not a denial). |
| Stage 1 Glue catalog enrichment | `tests/iac/test_iac_aws.py::TestAwsGlueCatalogEnrichment` (6) + `test_iac_tofu_validate.py[aws]` | ✅ — `aws_glue_catalog_table.description` + per-column comments + fluid_layer/fluid_product_type/fluid_domain/fluid_version/fluid_contract/forge.pii.<col> parameters, absorbed from the retired `GlueCatalogRegistrar`. **Zero new schema fields** — reads existing `description`/`metadata.description`/`metadata.layer`/`metadata.productType`/`domain`/`fluidVersion`/`column.tags[]`/`column.description`. |
| Stage 3 Glue catalog enrichment | `tests/iac/test_iac_aws_real_e2e.py::test_real_iceberg_on_glue_round_trip` | ✅ live — Description + Parameters + per-column Comments + `forge.pii.amount` verified via boto3 GetTable |
Expand Down Expand Up @@ -245,7 +246,7 @@ fields:
| Capability | Reads from |
|---|---|
| Cross-account S3 access | `binding.governance.lakeFormation.grants[].principal` (LF block); a bucket-policy statement is paired with each grantee in another account (`bucketPolicy: cross-account`, the default; `none` / `all-grantees` in fluid-schema 0.7.6) |
| Cross-project BQ access | `metadata.policies` (existing) → dataset `access[]` via `_bq_access_entries` — ⚠️ **but see the note below: `metadata.policies` does not pass `fluid validate`** |
| Cross-project BQ access | `metadata.policies` (existing) → `google_bigquery_dataset_iam_member` (was the dataset's authoritative `access[]`) — ⚠️ **but see the note below: `metadata.policies` does not pass `fluid validate`** |
| Glue catalog enrichment | `description` / `metadata.{description,layer,productType}` / `domain` / `fluidVersion` / `column.{description,tags}` |
| Snowflake catalog enrichment | same set as Glue |

Expand Down
53 changes: 52 additions & 1 deletion docs/apply.md
Original file line number Diff line number Diff line change
Expand Up @@ -192,7 +192,15 @@ builds:
non-parquet format is CSV, which is what that provider writes). An AWS
binding naming a `location.bucket` and `location.path` reads
`s3://<bucket>/<path>/*.<ext>`, the prefix the duckdb acquisition runner
writes into and the Glue table `fluid apply` declares for it. A warehouse table, a stream, or a GCS/Azure prefix is an
writes into and the Glue table `fluid apply` declares for it. A GCP
`bigquery_table` binding reads that table (`<project>.<dataset>.<table>`, the
one `fluid apply` created, whatever `gs://` path the binding also carries):
when the build runs, the table is read through the BigQuery API
(`tabledata.list` pages, streamed) into a Parquet file under the build's
`.fluid/staging/<build>/inputs/`, the view reads that file, and the file is
removed after the build. A BigQuery `TIMESTAMP` reads as a DuckDB `TIMESTAMP`
holding the UTC wall clock, which is what the same SQL reads on the local and
aws targets. Another warehouse's table, a stream, or a GCS/Azure prefix is an
`UnreadableBindingError` naming the platform. A `{{ env.X }}` in the upstream
binding with `X` unset is an error, not an empty string.
* **Explicit inputs win.** A `properties.parameters.inputs` entry whose `name`
Expand Down Expand Up @@ -222,6 +230,49 @@ unchanged. An expose declaring `policy.privacy.masking` is refused
(`MaskingNotAppliedError`): this path does not apply masking, and cleartext
must not land silently.

When the first expose's binding is a GCP `bigquery_table`, the result is
written as Parquet under `.fluid/staging/<build>/` and one load job moves it
into that table (`WRITE_TRUNCATE`, `CREATE_NEVER`, the table's own schema),
the load the duckdb acquisition runner performs; a `gs://` `location.path` on
the binding is never written to. A failed or short load fails the build. Any
other landing this path cannot write, a `gs://` or other non-S3 URI, a GCS
bucket, an Azure, Snowflake or Databricks binding, is refused
(`EmbeddedSqlLandingError`) before the SQL runs, instead of being written to a
local file of that name. The load is recorded as a run under
`.fluid/runs/<product>/<build>/runs/`, the way the acquisition load is, so
`fluid verify` holds the table's count to the rows it landed.

Also refused before anything is read (`EmbeddedSqlLandingError`):

* a landing that resolves to one of the build's own inputs: a BigQuery table
the build reads (names compared case-insensitively; a project left to the
client matches any), or an S3 object inside a prefix it reads. The load
replaces the table, so it would overwrite another product's rows, and no
`--allow-data-loss` is ever asked for a data write;
* a further expose named in the build's `outputs` and bound to a cloud store
or a warehouse (an aws, gcp, azure, snowflake or databricks binding, or any
remote URI): this path lands only the first expose. A further local expose
or output port is not written either, and the build prints a warning.

When the contract declares `sovereignty` and the build reads or loads a
BigQuery table, the locations those reads and the load actually use are held
to it by `fluid validate`'s rules (`EmbeddedSqlSovereigntyError`): every
BigQuery binding must name its region (without one it is `US`, the IaC's
default), the landing must be outside `deniedRegions`, inside
`allowedRegions` and in the declared `jurisdiction`, and, with `dataResidency`
and no `crossBorderTransfer` (the schema's defaults), every input must be in
the landing's jurisdiction. BigQuery's `EU` and `US` multi-regions count as
EU and US. `enforcementMode: strict` (the default) refuses, `advisory` warns,
`audit` logs.

BigQuery reads and loads authenticate with Application Default Credentials
(gcloud ADC, an attached service account, or a Workload Identity Federation
`external_account` file in `GOOGLE_APPLICATION_CREDENTIALS`). With
`BIGQUERY_EMULATOR_HOST` set they go to that emulator with anonymous
credentials, so no real token is sent to it. A load job that reports no row
count (the goccy emulator's never do) is checked by counting the table after
the load, never assumed.

### Skipped builds are not success

If every build in a build-augmented mode was skipped — a missing dbt
Expand Down
Loading
Loading